Yu Peng 0002

dblp:85/5580-2 · DBLP profile ↗
← Back
31ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-3315-5581ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 8 since 2021Artificial intelligence and machine learning · 8 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6Computer networks · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Resource-Optimized Time-Multiplexed Constant Multiplication via Adjacency Matrix Modeling
abstract
This article presents TmCM-AM, a new time-multiplexed constant multiplication framework based on adjacency matrix modeling. The TmCM-AM framework provides a universal and efficient approach for digital signal processing applications using much fewer resources and with greater adaptability based on conventional methods. By transforming adder graphs into adjacency matrices and using an optimization algorithm, the proposed framework minimizes the number of required adders and multiplexers to a large degree. In particular, three mathematical properties of adjacency matrices based on properties of adder graphs are presented. Meanwhile, the adjacency matrix is employed to model time-multiplexed adder graphs in detail, making hardware architecture analysis possible through matrix computation. Finally, heuristic algorithms are used to generate the best possible solution from matrices calculated. Experimental verification through FPGA and ASIC implementations further confirms the feasibility of TmCM-AM, presenting enormous reductions in area and power dissipation, as well as delay metrics across random data and various real-life coefficient sets.
Martin Kumm, Liansheng Liu, Zhixian Zhang, Yu Peng 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2026 Scaling-Free CORDIC Optimization Based on the Time-Multiplexed Constant Multiplication
abstract
Modern digital systems in communication, control, and sensing rely on fast and accurate hardware evaluation of trigonometric functions. The classic CORDIC algorithm provides a multiplier-less solution using iterative shifts and adds, but conventional CORDIC requires many iterations and a final scaling correction that increases latency and hardware cost. Scaling-Free CORDIC (SF-CORDIC) eliminates the scaling factor by adjusting micro-rotation angles to achieve a net gain of unity. However, this approach introduces fixed constant multiplications at each iteration, which can dominate hardware resources and critical path delay. Existing methods often restrict the set of rotation angles so that constants can be implemented with simple add-shift operations. Other approaches use high-radix or hybrid schemes, but these methods limit the angle accuracy, introduce extra complexity, and still rely on numerous add-shift networks. To address these issues, a new SF-CORDIC architecture is proposed to eliminate per-iteration multipliers by using the time-multiplexed constant multiplier units implemented across all iterations. Bit-width growth is controlled via shift propagation and a minor compensation step, maintaining precision without a final scaling stage. This compact, low-latency design significantly reduces hardware area and delay compared to prior SF-CORDIC implementations while preserving high output accuracy. In a six-stage FPGA pipeline, the design reaches a maximum clock frequency of 216.03 MHz with an end-to-end latency of 27.78 ns. It uses about 1200 LUTs and 300 registers. The root-mean-square error is 6.8 × 10−5for sine and 6.4 × 10−5for cosine.
Yu Peng 0002, Xuejing Wei, Liansheng Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2026 MultiSky: Dynamic Resource Allocation Framework for High-Throughput CGRA Multitask Execution
abstract
Coarse-grained reconfigurable arrays (CGRAs) offer a promising balance between high performance and flexibility, yet dynamic resource allocation in multi-task scenarios remains challenging due to unpredictable task creation/destruction. Existing static approaches lack flexibility, while dynamic methods suffer from high latency or limited applicability. This paper presents MultiSky, a framework for CGRA multi-task dynamic resource allocation, combining a hardware controller and a software pre-mapper. The hardware controller dynamically allocates resources within hundreds of cycles by calculating tile allocation for each task via weighted averaging, and generating tile shapes using a lightweight heuristic algorithm. The software pre-mapper employs incremental compilation to pre-generate configurations, avoiding online transformation overhead. Evaluations on a real-world multi-task scenario demonstrate that MultiSky achieves 1.72× higher throughput than baselines by maintaining 82.7% average resource utilization. The framework scales efficiently with larger CGRAs and task counts, with hardware overhead decreasing to 1% for 16×16 CGRAs. These results highlight MultiSky’s ability to balance flexibility, efficiency, and practicality in dynamic computing environments.
Chenhao Xie 0001, Rui Wang 0014, Liansheng Liu, Xiyuan Peng, Yu Peng 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 FexMo: Enabling Fuse Execution Mode for Multi-task CGRAs
Chenhao Xie 0001, Chuliang Guo, Liansheng Liu, Xiyuan Peng, Datong Liu, Yu Peng 0002
MICRO7
2025 Eyelet: A Cross-Mesh NoC-Based Fine-Grained Sparse CNN Accelerator for Spatio-Temporal Parallel Computing Optimization
abstract
Fine-grained sparse convolutional neural networks (CNNs) achieve a better trade-off between model accuracy and size than coarse-grained sparse CNNs. Due to irregular data structures and unbalanced computation loads, fine-grained sparse CNNs struggle to fully leverage the performance advantages of computation and storage on general-purpose edge hardware. However, existing custom sparse accelerators are designed from the perspective of emulating a balanced load by software or computational strategies, neglecting the exploration of the computing architecture’s adaptability and parallelism for fine-grained sparse models. To address these challenges, a cross-mesh NoC-based accelerator architecture is proposed. This architecture aligns with the irregular characteristics of fine-grained sparse CNN weights and enhances the spatio-temporal parallelism of fine-grained sparse CNNs. First, a sparse multiplier unit (SMU) array and an adder array are designed to enable parallel execution of convolution multiplication and accumulation operations. Then, element-wise unroll-based nonzero weight multiplication is mapped to the SMU array to provide more flexible spatial parallelism. A horizontal and vertical cross-mesh NoC is proposed for flexible dataflow scheduling between the SMU and adder arrays to further improve temporal parallelism. This architecture allows the multiplication and accumulation operations in convolution to be decoupled and pipelined with negligible latency. Finally, the proposed accelerator architecture is implemented on the ZU9EG platform. The experimental results show that the proposed accelerator achieves frame rates of 509.9, 249.3, 100.7, 48.4, and 168.9 frames per second (FPS) for AlexNet, VGG-16, ResNet-18, MobileNet-v2, and EfficientNet, respectively. Compared with related works, this accelerator achieves inference speed and energy efficiency improvements of$1.1\times \sim 36.1\times $and$2.4\times \sim 13.4\times $, respectively.
Liansheng Liu, Yu Peng 0002, Xiyuan Peng, Heming Liu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 PreTrans: Enabling Efficient CGRA Multi-Task Context Switch Through Config Pre-Mapping and Data Transceiving
abstract
Dynamic resource allocation guarantees the performance of CGRA multi-task, but incurs a wide range of incompatible contexts (config & data) to the CGRA architecture. However, traditional context switch approaches including online config transformation and data reloading may significantly block the task to process inputs under new resource allocation decisions, resulting in the limited task throughput. To address this issue, online config transformation can be avoided if compatible configs have been prepared through offline pre-mapping, but traditional CGRA mappers require days to achieve comprehensive pre-mapping with considerable quality. Besides, online data reloading can also be eliminated through memory sharing, but the traditional arbiter-based approach has the difficulty of trading off physical complexity and memory access parallelism. PreTrans is the first system design to achieve the efficient CGRA multi-task context switch. PreTrans first avoids the online config transformation through a software incremental pre-mapper, which re-utilizes the previously finished pre-mapping results to dramatically accelerate the pre-mapping of subsequent resource allocation decisions with negligible mapping quality loss. Secondly, PreTrans replaces the traditional arbiter with a hardware data transceiver to better support the memory sharing that eliminates data reloading, which allows each tile to possess an individual memory that maximizes the access parallelism without introducing significant physical overhead. The overall evaluation demonstrates that PreTrans achieves 1.13$\sim 2.46\times$throughput improvement on pipeline and parallel multi-task scenarios, and can reach the target throughput immediately after the new resource allocation decision takes effect. Ablation study further shows that the pre-mapper is more than 3 magnitudes faster than the traditional CGRA mapper while maintaining more than 99% of the optimal mapping quality, and the data transceiver only introduces 9.02% hardware area overhead under 16×16 CGRA.
Chenhao Xie 0001, Liansheng Liu, Xiyuan Peng, Yu Peng 0002, Hailong Yang 0002, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.5
2025 Resource Optimization in Polyphase-Filter STFT Based on Time-Multiplexed Constant Multiplication
Yu Peng 0002, Zhixian Zhang, Tongrui Zhang, Liansheng Liu
IEEE Trans. Very Large Scale Integr. Syst.2
2023 Probabilistic-Attention Fusion-Based Lithium-ion Battery Pack Multivariate Prediction Method
abstract
Lithium-ion battery packs are widely employed in various applications, such as electric vehicles, energy storage system solutions, and other practical uses. Inconsistency between battery cells is one of the main factors affecting the performance of the battery pack. However, existing studies have paid insufficient attention to the prediction of inconsistency evolution trends within a battery pack. They neglect the impact of local variables on the accuracy of the results in the prediction process, while the results lack the corresponding ability to express uncertainty. To address the aforementioned issues, this paper proposes a probabilistic-attention fusion-based method for multivariate prediction of lithium-ion battery pack performance. Firstly, it utilizes LSTM to identify correlations between parameters for multiparameter prediction. Then, an attention mechanism is introduced to enhance the model's focus on key information by assigning attention weights to the input data. Finally, probabilistic modeling is incorporated to provide the model with the capability to express uncertainties. The effectiveness of the proposed method is corroborated through rigorous testing using real-world battery data obtained from laboratory experiments.
Yuhang Du, Datong Liu, Yu Peng 0002
IECON4
2019 A Fast Recursive Collaboration Representation Anomaly Detector for Hyperspectral Image
abstract
Even though collaboration representation-based detector (CRD) performs well for hyperspectral image (HSI) anomaly detection, its computational cost is too high for the widely demanded real-time applications. To reduce the computational complexity, a recursive CRD is proposed in this letter. By constructing two elementary transformation matrices in accordance with the location of the pixels, a recursive update approach is derived by a matrix inversion lemma to speed up the detector. Experimental results on two real HSI data sets show that the proposed method saves over 30% processing time without accuracy loss.
Yu Peng 0002
IEEE Geosci. Remote. Sens. Lett.2
2019 Cross-Domain Noise Impact Evaluation for Black Box Two-Level Control CPS
abstract
Control Cyber-Physical Systems (CPSs) constitute a major category of CPS. In control CPSs, in addition to the well-studied noises within the physical subsystem, we are interested in evaluating the impact of cross-domain noise : the noise that comes from the physical subsystem, propagates through the cyber subsystem, and goes back to the physical subsystem. Impact of cross-domain noise is hard to evaluate when the cyber subsystem is a black box, which cannot be explicitly modeled. To address this challenge, this article focuses on the two-level control CPS, a widely adopted control CPS architecture, and proposes an emulation based evaluation methodology framework. The framework uses hybrid model reachability to quantify the cross-domain noise impact, and exploits Lyapunov stability theories to reduce the evaluation benchmark size. We validated the effectiveness and efficiency of our proposed framework on a representative control CPS testbed. Particularly, 24.1% of evaluation effort is saved using the proposed benchmark shrinking technology.
Liansheng Liu, Stefan Winter 0001, Qixin Wang 0001, Neeraj Suri, Lei Bu, Yu Peng 0002, Xue (Steve) Liu, Xiyuan Peng
ACM Trans. Cyber Phys. Syst.7
2017 Ideal Kernel-Based Multiple Kernel Learning for Spectral-Spatial Classification of Hyperspectral Image
abstract
Using multiple types of features can effectively improve the classification accuracy of hyperspectral image (HSI). Multiple kernel learning (MKL) provides a flexible framework to fuse different features in a very natural way. In this letter, a novel MKL algorithm is proposed to integrate multiple types of features [i.e., principal components of original spectrum, multistructure morphological profiles (MPs), and multiattribute MPs] for HSI classification. The basic kernels are constructed with each feature subset separately, and the weights of basic kernels are determined by solving an optimization problem with the objective of the ideal kernel. Then, linear programming (LP) and signal sparse representation are adopted to solve the optimization problem, thus leading to two variants of the proposed algorithm, ideal kernel MKL (IKMKL)-LP and IKMKL-sparse, respectively. Experiments carried out on real hyperspectral data show that the proposed algorithms outperform several state-of-the-art MKL algorithms and reveal the capability of ranking the relevant features.
Wei Gao 0005, Yu Peng 0002
IEEE Geosci. Remote. Sens. Lett.2
2017 MWPCA-ICURD: density-based clustering method discovering specific shape original features
Qinghua Luo, Yu Peng 0002, Junbao Li, Xiyuan Peng
Neural Comput. Appl.2
2017 Low-Cost Adaptive Lebesgue Sampling Particle Filtering Approach for Real-Time Li-Ion Battery Diagnosis and Prognosis
abstract
In the past decades, fault diagnosis and prognosis (FDP) approaches were developed in the Riemann sampling (RS) framework, in which samples are taken and algorithms are executed in periodic time intervals. With the increase of system complexity, a bottleneck of real-time implementation of RS-based FDP is limited calculation resources, especially for distributed applications. To overcome this problem, a Lebesgue sampling-based FDP (LS-FDP) is proposed. LS-FDP takes samples on the fault dimension axis and provides a need-based FDP philosophy in which the algorithm is executed only when necessary. In previous LS-FDP, the Lebesgue length is a constant. To accommodate the nonlinear fault dynamics, it is desirable to execute FDP algorithm more frequently when the fault growth is fast while less frequently when fault growth is slow. This requires to change the Lebesgue length adaptively and optimize the selection of Lebesgue length based on fault state and fault growth speed. The goal of this paper is to develop an improved LS-FDP method with adaptive Lebesgue length, which enables the FDP to be executed according to fault dynamics and has low cost in terms of computation and hardware resource needed. The design and implementation of adaptive LS-FDP (ALS-FDP) based on a particle filtering algorithm are illustrated with a case study of Li-ion batteries to verify the performances of the proposed approach. The experimental results show that ALS-FDP keeps close monitoring of fault growth and is accurate and time-efficient on long-term prognosis.
Wuzhao Yan, Bin Zhang 0008, Wan-Chun Dou, Datong Liu, Yu Peng 0002
IEEE Trans Autom. Sci. Eng.5
2017 Asynchronous and Selective Transmission for DeWiring of Building Management Systems
abstract
In this paper, we show a design and implementation of a (partial) wireless building management system (BMS). Compared to the existing wired BMS, a wireless system can be much cheaper and more flexible in deployment. There are existing studies on smart and wireless BMS. Our design differs from others as the latter usually takes a re-arch approach and develops a brand new suite of protocols. However, it can take a considerably long time for re-standardization and adoption by vendors. Our design does not intend to tear down the full suite of upper layer protocols. We thus face difficulties as we need to maintain the upper layer protocols in operation and support their data traffic. The key ideas of our approach are an asynchronous-response framework to maintain the control plane of the upper layer protocols intact, and a modular design to prioritize and schedule data flow to handle link quality and throughput variations. We implemented the proposed design into a real system and evaluated the system by comprehensive experiments with real BMS controllers and software. In addition, we conducted a field deployment by integrating our system with the BMS in FG-building of The Hong Kong Polytechnic University. The system operated smoothly during 5-h deployment.
Qinghua Luo, Abraham Hang-Yat Lam, Dan Wang 0002, Dawei Pan, Daniel Wai-Tin Chan, Yu Peng 0002, Xiyuan Peng
IEEE Trans. Ind. Informatics6
2016 DEDF: lightweight WSN distance estimation using RSSI data distribution-based fingerprinting
Qinghua Luo, Xiaozhen Yan, Junbao Li, Yu Peng 0002, Yumei Tang, Dan Wang 0002
Neural Comput. Appl.4
2016 A Microcoded Kernel Recursive Least Squares Processor Using FPGA Technology
abstract
Kernel methods utilize linear methods in a nonlinear feature space and combine the advantages of both. Online kernel methods, such as kernel recursive least squares (KRLS) and kernel normalized least mean squares (KNLMS), perform nonlinear regression in a recursive manner, with similar computational requirements to linear techniques. In this article, an architecture for a microcoded kernel method accelerator is described, and high-performance implementations of sliding-window KRLS, fixed-budget KRLS, and KNLMS are presented. The architecture utilizes pipelining and vectorization for performance, and microcoding for reusability. The design can be scaled to allow tradeoffs between capacity, performance, and area. The design is compared with a central processing unit (CPU), digital signal processor (DSP), and Altera OpenCL implementations. In different configurations on an Altera Arria 10 device, our SW-KRLS implementation delivers floating-point throughput of approximately 16 GFLOPs, latency of 5.5μ S , and energy consumption of 10 − 4 J, these being improvements over a CPU by factors of 12, 17, and 24, respectively.
Yeyong Pang, Yu Peng 0002, Xiyuan Peng, Nicholas J. Fraser, Philip H. W. Leong
ACM Trans. Reconfigurable Technol. Syst.3
2015 Locality structure preserving based feature selection for prognostics
abstract
Feature selection in data-driven modelling is an important research topic for prognostics. The performance of prediction model may vary considerably under different feature subsets. Hence it is important to devise a systematic feature selection method, which offers the guidance for choosing the mos t representative features for prognostics. Nowadays, feature selection algorithms in the field of prognostics are largely studied to the type of learning: supervised or unsupervised, which leads to poor generalization between different prognostics applications. In this paper, a unified feature selection method, called locality structure preserving based feature selection (LSPFS), is developed to improve the robustness and accuracy of prognostics under both unsupervised and supervised learning conditions. In LSPFS, the local structure of original data is constructed according to the similarity between data points, and the representative features are selected based on their ability to preserve the local structure. Moreover, by designing different local structure via local information and actual degradation information of the data, the introduced method can unify supervised and unsupervised feature selection, and enable their joint study under a general framework. Experiments on NASA turbofan engine simulation dataset and lithium-ion battery dataset are conducted to test and evaluate the proposed algorithm.
Yu Peng 0002, Datong Liu, Junbao Li
Intell. Data Anal.1
2015 A Health Indicator Extraction and Optimization Framework for Lithium-Ion Battery Degradation Modeling and Prognostics
abstract
Maximum releasable capacity and internal resistance are often used as the health indicators (HIs) of a lithium-ion battery for degradation modeling and estimation of remaining useful life (RUL). However, the maximum releasable capacity is usually difficult to estimate in online applications due to complex operating conditions in the field. Moreover, measuring the internal resistance is too expensive to be implemented on-line. In this paper, an HI extraction and optimization framework requiring only the operating parameters of lithium-ion batteries is proposed for battery degradation modeling and RUL estimation. The framework carries out raw HI extraction, transformation, correlation analysis, and verification and evaluation to achieve HI enhancement. In particular, the Box-Cox transformation is adopted to improve the correlation between the extracted HI and the battery's actual degradation state. To estimate the battery's RUL using the enhanced HI, an optimized relevance vector-machine algorithm is utilized, which can be performed in a flexible and agile way. Experimental studies using two different industrial testing data sets illustrate the high efficiency and adaptability of the proposed framework in lithium-ion battery degradation modeling and RUL estimation.
Datong Liu, Jianbao Zhou, Haitao Liao, Yu Peng 0002, Xiyuan Peng
IEEE Trans. Syst. Man Cybern. Syst.4
2014 High performance relevance vector machine on HMPSoC
abstract
Relevance Vector Machine (RVM) with the uncertainty expressing ability has spawned broad applications in Prognostic and Health Management (PHM). However computationally intensive intrinsic nature of RVM greatly limits its usage. This paper presents a software and hardware co-design approach based on HMPSoC technology, which efficiently exploited sequential and parallel nature of RVM. Multi-channel and pipelined hardware architecture for the acceleration of kernel formulation and intermediate values calculation is proposed. The hardware that wrapped with AXI-Stream interface is integrated into HMPSoC as an acceleration engine. We implement the design on an on-board PHM prototype platform with a Xilinx Zynq XC7Z020 AP SoC. The experiment results show 5.3× and 46.8× speed up in terms of the time cost than the RVM running on PC with a Xeon 5620 processor and ARM Cortex A9 processor. The energy consumption is reduced by 153.0× and 37.3×, respectively.
Yongfu He, Yu Peng 0002, Yeyong Pang, Jingyue Pang
FPT3
2014 Implementation of LS-SVM with HLS on Zynq
abstract
In recent years, implementing a complicated algorithm in an embedded system, especially in a heterogeneous computing system, has gained more and more attention in many fields. The problem is that the implementation needs amounts of coding and debugging work, even if the algorithm has been verified by high-level language in PC environment. Our demo presents a method which can reduce the time of developing an algorithm in an embedded and heterogeneous system by high level synthesis method. Least Square Support Vector Machine(LS-SVM) algorithm was realized on Zynq platform by translating high-level language to Hardware Description Language(HDL). Basing on the feature of the developed heterogeneous system and the theory of LS-SVM, three parts were implemented to realize LS-SVM which includes a generating Kernel Matrix module, a solving linear equations module and a forecasting module. The first and the third parts have been placed in ARM processor by C language. Moreover, considering that the second parts was compute-intensive, it has been realized in logic resource by using high-level language. To manage data communication and computing task, an SOPC system has been designed on Zynq platform which worked in PXI chassis. Experiments demonstrate that the design method is feasible and can be used for the implementation of other complicate algorithm. The precision and time consumption in computing are given at the end.
Yeyong Pang, Yu Peng 0002
FPT4
2014 Lithium-ion battery remaining useful life estimation based on fusion nonlinear degradation AR model and RPF algorithm
Datong Liu, Jie Liu 0015, Yu Peng 0002, Limeng Guo, Michael G. Pecht
Neural Comput. Appl.4
2014 A novel hybridization of echo state networks and multiplicative seasonal ARIMA model for mobile communication traffic series forecasting
Yu Peng 0002, Miao Lei, Junbao Li, Xiyuan Peng
Neural Comput. Appl.1
2013 A low latency kernel recursive least squares processor using FPGA technology
abstract
The kernel recursive least squares (KRLS) algorithm performs non-linear regression in an online manner, with similar computational requirements to linear techniques. In this paper, an implementation of the KRLS algorithm utilising pipelining and vectorisation for performance; and microcoding for reusability is described. The design can be scaled to allow tradeoffs between capacity, performance and area. Compared with a central processing unit (CPU) and digital signal processor (DSP), the processor improves on execution time, latency and energy consumption by factors of 5, 5 and 12 respectively.
Yeyong Pang, Yu Peng 0002, Nicholas J. Fraser, Philip H. W. Leong
FPT3
2013 Minimizing Building Electricity Costs in a Dynamic Power Market: Algorithms and Impact on Energy Conservation
abstract
Energy is a global concern and the electricity bills nowadays are leading to unprecedented costs. Electricity price is market-based and dynamic. In this paper, we investigate how to cut the electricity bills of commercial buildings in a dynamic power market. The building thermal systems (e.g., air-conditioning), which dominate electricity bills, has a special property of thermal storage, i.e., the energy will not immediately dissipate from thermal air/water. Intuitively, with storage, the energy can be "stored" in the thermal system, making it possible to purchase electricity in low price and use it at appropriate time. The building thermal supply and electricity purchasing surely depends on human activities that the building should support such as class and meeting schedules. To minimize electricity bills, we develop a holistic planning of electricity purchasing schedule with thermal storage management, and appropriate room assignment schedules for classes/meetings usage. The computing algorithms require inputs of physical modeling on energy consumption. We develop wireless sensing systems to collect fine-grained data which are used to assist the cross-disciplinary physical modeling. We conduct validation through real experiments. We formulate an optimization problem and show that it is NP-complete. Our primary focus is to minimize electricity bills, which matches the incentives of the commercial buildings. We show that this does not coincide with energy conservation. We further investigate the relationship of minimization of electricity bills and minimization of energy consumption. We develop algorithms for our problem and our evaluation shows that we can achieve a 40% cost reduction.
Dawei Pan, Dan Wang 0002, Jiannong Cao 0001, Yu Peng 0002, Xiyuan Peng
RTSS4
2013 Quasiconformal kernel common locality discriminant analysis with application to breast cancer diagnosis
Junbao Li, Yu Peng 0002, Datong Liu
Inf. Sci.2
2013 A study towards applying thermal inertia for energy conservation in rooms
abstract
We are in an age where people are paying increasing attention to energy conservation around the world. The heating and air-conditioning systems of buildings introduce one of the largest chunks of energy expenses. In this article, we make a key observation that after a meeting or a class ends in a room, the indoor temperature will not immediately increase to the outdoor temperature. We call this phenomenon thermal inertia . Thus, if we arrange subsequent meetings in the same room rather than in a room that has not been used for some time, we can take advantage of such undissipated cool or heated air and conserve energy. Though many existing energy conservation solutions for buildings can intelligently turn off facilities when people are absent, we believe that understanding thermal inertia can lead system designs to go beyond on-and-off-based solutions to a wider realm. We propose a framework for exploring thermal inertia in room management. Our framework contains two components. (1) The energy-temperature correlation model captures the relation between indoor temperature change and energy consumption. (2) The energy-aware scheduling algorithms: given information for the relation between energy and temperature change, energy-aware scheduling algorithms arrange meetings not only based on common restrictions, such as meeting time and room capacity requirement, but also energy consumptions. We identify the interface between these components so further works towards same on direction can make efforts on individual components. We develop a system to verify our framework. First, it has a wireless sensor network to collect indoor, outdoor temperature and electricity expenses of the heating or air-conditioning devices. Second, we build an energy-temperature correlation model for the energy expenses and the corresponding room temperature. Third, we develop room scheduling algorithms. In detail, we first extend the current sensor hardware so that it can record the electricity expenses in re-heating or re-cooling a room. As the sensor network needs to work unattendedly, we develop a hardware board for long-range communications so that the Imote2 can send data to a remote server without a computer relay close by. An efficient two-tiered sensor network is developed with our extended Imote2 and TelosB sensors. We apply laws of thermodynamics and build a correlation model of the energy needed to re-cool a room to a target temperature. Such model requires parameter calibration and uses the data collected from the sensor network for model refinement. Armed with the energy-temperature correlation model, we develop an optimal algorithm for a specified case, and we further develop two fast heuristics for different practical scenarios. Our demo system is validated with real deployment of a sensor network for data collection and thermodynamics model calibration. We conduct a comprehensive evaluation with synthetic room and meeting configurations, as well as real class schedules and classroom topologies of The Hong Kong Polytechnic University, academic calendar year of Spring 2011. We observe 20% energy savings as compared with the current schedules.
Yi Yuan 0005, Dawei Pan, Dan Wang 0002, Xiaohua Xu 0002, Yu Peng 0002, Xiyuan Peng, Peng-Jun Wan
ACM Trans. Sens. Networks5
2012 Thermal Inertia: Towards an energy conservation room management system
abstract
We are in an age where people are paying increasing attention to energy conservation around the world. The heating and air-conditioning systems of buildings introduce one of the largest chunk of energy expenses. In this paper, we make a key observation that after a meeting or a class ends in a room, the indoor temperature will not immediately increase to the outdoor temperature. We call this phenomenon Thermal Inertia. Thus, if we arrange subsequent meetings in the same room; than a room that has not been used for some time, we can take advantage of such un-dissipated cool or heated air and conserve energy. We develop a green room management system with three main components. First, it has a wireless sensor network to collect indoor, outdoor temperature and electricity expenses of the air-conditioning devices. Second, we build an energy-temperature correlation model for the energy expenses and the corresponding room temperature. Third, we develop room scheduling algorithms. Our system is validated with real deployment of a sensor network for data collection and thermodynamics model calibration. We conduct a comprehensive evaluation with synthetic room and meeting configurations. We observe a 30% energy saving as compared with the current schedules.
Dawei Pan, Yi Yuan 0005, Dan Wang 0002, Xiaohua Xu 0002, Yu Peng 0002, Xiyuan Peng, Peng-Jun Wan
INFOCOM5
2011 Accelerating on-line training of LS-SVM with run-time reconfiguration
abstract
Least Squares Support Vector Machines(LS-SVM), which is an efficient supervised learning tool, has been widely applied to real-time on-line data processing in many fields. However, the on-line training of LS-SVM always suffers from huge computation which greatly limits its practicability especially in embedded systems. By leveraging the flexibility and high degree parallelism offered by reconfigurable fabrics, we propose a Run-Time Reconfiguration(RTR) framework to accelerate the on-line training of LS-SVM. To realize maximum computational parallelism, we divide the training process into two parts, the kernel matrix formulation and the least-square problem solving. We dynamically load these two parts into FPGA with RTR under the control of the embedded PowerPC. In the kernel matrix formulation part, we design a piecewise linear interpolation method to realize the radial basis function. In the least-square problem solving part, the modified Cholesky Decomposition is introduced to avoid the latency caused by square roots operations. The whole design is tested on Virtex XC5VFX130T with a 150MHz clock. The experiments show appealing speed up which ranges from 6~218× over a Xeon CPU implementation on five different sized datasets. From time cost percentage analysis, our proposed architecture can be effectively applied to LS-SVM training in more than 1000 samples applications.
Yu Peng 0002, Guangquan Zhao, Xiyuan Peng
FPT2
2011 Analog Circuit Fault Diagnosis with Echo State Networks Based on Corresponding Clusters
Xiyuan Peng, Miao Lei, Yu Peng 0002
ISNN (1)4
2011 Anti Boundary Effect Wavelet Decomposition Echo State Networks
Jianmin Wang 0004, Yu Peng 0002, Xiyuan Peng
ISNN (1)2
2003 Virtual instrument parameter calibration with particle swarm optimization
abstract
In virtual instrument designs and applications, lots of functional parameters can be set through software methods. Currently, most parameter settings methods are lightly linked with the knowledge of instruments and basic principles related to specific applications. However, it is difficult for some end users to deal with those advanced operations. By adopting the particle swarm optimization (PSO) algorithm, the adaptive set and calibration of instrument parameters can be achieved by software with computational intelligence. Experiments and applications showed that the adaptive parameter calibration method based on the PSO can enhance the effectiveness of debugging and maintenance of virtual instrument and test system.
Yu Peng 0002, Xiyuan Peng, Shengwei Meng
SIS1