VLDB 2026 Research / reviewers in the wild / expert
Xiyuan Peng
dblp:91/988
· DBLP profile ↗
42ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0001-7424-1008ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 10 since 2021Computer networks · 7 · 1 since 2021Artificial intelligence and machine learning · 6 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MultiSky: Dynamic Resource Allocation Framework for High-Throughput CGRA Multitask ExecutionabstractCoarse-grained reconfigurable arrays (CGRAs) offer a promising balance between high performance and flexibility, yet dynamic resource allocation in multi-task scenarios remains challenging due to unpredictable task creation/destruction. Existing static approaches lack flexibility, while dynamic methods suffer from high latency or limited applicability. This paper presents MultiSky, a framework for CGRA multi-task dynamic resource allocation, combining a hardware controller and a software pre-mapper. The hardware controller dynamically allocates resources within hundreds of cycles by calculating tile allocation for each task via weighted averaging, and generating tile shapes using a lightweight heuristic algorithm. The software pre-mapper employs incremental compilation to pre-generate configurations, avoiding online transformation overhead. Evaluations on a real-world multi-task scenario demonstrate that MultiSky achieves 1.72× higher throughput than baselines by maintaining 82.7% average resource utilization. The framework scales efficiently with larger CGRAs and task counts, with hardware overhead decreasing to 1% for 16×16 CGRAs. These results highlight MultiSky’s ability to balance flexibility, efficiency, and practicality in dynamic computing environments. Chenhao Xie 0001, Rui Wang 0014, Liansheng Liu, Xiyuan Peng, Yu Peng 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | FexMo: Enabling Fuse Execution Mode for Multi-task CGRAs
Chenhao Xie 0001, Chuliang Guo, Liansheng Liu, Xiyuan Peng, Datong Liu, Yu Peng 0002 |
MICRO | 5 |
| 2025 | DynMap: A Heuristic Dynamic Mapper for CGRA Multitask Dynamic Resource AllocationabstractCoarse-grained reconfigurable architecture (CGRA) has received increasing attention in both industry and academia due to its comprehensive advantages of performance, energy efficiency, and flexibility. To improve the resource utilization and handle the mixing workloads in the real-world, multiple tasks sharing the whole CGRA has became an important technical trend, and the varying resource requirements throughout their life cycles also makes run-time dynamic resource allocation (DRA) necessary for higher-multitask throughput. As the key stage of DRA, dynamic mapping (DM) is responsible for mapping kernels within each task to the dynamically allocated CGRA resources. However, existing DM methods have difficulty to balance the mapping time and the mapping quality, resulting in a significant gap between the actual and the optimal task throughput. To address the challenge, we propose DynMap, a heuristic dynamic mapper for CGRA multitask DRA. With the support of specialized scheduling and routing schemes, DynMap heuristically references the placement tendency in the static mapping result to dramatically save the mapping time, while maintaining the high-mapping quality by minimizing the possibility of resource conflicts. Experimental evaluation demonstrates DynMap not only achieves the average 1.17 ms mapping time and average 98.33% of the optimal mapping quality on different CGRA architectures, but also reaches average 98.85% of the optimal task throughput expected by different CGRA multitask DRA scenarios, reducing the gap between actual and optimal task throughput average$31.75\times $smaller than that of the current methods. Chenhao Xie 0001, Liansheng Liu, Xiyuan Peng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | LVDE: A Lightweight Threshold Voltage Distribution Estimation Strategy for High-Performance 3-D Nand Flash MemoryabstractLow-density parity-check (LDPC) codes are now broadly employed as error correction code (ECC) solutions in NAND flash memories. With the rapid development of high-performance three-dimensional (3-D) NAND memories, mitigating the read latency caused by LDPC decoding attempts has emerged as an investigation focus. Existing investigations have demonstrated that efficient memory sensing is crucial to enhancing the efficiency of LDPC decoding attempts. The optimal memory sensing parameter options are determined by the threshold voltage distributions of NAND flash, which are difficult to be dynamically extracted during data retention. To address the issue, this article proposes an online lightweight strategy for threshold voltage distribution estimation, named LVDE. LVDE suggests a probe wordline design that uses one-read sampling to measure the threshold voltage distributions during long-term retention. Further, LVDE develops a methodology for the cross-layer estimation of threshold voltage distributions, thereby utilizing the measurements of the probe wordline to estimate those of each layer. LVDE is implemented on real flash chips to compare with numerous state-of-the-art strategies. Experimental results demonstrate that LVDE can significantly enhance the efficiency of LDPC decoding attempts and thus improve the read performance of 3-D NAND flash memory. Zhelong Piao, Debao Wei, Huqi Xiang, Liyan Qiao, Xiyuan Peng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | Eyelet: A Cross-Mesh NoC-Based Fine-Grained Sparse CNN Accelerator for Spatio-Temporal Parallel Computing OptimizationabstractFine-grained sparse convolutional neural networks (CNNs) achieve a better trade-off between model accuracy and size than coarse-grained sparse CNNs. Due to irregular data structures and unbalanced computation loads, fine-grained sparse CNNs struggle to fully leverage the performance advantages of computation and storage on general-purpose edge hardware. However, existing custom sparse accelerators are designed from the perspective of emulating a balanced load by software or computational strategies, neglecting the exploration of the computing architecture’s adaptability and parallelism for fine-grained sparse models. To address these challenges, a cross-mesh NoC-based accelerator architecture is proposed. This architecture aligns with the irregular characteristics of fine-grained sparse CNN weights and enhances the spatio-temporal parallelism of fine-grained sparse CNNs. First, a sparse multiplier unit (SMU) array and an adder array are designed to enable parallel execution of convolution multiplication and accumulation operations. Then, element-wise unroll-based nonzero weight multiplication is mapped to the SMU array to provide more flexible spatial parallelism. A horizontal and vertical cross-mesh NoC is proposed for flexible dataflow scheduling between the SMU and adder arrays to further improve temporal parallelism. This architecture allows the multiplication and accumulation operations in convolution to be decoupled and pipelined with negligible latency. Finally, the proposed accelerator architecture is implemented on the ZU9EG platform. The experimental results show that the proposed accelerator achieves frame rates of 509.9, 249.3, 100.7, 48.4, and 168.9 frames per second (FPS) for AlexNet, VGG-16, ResNet-18, MobileNet-v2, and EfficientNet, respectively. Compared with related works, this accelerator achieves inference speed and energy efficiency improvements of$1.1\times \sim 36.1\times $and$2.4\times \sim 13.4\times $, respectively. Liansheng Liu, Yu Peng 0002, Xiyuan Peng, Heming Liu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | PreTrans: Enabling Efficient CGRA Multi-Task Context Switch Through Config Pre-Mapping and Data TransceivingabstractDynamic resource allocation guarantees the performance of CGRA multi-task, but incurs a wide range of incompatible contexts (config & data) to the CGRA architecture. However, traditional context switch approaches including online config transformation and data reloading may significantly block the task to process inputs under new resource allocation decisions, resulting in the limited task throughput. To address this issue, online config transformation can be avoided if compatible configs have been prepared through offline pre-mapping, but traditional CGRA mappers require days to achieve comprehensive pre-mapping with considerable quality. Besides, online data reloading can also be eliminated through memory sharing, but the traditional arbiter-based approach has the difficulty of trading off physical complexity and memory access parallelism. PreTrans is the first system design to achieve the efficient CGRA multi-task context switch. PreTrans first avoids the online config transformation through a software incremental pre-mapper, which re-utilizes the previously finished pre-mapping results to dramatically accelerate the pre-mapping of subsequent resource allocation decisions with negligible mapping quality loss. Secondly, PreTrans replaces the traditional arbiter with a hardware data transceiver to better support the memory sharing that eliminates data reloading, which allows each tile to possess an individual memory that maximizes the access parallelism without introducing significant physical overhead. The overall evaluation demonstrates that PreTrans achieves 1.13$\sim 2.46\times$throughput improvement on pipeline and parallel multi-task scenarios, and can reach the target throughput immediately after the new resource allocation decision takes effect. Ablation study further shows that the pre-mapper is more than 3 magnitudes faster than the traditional CGRA mapper while maintaining more than 99% of the optimal mapping quality, and the data transceiver only introduces 9.02% hardware area overhead under 16×16 CGRA. Chenhao Xie 0001, Liansheng Liu, Xiyuan Peng, Yu Peng 0002, Hailong Yang 0002, Depei Qian 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2024 | Sub-Nyquist Frequency and DOA Estimation With DOF Extension Arrayed MWCabstractWe conducted frequency and direction of arrival estimations when more signals than sensors were present using a sub-Nyquist sampling system. We proposed a degrees of freedom (DOF) extension arrayed modulated wideband converter method based on a novel joint space-frequency sparse model, which can increase the DOF of the receiving array by appropriately increasing the sub-Nyquist sampling rate, was proposed. Our method does not require any special sparse array arrangement or an increase in the number of channels, thereby reducing cost and simplifying implementation. Coherent signals and source ambiguity were theoretically analyzed, and rigorous reconstruction conditions were provided for each case. Numerical results were provided to verify the effectiveness of the proposed method. Siyi Jiang, Zhiliang Wei, Ning Fu, Liyan Qiao, Xiyuan Peng |
IEEE Signal Process. Lett. | 5 |
| 2024 | Decode-and-forward cooperative transmission in wireless sensor networks based on physical-layer network coding
Bo Li 0034, Gongliang Liu, Ruofei Ma, Xiyuan Peng |
Wirel. Networks | 6 |
| 2022 | Experimental Verification and Analysis of the Acceleration Factor Model for 3-D nand Flash MemoryabstractWith the increasing density of flash memory, its service life is declining. As the most important technical index of a flash device, flash retention parameters often need to consume a substantial amount of time to be tested. With a view toward reducing the test period, the most popular method currently is to speedup the retention loss of flash memory at high temperatures. However, since the retention loss of flash memory is the result of a mixture of multiple failure mechanisms, the Arrhenius model with a single apparent activation energy ($E_{aa}$) has significant limitations in calculating the high-temperature acceleration factor for the retention loss of flash memory. In this article, based on a large number of real experiments, we have studied the retention characteristics of the flash memory in high-temperature environments and closely tracked the change of$E_{aa}$at different temperatures, retention time, and program/erase (P/E) cycle count. The experimental results show that the equivalent acceleration factor (AF) grows rapidly with retention time at the beginning of flash retention and that flash with a high P/E cycle count has a faster AF growth rate. Debao Wei, Liyan Qiao, Xiyuan Peng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | TCSE: A Target Cell States Elimination Coding Strategy for Highly Reliable Data Storage Based on 3-D nand Flash MemoryabstractWith the wide application of NAND flash storage systems in read-intensive memory, the corresponding reliability enhancement strategies for mitigating read disturb become the focus of investigations in recent years. The prior investigations have reported the strategies based on data modulation and verified their effectiveness in mitigating retention loss. The most significant advantage of these strategies is that they can often achieve significant reliability enhancement effect with great read performance, since they are usually based on asymmetric coding. This feature means that they have the potential to be ideal reliability enhancement strategies for read-intensive memory applications. To propose a universal and highly reliable data storage strategy, this article first observes the error modes at the cell-state level under the reliability stresses of retention loss and read disturb with floating-gate (FG) 3-D triple-level (TLC) NAND flash, and then proposes a target cell states elimination (TCSE) coding strategy for further restraining bit errors. In addition, this article for the first time reports the extra bit errors generated in the decoding process of the storage strategies based on data modulation, and defines the concept of transfer factor (TF) for evaluation. By using the proposed TCSE, the experimental results show that the overall bit error rate (BER) can be reduced by 80%–90% on average, compared with the raw random data pattern. Debao Wei, Zhelong Piao, Liyan Qiao, Xiyuan Peng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | An Online Noninvasive Estimation Method of Electrolytic Capacitor for Boost ConvertersabstractThe aluminum electrolytic capacitor (AEC) plays a key role in power electronic converter, but it is the most vulnerable component in the DC-DC converter. Throughout its short lifespan, the equivalent series resistance (ESR) will increase with the AEC aging. Therefore, it is important to monitor the ESR for ensuring the normal operation of a DC-DC converter. For a boost converter operating on discontinuous conduction mode and continuous conduction mode, an online noninvasive ESR estimation method was proposed in this paper. Based on the expression derivation of the output voltage, the ESR can be estimated by sampling the output ripple voltage at three specific moments. The proposed method is a noninvasive method, which does not require the current sensors. This method can be used in the discontinuous conduction mode and critical conduction mode. The simulation results verify the effectiveness of the proposed method. Chuanfeng Li, Yang Yu 0015, Qingxin Liu, Xiyuan Peng |
IECON | 4 |
| 2021 | TSV Fault Modeling and A BIST Solution for TSV Pre-bond TestabstractAs semiconductor technology develops, three-dimensions integrated circuits (3-D ICs) is thought as a viable solution to further improve the performance of ICs. Throughsilicon-vias (TSVs) are the most important device in 3-D ICs. Therefore, TSVs testing is a very critical flow in 3-D ICs manufacturing. A new pre-bond TSVs test solution is proposed in this paper. The new test structure includes an improved ring oscillator and a voltage-divider structure. Compared with conventional TSV test structure, the output of new test structure is more sensitive to TSV faults. We prove the effectiveness of the new test structure by HSPICE simulation. Yang Yu 0015, Xiyuan Peng |
VTS | 3 |
| 2020 | An Adaptive Carrier and Symbol Synchronization Approach for QPSK SignalabstractCarrier and symbol synchronization are the key prerequisite for the accurate reception of the Quadrature Phase Shift Keying (QPSK) signal. The traditional synchronization methods have some intrinsic problems, such as slow speed, difficult to design the coefficients for the loop filter. In this paper, an adaptive carrier and symbol joint synchronization method is proposed, which applies the idea of segmented synchronization. The proposed approach divides the synchronization process into two stages: coarse stage and fine stage. In the coarse stage, loop filters with large bandwidth are used to fast capture the range of the carrier frequency offset (CFO) and symbol timing offset (STO). While, in the fine stage, filters with small bandwidth are adopted to accurately calculate and track the offsets. In addition, the proposed method can track the variation of the offsets dynamically, thus, it can adaptively identify the synchronization status and select the suitable structure and coefficients for different synchronization stages. Simulation results show that, compared with the traditional joint method, our proposed method can take into account the synchronization speed and stability simultaneously, and the speed and stability can be increased by 1.2 times and 21.77%, respectively, when the CFO is 1.5% of the symbol rate and STO is 0.5% of the symbol rate. Our proposed method can be used to solve the synchronization problem for the harsh environment, especially for the field of satellite communications, which has extremely large frequency and timing offsets. Zhiming Yang 0001, Xin Ji, Zhiyi Song, Liyan Qiao, Yang Yu 0015, Xiyuan Peng |
IWCMC | 6 |
| 2020 | A Post-Bond TSV Test Method Based on RGC Parameters MeasurementabstractIn this paper, we propose a novel post-bond through-silicon via (TSV) test method for post-bond TSV manufacture test. The proposed method can measure the RGC parameters of TSVs using switched-capacitor circuits and calibration method and detect TSV leakage fault, open fault, high-resistance fault, and delay fault based on the measured RGC parameters. The test process has a high resolution and is robust. The fault-detection effectiveness is evaluated via HSPICE simulations. The results of the TSV capacitance measurement have a relative error within 1%, the TSV resistance measurement results have a relative error within 5%, and the TSV conductance measurement results have a relative error within 8.5%. We also perform an Monte Carlo simulation to prove that even under process variation, the proposed method can obtain a TSV capacitance within an absolute error of ±3 fF, obtain a TSV resistance with a relative error of no more than 18.4%, and obtain a TSV conductance of less than 100 ${\mu \Omega }^{{-1}}$ with a relative error of no more than 13%. Yang Yu 0015, Xu Fang 0002, Xiyuan Peng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | Cross-Domain Noise Impact Evaluation for Black Box Two-Level Control CPSabstractControl Cyber-Physical Systems (CPSs) constitute a major category of CPS. In control CPSs, in addition to the well-studied noises within the physical subsystem, we are interested in evaluating the impact of cross-domain noise : the noise that comes from the physical subsystem, propagates through the cyber subsystem, and goes back to the physical subsystem. Impact of cross-domain noise is hard to evaluate when the cyber subsystem is a black box, which cannot be explicitly modeled. To address this challenge, this article focuses on the two-level control CPS, a widely adopted control CPS architecture, and proposes an emulation based evaluation methodology framework. The framework uses hybrid model reachability to quantify the cross-domain noise impact, and exploits Lyapunov stability theories to reduce the evaluation benchmark size. We validated the effectiveness and efficiency of our proposed framework on a representative control CPS testbed. Particularly, 24.1% of evaluation effort is saved using the proposed benchmark shrinking technology. Liansheng Liu, Stefan Winter 0001, Qixin Wang 0001, Neeraj Suri, Lei Bu, Yu Peng 0002, Xue (Steve) Liu, Xiyuan Peng |
ACM Trans. Cyber Phys. Syst. | 9 |
| 2019 | TSV Prebond Test Method Based on Switched CapacitorsabstractAs an integrated circuit (IC) technology develops, 3-D stacked ICs based on through-silicon-via (TSV) technology have attracted the attention of the semiconductor industry. In 3-D stacked ICs, the TSV forms signal paths that vertically connect different ICs. The signal integrity largely depends on the quality of the TSV. A key challenge to the commercial viability of 3-D ICs is developing a method of accomplishing prebond TSV tests. To address this problem, we propose a prebond TSV test method with easily applied Design for Testability (DfT) structures and a simple fault-detection process. The proposed method is based on switched-capacitor circuits, which can detect TSV leakage faults, open faults, and high-resistance faults with high resolution. The fault-detection effectiveness is evaluated via HSPICE simulations. We also present an overall assessment of the test resolution, test time, and DfT area cost. Xu Fang 0002, Yang Yu 0015, Xiyuan Peng |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | DNF-SC-PNC: a new physical-layer network coding scheme for two-way relay channels with asymmetric data length
Bo Li 0034, Gongliang Liu, Xin Liu 0009, Xiyuan Peng |
Wirel. Networks | 5 |
| 2018 | NBTI and Power Reduction Using a Workload-Aware Supply Voltage Assignment Approach
Yang Yu 0015, Zhiming Yang 0001, Xiyuan Peng |
J. Electron. Test. | 4 |
| 2018 | Physical-Layer Network Coding Scheme over Asymmetric Rayleigh Fading Two-Way Relay Channels
Bo Li 0034, Xuesong Ding, Gongliang Liu, Xiyuan Peng |
Mob. Networks Appl. | 5 |
| 2018 | The Performance of Physical-Layer Network Coding in Asymmetric Rayleigh Fading Two-Way Relay Channels
Bo Li 0034, Gongliang Liu, Xiyuan Peng |
Mob. Networks Appl. | 4 |
| 2017 | MWPCA-ICURD: density-based clustering method discovering specific shape original features
Qinghua Luo, Yu Peng 0002, Junbao Li, Xiyuan Peng |
Neural Comput. Appl. | 4 |
| 2017 | Asynchronous and Selective Transmission for DeWiring of Building Management SystemsabstractIn this paper, we show a design and implementation of a (partial) wireless building management system (BMS). Compared to the existing wired BMS, a wireless system can be much cheaper and more flexible in deployment. There are existing studies on smart and wireless BMS. Our design differs from others as the latter usually takes a re-arch approach and develops a brand new suite of protocols. However, it can take a considerably long time for re-standardization and adoption by vendors. Our design does not intend to tear down the full suite of upper layer protocols. We thus face difficulties as we need to maintain the upper layer protocols in operation and support their data traffic. The key ideas of our approach are an asynchronous-response framework to maintain the control plane of the upper layer protocols intact, and a modular design to prioritize and schedule data flow to handle link quality and throughput variations. We implemented the proposed design into a real system and evaluated the system by comprehensive experiments with real BMS controllers and software. In addition, we conducted a field deployment by integrating our system with the BMS in FG-building of The Hong Kong Polytechnic University. The system operated smoothly during 5-h deployment. Qinghua Luo, Abraham Hang-Yat Lam, Dan Wang 0002, Dawei Pan, Daniel Wai-Tin Chan, Yu Peng 0002, Xiyuan Peng |
IEEE Trans. Ind. Informatics | 7 |
| 2016 | NBTI-aware adaptive minimum leakage vector selection using a linear programming approach
Zhiming Yang 0001, Yang Yu 0015, Xiyuan Peng |
Integr. | 4 |
| 2016 | A page-granularity wear-leveling (PGWL) strategy for NAND flash memory-based sink nodes in wireless sensor networks
Debao Wei, Libao Deng, Liyan Qiao, Xiyuan Peng |
J. Netw. Comput. Appl. | 5 |
| 2016 | Window-Based Three-Dimensional Aggregation for Stereo MatchingabstractThis paper presents a window-based three-dimens-ional (3-D) aggregation technique, which can approximate the surfaces of all kinds of objects, in stereo matching. The 3-D aggregation, which means to aggregate in 3-D surfaces, is implemented by decomposing the adaptive support window into horizontal segments; we allow the disparity to change smoothly in or between segments. Compared to traditional local stereo methods, the 3-D aggregation greatly improves the accuracy of results in slanted surfaces and occlusion areas while keeping excellent performance near depth discontinuities. We also propose an acceleration method that surprisingly improves the accuracy at the same time. The evaluation experiments confirm our achievements. Xiyuan Peng, Liyan Qiao |
IEEE Signal Process. Lett. | 2 |
| 2016 | NRC: A Nibble Remapping Coding Strategy for NAND Flash Reliability ExtensionabstractAccording to the appearance frequency of nibble short code, this paper presents a novel nibble remapping coding (NRC) strategy to increase the ratio of “1”s in the programming data of NAND flash memory. Because the NRC strategy does not change the length of data during encoding and decoding process, it does not consume any extra user data area of NAND flash. In addition, it can be transparently fit within the flash translation layer algorithms. The experimental results show that the NRC strategy reduces the program disturb errors by 56.9% and decreases the data retention errors by 52.5%, while experiencing only 2.2% degradation in writing speed and 3.2% degradation in reading speed. Compared with conventional BCH code, NRC can significantly enhance the efficacy of the out-of-band space to improve the storage reliability. Debao Wei, Libao Deng, Liyan Qiao, Xiyuan Peng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2016 | A Microcoded Kernel Recursive Least Squares Processor Using FPGA TechnologyabstractKernel methods utilize linear methods in a nonlinear feature space and combine the advantages of both. Online kernel methods, such as kernel recursive least squares (KRLS) and kernel normalized least mean squares (KNLMS), perform nonlinear regression in a recursive manner, with similar computational requirements to linear techniques. In this article, an architecture for a microcoded kernel method accelerator is described, and high-performance implementations of sliding-window KRLS, fixed-budget KRLS, and KNLMS are presented. The architecture utilizes pipelining and vectorization for performance, and microcoding for reusability. The design can be scaled to allow tradeoffs between capacity, performance, and area. The design is compared with a central processing unit (CPU), digital signal processor (DSP), and Altera OpenCL implementations. In different configurations on an Altera Arria 10 device, our SW-KRLS implementation delivers floating-point throughput of approximately 16 GFLOPs, latency of 5.5μ S , and energy consumption of 10 − 4 J, these being improvements over a CPU by factors of 12, 17, and 24, respectively. Yeyong Pang, Yu Peng 0002, Xiyuan Peng, Nicholas J. Fraser, Philip H. W. Leong |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2016 | PEVA: A Page Endurance Variance Aware Strategy for the Lifetime Extension of NAND FlashabstractWith aggressive scaling and multilevel cell technology, the reliability of NAND flash continuously degrades. The lifetime of NAND flash is highly restricted by the bit error rate (BER), and error-correcting codes (ECCs) can provide only limited error correction capability to tolerate increasing bit errors. To cope with this issue, a novel page endurance variance aware (PEVA) strategy is proposed to extend the lifetime of NAND flash based on the experimental observations from our hardware-software codesigned experimental platform. The experimental observations indicate that the BER distribution of retention error shows distinct variances in different pages. The key purpose of PEVA is to exploit the lifetime potency of every page in a block by introducing fine-grained bad page management instead of coarse-grained bad block management (BBM). The experimental results show that the PEVA can extend the lifetime of 2×-nm NAND flash by 9.8× compared with the conventional BBM and that there is at most an 8.7% degradation in writing speed compared with the traditional sector mapping technology. In addition, the maximum writing response time increased by at most 5.9% during the operation of the PEVA strategy. Debao Wei, Libao Deng, Liyan Qiao, Xiyuan Peng |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2016 | Fixation Ratio of Error Location-Aware Strategy for Increased Reliable Retention Time of Flash MemoryabstractThe lifetime of NAND flash is highly restricted by the bit error rate (BER), while error-correcting codes (ECC) can provide only limited error correction capability to tolerate increasing bit errors. In this paper, a novel fixation ratio of error locations aware (FRELA) strategy is proposed to prolong the reliable retention time of flash memory. The concept of FRELA is motivated by following two observations obtained from our real hardware experimental platform: first, the increasing trend of average BER with the retention time and the program/erase cycles can be easily modeled; second, there is a great possibility that an error bit of flash memory remains wrong after a certain retention time. The key purpose of FRELA is to flip the corresponding data bits according to the locations recording of error bits prior to operating ECC. By virtue of the BER prediction model, the required storage space of FRELA is effectively reduced because it avoids updating the information of locations frequently. The experimental results show that FRELA can prolong the reliable retention time of 2×-nm NAND flash by more than 60% without stronger ECC, while experiencing at most 1.03% degradation in writing speed and 0.79% degradation in reading speed, respectively. Debao Wei, Liyan Qiao, Xiyuan Peng |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2015 | Point set registration based on implicit surface fitting with equivalent distanceabstractPoint set registration can be reformulated as the problem of points-to-model alignment with implicit surface fitting. This formulation avoids the correspondence search that is time-consuming. However, the minimization of the sum of the squared “Approximate Distance” between the point set and the model under rigid transformation can be easily trapped into false positions. In this paper, we explicitly derive the detailed formulation of Levenberg-Marquadt algorithm (LMA) with the Approximate Distance for nonlinear least square optimization of registration in 3D case. Based on the analysis for the defect of the Approximate Distance, we propose a novel metric called “Equivalent Distance” and give the full solution for the nonlinear least square optimization of the rigid transformation parameters with the Equivalent Distance. Contrary to the method with the Approximate Distance, the LMA with the Equivalent Distance can converge into optimal positions with much wider convergence range and lead to more accurate transformation parameters. Experimental results and comparisons in 3D cases demonstrate the speed, the accuracy and the convergence performance of the proposed approach. Liyan Qiao, Tiannan Luo, Xiyuan Peng |
ICIP | 5 |
| 2015 | A Health Indicator Extraction and Optimization Framework for Lithium-Ion Battery Degradation Modeling and PrognosticsabstractMaximum releasable capacity and internal resistance are often used as the health indicators (HIs) of a lithium-ion battery for degradation modeling and estimation of remaining useful life (RUL). However, the maximum releasable capacity is usually difficult to estimate in online applications due to complex operating conditions in the field. Moreover, measuring the internal resistance is too expensive to be implemented on-line. In this paper, an HI extraction and optimization framework requiring only the operating parameters of lithium-ion batteries is proposed for battery degradation modeling and RUL estimation. The framework carries out raw HI extraction, transformation, correlation analysis, and verification and evaluation to achieve HI enhancement. In particular, the Box-Cox transformation is adopted to improve the correlation between the extracted HI and the battery's actual degradation state. To estimate the battery's RUL using the enhanced HI, an optimized relevance vector-machine algorithm is utilized, which can be performed in a flexible and agile way. Experimental studies using two different industrial testing data sets illustrate the high efficiency and adaptability of the proposed framework in lithium-ion battery degradation modeling and RUL estimation. Datong Liu, Jianbao Zhou, Haitao Liao, Yu Peng 0002, Xiyuan Peng |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2014 | A novel hybridization of echo state networks and multiplicative seasonal ARIMA model for mobile communication traffic series forecasting
Yu Peng 0002, Miao Lei, Junbao Li, Xiyuan Peng |
Neural Comput. Appl. | 4 |
| 2014 | Compressive Circulant Matrix Based Analog to Information ConversionabstractCompressive Sampling is an attractive way implementing analog to information conversion (AIC), of which the most successful hardware architecture is modulated wideband converter (MWC). Unfortunately, the MWC has high hardware complexity owing to high degree of freedom of the random waveforms constructing the measurement matrix. To reduce the complexity, in this letter, we present a novel Compressive Circulant Matrix based AIC (CCM-AIC) generating random waveforms by cyclic shift of a special sequence with unit amplitude and random phase in frequency domain. Theoretical analysis shows this scheme is optimal for signals sparse in frequency. CCM-AIC outperforms MWC and is more robust. Simulations classify the above analysis. Jingchao Zhang, Ning Fu, Xiyuan Peng |
IEEE Signal Process. Lett. | 3 |
| 2013 | Local stereo matching using binary weighted normalized cross-correlationabstractSignificant achievements have been attained in the field of dense stereo correspondence by local algorithms since the emergence of adaptive support weight by Yoon [1]. However, most algorithms suffer from photometric distortions and low-texture areas. In this paper, we present a novel stereo matching algorithm that can be sensitive to low-texture changes within support windows while keep insensitive to radiometric variations between left and right images. The algorithm performs Normalized Cross-Correlation with Binary Weighted support window (BWNCC) using k-nearest neighbors algorithm to resolve boundary problems. And, the proposed algorithm can be accelerated with transform domain convolution. We also propose to accelerate the BWNCC with transform domain computation. Experiment results confirm that the proposed method is robust, and has the comparable accuracy as the state-of-the-art. Liyan Qiao, Xiyuan Peng |
ICMV | 3 |
| 2013 | Minimizing Building Electricity Costs in a Dynamic Power Market: Algorithms and Impact on Energy ConservationabstractEnergy is a global concern and the electricity bills nowadays are leading to unprecedented costs. Electricity price is market-based and dynamic. In this paper, we investigate how to cut the electricity bills of commercial buildings in a dynamic power market. The building thermal systems (e.g., air-conditioning), which dominate electricity bills, has a special property of thermal storage, i.e., the energy will not immediately dissipate from thermal air/water. Intuitively, with storage, the energy can be "stored" in the thermal system, making it possible to purchase electricity in low price and use it at appropriate time. The building thermal supply and electricity purchasing surely depends on human activities that the building should support such as class and meeting schedules. To minimize electricity bills, we develop a holistic planning of electricity purchasing schedule with thermal storage management, and appropriate room assignment schedules for classes/meetings usage. The computing algorithms require inputs of physical modeling on energy consumption. We develop wireless sensing systems to collect fine-grained data which are used to assist the cross-disciplinary physical modeling. We conduct validation through real experiments. We formulate an optimization problem and show that it is NP-complete. Our primary focus is to minimize electricity bills, which matches the incentives of the commercial buildings. We show that this does not coincide with energy conservation. We further investigate the relationship of minimization of electricity bills and minimization of energy consumption. We develop algorithms for our problem and our evaluation shows that we can achieve a 40% cost reduction. Dawei Pan, Dan Wang 0002, Jiannong Cao 0001, Yu Peng 0002, Xiyuan Peng |
RTSS | 5 |
| 2013 | A study towards applying thermal inertia for energy conservation in roomsabstractWe are in an age where people are paying increasing attention to energy conservation around the world. The heating and air-conditioning systems of buildings introduce one of the largest chunks of energy expenses. In this article, we make a key observation that after a meeting or a class ends in a room, the indoor temperature will not immediately increase to the outdoor temperature. We call this phenomenon thermal inertia . Thus, if we arrange subsequent meetings in the same room rather than in a room that has not been used for some time, we can take advantage of such undissipated cool or heated air and conserve energy. Though many existing energy conservation solutions for buildings can intelligently turn off facilities when people are absent, we believe that understanding thermal inertia can lead system designs to go beyond on-and-off-based solutions to a wider realm. We propose a framework for exploring thermal inertia in room management. Our framework contains two components. (1) The energy-temperature correlation model captures the relation between indoor temperature change and energy consumption. (2) The energy-aware scheduling algorithms: given information for the relation between energy and temperature change, energy-aware scheduling algorithms arrange meetings not only based on common restrictions, such as meeting time and room capacity requirement, but also energy consumptions. We identify the interface between these components so further works towards same on direction can make efforts on individual components. We develop a system to verify our framework. First, it has a wireless sensor network to collect indoor, outdoor temperature and electricity expenses of the heating or air-conditioning devices. Second, we build an energy-temperature correlation model for the energy expenses and the corresponding room temperature. Third, we develop room scheduling algorithms. In detail, we first extend the current sensor hardware so that it can record the electricity expenses in re-heating or re-cooling a room. As the sensor network needs to work unattendedly, we develop a hardware board for long-range communications so that the Imote2 can send data to a remote server without a computer relay close by. An efficient two-tiered sensor network is developed with our extended Imote2 and TelosB sensors. We apply laws of thermodynamics and build a correlation model of the energy needed to re-cool a room to a target temperature. Such model requires parameter calibration and uses the data collected from the sensor network for model refinement. Armed with the energy-temperature correlation model, we develop an optimal algorithm for a specified case, and we further develop two fast heuristics for different practical scenarios. Our demo system is validated with real deployment of a sensor network for data collection and thermodynamics model calibration. We conduct a comprehensive evaluation with synthetic room and meeting configurations, as well as real class schedules and classroom topologies of The Hong Kong Polytechnic University, academic calendar year of Spring 2011. We observe 20% energy savings as compared with the current schedules. Yi Yuan 0005, Dawei Pan, Dan Wang 0002, Xiaohua Xu 0002, Yu Peng 0002, Xiyuan Peng, Peng-Jun Wan |
ACM Trans. Sens. Networks | 6 |
| 2012 | Thermal Inertia: Towards an energy conservation room management systemabstractWe are in an age where people are paying increasing attention to energy conservation around the world. The heating and air-conditioning systems of buildings introduce one of the largest chunk of energy expenses. In this paper, we make a key observation that after a meeting or a class ends in a room, the indoor temperature will not immediately increase to the outdoor temperature. We call this phenomenon Thermal Inertia. Thus, if we arrange subsequent meetings in the same room; than a room that has not been used for some time, we can take advantage of such un-dissipated cool or heated air and conserve energy. We develop a green room management system with three main components. First, it has a wireless sensor network to collect indoor, outdoor temperature and electricity expenses of the air-conditioning devices. Second, we build an energy-temperature correlation model for the energy expenses and the corresponding room temperature. Third, we develop room scheduling algorithms. Our system is validated with real deployment of a sensor network for data collection and thermodynamics model calibration. We conduct a comprehensive evaluation with synthetic room and meeting configurations. We observe a 30% energy saving as compared with the current schedules. Dawei Pan, Yi Yuan 0005, Dan Wang 0002, Xiaohua Xu 0002, Yu Peng 0002, Xiyuan Peng, Peng-Jun Wan |
INFOCOM | 6 |
| 2012 | Oscillation property for fuzzy delay differential equations
Mengshu Guo, Xiyuan Peng |
Fuzzy Sets Syst. | 2 |
| 2011 | Accelerating on-line training of LS-SVM with run-time reconfigurationabstractLeast Squares Support Vector Machines(LS-SVM), which is an efficient supervised learning tool, has been widely applied to real-time on-line data processing in many fields. However, the on-line training of LS-SVM always suffers from huge computation which greatly limits its practicability especially in embedded systems. By leveraging the flexibility and high degree parallelism offered by reconfigurable fabrics, we propose a Run-Time Reconfiguration(RTR) framework to accelerate the on-line training of LS-SVM. To realize maximum computational parallelism, we divide the training process into two parts, the kernel matrix formulation and the least-square problem solving. We dynamically load these two parts into FPGA with RTR under the control of the embedded PowerPC. In the kernel matrix formulation part, we design a piecewise linear interpolation method to realize the radial basis function. In the least-square problem solving part, the modified Cholesky Decomposition is introduced to avoid the latency caused by square roots operations. The whole design is tested on Virtex XC5VFX130T with a 150MHz clock. The experiments show appealing speed up which ranges from 6~218× over a Xeon CPU implementation on five different sized datasets. From time cost percentage analysis, our proposed architecture can be effectively applied to LS-SVM training in more than 1000 samples applications. Yu Peng 0002, Guangquan Zhao, Xiyuan Peng |
FPT | 4 |
| 2011 | Analog Circuit Fault Diagnosis with Echo State Networks Based on Corresponding Clusters
Xiyuan Peng, Miao Lei, Yu Peng 0002 |
ISNN (1) | 1 |
| 2011 | Anti Boundary Effect Wavelet Decomposition Echo State Networks
Jianmin Wang 0004, Yu Peng 0002, Xiyuan Peng |
ISNN (1) | 3 |
| 2003 | Virtual instrument parameter calibration with particle swarm optimizationabstractIn virtual instrument designs and applications, lots of functional parameters can be set through software methods. Currently, most parameter settings methods are lightly linked with the knowledge of instruments and basic principles related to specific applications. However, it is difficult for some end users to deal with those advanced operations. By adopting the particle swarm optimization (PSO) algorithm, the adaptive set and calibration of instrument parameters can be achieved by software with computational intelligence. Experiments and applications showed that the adaptive parameter calibration method based on the PSO can enhance the effectiveness of debugging and maintenance of virtual instrument and test system. Yu Peng 0002, Xiyuan Peng, Shengwei Meng |
SIS | 2 |