EDBT 2026 Demo / reviewers in the wild / expert
Yongxin Zhu 0001
dblp:27/3343-1
· DBLP profile ↗
66ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-1813-1792ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 37 · 1 first-author · 11 since 2021Security and privacy · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 since 2021Computer networks · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PipeInfer: A pipelined inference framework for U-Net like neural networks
Chujun Feng, Congyu Lin, Yongxin Zhu 0001 |
ISCAS | 4 |
| 2026 | PacViT: An Efficient ViT Accelerator with Native Dynamic Pruning and Sliding Cache Attention
Jun Gong 0003, Zixuan Zhu 0001, Yongxin Zhu 0001 |
ISCAS | 4 |
| 2026 | Accelerating string search: A microarchitecture-aware index for variable-length data
Yuanliang Cui, Chujun Feng, Yongxin Zhu 0001 |
J. Syst. Archit. | 5 |
| 2026 | RapidFS: A hybrid user-kernel PMEM file system for high-throughput streaming writes in large-scale facility data pipelines
Chujun Feng, Xiangcong Kong, Yuanliang Cui, Congyu Lin, Yongxin Zhu 0001 |
J. Syst. Archit. | 6 |
| 2026 | A 50 μW/Gbps/Lane Power-Efficient MIPI D-PHY Receiver With Architecture-Level Adaptive and Structural Optimizations for Micro-DisplaysabstractAchieving high power efficiency in Mobile Industry Processor Interface (MIPI) D-PHY receivers is crucial for micro-display chips in AR/VR systems, where stringent power constraints exist. However, existing designs often sacrifice power efficiency for higher data rates due to architectural limitations, neglecting optimization for low-power applications. To address this issue, we propose a receiver architecture that substantially enhances power efficiency through three key techniques. First, we improve the gain-bandwidth product (GBW) by employing an autonomous gain scheduling analog front-end (AFE) that dynamically tunes the gain while reducing drive current. Second, we reduce clocking overhead by introducing a self-monitoring interferometric deserializer that enables clock-free pre-scaling and halves the DDR sampling frequency. Third, we increase transition speed and minimize short-circuit power by utilizing a chaotic topological flow actuator (CTFA) with multi-path current feedthrough. Compared to prior state-of-the-art designs, the proposed receiver achieves a power efficiency of$50~\mu $W/Gbps/lane ($42~\mu $A/Gbps/lane), reducing power and current consumption by 46% and 45%, respectively, using a standard 180-nm process. Haoran Zeng, Yingqi Feng, Tianai Li, Hang Ye 0007, Zunkai Huang, Hui Wang 0036, Yongxin Zhu 0001, Qiliang Li, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | Accelerating Elliptic Curve Digital Signature Verification on FPGA for Secure Communication
Chujun Feng, Ning Ni 0001, Congyu Lin, Yongxin Zhu 0001, Hui Wang 0036 |
WASA (1) | 4 |
| 2025 | An Open X-Ray Spectrometric Dataset for Deep Learning-Based Pile-up Correction
Congyu Lin, Chujun Feng, Songqi Gu, Yongxin Zhu 0001, Thomas Trigano, Dima Bykhovsky |
WASA (3) | 6 |
| 2025 | KARMA: A Multilevel Decomposition Hybrid Mamba Framework for Multivariate Long-Term Time Series Forecasting
Hang Ye 0007, Gaoxiang Duan, Haoran Zeng, Yangxin Zhu, Lingxue Meng, Yongxin Zhu 0001 |
WASA (3) | 7 |
| 2025 | Bit-Sparsity Aware Acceleration With Compact CSD Code on Generic Matrix MultiplicationabstractThe ever-increasing demand for matrix multiplication in artificial intelligence (AI) and generic computing emphasizes the necessity of efficient computing power accommodating both floating-point (FP) and quantized integer (QINT). While state-of-the-art bit-sparsity-aware acceleration techniques have demonstrated impressive performance and efficiency in neural networks through software-driven methods such as pruning and quantization, these approaches are not always feasible in typical generic computing scenarios. In this paper, we propose Bit-Cigma, a hardware-centric architecture that leverages bit-sparsity to accelerate generic matrix multiplication. Bit-Cigma features (1) CCSD encoding, an optimized on-chip sparsification technique based on canonical signed digit (CSD) representation; (2) segmented dot product, a multi-stage exponent matching technique for long FP vectors; and (3) the versatility to efficiently process both FP and QINT data types. CCSD encoding halves the cost of CSD encoding while achieving optimal bit-sparsity, and segmented dot product improves both accuracy and throughput. Bit-Cigma cores are implemented using 65 nm technology at 1 GHz, demonstrating substantial gains in performance and efficiency for both FP and QINT configurations. Compared to state-of-the-art Bitlet, Bit-Cigma achieves 3.2$\boldsymbol{\times}$performance, 6.1$\boldsymbol{\times}$area efficiency, and 15.3$\boldsymbol{\times}$energy efficiency when processing FP32 data while ensuring zero computing error. Zixuan Zhu 0001, Chundong Wang 0001, Zunkai Huang, Yongxin Zhu 0001 |
IEEE Trans. Computers | 6 |
| 2024 | Liver Fibrosis Classification based on Multimodal Imaging Feature FusionabstractThis study introduces a liver fibrosis staging method based on the fusion of multi-modal imaging features. By leveraging ultrasound gray-scale images and ultrasound shear wave elastography (SWE), the method employs an attention-based weighted strategy to effectively fuse different feature maps, integrating information from diverse modalities. Additionally, it integrates multi-modal network branches and designs a multi- modal integrated loss function to update network parameters, thereby enhancing the model's generalization and anti- interference capabilities. The experimental results demonstrate that the proposed fusion network achieves a high AUC of 95.01 and an accuracy of 76.21%. Compared to existing liver fibrosis classification methods for the five-class classification task, which integrate multi-modal features from various liver imaging modalities, our approach shows a significant improvement of 6% in accuracy. With its lightweight model architecture and low computational resource consumption, the proposed method effectively performs liver fibrosis staging, holding significant promise for clinical auxiliary diagnosis. Xinyan Jiang, Xinping Ren, Yongxin Zhu 0001, Yueying Zhou |
CSCloud | 3 |
| 2024 | Deep Fingerprinting Data Learning Based on Federated Differential Privacy for Resource-Constrained Intelligent IoT SystemsabstractWith the rapid integration of Internet of Things (IoT) devices and artificial intelligence (AI) function, the data management and privacy issue has drawn great attentions in intelligent IoT systems where communication infrastructures frequently exchange open data flows over the air. Therefore, lightweight and private access over radio communication pipes becomes a critical but challengeable need for resource-constrained IoT devices due to the limited memory capacity, computing, and energy consumption. In this article, we develop the concept of deep federated scattering fingerprinting aided by differential privacy (DFSF-DP) in which a deep fingerprinting data learning network exploits fingerprinting data to realize lightweight intelligent access and incorporates federated learning with differential privacy to guarantee the data privacy in a way of distributed training. Particularly, first, we employ a wavelet scattering network for the efficient radio frequency fingerprinting (RFF) feature extraction and construct a high information density database. Subsequently, the implementation of distributed learning minimizes the demand for computing resources, by exploiting the full potential of edge and cloud nodes to aggregate the global model. To bolster the data privacy and security, adaptive clipping and gradient noising are incorporated into DFSF-DP. Experimental results demonstrate that DFSF-DP obtains outstanding performance and achieves equivalent advancements while utilizing a mere 25% of the original data set. Moreover, it attains a 93% identification accuracy with 0.1 noise multiplier which confirms the remarkable performance of DFSF-DP while upholding privacy and security considerations. Dongyang Xu 0003, Pandi Vijayakumar, Yongxin Zhu 0001, Amr Tolba |
IEEE Internet Things J. | 5 |
| 2023 | Enabling zero knowledge proof by accelerating zk-SNARK kernels on GPU
Ning Ni 0001, Yongxin Zhu 0001 |
J. Parallel Distributed Comput. | 2 |
| 2023 | Optimized CPU-GPU collaborative acceleration of zero-knowledge proof for confidential transactions
Yongxin Zhu 0001 |
J. Syst. Archit. | 3 |
| 2023 | I/O-efficient GPU-based acceleration of coherent dedispersion for pulsar observation
Xiangcong Kong, Yongxin Zhu 0001, Gaoxiang Duan |
J. Syst. Archit. | 3 |
| 2023 | Federated-Learning-Based Synchrotron X-Ray Microdiffraction Image Screening for Industry MaterialsabstractSynchrotron X-ray microdiffraction (μXRD) services are conducted for industrial minerals to identify their crystal impurities in terms of crystallinity and potential impurities.μXRD services generate huge loads of images that have to be screened before further processing and storage. However, there are insufficient effective labeled samples to train a screening model since service consumers are unwilling to share their original experimental images. In this article, we propose a physics law-informed federated learning (FL) basedμXRD image screening method to improve the screening while protecting data privacy. In our method, we handle the unbalanced data distribution challenge incurred by service consumers with different categories and amounts of samples with novel client sampling algorithms. We also propose hybrid training schemes to handle asynchronous data communications between FL clients and servers. The experiments show that our method can ensure effective screening for industrial users conducting industrial material testing while keeping commercially confidential information. Yongxin Zhu 0001, Victor Chang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | A Universal RRAM-Based DNN Accelerator With Programmable Crossbars Beyond MVM OperatorabstractResistive-RAM (RRAM)-based deep neural network (DNN) accelerator has shown a great potential as it is good at the matrix–vector multiplication (MVM) operator. However, it does not benefit non-MVM operators, such as transcendental activation or elementwise operations, which often require customized CMOS circuits in conventional DNN accelerator designs. In this article, we propose a new RRAM-based DNN inference accelerator, which leverages the proposed RRAM-CORDIC and RRAM-MLP algorithms to make the transcendental and elementwise operators calculable in the RRAM crossbar just like MVM. Both algorithms can exploit the higher multiply-and-accumulation (MAC) parallelism that is traditionally expensive in CMOS but now efficient in the RRAM crossbar. Then, we further propose an intercrossbar pipelining scheme, which can balance the number of crossbars for MVM and non-MVM operations and orchestrate them in pursuing higher DNN computing throughput. The experimental results show that both algorithms can sustain a high arithmetic accuracy and deliver less than 1% DNN accuracy loss on typical inference workloads. The elimination of expensive CMOS circuits, in turn, can trade more crossbar resources in the same area to speed up the performance by$1.16\times $to$2.33\times $. With the extended operators, the RRAM-based DNN accelerator can switch crossbar functions at will, and apply for a diverse of DNN models in a unified in-memory accelerator architecture. Jianfei Jiang 0001, Yongxin Zhu 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | An Improved Federated Learning Algorithm for Privacy Preserving in Cybertwin-Driven 6G SystemabstractWith the expected explosive use of the Internet of Everything in sixth generation (6G), the cybertwin network is able to convert user information to digital assets and provide extensive services. However, protecting and enhancing privacy of the processed and transmitted data in cybertwin-driven 6G is still in its infancy. Federated learning (FL) is a nascent distributed machine learning paradigm that is able to facilitate privacy protection in cybertwin networks. In a cybertwin network, imbalanced data distribution of the clients can increase the bias of the global model and sacrifice the performance of the FL model. Prior research work dealing with imbalanced data requires extra data information exchanged between clients and the server, which increases the risk of privacy leakage. To avoid privacy leakage, we design an estimation algorithm to determine the distribution of local data collected at the clients without the awareness of specific raw data. We consider two scenarios in FL: 1) the server could receive the individual trained model for each selected device and 2) the server could receive the aggregated model from the selected clients. We formulate two device selection problems to improve the training performance of the aforementioned scenarios. We develop two online learning algorithms to tackle the selection problems for both individual model uploading and aggregated model uploading. The proposed algorithms are conducted on the server, thereby avoiding privacy leakage and extra computation at the clients. We validate the effectiveness of the proposed client selection algorithms with sufficient experiments in cybertwin-driven 6G networks. Ximin Wang, Hua Qian, Yongxin Zhu 0001, Hongbin Zhu, Mohsen Guizani, Victor Chang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2021 | SparkNoC: An energy-efficiency FPGA-based accelerator using optimized lightweight CNN for edge computing
Zunkai Huang, Hui Wang 0036, Victor Chang 0001, Yongxin Zhu 0001, Songlin Feng |
J. Syst. Archit. | 6 |
| 2020 | Anomaly Detection Based on RBM-LSTM Neural Network for CPS in Advanced Driver Assistance SystemabstractAdvanced Driver Assistance System (ADAS) is a typical Cyber Physical System (CPS) application for human–computer interaction. In the process of vehicle driving, we use the information from CPS on ADAS to not only help us understand the driving condition of the car but also help us change the driving strategies to drive in a better and safer way. After getting the information, the driver can evaluate the feedback information of the vehicle, so as to enhance the ability to assist in driving of the ADAS system. This completes a complete human–computer interaction process. However, the data obtained during the interaction usually form a large dimension, and irrelevant features sometimes hide the occurrence of anomalies, which poses a significant challenge to us to better understand the driving states of the car. To solve this problem, we propose an anomaly detection framework based on RBM-LSTM. In this hybrid framework, RBM is trained to extract general underlying features from data collected by CPS, and LSTM is trained from the features learned by RBM. This framework can effectively improve the prediction speed and present a good prediction accuracy to show vehicle driving condition. Besides, drivers are allowed to evaluate the prediction results, so as to improve the accuracy of prediction. Through the experimental results, we can find that the proposed framework not only simplifies the training of the entire neural network and increases the training speed but also greatly improves the accuracy of the interaction-driven data analysis. It is a valid method to analyze the data generated during the human interaction. Hanlin Zhu, Yongxin Zhu 0001, Victor Chang 0001, Cong He, Ching-Hsien Hsu, Hui Wang 0036, Songlin Feng, Zunkai Huang |
ACM Trans. Cyber Phys. Syst. | 3 |
| 2019 | Efficient Weight Reuse for Large LSTMsabstractLong Short-Term Memory (LSTM) networks have been deployed in speech recognition, natural language processing and financial calculations in recent years, and are beginning to be used in systems where low latency and low power are required. To meet such requirements, we propose a stall-free hardware architecture by reorganising the order of operations in an LSTM system and develop a unique blocking-batching strategy to reuse the LSTM weights fetched from external memory to optimise the benefits of on-chip memory with a limited size for a large machine learning model. Evaluation results show that our architecture can achieve up to 20.8 GOPS/W, which would be among the highest for FPGA designs targeting LSTM systems with weights stored in off-chip memory. Comparing to the state-of-the-art design using off-chip memory to store the weights, we achieve 1.65 times higher performance-per-watt efficiency and 1.60 times higher performance-per-DSP efficiency. When compared with CPU and GPU implementation, our novel hardware architecture is 23.7 and 1.3 times faster while consuming 208 and 19.2 times lower energy respectively, which shows that our approach contributes to high performance and low power FPGA-based LSTM systems. Zhiqiang Que, Thomas Nugent, Shuanglong Liu, Xinyu Niu, Yongxin Zhu 0001, Wayne Luk |
ASAP | 6 |
| 2019 | Design and implementation of reconfigurable acceleration for in-memory distributed big data computing
Junjie Hou, Yongxin Zhu 0001, Shijin Song |
Future Gener. Comput. Syst. | 2 |
| 2019 | Secure big data communication for energy efficient intra-cluster in WSNs
Anxi Wang, Jian Shen 0001, Pandi Vijayakumar, Yongxin Zhu 0001 |
Inf. Sci. | 4 |
| 2019 | A Dependable Time Series Analytic Framework for Cyber-Physical Systems of IoT-based Smart GridabstractWith the emergence of cyber-physical systems (CPS), we are now at the brink of next computing revolution. The Smart Grid (SG) built on top of IoT (Internet of Things) is one of the foundations of this CPS revolution, which involves a large number of smart objects connected by networks. The volume of time series of SG equipment is tremendous and the raw time series are very likely to contain missing values because of undependable network transferring. The problem of storing a tremendous volume of raw time series thereby providing a solid support for precise time series analytics now becomes tricky. In this article, we propose a dependable time series analytics (DTSA) framework for IoT-based SG. Our proposed DTSA framework is capable of providing a dependable data transforming from CPS to the target database with an extraction engine to preliminary refining raw data and further cleansing the data with a correction engine built on top of a sensor-network-regularization-based matrix factorization method. The experimental results reveal that our proposed DTSA framework is capable of effectively increasing the dependability of raw time series transforming between CPS and the target database system through the online lightweight extraction engine and the offline correction engine. Our proposed DTSA framework would be useful for other industrial big data practices. Chang Wang 0003, Yongxin Zhu 0001, Weiwei Shi 0002, Victor Chang 0001, Pandi Vijayakumar, Yishu Mao 0001, Yiping Fan |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2019 | A Novel Resistive Memory-based Process-in-memory Architecture for Efficient Logic and Add OperationsabstractThe coming era of big data revives the Processing-in-memory (PIM) architecture to relieve the memory wall problem that embarrasses the modern computing system. However, most existing PIM designs just put computing units closer to memory, rather than a complete integration of them due to their incompatibility in CMOS manufacturing. Fortunately, the emerging Resistive-RAM (ReRAM) offers new hope to this dilemma owing to its inherent memory and computing capability using the same device. In this article, we propose a ReRAM memory structure with efficient PIM capability of both logic and add operations. It first leverages non-linearity to suppress sneak current and thus sustains high memory density. Using a differential bit cell, it also enables efficient processing of arbitrary logic functions using the same memory cells with non-destructive operations. Then, a novel PIM adder is proposed, which customizes a sneak current path as the carry-chain for fast carry propagation and improves adder performance significantly. In the experiment, the proposed PIM demonstrates higher efficiency in both computing area and performance for logic and addition, which greatly increases the ReRAM PIM applicability for future computable architectures. Taozhong Li, Qin Wang 0009, Yongxin Zhu 0001, Jianfei Jiang 0001, Guanghui He 0002, Jing Jin 0005, Zhigang Mao, Naifeng Jing |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2019 | Sustainable Computing Based Deep Learning Framework for Writing Research ManuscriptsabstractWriting research manuscripts is always a tough task at the eleventh hour. Often researchers do not find time to rewrite the manuscript to satisfaction, which is not quantifiable though. This paper proposes a sustainable computing based deep learning framework for iterated accumulation of ideas while writing research manuscripts. The framework suggests Deep Author Topic Models (DATM) where every author of the manuscript is modeled. For this, we have assumed time based sustainable computing as a measure of evaluation for research manuscript effectiveness. Using respective DATM, the region contributed by every author in the manuscriptis analyzed and fine-tuned semantically such that the manuscript is made to perfection in least time. G. S. Mahalakshmi 0001, G. Muthu Selvi, Sendhilkumar Selvaraju, Pandi Vijayakumar, Yongxin Zhu 0001, Victor Chang 0001 |
IEEE Trans. Sustain. Comput. | 5 |
| 2018 | Effective Prediction of Missing Data on Apache Spark over Multivariable Time SeriesabstractMore massive volume of data are generated in many areas than ever before. However, the missing of some values in collected data always occurs in practice and challenges extracting maximal value from these large scale data sets. Nevertheless, in multivariable time series, most of the existing methods either might be infeasible or could be inefficient to predict the missing data. In this paper, we have taken up the challenge of missing data prediction in multivariable time series by employing improved matrix factorization techniques. Our approaches are optimally designed to largely utilize both the internal patterns of each time series and the information of time series across multiple sources. Based on the idea, we have imposed three different regularization terms to constrain the objective functions of matrix factorization and built five corresponding models. Extensive experiments on real-world data sets and synthetic data set demonstrate that the proposed approaches can effectively improve the performance of missing data prediction in multivariable time series. Furthermore, we have also demonstrated how to take advantage of the high processing power of Apache Spark to perform missing data prediction in large scale multivariable time series. Weiwei Shi 0002, Yongxin Zhu 0001, Philip S. Yu, Jiawei Zhang 0001, Tian Huang, Chang Wang 0003 |
IEEE Trans. Big Data | 2 |
| 2018 | Statistical Learning for Anomaly Detection in Cloud Server Systems: A Multi-Order Markov Chain FrameworkabstractAs a major strategy to ensure the safety of IT infrastructure, anomaly detection plays a more important role in cloud computing platform which hosts the entire applications and data. On top of the classic Markov chain model, we proposed in this paper a feasible multi-order Markov chain based framework for anomaly detection. In this approach, both the high-order Markov chain and multivariate time series are adopted to compose a scheme described in algorithms along with the training procedure in the form of statistical learning framework. To curb time and space complexity, the algorithms are designed and implemented with non-zero value table and logarithm values in initial and transition matrices. For validation, the series of system calls and the corresponding return values are extracted from classic Defense Advanced Research Projects Agency (DARPA) intrusion detection evaluation data set to form a two-dimensional test input set. The testing results show that the multi-order approach is able to produce more effective indicators: in addition to the absolute values given by an individual single-order model, the changes in ranking positions of outputs from different-order ones also correlate closely with abnormal behaviours. Wenyao Sha, Yongxin Zhu 0001, Min Chen 0003, Tian Huang |
IEEE Trans. Cloud Comput. | 2 |
| 2018 | A Hardware Pipeline with High Energy and Resource Efficiency for FMM AccelerationabstractThe fast multipole method (FMM) is a promising mathematical technique that accelerates the calculation of long-ranged forces in the large-sized n-body problem. Existing implementations of the FMM on general-purpose processors are energy and resource inefficient. To mitigate these issues, we propose a hardware pipeline that accelerates three key FMM steps. The pipeline improves energy efficiency by exploiting fine-granularity parallelism of the FMM. We reuse the pipeline for different FMM steps to reduce resource usage by 66%. Compared to the state-of-the-art implementations on CPUs and GPUs, our implementation requires 15% less energy and delivers 2.61 times more floating-point operations. Tian Huang, Yongxin Zhu 0001, Yajun Ha, Xu Wang 0010, Meikang Qiu |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2017 | Enhancing Precision and Bandwidth in Cloud Computing: Implementation of a Novel Floating-Point Format on FPGAabstractCloud computing is a type of Internet-based service computing that provides computing, storage and networking services to multiple users. With the increase of data size, computing capacity runs out quickly in cloud computing services. To fill the shortage of computation capacity, we propose to adopt variable precision by implementing unum (universal number), which is a number format different from IEEE Standard for Floating-Point Arithmetic - IEEE 754 floats. Compared with IEEE 754 floats, the outstanding features of unum are clearance of rounding errors, high information-per-bit and variable precision. As a candidate replacement of IEEE 754 floats, the application of unum can improve the precision in computing, decrease the bit width for high precision numbers. However, unum was only implemented in software model before due to technical complexity, in order to validate the performance on chip, we implement this arithmetic on FPGA for the first time. We also implement an unum based 16-point FFT on FPGA. We validate the design and compare the bit width in computing with IEEE 754 floats, evaluate the power dissipation on FPGA. The experimental results of comparison show that unum arithmetic can ensure correctness even in some extreme arithmetic cases in which IEEE 754 floats cannot work properly, furthermore the bit width of unum is much less than IEEE 754 floats in the same precision. Junjie Hou, Yongxin Zhu 0001, Yulan Shen |
CSCloud | 2 |
| 2017 | Exploring High Efficiency Hardware Accelerator for the Key Algorithm of Square Kilometer Array Telescope Data ProcessingabstractThe SKA (Square Kilometer Array) radio telescope under construction will become the largest telescope in the world by integrating the sampled data from a huge number of small antenna nodes in the array to emulate a giant antenna. Due to the limited storage space, the SKA needs to process massive data in real-time, which makes the SKA scientific data processing become a bottleneck of the computational performance. However, existing off-the-shelf high performance computing solutions cannot meet the computation requirements (5 times more than top 1 supercomputer) as well as low power budget (1/3 of the power of the top 1 supercomputer). In this paper, we explore high efficiency solution design based on FPGA by addressing the most representative key algorithm in SKA data processing, i.e. Gridding, which is the most time and memory consuming. We propose an efficient hardware accelerator design of Gridding algorithm on FPGA, which would the first FPGA based design of Gridding algorithm in this community. In our design, we unfold the third loop in the Gridding algorithm and design corresponding hardware pipeline stages to achieve the high efficiency hardware acceleration. The functionality and performance of our design is verified in both simulation and FPGA prototyping board, whose results show that our proposed hardware implementation achieved great improvement in performance compared with software implementation running on generic CPUs. We believe our design would be a strong candidate design to solve the bottleneck in SKA data processing. Yongxin Zhu 0001, Xu Wang 0010, Junjie Hou, Ali Masoumi |
FCCM | 2 |
| 2017 | Synergistic design of an application-oriented sparse directory on many-core embedded systems
Chang Wang 0003, Yongxin Zhu 0001, Victor Chang 0001, Han Song |
J. Syst. Archit. | 2 |
| 2017 | An energy-efficient system on a programmable chip platform for cloud applications
Xu Wang 0010, Yongxin Zhu 0001, Yajun Ha, Meikang Qiu, Tian Huang, Xueming Si |
J. Syst. Archit. | 2 |
| 2017 | Dynamic application allocation with resource balancing on NoC based many-core embedded systems
Chang Wang 0003, Yongxin Zhu 0001, Meikang Qiu, Xu Wang 0010 |
J. Syst. Archit. | 2 |
| 2016 | Combining the histogram method and the ultrafast segmented model identification of linearity errors algorithm for ADC linearity testingabstractIn this paper, we propose a novel method termed as the histogram and ultrafast segmented model identification of linearity errors (H-USMILE), in order to implement an efficient linearity test of ADCs. We use the standard histogram method with only a small number of test samples to obtain coarse INL results, along with the segmented non-parametric model of the ultrafast segmented model identification of linearity errors (USMILE) algorithm employed to convert the coarse INL results into precise ones. Experimental results show that, by using the proposed method, we can significantly reduce the required test data and test time while maintaining or even improving the precision of the standard histogram method. Weida Chen, Yongxin Zhu 0001, Dongyu Ou |
ETS | 2 |
| 2016 | Evaluation of variable precision computing with variable precision FFT implementation on FPGAabstractIn the workflow of SKA-SDP (Square Kilometer Array Radio Telescope-Scientific Data Processing), FFT (Fast Fourier Transform) calculation takes a significant proportion of computation overhead. Moreover, FFT computation has to be done within the tight power budget, which existing generic high performance computing architectures cannot meet. To explore power efficiency of FFT computation, this study is designed with initial evaluation of FFT implementation in variable precision on Xilinx ML605 FPGA (field programmable gate array). The FPGA-based implementation of FFT includes a Xilinx IP Core version 7.1, which supports fixed-point and floating-point computation in single precision. The input data width and phase factor width of fixed-point computation can be varied from 8 bits to 34 bits, allowing that calculation accuracy can be adjusted by setting the width. Since single precision is redundant to accuracy requirement of SKA-SDP, fixed-point calculation is designed to emulate single-precision floating point computation. Calculational power dissipation and throughput with different phase factors on FPGA was measured respectively. The final result demonstrates that the implementation on FPGA ensures sufficient precision at a much less power cost compared with floating-point FFT. In other words, this study indicates variable precision computation would be an efficient way to improve power efficiency. Yongxin Zhu 0001, Xu Wang 0010, Tian Huang, Weida Chen, Yishu Mao 0001 |
FPT | 2 |
| 2016 | Parallel Discord Discovery
Tian Huang, Yongxin Zhu 0001, Yishu Mao 0001, Yajun Ha, Gillian Dobbie |
PAKDD (2) | 2 |
| 2016 | Anomaly detection and identification scheme for VM live migration in cloud infrastructure
Tian Huang, Yongxin Zhu 0001, Stéphane Bressan, Gillian Dobbie |
Future Gener. Comput. Syst. | 2 |
| 2016 | A comprehensive reconfigurable computing approach to memory wall problem of large graph computation
Xu Wang 0010, Yongxin Zhu 0001, Linan Huang |
J. Syst. Archit. | 2 |
| 2016 | Intrusion detection techniques for mobile cloud computing in heterogeneous 5GabstractMobile cloud computing is applied in multiple industries to obtain cloud-based services by leveraging mobile technologies. With the development of the wireless networks, defending threats from wireless communications have been playing a remarkable role in the Web security domain. Intrusion detection system IDS is an efficient approach for protecting wireless communications in the Fifth Generation 5G context. In this paper, we identify and summarize the main techniques being implemented in IDSs and mobile cloud computing with an analysis of the challenges for each technique. Addressing the security issue, we propose a higher level framework of implementing secure mobile cloud computing by adopting IDS techniques for applying mobile cloud-based solutions in 5G networks. On the basis of the reviews and synthesis, we conclude that the implementation of mobile cloud computing can be secured by the proposed framework because it will provide well-protected Web services and adaptable IDSs in the complicated heterogeneous 5G environment. Copyright © 2015John Wiley & Sons, Ltd. Keke Gai, Meikang Qiu, Lixin Tao, Yongxin Zhu 0001 |
Secur. Commun. Networks | 4 |
| 2016 | A Real-Time FPGA-Based Accelerator for ECG Analysis and Diagnosis Using Association-Rule MiningabstractTelemedicine provides health care services at a distance using information and communication technologies, which intends to be a solution to the challenges faced by current health care systems with growing numbers of population, increased demands from patients, and shortages in human resources. Recent advances in telemedicine, especially in wearable electrocardiogram (ECG) monitors, call for more intelligent and efficient automatic ECG analysis and diagnostic systems. We present a streaming architecture implemented on Field-Programmable Gate Arrays (FPGAs) to accelerate real-time ECG signal analysis and diagnosis in a pipelining and parallel way. Association-rule mining is employed to generate early diagnostic results by matching features of ECG with generated association rules. To improve performance of the processing, we propose a hardware-oriented data-mining algorithm named Bit_Q_Apriori . The corresponding hardware implementation indicates a good scalability and outperforms other hardware designs in terms of performance, throughput, and hardware cost. Xiaoqi Gu, Yongxin Zhu 0001, Shengyan Zhou, Chaojun Wang, Meikang Qiu, Guoxing Wang |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2015 | Designing ARINC653 Partition Constrained Scheduling for Secure Real Time Embedded AvionicsabstractBeing a high end embedded system, an avionic system calls for stringent real time constraints as well as secure guarantees. In terms of logical architecture, avionic systems have recently grown into the form of Integrated Modular Avionics (IMA) from the traditional federated avionics system whose redundancy level is overwhelming for modern large aircrafts. The key idea of IMA system lies in the rules of time and space partitioning, which guarantees system predictability and reliability. However, existing industrial practices of IMA partition and priority settings usually incur significant waste of resources, which would eventually lower the performance of IMA tasks in terms of latency or throughput. This issue was not properly addressed by previous researchers who assumed settings of priority variances and fixed partitions, which differ from practical applications. In this paper, a secure real time scheduling scheme with partition readjustment is proposed with inputs of features exhibited by tasks under partition. In our scheme, the resource costs are reduced by merging and restructuring partitions without compromising hard real time constraints. The simulation results of actual flight missions show that significant improvement by our method in terms of the average response time of tasks as well as number of partitions. Xiang Tao, Yongxin Zhu 0001, Yishu Mao 0001, Han Song, Weiguang Sheng, Weiwei Shi 0002 |
CSCloud | 2 |
| 2014 | An FPGA-Assisted Cloud Framework for Massive ECG Signal ProcessingabstractCurrent aging society has seen huge increases in portable devices sending massive volume of signals to medical servers. To address inefficient and unscalable signal processing on generic servers and clients, in this paper, we present an FPGA(field programmable gate array)-assisted cloud system providing an efficient framework for electrocardiogram (ECG) telemedicine including real-time data acquisition, transmission and analyzing over the Internet. We explore the requirements of massive cloud signal processing supporting a large number of connections or channels from cloud clients. A prototype system was composed of a client-server platform and an FPGA hardware system with PCI-E endpoint and ECG R-peak detection modules. A streaming micro-architecture for hardware system is proposed to detect ECG pattern from scalable number of channels. Evaluation results show that our streaming system design has good performance in terms of real-time, scalability and latency. Shengyan Zhou, Yongxin Zhu 0001, Chaojun Wang, Xiaoqi Gu, Guoguang Rong |
DASC | 2 |
| 2013 | An FPGA Based PCI-E Root Complex Architecture for Standalone SOPCsabstractWe present an FPGA (field programmable gate array) based PCI-E (PCI-Express) root complex architecture for SOPCs (System-on-a-Programmable-Chip) in this paper. In our work, the system on the FPGA serves as a PCIE master device rather than a PCIE endpoint, which is usually a common practice as a co-processing device driven by a desktop computer or a server. We use this system to control a PCIE endpoint, which is also an FPGA based endpoint implemented on another FPGA board. This architecture requires only IP cores free of charge. We also provide basic software driver so that specific device driver can be developed on it to control popular PCIE device in the future, i.e. ethernet card or graphic card. The whole architecture has been implemented on Xilinx Virtex-6 FPGAs to indicate that this architecture is a feasible approach to standalone SOPCs, which has better efficiencies than those with additional generic controlling processors. Yingjie Cao, Yongxin Zhu 0001, Xu Wang 0010, Meikang Qiu |
FCCM | 2 |
| 2013 | A Reconfigurable Architecture for 1-D and 2-D Discrete Wavelet TransformabstractIn this paper, we propose a novel architecture for DWT that can be reconfigured to be adapted to different kinds of filter banks and different sizes of inputs. High flexibility and generality are achieved by using the MAC loop based filter(MLBF). Classic methods, such as polyphase structure and fragment-based sample consumption, are used to enhance the parallelism of the system. The architecture can be reconfigured to 3 modes to deal with 1-D or 2-D DWT with different bandwidth and throughput requirements. Qing Sun 0002, Yongxin Zhu 0001, Yuzhuo Fu |
FCCM | 3 |
| 2013 | An Efficient Power-Aware Resource Scheduling Strategy in Virtualized DatacentersabstractIn the era of cloud computing, data centers are well-known to be bounded by the power wall issue. This issue lowers the profit of service providers and obstructs the expansions of data center's scale. As virtual machine's behavior was not explored sufficiently in classic data center's power-saving strategies, in this paper we address the power consumption issue in the setting of a virtualized data center. We propose an efficient power-aware resource scheduling strategy that reduces data center's power consumption effectively based on VM live migration which is a key technical feature of cloud computing. Our scheduling algorithm leverages the Xen platform and consolidates VM workloads periodically to reduce the number of running servers. To satisfy each VM's service level agreements, our strategy keeps adjusting VM placements between scheduling rounds. We developed a power-aware data center simulator to test our algorithm. The simulator runs in time domain and includes server's segmented linear power model. We validated our simulator using measured server power trace. Our simulation shows that compared with event-driven schedulers, our strategy improves data center power budget by 35% for random workloads resembling web-requests, and improve data center power budget by 22.7% for workloads exhibiting stable resource requirements like ScaLAPACK. Yazhou Zu, Tian Huang, Yongxin Zhu 0001 |
ICPADS | 3 |
| 2013 | An Intelligent Anomaly Detection and Reasoning Scheme for VM Live Migration via Cloud Data MiningabstractCloud computing operators provide flexible, convenient, and affordable means to access public and private services. Virtual machine (VM) live migration, as an important feature of virtualization technique in cloud computing, ensures high efficiency and performance of computing infrastructure, while it stays transparent to clients. However, VM live migration is observed to cover anomalies due to their statistical similarity. To tackle the critical security issue, in this work, we propose an intelligent scheme to mine statistical data from cloud infrastructure to detect anomalies even if VMs are migrated to a new host with different infrastructure settings. In addition to detection of the existence of anomalies, our scheme is capable of identifying the possible sources of anomalies, which gives administrators clues to pinpoint and clear the anomalies. Tian Huang, Yongxin Zhu 0001 |
ICTAI | 4 |
| 2013 | Clustering scheduling for hardware tasks in reconfigurable computing systems
Zhi Chen 0008, Meikang Qiu, Zhong Ming 0001, Laurence T. Yang, Yongxin Zhu 0001 |
J. Syst. Archit. | 5 |
| 2013 | Informer homed routing fault tolerance mechanism for wireless sensor networks
Meikang Qiu, Zhong Ming 0001, Jianning Liu, Gang Quan, Yongxin Zhu 0001 |
J. Syst. Archit. | 6 |
| 2013 | Thermal-aware task scheduling in 3D chip multiprocessor with real-time constrained workloadsabstractChip multiprocessor (CMP) techniques have been implemented in embedded systems due to tremendous computation requirements. Three-dimension (3D) CMP architecture has been studied recently for integrating more functionalities and providing higher performance. The high temperature on chip is a critical issue for the 3D architecture. In this article, we propose an online thermal prediction model for 3D chips. Using this model, we propose novel task scheduling algorithms based on rotation scheduling to reduce the peak temperature on chip. We consider data dependencies, especially inter-iteration dependencies that are not well considered in most of the current thermal-aware task scheduling algorithms. Our simulation results show that our algorithms can efficiently reduce the peak temperature up to 8.1 ˆ C. Meikang Qiu, Jianwei Niu 0002, Laurence T. Yang, Yongxin Zhu 0001, Zhong Ming 0001 |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2013 | Extending Amdahl's law and Gustafson's law by evaluating interconnections on multi-core processors
Tian Huang, Yongxin Zhu 0001, Meikang Qiu, Xu Wang 0010 |
J. Supercomput. | 2 |
| 2012 | Prototyping Efficient Desktop-as-a-Service for FPGA Based Cloud Computing ArchitectureabstractCloud computing, a delivery of computing as a service mainly implying how to use utilities in our context, can be provided either at infrastructure, platform or software levels. The Desktop-as-a-Service (DaaS) paradigm, derives from the software level Software-as-a-Service (SaaS) paradigm, is drawing increasing interest because of its transformation from desktops into a cost-efficient, scalable and comfortable subscription service. Unlike most existing solutions delivering service with various protocols based on image transmitting in PC dominating environment, we present a DaaS with cloud server technologies on FPGA to address the problem of high power consumption and heavy network traffic. With the booming of mobile cloud computing, users can access the service on demand with smart phones or other portable devices like iPad or Amazon kindle as well as PC. Our system provides virtual desktop web pages written in HTML/JavaScript to avoid frequent image transmissions and reduce network traffic. To build the cloud prototype system, we combine Lightweight TCP/IP stack (LwIP) and Java Optimized Processor (JOP) to build a web server enabling dynamic web page interactions. Our system significantly saves volumes of data in transmission and network bandwidth. Analytical performance evaluation shows that on average, our system suffers only 25% transmitting latency and saves 46% of energy efficiency in comparison to other solutions. Our efficient DaaS based on FPGA explores new application of embedded web server in green cloud computing as well as new service paradigm of mobile cloud computing. Shi Shu, Yongxin Zhu 0001, Tian Huang, Shunqing Yan |
IEEE CLOUD | 3 |
| 2012 | Optimizing Scheduling in Embedded CMP Systems with Phase Change MemoryabstractPhase Change Memory (PCM) is emerging as one of the most promising alternative technology to the Dynamic RAM (DRAM) when building large-scale main memory systems. Even though the PCM is easy to scale, it encounters serious endurance problems. Writes are the primary wear mechanism in the PCM. The PCM can perform 108to 109times of writes before it cannot be programmed reliably. In addition, the PCM has high write latency. To prolong the lifetime of the PCM as the main memory and enhance the performance, we propose a Scratch Pad Memory (SPM) based memory mechanism and an Integer Linear Programming (ILP) memory activities scheduling algorithm to reduce the redundant write operations in the PCM. The idea of our approach is to share the data copies among the SPMs, instead of writing back to the PCM main memory each time a modify occurs. Our experimental results show that the ILP scheduling can generate the optimal schedule of memory activities with minimum write operations, reducing the number of write by up to 61%. Meikang Qiu, Gang Quan, Yongxin Zhu 0001 |
ICPADS | 6 |
| 2011 | Efficient Pattern Detection for Embedded Optical Bio-sensing SystemabstractTo enable pattern detection in optical signals from a novel optical biosensor used in medical embedded system, we propose a set of efficient algorithms and their corresponding implementation on FPGA (field programmable gate array). The optical biosensor is a porous silicon micro cavity membrane, which can generate different optical reflectance spectra for varieties of molecule solutions. In measured reflectance spectra of the membrane, our design is able to detect the shift of the resonant dip which is considered as the pattern to distinguish target molecule solution of different concentration. According to measured results, besides the much higher sensitivity of the novel optical sensor than classic electrochemical methods, our FPGA based implementation of our detection algorithms also shows significant speedup over software implementation on PC. The small chip area cost of FPGA implementation of our detection algorithms further ensures feasibility of ASIC (application specific integrated circuits) to incorporate both sensors and signal processing in near future. Yingjie Cao, Yongxin Zhu 0001, Guoguang Rong, Meikang Qiu |
DASC | 2 |
| 2011 | Efficient Implementation of Thermal-Aware Scheduler on a Quad-core ProcessorabstractDue to power wall and slow performance improvement in a single core micro-architecture, multiple even many cores based processors rose as the main stream processor. Nevertheless, thermal threats regarding reliability and lifetime of processors are still among the major concerns which received much attention in terms of algorithms and hardware design to reduce processor temperature and keep application performance in recent years. In this paper, we propose and implement a thermal-aware Round-Robin scheduling algorithm for process migration in the Linux environment on a quad-core processor. Bearing designer's goals in mind, such as performance, load-balancing, and reliability, we managed to achieve much bigger temperature fall than previous results of Round-Robin scheduler on a dual-core processor as well as baseline Linux scheduler on a quad-core processor. Moreover, the performance loss due to scheduling overhead is modest in our approach. Our results indicate that thermal-aware scheduling is a valid approach to tackling thermal issues on multi-core processors. There will be increasing demand for thermal-aware scheduling as the number of cores on a single processor increases. Yongxin Zhu 0001, Jingwei Ye, Tian Huang, Yuzhuo Fu, Meikang Qiu |
TrustCom | 2 |
| 2010 | Design and Implementation of 3D Positioning Algorithms Based on RF Signal Radiation Patterns for In Vivo Micro-robotabstractRecent popularity of capsule endoscopes as a medical means to investigate lesions along digestive tract has prompted complaints by doctors about the in sufficient positioning system for the capsules. The possible major reason would be that current positioning system is based on time difference of arrival whose precision is hard to improve for short flight distances. To enable better control of active capsules, we propose to exploit the radiation patterns of radio frequency (RF) signals from the capsules. The symmetric RF signals can be sensed by receivers in similar areas of external receiver arrays. The similarity, position, shape and size of those areas are considered to make back projection calculation and identify accurate 3D positions of the signal source. Our positioning algorithms are scalable to achieve better precision than the density of receivers. Yongxin Zhu 0001, Tingting Mo, Jinlong Hou, Guoguang Rong |
BSN | 2 |
| 2010 | Real-Time Constrained Task Scheduling in 3D Chip Multiprocessor to Reduce Peak TemperatureabstractChip multiprocessor technique has been implemented in embedded systems due to the tremendous computation requirements. Three dimension chip multiprocessor architecture has been studied recently for integrating more functionalities and providing higher performance. The high temperature on chip is a critical issue for the 3D architecture. In this paper, we propose an online thermal prediction model for 3D chip. Using this model, we present a task scheduling algorithm based on rotation scheduling to reduce the peak temperature on chip. We consider the data dependencies, especially the inter-iteration dependencies which are not well considered in most of the current thermal-aware task scheduling algorithms. Our simulation result shows that our algorithm can efficiently reduce the peak temperature up to 10°C. Meikang Qiu, Jianwei Niu 0002, Tianzhou Chen, Yongxin Zhu 0001 |
EUC | 5 |
| 2010 | Task Allocation and Optimization of Distributed Embedded Systems with Simulated Annealing and Geometric ProgrammingabstractWe consider the task model of periodic tasks running on a network of processor nodes connected by a bus based on the time-triggered protocol, an industry-standard bus protocol designed for safety-critical automotive and avionics distributed embedded systems, and present an integrated optimization framework that jointly considers one or more of the following attributes: task-to-processor allocation, task priority assignment, task period assignment and bus access configuration. We adopt a hierarchical optimization framework, where each possible task allocation and priority assignment is treated as one top-level coarse-grained state, which may contain many lower-level fine-grained states defined by different task period assignments and bus access configurations. Simulated annealing is used to explore the top-level states, which calls a geometric programming solver as a subroutine to explore the lower-level states contained within a given top-level state. Performance evaluation shows that our framework has good performance in terms of solution quality and scalability. Xiuqiang He 0001, Zonghua Gu 0001, Yongxin Zhu 0001 |
Comput. J. | 3 |
| 2010 | Implementing a Thermal-Aware Scheduler in Linux Kernel on a Multi-Core ProcessorabstractAs power dissipation causes thermal issues in cooling costs, lifetime and reliability, thermal management has become an important issue in today's OS and processor design. Early OS-level thermal management schemes were proposed and evaluated mainly with simulators or analytical models. In this paper, we implement a thermal-aware round-robin scheduling algorithm in the Linux kernel, and compare its performance with the ‘Heat-and-Run’ algorithm and the default Linux baseline scheduler on an Intel Core 2 Duo processor using representative benchmarks from SPEC2000, MiBench and NetBench. Our results indicate that the current Linux scheduler can easily be enhanced with thermal-awareness to show improved performance in terms of both the on-chip temperature condition and application throughput. Yongxin Zhu 0001, Jingwei Ye, Zonghua Gu 0001 |
Comput. J. | 2 |
| 2010 | Statistical estimation and evaluation for communication mapping in Network-on-Chip
Naifeng Jing, Weifeng He, Yongxin Zhu 0001, Zhigang Mao |
Integr. | 3 |
| 2009 | Design and Implementation of a High Resolution Localization System for In-Vivo Capsule EndoscopyabstractCapsule endoscopes have been proven useful to diagnose lesions along digestive tracts in recent years. However,the capsule moving through the GI tract can’t be externally controlled, thus it may miss capturing important lesions. Locomotion is in urgent need for future in-vivo capsule endoscopy, and location is one of the most needed information to guide the movement of capsule inside the body. We present a novel method and implementation of a high resolution localization system based on UHF band RFID. It consists of an RFID reader with 3D antenna array, a tag inside the capsule, and a data processing module. We also propose a location estimation algorithm by calculating center of gravity of antennas which have detected the tag, and the results show good estimation accuracy. Furthermore, some RF parts of the system are simulated with Agilent ADS and ADIsimPLL tools. We believe that the experience gained in the process of our design would serve as an important reference for the future in-vivo capsule endoscope. Jinlong Hou, Yongxin Zhu 0001, Yuzhuo Fu, Guoguang Rong |
DASC | 2 |
| 2009 | Design of 3D Positioning Algorithm Based on RFID Receiver Array for In Vivo Micro-RobotabstractThe clinical applications of capsule endoscopes have been increasing consistently since the invention of a passive capsule endoscope was made. Though the capsule endoscopes are effective in detecting large lesions along human digestive tract, they cannot meet doctors' requirements of active control of the capsules to carefully examine small lesions. The poor precision of positioning system is one of the major hurdles blocking robots from approaching the accurate position of suspected areas. In this paper, we propose a novel algorithm to exploit the radiation pattern of a radio frequency (RF) tag inside the capsule, which forms shadows or traces on a set of receiver arrays. According to the shape of the traces and the radiation pattern, the position of the radiation source, i.e. the tag inside the capsule can be calculated with our algorithm. The details of our algorithm are presented with the simulation results in the paper. With the settings under medical constraints, the degree of positioning precision in simulation is improved to less than 1 cm horizontally and 2 cm vertically. Yongxin Zhu 0001, Tingting Mo, Jinlong Hou |
DASC | 2 |
| 2008 | A Hybrid Anti-Collision Algorithm for RFID with Enhanced Throughput and Reduced Memory ConsumptionabstractIn order to solve the transponder collision problem in a RFID system, this paper proposes a hybrid algorithm which combines strengths of existing high-performance algorithms while avoiding their major drawbacks. Experiments show this hybrid algorithm uses fewer time slots and less total communication time compared to adaptive slot-count algorithm based on ALOHA as well as enhanced anti-collision algorithm based on binary tree, both of which are superior and well accepted algorithms in literature. Meanwhile, high computational intensity of adaptive slot-count algorithm is significantly relieved, and memory consumption of enhanced anti-collision algorithm is reduced to a negligible level. Results from simulations under large number of transponders and long transponder IDs, which is the case in real RFID applications, suggest the advantage of this hybrid algorithm is prominent. Majun Zheng, Jing Xie 0010, Zhigang Mao, Yongxin Zhu 0001 |
EUC (1) | 4 |
| 2005 | Design of clocked circuits using UMLabstractProceedings of the Asia and South Pacific Design Automation Conference, ASP-DAC Zhenxin Sun, Weng-Fai Wong, Yongxin Zhu 0001, Santhosh Kumar Pilakkat |
ASP-DAC | 3 |
| 2005 | An integrated performance and power model for superscalar processor designsabstractOn current superscalar processors, performance and power issues cannot be decoupled for designers. Extensive simulations are usually required to meet both power and performance constraints. This paper describes an integrated performance and power analytical model. The model's performance and power results are in good agreement with detailed simulations, previous models and physically measured results. For designers, the model enables quick and flexible explorations into a subset of even entire huge parameter space of more than 15 workload and architectural parameters plus leakage power, feature sizes, clock and voltage. Yongxin Zhu 0001, Weng-Fai Wong, Stefan Andrei |
ASP-DAC | 1 |
| 2005 | Runtime-Coordinated Scalable Incremental Checksum Testing of Combinational CircuitsabstractCircuit testing is the most significant cost in modern chip design and production. Due to the complexity in terms of millions of gates, manufacturers often have to truncate test patterns to make the testing feasible on ATEs with limited capacities. In this paper, we present a novel approach to this challenge by run-time coordinating the algorithm and ATE. A unique combination of a #SAT solver, checksum computation and frame testing enables the efficient incremental testing. Unlike checksums from the communication domain which can only detect the existence of stuck-at faults, our approach differentiates by also locating them. In our experimental results, our method further demonstrates a shorter testing time. Stefan Andrei, Wei-Ngan Chin, Albert Mo Kim Cheng, Yongxin Zhu 0001 |
RTCSA | 4 |
| 2005 | Using UML 2.0 for System Level Design of Real Time SoC Platforms for Stream ProcessingabstractWhile enabling fast implementation and reconfiguration of stream applications, programmable stream processors expose issues of incompatibility and lack of adoption in existing stream modeling languages. To address them we describe a design approach in which specifications are captured in UML 2.0, and automatically translated into SystemC models consisting of simulators and synthesizable code under proper style constraints. As an application case, we explain real time stream processor specifications using new UML 2.0 notations. Then we expound how our translator generates SystemC models and includes additional hardware details. Verifications are made during UML execution as well as assertions in SystemC. The case study demonstrates the feasibility of fast specifications, modifications and generation of real time stream processor designs. Yongxin Zhu 0001, Zhenxin Sun, Alexander Maxiaguine, Weng-Fai Wong |
RTCSA | 1 |