Yan Luo 0001

dblp:25/2208-1 · DBLP profile ↗
← Back
65ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-5301-5092ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 4 first-author · 2 since 2021Computer networks · 23 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Software engineering, systems software and programming languages · 3Applied, interdisciplinary, general and emerging computing · 3Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Poster: Networked Multimodal Sensor Framework for Shrimp Health and Behavior Analysis
abstract
This study presents a multimodal sensing and deep learning framework to enhance monitoring shrimp health and behavior. By integrating real-time water quality sensors, acoustic monitoring, and computer vision, we track key parameters closely tied to feed consumption in healthy and diseased conditions in shrimp. Over a 5-week growth period, we analyzed shrimp feeding behavior through acoustic signals, and performed a controlled disease challenge to identify unique patterns associated with healthy and diseased condition in shrimp. The resulting data can help offer actionable insights for farm operators to develop cost effective feeding strategies and reduce cost associated with aquafeed use for both commercial applications and small individual farm owners.
Calvin Alexander Ng, Sage Lyon, Sheree Pagsuyoin, Paul J. Schofield, Arun K. Dhar, Yan Luo 0001
MobiCom6
2023 A Comprehensive Study on Efficient and Accurate Machine Learning-Based Malicious PE Detection
abstract
For safe and trustworthy digital services, fast and accurate malware detection is critical. Because of the financial rewards, ransomware assaults are one of the most commonly employed mal ware variants by cyber criminals. Because of the dynamic environment in which new malware variants arise on a regular basis, it is critical to maintain databases up-to-date in order to protect the digital world from ransomware threats. In this study, we curated the Ransomary dataset containing 2871 ransomware and 4208 benign PE files to allow researchers to use their own algorithms to accomplish fast and precise detection. We examined the Ransomary dataset and compared feature extraction and raw data techniques of static malware analysis. In the EMBER, DeepDetectNet, and Ransomary datasets, we found that effective feature selection with the LightGBM model can yield more than 0.99 AUC. Finally, we demonstrate that using raw data from the first 1KB of PE files may result in an accurate and extremely rapid response time. We intend to continuously expand Ransomary dataset and encourage more researchers to use static, dynamic, or hybrid analysis to identify ransomware more quickly and accurately.
Onur Barut, Yan Luo 0001
CCNC3
2023 R1DIT: Privacy-Preserving Malware Traffic Classification With Attention-Based Neural Networks
abstract
With the advances in deep learning techniques and the increase in the volume of network traffic data, deep neural networks trained directly with the raw traffic data have become more popular and successful for malware traffic classification without explicit feature extraction. However, most of the existing studies raises privacy concerns when using the payload data and ignore the generalization of the model to the newly emerged traffic such as DDoS detection on TLS 1.3. To overcome these limitations, we introduce a malware traffic classification system, Residual 1-D Image Transformer (R1DIT) model. We first leverage network domain knowledge by carefully parsing IP, HTTP, DNS, and unencrypted TLS record headers as sequences of bytes for input without interfering with IP addresses, port numbers and the payload. Then, we apply raw data transform and attention-based modules in our deep model to classify different malware types and benign traffic. Our results on NetML dataset show that the proposed model delivers 0.972 F1 score, nearly 0.3 higher than the feature-based methods and outperforms state-of-the-art models with 0.9999 F1 score for multi-class malware classification task using CICIDS2017 dataset. The generalization of this model has been proven using the TLS 1.3 traffic obtained from CICDDoS2019 dataset with the detection rate 0.9897 using meta-learning.
Onur Barut, Yan Luo 0001
IEEE Trans. Netw. Serv. Manag.2
2023 Flow-Based Encrypted Network Traffic Classification With Graph Neural Networks
abstract
Classifying encrypted traffic from emerging applications is important but challenging as many conventional traffic classification approaches are ineffective, thus calling for novel methods for identifying encrypted network flows. Recent machine learning and deep learning-based approaches are severely limited by their feature selection and inherent neural network architecture. More importantly, they overlook the opportunity to capture latent information in the temporal dimension of packets. As network data by nature are of non-Euclidean distance space and carry abundant chronological and temporal relations, we are inspired to utilize geometric deep learning that simultaneously takes into account packet raw bytes, metadata and packet relations for classifying encrypted network traffic. Our proposed graph neural network (GNN) model outperforms the two reference methods, convolutional neural networks (CNN) and recurrent neural networks (RNN) quantitatively as indicated by three metrics: sensitivity, precision and F1 score.
Ting-Li Huoh, Yan Luo 0001
IEEE Trans. Netw. Serv. Manag.2
2022 BBS: A Blockchain Big-Data Sharing System
abstract
Chain of custody is needed to document the sequence of custody of sensitive big data. In this paper, we design a blockchain big-data sharing system (BBS) based on Hyperledger Fabric. We denote the data stored outside of a ledger for sharing as "off-state" and "big data" (referring to extremely large data) is in this category. In our off-state sharing protocol, a sender registers a file with BBS for sharing. To acquire the file, an authenticated and authorized receiver has to use transactions and interacts with BBS in four phases, including the file transfer request, encrypted file transfer, key retrieval, and file decryption. The corresponding transactions are recorded in the ledger and serve as chain of custody to document the trail of the data. Compared with related work, BBS can perform the four phases autonomously. It utilizes the permissioned blockchain, i.e. Hyperledger Fabric, for access control and can defeat dishonest receivers. We design and implement a prototype of BBS for big file sharing. Extensive experiments were performed to validate its feasibility and performance.
Shan Wang 0008, Ming Yang 0001, Tingjian Ge, Yan Luo 0001, Xinwen Fu
ICC4
2022 A Spatio-temporal Learning for Music Conditioned Dance Generation
abstract
The music-conditioned dance generation, i.e., dancing to music, is a usage scenario of multi-modality human motion synthesis. Typically, it is a challenge to choreograph continuous motions coinciding with the melody and rhythm of the music. This paper proposes a position-wise encoding-decoding framework for spatio-temporal learning of motions and long-term skeleton-based dance generation oriented on music. Given the positional embedding of the frames in 1-minute video clips, firstly, we modularize a regional attention-based feed-forward mechanism to encode the music features. Secondly, based on the skeleton of each frame and the joint trajectories across motion frames, we formalize a graph topology to represent each dance sequence’s spatial and temporal knowledge. Specifically, we propose a graph convolutional network (GCN) based blocks to process long-term dependencies of motions and leverage the spatial and temporal features. Both music and motion paths are learned fully in positional embedding schemes and constructed by repeating the corresponding blocks. Finally, as the task of dance generation is inherently the consistency between music and motions, we proposed a cross-modality feature fusion for multimodal interaction and music-conditioned dance generation. Experimental results demonstrate that our method outperforms state-of-art methods in motion quality and motion-music correlation metrics.
Li Zhou 0014, Yan Luo 0001
ICMI2
2021 Multi-Task Hierarchical Learning Based Network Traffic Analytics
abstract
Classifying network traffic is the basis for important network applications. Prior research in this area has faced challenges on the availability of representative datasets, and many of the results cannot be readily reproduced. Such a problem is exacerbated by emerging data-driven machine learning based approaches. To address this issue, we present (Net)2database with three open datasets containing nearly 1.3M labeled flows in total, with a comprehensive list of flow features, for the research community1. We focus on broad aspects in network traffic analysis, including both malware detection and application classification. As we continue to grow them, we expect the datasets to serve as a common ground for AI driven, reproducible research on network flow analytics. We release the datasets publicly and also introduce a Multi-Task Hierarchical Learning (MTHL) model to perform all tasks in a single model. Our results show that MTHL is capable of accurately performing multiple tasks with hierarchical labeling with a dramatic reduction in training time.
Onur Barut, Yan Luo 0001, Weigang Li 0002
ICC2
2021 On Private Data Collection of Hyperledger Fabric
abstract
Hyperledger Fabric is a popular permissioned Blockchain framework for a consortium of organizations to develop Blockchain based applications and transact within the consortium. Hyperledger Fabric introduces a fine-grained access control mechanism called the private data collection (PDC), which allows private data to be shared by only a subset of participants. In this paper, we analyze PDC and show three classes of use cases in which misuse of Hyperledger Fabric features may endanger implemented Hyperledger Fabric systems. We present two groups of potential attacks including fake PDC results injection and PDC leakage against the misuse of the policy based consensus protocol. We use prototype systems to validate the discovered attacks. We also collected 6392 Hyprledger Fabric projects on GitHub and built a tool to statically analyse them. We find that 86.51% of the PDC related projects are potentially vulnerable to the fake PDC results injection attacks, and 91.67% have PDC leakage issues. We design new features for the Hyper-ledger Fabric framework to mitigate the attacks and show that the new features have minor impact on the system performance.
Shan Wang 0008, Ming Yang 0001, Yue Zhang 0025, Yan Luo 0001, Tingjian Ge, Xinwen Fu, Wei Zhao 0001
ICDCS4
2021 Deep Features Fusion with Mutual Attention Transformer for Skin Lesion Diagnosis
abstract
Early skin lesion diagnosis is crucial to prevent skin cancer, and deep learning (DL) based methods are well exploited to support dermatologists’ diagnosis. The data for the diagnosis tasks include dermoscopic lesion images and textual information. It is a challenge to learn features from the multimodal data to improve diagnostic quality. Inspired by the vision and language integration models in Visual Question Answer (VQA), we present an end-to-end neural network model for skin lesion diagnosis using both images and textual information simultaneously. Specifically, we fine-grained features from the two modalities (image and text) of the dataset by the pre-trained DL models. We propose a novel approach named Mutual Attention Transformer (MAT), which consists of self-attention blocks and guided-attention blocks, to enable the interactions between the features from both modalities concurrently. We then develop a fusion mechanism to integrate the represented features before the final classification output layer. The experimental results on the HAM10000 dataset demonstrate that the proposed method outperforms the state-of-art methods for skin lesion diagnosis.
Li Zhou 0014, Yan Luo 0001
ICIP2
2021 Lower Body Rehabilitation Dataset and Model Optimization
abstract
Human pose estimation has enabled numerous applications by classifying and tracking body movements. Although a few open datasets have emerged to facilitate the evaluation of pose detection methods, they are too generic to benefit do-main specific applications such as physical therapy which has quantitative clinical metrics and requires precise differentiation and measurement. To address this issue, we construct the first human keypoints detection dataset for physical therapy, in particular lower body rehabilitation. The dataset consists of 1,885,637 distinctive human poses for 31 lower body rehab exercises, which are performed by 20 actors under the guidance of a licensed physical therapist. Their motion are captured with state of art of motion tracking system in both 3D and 2D to establish the ground truth. Furthermore, to optimize a number of deep learning models applied on this unique dataset, we utilize Extremely Efficient Spatial Pyramid (EESP) and attention mechanism to reduce the models’ computational complexity. Our experiment results show that the optimized models achieve comparable performance with nearly 4x reduction in complexity.
Ying Li 0133, Zinan Xiong, Yan Luo 0001, Yu Cao 0002
ICME4
2021 Flow Scheduling in a Heterogeneous NFV Environment using Reinforcement Learning
abstract
Network function virtualization (NFV) allows net-work functions executed on general-purpose servers or virtual machines (VMs) instead of proprietary hardware, greatly improving the flexibility and scalability of network services. Recent trends in using programmable accelerators to speed up NFV performance introduce challenges in flow scheduling in a dynamic NFV environment. Reinforcement learning (RL) trains machine learning models for decision making to maximize returns in uncertain environments such as NFV. In this paper, we study the allocation of heterogeneous processors (CPUs and FPGAs) to minimize the delays of flows in the system. We conduct extensive simulations to evaluate the performance of reinforcement learning based scheduling algorithms such as Advantage Actor Critic (A2C), Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), and compare with greedy policies. The results show that RL based schedulers can effectively learn from past experiences and converge to the optimal greedy policy. We also analyze in-depth how the policies lead to different processor utilization and flow processing time, and provide insights into these policies.
Chun Jen Lin, Yan Luo 0001, Liang-Min Wang 0002, Li-De Chen
NAS2
2021 Prediction of pediatric activity intensity with wearable sensors and bi-directional LSTM models
Li Zhou 0014, Xiao Qu, Hongyan Guan, Yan Luo 0001
Pattern Recognit. Lett.7
2020 Blockchain-Based Secure and Privacy-Preserving Clinical Data Sharing and Integration
Yan Luo 0001
ICA3PP (3)3
2020 ACETA: Accelerating Encrypted Traffic Analytics on Network Edge
abstract
Applying machine learning techniques to detect malicious encrypted network traffic has become a challenging research topic. Traditional approaches based on studying network patterns fail to operate on encrypted data, especially without compromising the integrity of encryption. In addition, the requirement of rendering network-wide intelligent protection in a timely manner further exacerbates the problem. In this paper, we propose to leverage ×86 multicore platforms provisioned at enterprises' network edge with the software accelerators to design an encrypted traffic analytics (ETA) system with accelerated speed. Specifically, we explore a suite of data features and machine learning models with an open dataset. Then we show that by using Intel DAAL and OpenVINO libraries in model training and inference, we are able to reduce the training and inference time by a maximum order of 31× and 46× respectively while retaining the model accuracy.
Derek Manning, Xiaoban Wu, Yan Luo 0001, Weigang Li 0002
ICC4
2020 Human pose estimation based in-home lower body rehabilitation system
abstract
In this paper, we design, develop and evaluate an in-home lower body rehabilitation system based on a novel lightweight human pose estimation model. To achieve that, we first create a lower body rehabilitation dataset of 500,000 images with each image annotated with the ground truth joint point locations. The dataset consists of 31 different types of lower body rehabilitation activities from twenty volunteers. After that, we design a lightweight but powerful neural network model, which runs on a smartphone, to estimate human pose. Furthermore, we develop a series of principles for evaluating in-home rehabilitation activities of patients in terms of the range of motion and duration of activities. For the concern of privacy, all the data collected from patients are encrypted, stored and processed locally on patients' own smartphones. Only the sanitized evaluation reports are uploaded and shared with the patients' primary doctors. Our model achieves 70.8 in AP score on the COCO val2017 set with only 4.7M parameters and 1.0 GFLOPs. Using our system, patients can perform lower body rehabilitation activities at home and obtain evaluation report without the presence of physical therapists. We believe our system can greatly facilitate in-home rehabilitation and reduce the cost for patients.
Ying Li 0133, Yu Cao 0002, Benyuan Liu, Joanna Tan, Yan Luo 0001
IJCNN6
2020 Multitask LSTM Model for Human Activity Recognition and Intensity Estimation Using Wearable Sensor Data
abstract
Human activity recognition (HAR) and measuring the intensity of activity are increasingly important for healthcare applications, such as fitness tracking and patient monitoring. However, these two tasks have been performed separately, leading to delays and expensive implementations. In this article, we propose a holistic approach to achieve both HAR and intensity estimation simultaneously. We introduce a new data set with both activity types and activity intensities, and a multitask long short-term memory (LSTM) model to accurately classify the activity types and estimate the intensity of each activity. In addition, we evaluate our proposed neural network with other publicly available data sets and show that including activity intensities in the data set help multitask models perform comparably to two separate single-task models.
Onur Barut, Li Zhou 0014, Yan Luo 0001
IEEE Internet Things J.3
2020 AMACS: Automated Mobile Application Content Sensing
abstract
After a decade of rapid development, mobile devices have become an essential part of people's daily life and mobile applications generate rich contents and provide important information access. Contents inside mobile applications can provide valuable insights for social sensing and application development. However, unlike Web search, there is a lack of efficient methods to sense and index information from mobile applications. This article proposes an automated mobile application content sensing (AMACS) framework, which can be the fundamental tool to extract effectively the contents and conducting measurements in mobile applications. AMACS uses a content-aware discovery model to extract effectively the contents from a variety of mobile applications without any manual intervention. Its modular design with independent building blocks can be readily expanded into a distributed system for large-scale deployment. AMACS is compared with the existing relevant mobile automated analysis methods to evaluate its performance. The results show that the crawler can cover more contents efficiently in mobile applications with low overheads, and it can be used to extract contents for application scenarios like social sensing and network measurements.
Zexun Jiang, Yan Luo 0001, Jiaying Gong
IEEE Trans. Comput. Soc. Syst.3
2019 Dr. BFS: Data Centric Breadth-First Search on FPGAs
abstract
The flexible architectures of Field Programmable Gate Arrays (FPGAs) lend themselves to an array of data analytical applications, among which Breadth-First Search (BFS), due to its vital importance, draws particular attention. Recent attempts that offload BFS on FPGAs either simply imitate the existing CPU- or Graphics Processing Units (GPU)- based mechanisms or suffer from scalability issues. To this end, we introduce a novel data centric design which extensively extracts the potential of FPGAs for BFS with the following two techniques. First, we advocate to partition and compress the BFS algorithmic metadata in order to buffer them in fast on-chip memory and circumvent the expensive metadata access. Second, we propose a hierarchical coalescing method to improve the throughput of graph data access. Taken together, our evaluation demonstrates that the proposed design achieves, on average, 1.6× and 2.2× speedups over the state-of-the-art FPGA designs TorusBFS and Umuroglu, respectively, across a collection of graph datasets.
Eric Finnerty, Zachary Sherer, Hang Liu 0001, Yan Luo 0001
DAC4
2019 Software Hardware Co-Optimized BFS on FPGAs
abstract
No abstract available.
Zachary Sherer, Eric Finnerty, Yan Luo 0001, Hang Liu 0001
FPGA3
2019 Toward Secure, Privacy-Preserving, and Interoperable Medical Data Sharing via Blockchain
abstract
In the era of cloud computing and big data analysis, how to efficiently share and utilize medical information scattered across various care providers has become a critical problem. This paper proposes a new framework for sharing medical data in a secure and privacy-preserving way. This framework holistically integrates multi-authority attribute based encryption, blockchain and smart contract, as well as software defined networking to define and enforce sharing policies. Specifically in our framework, patients' medical records are encrypted and stored in hospital databases, where strict access controls are enforced with attribute based encryption coupled with privacy level classification. Our framework leverages blockchain technology to connect scattered private databases from participating hospitals for efficient and secure data provision, smart contracts to enable the business logic of clinical data usage, and software defined networking to revoke sharing privileges. The performance evaluation of our prototype demonstrates that the associated computation costs are reasonable in practice.
Yan Luo 0001, Yu Cao 0002, Jomol Mathew
ICPADS3
2019 Edison: Event-driven Distributed System of Network Measurement
Xiaoban Wu, Timothy Miskell, Yan Luo 0001, Liang-Min Wang 0002, Li-De Chen
IM3
2019 Quantitative Analysis of Mobile Application User Interface Design
abstract
With the development and proliferation of mobile devices, the industry of mobile applications has grown into a massive market. Well-designed user interfaces are essential for a high-quality and popular mobile application. To understand the basic principles of UI designing, this work employs a data-driven quantitative approach. In this work, we build a dataset of mobile UI and propose a systematic approach for quantitative analysis of UI design using five measurable metrics. Our dataset and analytical methods make it possible to correlate the quantitative metrics with the qualitative features so that the quality of UI designs can be reliably studied and optimized. The evaluation shows that these proposed metrics can be correlated with the qualitative features, and some UI designing suggestions or insights are generated.
Zexun Jiang, Yan Luo 0001, Jiaying Gong, Yuannan Yang, Manshan Lin
IPCCC3
2019 Ares: A Scalable High-Performance Passive Measurement Tool Using a Multicore System
abstract
Network measurement tools must support the collection of fine-grain flow statistics and scale well to the increasing line rates. However, conventional network measurement software tools are inadequate in high-speed network at the current scale. In this paper, we present Ares, a scalable high-performance passive network measurement tool to collect accurate per-flow metrics. Ares is built on a multicore platform, consisting of an effective hierarchical core assignment strategy, an efficient hash table for keeping flow statistics, a novel lockless flow statistics management scheme, as well as cache friendly prefetching. Our extensive performance evaluation shows that Ares brings about 19x speedup for 64-byte packets over existing approaches and can sustain up to a line rate of 100Gbps, while delivering the same level of fine-grained flow metrics.
Xiaoban Wu, Yan Luo 0001, Jeronimo Bezerra, Liang-Min Wang 0002
NAS2
2018 Do Bitcoin Users Really Care About Anonymity? An Analysis of the Bitcoin Transaction Graph
abstract
The pseudonymous nature of Bitcoin has sparked the twin rivaling researches in Bitcoin community, that is, either protecting or attacking anonymity. In spite of this intense battle, the answer to a primary question is absent – Do Bitcoin users themselves care about anonymity? This paper demystifies this doubt via analyzing the Bitcoin transaction graphs with the following three contributions: 1). We outline three representative metrics that can signify whether users concern about anonymity. 2). We examine the collective trend of anonymity concerns from a macroscope. 3). We pay particular attention on critical addresses in a microscope to unveil their anonymity concerns.This paper arrives at both expected conclusions and unexpected surprises. In particular, the expected ones are: rich addresses concern more about anonymity than poor ones. Miner addresses start caring about anonymity when exchange rate soars. Stock addresses never hide their intent of jump-and-dump. The surprises are: the majority of the users show weak concerns on anonymity. One can easily find both hot and cold wallet addresses owned by big organizations.
Anil Gaihre, Yan Luo 0001, Hang Liu 0001
IEEE BigData2
2018 UKSM: Swift Memory Deduplication via Hierarchical and Adaptive Memory Region Distilling
Nai Xia, Chen Tian 0001, Yan Luo 0001, Hang Liu 0001, Xiaoliang Wang 0001
FAST3
2018 EQuery: Enable event-driven declarative queries in programmable network measurement
abstract
Network measurement is critical in network management such as performance monitoring, diagnosis, and traffic engineering. However, conventional network measurement solutions are limited by simple and fixed functionalities as well as coarse-grained statistics which often fail to precisely illustrate network conditions. In this paper, we propose an event-driven declarative query language, EQuery, for programmable network management in order to design sophisticated measurement tasks and enable event mechanism to avoid human intervene. Furthermore, we design a compiler to support the query language on the EQuery Controller, which drives the chaining query workflow with nondeterministic finite automaton (NFA), and translates measurement jobs into low-level rules/states on the physical devices. Finally, we evaluate the effectiveness of our EQuery framework on a nation-wide operational network with real-time network statistics.
Yongyi Ran, Xiaoban Wu, Yan Luo 0001, Liang-Min Wang 0002
NOMS5
2018 Network measurement for 100 GbE network links using multicore processors
Xiaoban Wu, Yongyi Ran, Yan Luo 0001
Future Gener. Comput. Syst.4
2018 A programmable policy engine to facilitate time-efficient science DMZ management
Yan Luo 0001
Future Gener. Comput. Syst.3
2018 Immunization-based redundancy elimination in Mobile Opportunistic Networks-Generated big data
Junbao Zhang, Haojun Huang, Yan Luo 0001, Yinting Fan, Guan Yang
Future Gener. Comput. Syst.3
2018 A framework with data-centric accountability and auditability for cloud storage
Ke Zhou 0001, Yan Luo 0001
J. Supercomput.3
2018 A New Deep Learning-Based Food Recognition System for Dietary Assessment on An Edge Computing Service Infrastructure
abstract
Literature has indicated that accurate dietary assessment is very important for assessing the effectiveness of weight loss interventions. However, most of the existing dietary assessment methods rely on memory. With the help of pervasive mobile devices and rich cloud services, it is now possible to develop new computer-aided food recognition system for accurate dietary assessment. However, enabling this future Internet of Things-based dietary assessment imposes several fundamental challenges on algorithm development and system design. In this paper, we set to address these issues from the following two aspects: (1) to develop novel deep learning-based visual food recognition algorithms to achieve the best-in-class recognition accuracy; (2) to design a food recognition system employing edge computing-based service computing paradigm to overcome some inherent problems of traditional mobile cloud computing paradigm, such as unacceptable system latency and low battery life of mobile devices. We have conducted extensive experiments with real-world data. Our results have shown that the proposed system achieved three objectives: (1) outperforming existing work in terms of food recognition accuracy; (2) reducing response time that is equivalent to the minimum of the existing approaches; and (3) lowering energy consumption which is close to the minimum of the state-of-the-art.
Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma, Songqing Chen
IEEE Trans. Serv. Comput.3
2017 Designing Virtual Network Functions for 100 GbE Network Using Multicore Processors
abstract
Network function virtualization (NFV) introduces great flexibility in designing software-based network appliances to reduce cost and accelerate service deployment for network operators. However, with the fast development of high speed network of 100 GbE and beyond, how to efficiently design virtual network functions (VNF) on commodity servers has become a challenging problem. Although the advances in network hardware and software have facilitated the design of high speed network applications with hardware acceleration and kernel/driver optimization, how to leverage the existing techniques to design optimized high-performance VNFs still remains vague. In this study, we focus on the design and evaluation of four widely used VNFs covering the domains of network switching/routing, access control, measurement and security, by using a multicore platform supported by Intel DPDK fast packet I/O library. We describe the versatile network packet receiving and processing design options available for implementing such a programmable NFV platform for 100 Gbps network speed. With extensive experiments, we evaluate the performance of the each VNF and its design options in terms of packet drop rate, processing time per packet and delay per packet. Based on the evaluation over the collected data, we propose the optimal design with the given hardware resources to sustain the line rate while achieving the highest level of programmability.
Xiaoban Wu, Yongyi Ran, Yan Luo 0001
ANCS4
2017 Dynamic Virtual Measurement Function scheduling in software-oriented measurement environment
abstract
Network function virtualization (NFV) allows software-oriented network functions executed on general-purpose servers or virtual machines (VMs) instead of dedicated hardware, greatly improving the flexibility and scalability of network services. Consequently, NFV would facilitate versatile and dynamic network measurement services to meet the increasingly diversified measurement demands. However, it is challenging to provision the measurement services in a virtualized environment due to the stochastic nature in measurement demand and the special requirements of measurement functions such as location constraint and execution time. In this paper, we compose a measurement service chain as a Virtual Measurement Function (VMF) graph, and then propose a dynamic VMF scheduling algorithm for a software-oriented measurement system using Lyapunov optimization technique to maximize the revenue of the system while guaranteeing the Quality of Service (QoS). The scheduling algorithm decides whether to accept a measurement service request and which measurement nodes (MNs) instantiate the VMF graph. Finally, the performance of our proposed algorithm is verified through theoretical analysis and numerical evaluation. The simulation results show that the proposed algorithm can increase the total avenue by up to 10%, reduce the service average queue delay by 24%, and decrease the service reject rate by up to 10%, comparing with a heuristic algorithm.
Yongyi Ran, Xiaoban Wu, Yan Luo 0001
ICC4
2017 Improved Multimodal Representation Learning with Skip Connections
abstract
Multimodal Deep Boltzmann Machines (DBMs) have demonstrated huge successes in multimodal representation learning tasks. During inference, DBMs function as Recurrent Neural Nets (RNNs) because of the intractable distributions. To learn the parameters, optimizations can alternatively be operated on these surrogate RNNs with "truncated message passing". As a consequence, the gradient will propagate through a long chain without any local guidance which can potentially affects the optimization procedure. In this paper, we address this problem by adding skip connections during back-propagation while keeping the forward propagation (inference) untouched. With skip connections, we implicitly assign local "targets" for the states of intermediate inference loops to approach. Applied to different training criteria on different data sets, we demonstrate the proposed algorithms can consistently help to train better models while at a lower cost of training time. Experimental results show that our algorithms can achieve state-of-the-art performance on the Multimedia Information Retrieval (MIR) Flickr data set.
Yu Cao 0002, Benyuan Liu, Yan Luo 0001
ACM Multimedia4
2017 Tradeoffs Between Cost and Performance for CDN Provisioning Based on Coordinate Transformation
abstract
Today's content delivery is characterized by key trends such as converged media delivery over HTTP, increasing volumes of multimedia content delivered over IP, and elevated user expectations on quality-of-experience. In this respect, server provisioning is a critical phase of CDN management, which affects both incumbent and entrant CDN operators as well as internet service providers. However, existing tools and approaches to solve server placement problems have serious shortcomings: they offer only coarse tuning knobs and limit servers to a set of candidate sites givena priori. Our conversations with CDN operators reveal that a new provisioning mechanism is necessary to take advantage of emerging opportunities such as faster speed to roll out new locations and more access networks. In this paper, we present the design of DISC, a decision support system to help CDN operators systematically investigate different design tradeoffs and evaluate what-if scenarios. The key enabler underlying DISC is a network coordinate-based data analysis workflow that can flexibly embed different cost, performance, and workload characteristics without sacrificing the fidelity. We describe practical use cases and experiences in applying DISC to a large country-wide deployment. The results show that DISC significantly reduces average latency, deployment cost, and interdomain traffic.
Xu Zhang 0006, Shuoyao Zhao, Yan Luo 0001, Chen Tian 0001, Vyas Sekar
IEEE Trans. Multim.4
2017 Edge Provisioning with Flexible Server Placement
abstract
We present$\sf {Tentacle}$, a decision support framework to provision edge servers for online services providers (OSPs).$\sf {Tentacle}$takes advantage of the increasingly flexible edge server placement, which is enabled by new technologies such as edge computing platforms, cloudlets and network function virtualization, to optimize the overall performance and cost of edge infrastructures. The key difference between$\sf {Tentacle}$and traditional server placement approaches lies on that$\sf {Tentacle}$can discover proper unforeseen edge locations which significantly improve the efficiency and reduce the cost of edge provisioning. We show how$\sf {Tentacle}$effectively identifies promising edge locations which are close to a collection of users merely with inaccurate network distance estimation methods, e.g., geographic coordinate (GC) and network coordinate systems (NC). We also show how$\sf {Tentacle}$comprehensively considers various pragmatic concerns in edge provisioning, such as traffic limits by law or ISP policy, edge site deployment and resource usage cost, over-provisioning for fault tolerance, etc., with a simple optimization model. We simulate$\sf {Tentacle}$using real network data at global and county-wide scales. Measurement-driven simulations show that with a given cost budget$\sf {Tentacle}$can improve user performance by around 10-45 percent at global scale networks and 15-35 percent at a country-wide scale network.
Xu Zhang 0006, Hongqiang Harry Liu, Yan Luo 0001, Chen Tian 0001, Shuoyao Zhao
IEEE Trans. Parallel Distributed Syst.4
2016 P4GPU: Accelerate Packet Processing of a P4 Program with a CPU-GPU Heterogeneous Architecture
abstract
The P4 language is an emerging domain-specific language for describing the data plane processing at a network device. P4 has been mapped to a wide range of forwarding devices including NPUs, programmable NICs and FPGAs, except for General Purpose Graphics Processing Unit (GPGPU) which is a salient parallel architecture for processing network flows. In this work, we design a heterogeneous architecture with both CPU and GPU as a P4 programming target, and present a toolset to map a P4 program onto the proposed architecture. Our evaluation reveals that a P4 program can render promising performance on such architecture by parallelizing its "match+action" engine with the GPGPU accelerator. The experiment results show that the auto-configured GPU kernels achieve scalable lookup and classification speeds: the prototype system can reach up to 580 Gbps for IP lookups (64-byte packets) and 60 million classifications per second for 4k firewall rules, respectively.
Yan Luo 0001
ANCS2
2016 P4GPU: Acceleration of programmable data plane using a CPU-GPU heterogeneous architecture
abstract
The programmability of the network data plane has become one of the most desirable features within the context of software defined networks, with P4 serving as a domain-specific language for defining data plane processing. In this work, we are motivated to address the challenges of mapping a P4 defined data plane to a heterogeneous programmable hardware architecture consisting of both a CPU and a GPU, which includes a salient parallel SIMD architecture for processing network flows. We first design a toolset that can be used to map a P4 program onto the proposed architecture. We then optimize the GPU kernel designs for “match-action” primitives and present latency-hiding techniques to reduce the overheads of CPU/GPU communication. In addition, load balancing is investigated to maximize the utilization of CPU and GPU resources. Our toolset and optimizations allow a P4 program to render promising performance on the given heterogeneous architecture. Specifically, the experimental results collected on our prototype systems show that the automatically configured GPU kernels achieve scalable lookup and classification speeds with 420 million IP lookups per second, and more than 60 million classifications per second (for 4K firewall rules).
Yan Luo 0001
HPSR2
2016 DeepFood: Deep Learning-Based Food Image Recognition for Computer-Aided Dietary Assessment
Chang Liu 0033, Yu Cao 0002, Yan Luo 0001, Vinod Vokkarane, Yunsheng Ma
ICOST3
2016 Minimizing Content Reorganization and Tolerating Imperfect Workload Prediction for Cloud-Based Video-on-Demand Services
abstract
Video-on-demand (VoD) services historically rely on commercial content distribution networks (CDNs) for on-demand capacity provisioning. Content providers gradually prefer a self-managed content infrastructure because of its full control and customization. However, such a dedicated physical infrastructure could be costly in initial capital investment, and complex in management. It has become a promising alternative to host VoD services on pay-as-you-go cloud platforms, on which using dynamic server provisioning to reduce server rental cost is the key objective of content providers. In this paper we address two major challenges to reducing cost: to minimize content reorganization and to tolerate imperfect workload prediction. We first present a practical VoD servicing system design based on a pay-as-you-go cloud. We prove that previous works, focusing exclusively on cost savings, cause significant content reorganization and are vulnerable to imperfect workload prediction. To address such issues, we propose a novel idea called workload absorber, and design a provisioning algorithm called Absorb Window based on the idea. Workload absorbers eliminate the bandwidth wastage and significantly reduce content reorganization. We conduct extensive evaluations with real VoD access traces, and demonstrate the superior scalability of the proposed algorithm by producing highly optimized provisioning in seconds for thousands of servers.
Chen Tian 0001, Yi Wang 0049, Yan Luo 0001, Hongbo Jiang 0001, Wenyu Liu 0001, Jie Wu 0003
IEEE Trans. Serv. Comput.3
2015 HeteroSpark: A heterogeneous CPU/GPU Spark platform for machine learning algorithms
abstract
Analytics algorithms on big data sets require tremendous computational capabilities. Spark is a recent development that addresses big data challenges with data and computation distribution and in-memory caching. However, as a CPU only framework, Spark cannot leverage GPUs and a growing set of GPU libraries to achieve better performance and energy efficiency. We present HeteroSpark, a GPU-accelerated heterogeneous architecture integrated with Spark, which combines the massive compute power of GPUs and scalability of CPUs and system memory resources for applications that are both data and compute intensive. We make the following contributions in this work: (1) we integrate the GPU accelerator into current Spark framework to further leverage data parallelism and achieve algorithm acceleration; (2) we provide a plug-n-play design by augmenting Spark platform so that current Spark applications can choose to enable/disable GPU acceleration; (3) application acceleration is transparent to developers, therefore existing Spark applications can be easily ported to this heterogeneous platform without code modifications. The evaluation of HeteroSpark demonstrates up to 18× speedup on a number of machine learning applications.
Yan Luo 0001, Yu Cao 0002
NAS2
2015 Optimal bandwidth allocation for hybrid Video-on-Demand streaming with a distributed max flow algorithm
Chen Tian 0001, Jingdong Sun, Weimin Wu 0003, Yan Luo 0001
Comput. Networks4
2015 Demystifying commercial content delivery networks in China
abstract
Summary Over the past decade, content delivery networks (CDNs) have attracted substantial Internet traffic and improved quality of experience for Internet users. However, the evolution of the Internet ecosystem, which is driven by underlying economic incentives and ever emerging technologies, posts great challenges to the existing commercial CDNs (CCDNs). Thoroughly understanding the CDN industry from different aspects including market choice, technology, performance, tendency and infrastructure is indispensable to future Internet. In this paper, we conduct the first comprehensive study of China's CDNs using continuous, at‐scale, content‐driven measurements. Based on the massive amount of measurement data with multidimensional properties, we demystify the CCDNs in China and answer two important questions: (1) what is the development trend of CCDNs in China and (2) what are their unique characteristics. The answers to these questions have significant implications on CDN providers and users. Copyright © 2015 John Wiley & Sons, Ltd.
Bo Qiao 0007, Yan Luo 0001, Chen Tian 0001, Yang Richard Yang
Concurr. Comput. Pract. Exp.3
2015 Transformer: Run-time reprogrammable heterogeneous architecture for transparent acceleration of dynamic workloads
Yan Luo 0001, Jun Yang 0002
J. Parallel Distributed Comput.2
2014 Accelerator of Stacked Convolutional Independent Subspace Analysis for Deep Learning-Based Action Recognition
abstract
Action recognition has been a research challenge in multimedia computing and machine vision. Recent advances in deep learning combined with stacked convolutional Independent Subspace Analysis (ISA) has achieved a better performance superior to all previously published results on several public available data sets. Unfortunately, one major issue in large-scale deployment of this new deep learning-based approach is the unacceptable latency of training with high-dimension data. In this paper, we propose a new hardware accelerator that can reduce the training time substantially for deep learning-based action recognition. Specifically, our proposed approach focuses on accelerating the convolutional stacked ISA algorithm, the core components of the deep learning-based action recognition algorithms. We design parallel pipelines, data parallelisms and look-up table to speed up the algorithm. With an embedded heterogeneous platform consisting of a general purpose processor and a FPGA, we are able to achieve up to 10X speedup for stacked ISA training compared to a software-only implementation.
Yan Luo 0001, Yu Cao 0002
FCCM2
2012 Virtual network embedding through topology awareness and optimization
Xiang Cheng 0003, Sen Su, Zhongbao Zhang, Kai Shuang, Fangchun Yang, Yan Luo 0001, Jie Wang 0002
Comput. Networks6
2011 Improving IPS by network processors
Pablo Cascón, Julio Ortega 0001, Yan Luo 0001, Eric Murray, Antonio F. Díaz, Ignacio Rojas
J. Supercomput.3
2009 Accelerating OpenFlow switching with network processors
abstract
OpenFlow switching enables flexible management of enterprise network switches and experiments on regular network traffic. We present in this paper a complementary design to OpenFlow's existing reference designs. We apply network processor based acceleration cards to perform OpenFlow switching. We describe the design options and report our experiment results that show a 20% reduction on packet delay and the comparable packet forwarding throughput compared to conventional designs.
Yan Luo 0001, Pablo Cascón, Eric Murray, Julio Ortega 0001
ANCS1
2009 Discernibility Analysis and Accuracy Improvement of Machine Learning Algorithms for Network Intrusion Detection
abstract
Network intrusion detection based on machine learning algorithms has demonstrated high performance in execution time and overall classification accuracy. However, very poor identification skill is showed for certain specific attack types, especially for the unknown attack types appeared in the test data only. We use the Parallel Coordinates Plot (PCP), one kind of visualization technique for multi-dimension data analysis, to comparatively analyze the data distribution characteristic for both training and test datasets. On the other hand, we make use of rough sets theory to investigate the discernibility in respect of whole training dataset, randomly sampled dataset and reduct attributes set. Furthermore, based on the higher classification accuracy for data with unknown attack types by using rough sets method, the decision rules extracted from both C4.5 and rough sets method are combined to improve the detection capability of classification model.
Sanping Li, Yan Luo 0001
ICC2
2009 Distributed Intrusion Detection with Intelligent Network Interfaces for Future Networks
abstract
Intrusion detection remains an important and challenging task in current and next generation networks (NGN). Emerging technologies such as multi-core processors and virtualization have changed the architecture of the building elements of NGN significantly, thus call for rethinking of how network processing is done. In this paper, we propose distributed intrusion detection using intelligent network interfaces where additional processing capabilities are available. We design and implement a prototype to perform pattern matching using network processors since pattern matching is one of the important workloads in intrusion detection. Through the experimental results, we show the feasibility and performance of distributed intrusion detection in next generation networks.
Yan Luo 0001, Ke Xiang, Jie Fan 0002
ICC1
2008 Acceleration of decision tree searching for IP traffic classification
abstract
Traffic classification remains a hot research problem, especially when facing new traffic trends and new hardware architectures. We propose a classification tree search method called explicit range search, motivated by the characteristics of machine learning based classification approaches. Our method differs from previously known algorithms such as HiCut and HyperCut in how to cut the ranges within a dimension and how to search within the ranges. By storing explicit marks and performing hardware supported parallel comparison, the explicit range search can reduce the worst-case number of memory accesses from 26 to 5 on a number of realistic rule sets generated from a well-known machine learning algorithm (C4.5). We also describe in this paper the proposed design based on FPGA devices.
Yan Luo 0001, Ke Xiang, Sanping Li
ANCS1
2008 Design of high performance pattern matching engine through compact deterministic finite automata
abstract
Pattern matching relies on deterministic finite automata (DFA) to search for predefined patterns. While a bit-DFA method is recently proposed to exploit the parallelism in pattern matching, we identify its limitations and present two schemes, Label Translation Table (LTT) and CAM-based Lookup Table (CLT), to reduce the DFA memory size by 85%, and simplify the design by requiring only four processing elements of bit-DFA instead of thousands.
Piti Piyachon, Yan Luo 0001
DAC2
2008 The Design of a Programmable Edge Node with Hybrid Multi-Core Processors for Virtual Networks
abstract
Emerging new network applications challenge the existing Internet infrastructure and call for flexible and open network facilities that can provide virtual networks for conducting experiments or deploying new services. In this paper, we propose to use hybrid multi-core processors including general purpose ones and network processors as the main processing elements of a programmable edge node (PEN), an integral part of the virtualizable network substrate. We describe the design of the PEN and compare it with two newly proposed architectures. We present experiment results on the performance limitations of an existing design and show the potential of network processors in forwarding packets, isolating experiments and measurements, and offloading complex packet processing.
Yan Luo 0001
ICCCN1
2008 Fault tolerant practices on network processors for dependable network processing
abstract
In this paper, we study how to provide dependable network processing through multi-core based network processors (NPs). We present the performance analysis results of an NP based network system and motivate our research. We propose to use the redundant cores available in an NP to handle faults. We outline a fault-tolerant (FT) task model specifically for NP based applications and describe our implementation taking advantage of hardware features of an Intel IXP2xxx NP. A set of experiments are conducted to evaluate the performance and effectiveness of a FT-enabled NP system. The experiment results show that our fault- tolerant design can effectively improve the schedulability of the system.
Yan Luo 0001, Jie Fan 0002
IPDPS1
2007 DPICO: a high speed deep packet inspection engine using compact finite automata
abstract
Deep Packet Inspection (DPI)has been widely adopted in detecting network threats such as intrusion, viruses and spam. It is challenging, however, to achieve high speed DPI due to the expanding rule sets and ever increasing line rates. A key issue is that the size of the finite automata falls beyond the capacity of on-chip memory thus incurring expensive off-chip accesses. In this paper we present DPICO a hardware based DPI engine that utilizes novel techniques to minimize the storage requirements for finite automata. The techniques proposed are modified content addressable memory (mCAM), interleaved memory banks, and data packing. The experiment results show the scalable performance of DPICO can achieve up to 17.7 Gbps throughput using a contemporary FPGA chip. Experiment data also show that a DPICO based accelerator can improve the pattern matching performance of a DPI server by up to 10 times.
Christopher L. Hayes, Yan Luo 0001
ANCS2
2007 Compact State Machines for High Performance Pattern Matching
abstract
Pattern matching is essential to a wide range of applications such as network intrusion detection, virus scanning, etc. Pattern matching algorithms normally rely on state machines to detect predefined patterns. Recently, parallel pattern matching engines, based on ASICs, FPGAs or network processors, perform matching with multiple state machines. The state migration in the matching procedure incurs intensive memory accesses. Thus, it is critical to minimize the storage of state machines such that they can be fit in on-chip or other fast memory modules to achieve high-speed pattern matching. This paper proposes novel optimization techniques, namely state re-labeling and memory partition, to reduce state machine storage. The paper also presents architectural designs based on the optimization strategy. We evaluate our design using realistic pattern sets, and the results show state machine memory reduction up to 80.1%.
Piti Piyachon, Yan Luo 0001
DAC2
2007 Conserving network processor power consumption by exploiting traffic variability
abstract
Network processors (NPs) have emerged as successful platforms for providing both high performance and flexibility in building powerful routers. Typical NPs incorporate multiprocessing and multithreading to achieve maximum parallel processing capabilities. We observed that under low incoming traffic rates, processing elements (PEs) in an NP are idle for most of the time but still consume dynamic power. This paper develops a low-power technique to reduce the activities of PEs in accordance with the varying traffic volume. We propose to monitor the average number of idle threads in a time window, and gate off the clock signals to unnecessary PEs when a subset of PEs is enough to handle the network traffic. We solve the difficulties arising from clock gating the PEs, such as redirecting network packets, determining the thresholds of turning on/off PEs, and avoiding unnecessary packet loss. Our technique brings significant reduction in power consumption of NPs with no packet loss and little impact on overall throughput.
Yan Luo 0001, Jia Yu 0008, Jun Yang 0002, Laxmi N. Bhuyan
ACM Trans. Archit. Code Optim.1
2006 Efficient memory utilization on network processors for deep packet inspection
abstract
Deep Packet Inspection (DPI) refers to examining both packet header and payload to look for predefined patterns, which is essential for network security, intrusion detection and content-aware switch etc. The increasing line speed and expanding pattern sets make DPI a challenging task. Network Processors (NPs) are chosen to perform DPI due to their packet processing performance and programmability. In this paper, we focus on achieving high performance DPI through exploitation of NP's on-chip resources (particularly memory) and inherent parallel processing capability. We study the parallelism in classical DPI algorithms and construct a memory model for different parallel matching methods. Based on the model, we find the optimal organization of state machines that requires minimal on-chip memory space and guides us to high performance NP architectures for DPI. The performance evaluation experiments show that our method can reduce the memory usage by up to 86%. With an Intel IXP28xx NP simulator, we observe that the estimated DPI throughput reaches up to 5 Gbps.
Piti Piyachon, Yan Luo 0001
ANCS2
2005 SpliceNP: a TCP splicer using a network processor
abstract
TCP Splicing can be used in content-aware switches to tremendously reduce overall request latency. In order to reduce the processing latency further, we propose to offload the protocol processing onto network processors (NPs). An NP consists of a multithreaded multiprocessor architecture that can provide high throughput for packet processing or forwarding. However, offloading any protocol software to an NP needs to be carefully designed due to its low-level programming and limited control memory size.In this paper, we first analyze the operation of TCP Splicing in detail and evaluate its performance through measurements on a Linux-based switch. Then various possibilities of workload allocation among different computation resources in an NP are presented, and the design tradeoffs are discussed. A content aware switch is implemented using IXP 2400 NP and evaluated for performance comparison. The measurement results demonstrate that our NP-based switch can reduce the http processing latency by an average of 83.3% for a 1K byte web page. The amount of reduction increases with larger file sizes. It is also shown that the packet throughput can be improved by up to 5.7x across a range of files by taking advantage of multithreading and multiprocessing, available in the NP.
Li Zhao 0002, Yan Luo 0001, Laxmi N. Bhuyan, Ravi R. Iyer 0001
ANCS2
2005 Low power network processor design using clock gating
abstract
Network processors (NPs) have emerged as successful platforms to providing both high performance and flexibility in building powerful routers. Typical NPs incorporate multiprocessing and multi-threading to achieve maximum parallel processing capabilities. We observed that under low incoming traffic rates, most processing elements (PEs) in NPs are nearly idle and yet still consume dynamic power. This paper develops a low power technique to reduce the activities of PEs according to the varying traffic volume. We propose to monitor the average number of idle threads in a time window, and gate off the clock network of unused PEs when a subset of PEs is enough to handle the network traffic. We show that our technique brings significant reduction in power consumption (up to 30%) of NPs with no packet loss and little impact to the overall throughput.
Yan Luo 0001, Jia Yu 0008, Jun Yang 0002, Laxmi N. Bhuyan
DAC1
2005 Optimal network processor topologies for efficient packet processing
abstract
In this paper, we propose a novel strategy to determine the optimal network processor (NP) topology for the target application tasks. We partition network applications into different stages with the consideration of limited instruction memory of the processing elements (PEs). We develop a theoretical approach to determine an optimal topology of the PEs via multiple pipelines. The idea of multiple pipelining is to exploit the task/packet level parallelism and the pipelines are further optimized to achieve the maximum throughput and resource utilization. Simulation results verify our analytical model and demonstrate the robustness of our approach in different NP configurations.
Jingnan Yao, Yan Luo 0001, Laxmi N. Bhuyan, Ravi R. Iyer 0001
GLOBECOM2
2005 Enhancing Network Processor Simulation Speed with Statistical Input Sampling
Jia Yu 0008, Jun Yang 0002, Shaojie Chen, Yan Luo 0001, Laxmi N. Bhuyan
HiPEAC4
2004 Utilizing Formal Assertions for System Design of Network Processors
abstract
System level modeling with executable languages such as C/C++ has been crucial in the development of large electronic systems from general processors to application specific designs. To make sure that the executable models behave as they should, the designers often have to "eye-ball" the simulation traces and at best, apply simple "assert" statements or write simple trace checkers in some scripting languages. The problem is the lack of a concise and formal method to specify and check desired properties, whether they be functional or performance in nature. In this paper, we apply assertion checking methodology to the system design of network processors. Functional and performance assertions, based on linear temporal logic and logic of constraints, are written during the design process. Trace checkers and simulation monitors are automatically generated to validate particular simulation runs or to analyze their performance characteristics. Several categories of assertions are checked throughout the design process, such as equivalence, functionality, transaction, and performance. We demonstrate that the assertion-based methodology is very useful for both system level verification and design exploration.
Xi Chen 0024, Yan Luo 0001, Harry Hsieh, Laxmi N. Bhuyan, Felice Balarin
DATE2
2003 Shared memory multiprocessor architectures for software IP routers
abstract
We propose new shared memory multiprocessor architectures and evaluate their performance for future Internet protocol (IP) routers based on symmetric multiprocessor (SMP) and cache coherent nonuniform memory access (CC-NUMA) paradigms. We also propose a benchmark application suite, RouterBench, which consists of four categories of applications representing key functions on the time-critical path of packet processing in routers. An execution driven simulation environment is created to evaluate SMP and CC-NUMA router architectures using this RouterBench. The execution driven simulation can produce accurate cycle-level execution time prediction and reveal the impact of various architectural parameters on the performance of routers. We port the FUNET trace and its routing table for use in our experiments. We find that the CC-NUMA architecture provides an excellent scalability for design of high-performance IP routers. Results also show that the CC-NUMA architecture can sustain good lookup performance, even at a high frequency of route updates.
Yan Luo 0001, Laxmi N. Bhuyan, Xi Chen 0024
IEEE Trans. Parallel Distributed Syst.1
2001 Experiences with Oasis+: A Fault Tolerant Storage System
abstract
The Oasis+ distributed storage system is a reliable memory store for small scale computing clusters. It is implemented entirely in-memory using Distributed Shared Memory (DSM) and was built to operate as a backbone service for a computing cluster that supports mobile workstations or remote clients needing fast access to storage. The system can store data quickly in a dependable manner in part by using a highperformance, high-availability page-based protocol called BR [1]. BR guarantees robust functionality despite multiple site failures that could occur. By integrating address range locking and eager release consistency (ERC) [2] Oasis+ provides a flexible and efficient platform for the development of distributed services. Reliability is achieved by replication and corrective cleanup recovery actions once failures arise.
Yan Luo 0001, Brett D. Fleisch
CLUSTER2