VLDB 2026 Research / reviewers in the wild / expert
Yao Xin
dblp:132/2735
· DBLP profile ↗
22ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-6495-081XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 4 first-author · 11 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MegaTurbo: A Scalable FPGA-based Engine for MegaFlow Classifier in Open vSwitchabstractOpen vSwitch (OVS) is a key component in cloud and data center networks, yet its MegaFlow classifier imposes significant CPU overhead. Existing SmartNIC-based acceleration approaches for the MegaFlow classifier typically employ simplistic hardware offloading techniques, which exhibit limited scalability for dynamic, large-scale flow tables. Motivated by these challenges, we argue that a hardware accelerator specifically tailored for the MegaFlow classifier is necessary, forming the basis of our FPGA-based solution, MegaTurbo. The core innovations of MegaTurbo are threefold: (1) a scalable and hardware-friendly decision-tree based packet classification algorithm, specifically optimized for the structure of MegaFlow rules; (2) a novel hardware architecture incorporating multiple pipelined matching engines, designed to process multiple decision trees generated by the software algorithm in parallel; and (3) a heterogeneous framework composed of CPU and FPGA, which can work together to support online rule updates, with little and bounded impact on rule searching. Experimental results on a Xilinx Virtex UltraScale+ FPGA demonstrate that MegaTurbo achieves a sustained classification throughput of 500 MPPS while supporting dynamic rule updates at 300-500 KUPS on 100K-scale rulesets. These results not only validate the effectiveness of our domain-specific co-design approach, but also highlight the potential of FPGA-based SmartNICs to address the performance bottlenecks of software switches in large-scale cloud and data center networks. Zhongxian Liang, Wenjun Li 0004, Yao Xin, Ying Wan 0001, Hui Li 0022, Weizhe Zhang |
FPGA | 5 |
| 2026 | FlowTurbo: From Best-Effort to Hit-Driven MegaFlow Hardware Offloading in Open vSwitchabstractOffloading fast-path MegaFlows in Open vSwitch to hardware accelerators is a common approach for accelerating packet forwarding in modern cloud data centers. However, due to the limited capabilities of current hardware accelerators, existing solutions still rely on coarse-grained, best-effort offloading, which struggles with dynamic, large-scale traffic and results in inefficient resource utilization and limited performance gains. We present FlowTurbo, a self-adaptive, system-level offloading approach that implements hit-driven MegaFlow hardware offloading by jointly optimizing software rule scheduling and hardware rule lookup. The core innovations of FlowTurbo are threefold: (1) a traffic-aware, hit-driven MegaFlow offloading framework that selectively migrates hotspot wildcard rules to hardware; (2) a domain-specific, hardware-friendly sketch for MegaFlow rules that tracks rule hotness and enables the scheduler to make timely and precise offloading decisions; and (3) a domain-specific, algorithm-hardware co-designed packet classification accelerator that supports both line-rate rule matching and online rule updates. We implemented FlowTurbo on Open vSwitch and prototyped its hardware accelerator on a Xilinx Alveo U200. Evaluation using multiple real-world traffic traces shows that FlowTurbo achieves an average acceleration coverage of 89.4%, and the hardware accelerator delivers a maximum throughput of 400 MOPS while consuming only 3.3% of FPGA logic resources. Zhongxian Liang, Wenjun Li 0004, Yao Xin, Tong Yang 0003, Gaogang Xie, Weizhe Zhang |
SIGCOMM | 6 |
| 2026 | Welford-Sketch: Finding Steady Heavy Flows in Data Streams
Yao Xin, Lingfeng Qu, Qingfeng Tan |
IEEE Internet Things J. | 1 |
| 2026 | A Multimechanism Fuzzy Logic Control Scheme for Autonomous Vehicle System Security: Machine Learning-Supervised Data Compression With Integrated Performance EstimationabstractThis paper investigates the resilience of autonomous vehicle (AV) systems against denial-of-service (DoS) attacks in networked environments. While existing fuzzy logic and event-triggered control strategies have been explored, they often fail to address the combined impacts of DoS-induced network congestion, data compression errors, and the lack of real-time adaptive learning to maintain stability. To address this gap, this work introduces a unified, resilient fuzzy logic control framework. First, a Performance Error Estimation (PEE) framework is developed to quantify control performance degradation under cyber threats, informing the design of the resilient controller. Second, an efficient data compression scheme is designed to mitigate DoS-induced network overload, ensuring reliable real-time data transmission under constrained bandwidth. Third, an Intelligent Event-Triggered Fuzzy Controller (IETFC) is proposed, which adaptively optimizes its triggering threshold via a mini-batch machine learning algorithm, effectively balancing communication efficiency with system robustness. Validation on the CarSim–Simulink platform demonstrates that the proposed framework successfully mitigates DoS-induced performance degradation, enhances triggering sparsity, and improves computational efficiency. By embedding learning-based adaptation into fuzzy control, this work provides a scalable and secure solution for AV systems operating under adversarial conditions. Yanbin Sun, Yao Xin, Kaibo Shi, Huaicheng Yan 0001, Shiping Wen 0001, Zhihong Tian 0001 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2025 | ReCQF: Enhancing CQF Redundancy with Delay Alignment Scheduling in TSNabstractIntegrating Frame Replication and Elimination for Reliability (FRER) with Cyclic Queuing and Forwarding (CQF) in Time-Sensitive Networks (TSN) encounters redundancy failures and resource reservation inefficiencies due to length disparities across redundant paths. To address these challenges, we propose ReCQF, a Reliability-Enhanced CQF scheduling framework built on Multi-Instance CQF. ReCQF adaptively assigns redundant flows to multiple CQF queue pairs with specific cycles, effectively aligning transmission delays across redundant paths to ensure low delay and inter-path delay differences while significantly reducing resource reservations. Yan Liu 0062, Zhuyun Qi, Xingbo Feng, Shuangping Zhan, Yao Xin, Jiashuo Lin, Chenxi Ling, Ruide Cao, Weichao Li 0001, Yi Wang 0004 |
IWQoS | 5 |
| 2025 | WisePIFinder: Efficient and Accurate Detection of Persistent and Infrequent FlowsabstractIn large-scale data stream analytics, accurate identification of Persistent and Infrequent (PI) flows is of great significance for monitoring and protecting against network attacks such as Advanced Persistent Threats (APT). However, existing research focuses mainly on detecting frequent flows or persistent flows, with insufficient studies on the characterization and detection methods for PI flows. Based on the analysis of sufficient APT flows, we propose a method that combines global and local features to effectively characterize PI flows. Further, we propose a novel sketch algorithm called WisePIFinder, which aims to detect PI flows more accurately and efficiently in realtime. The key idea is to continuously filter out non-PI flows while detecting flow persistence, to achieve accurate statistics on PI flows. Experimental results show that WisePIFinder improves the F1 Score by at least 20 % and insertion throughput by at least 60 % compared to the state-of-the-art solution for detecting PI flows. All related codes have been open-sourced on GitHub. Zengxie Ma, Yao Xin, Zhuochen Fan, Tong Li 0014, Qing Liao 0001, Yi Zhao 0011, Feng Zhang 0007 |
IWQoS | 2 |
| 2025 | REFS: a novel framework for accelerated receive encrypted flow steeringabstractAbstract In virtual private network (VPN) tunnel mode, the entire original packet, including the header’s five-tuple information, is encrypted, which prevents traditional scheduling algorithms from evenly distributing packets to central processing unit (CPU) cores based on packet header information. To address the need for data security and encrypted packet scheduling, we propose a novel framework, named REFS (receive encrypted flow steering), for accelerated receive encrypted flow steering. This work creatively adopts a new method that allows encrypted packets to be distributed across CPU cores without decrypting them, overcoming limitations of traditional scheduling approaches. It efficiently distributes encrypted packets across CPU cores, enabling dynamic allocation of CPU resources. A key feature of REFS is its ability to perform this distribution without decrypting the packets, which enhances dynamic load balancing and improves system responsiveness. When integrated into the Linux kernel’s VPN functionality, REFS can potentially increase throughput by up to 50% compared to WireGuard, which is a benchmark for kernel-based VPN performance. Upon integration of REFS into userspace, network performance shows significant improvements: throughput doubles, while latency is reduced by 80%. Zengxie Ma, Yao Xin, Tong Li 0014, Feng Zhang 0007 |
Comput. J. | 2 |
| 2025 | Reversible data hiding in Redundancy-Free cipher images through pixel rotation and multi-MSB replacement
Lingfeng Qu, Xu Wang 0027, Yuan Yuan 0038, Yao Xin |
J. Inf. Secur. Appl. | 5 |
| 2025 | A Heterogeneous and Adaptive Architecture for Decision-Tree-Based ACL Engine on FPGAabstractAccess Control Lists (ACLs) are crucial for ensuring the security and integrity of modern cloud and carrier networks by regulating access to sensitive information and resources. However, previous software and hardware implementations no longer meet the requirements of modern datacenters. The emergence of FPGA-based SmartNICs presents an opportunity to offload ACL functions from the host CPU, leading to improved network performance in datacenter applications. However, previous FPGA-based ACL designs lacked the necessary flexibility to support different rulesets without hardware reconfiguration while maintaining high performance. In this paper, we propose HACL, a heterogeneous and adaptive architecture for decision-tree-based ACL engine on FPGA. By employing techniques such as tree decomposition and recirculated pipeline scheduling, HACL can accommodate various rulesets without reconfiguring the underlying architecture. To facilitate the efficient mapping of different decision trees to memory and optimize the throughput of a ruleset, we also introduce a heterogeneous framework with a compiler in CPU platform for HACL. We implement HACL on a typical SmartNIC and evaluate its performance. The results demonstrate that HACL achieves a throughput exceeding 260 Mpps when processing 100K-scale ACL rulesets, with low hardware resource utilization. By integrating more engines, HACL can achieve even higher throughput and support larger rulesets. Yao Xin, Chengjun Jia, Wenjun Li 0004, Ori Rottenstreich, Yang Xu 0010, Gaogang Xie, Zhihong Tian 0001, Jun Li 0002 |
IEEE Trans. Computers | 1 |
| 2025 | ERPC: Efficient Rule Partitioning Through Community Detection for Packet ClassificationabstractPacket classification is crucial for network security, traffic management, and quality of service by enabling efficient identification and handling of data packets. Decision tree-based rule partitioning has emerged as a prominent method in recent research. A significant challenge for decision tree algorithms is rule replication, which occurs when rules span multiple subspaces, leading to substantial memory consumption increases. Rule partitioning can effectively mitigate or eliminate this replication by separating overlapping rules. However, existing partitioning techniques heavily rely on manual parameter tuning across a wide range of possible values, making optimal solution discovery challenging. Furthermore, due to the lack of global optimization, these approaches face a critical trade-off: either the number of subsets becomes uncontrollable, resulting in diminished query speed, or rule replication becomes severe, causing substantial memory overhead. To bridge these gaps and achieve high-performance adaptive partitioning, we propose ERPC, a novel algorithm with the following key features: First, ERPC leverages graph theory to model rule sets, enabling global optimization that balances intra-group rule replication against the total number of groups. Second, ERPC advances rule set partitioning by modifying traditional community detection algorithms, strategically shifting the optimization objective from positive to negative modularity. Third, ERPC allows the rule set itself to determine the optimal number of groups, thus eliminating the need for manual parameter tuning. Experimental results demonstrate the efficacy of ERPC when applied to CutSplit, a state-of-the-art multi-tree method. It preserves 88% of CutSplit’s average classification throughput while reducing tree-building time by 89% and memory consumption by 77%. Furthermore, ERPC exhibits strong scalability, being adaptable to mainstream decision tree methods. Jinshui Wang, Yao Xin, Chongwu Dong, Lingfeng Qu |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Bubble Sketch: A High-performance and Memory-efficient Sketch for Finding Top-k Items in Data StreamsabstractSketch algorithms are crucial for identifying top-k items in large-scale data streams. Existing methods often compromise between performance and accuracy, unable to efficiently handle increasing data volumes with limited memory. We present Bubble Sketch, a compact algorithm that excels in both performance and accuracy. Bubble Sketch achieves this by (1) Recording only full keys of hot items, significantly reducing memory usage, and (2) Using threshold relocation to resolve conflicts, enhancing detection accuracy. Unlike traditional methods, Bubble Sketch eliminates the need for a Min-Heap, ensuring fast processing speeds. Experiments show Bubble Sketch outperforms the other seven algorithms compared, with the highest throughput and precision, and surpasses HeavyKeeper in accuracy by up to two orders of magnitude. Qilong Shi, Yuxi Liu 0017, Hanyue Zheng, Yao Xin, Wenjun Li 0004, Tong Yang 0003, Yangyang Wang 0001, Yang Xu 0010, Weizhe Zhang, Mingwei Xu 0001 |
CIKM | 5 |
| 2024 | A novel low-latency scheduling approach of TSN for multi-link rate networkingabstractTime Sensitive Networks (TSN), as an important representative of deterministic networks, provide low-latency and highly reliable communication services for the growing network applications that have strict requirements. Cyclic Queuing and Forwarding (CQF) is a well-known mechanism proposed by IEEE 802.1Qch for low-latency flow control of time-sensitive networks. It achieves bounded end-to-end delay and jitter transmission through a set of queues without complicated queue gating. However, most of the current work overlooks the widespread existence of multi-link rate networks in LANs and WANs, and the single-cycle CQF is unable to adjust different link rates, resulting in low bandwidth utilization and high latency. In this paper, we propose a novel scheduling approach named Multi-Cycle CQF (MCCQF) to solve the transmission problem in multi-link rate networks, aiming to reduce deterministic end-to-end delay and improve link bandwidth utilization. In addition, we formulate the scheduling constraints, being of guiding significance for designing the transmission of multi-link-rate networks, and we design an online scheduling algorithm based on it. We compare the proposed scheme with the single-cycle CQF online scheduling algorithm in hierarchical multi-link-rate networking scenarios, and the evaluation shows that our algorithm achieves better end-to-end ultra-low latency (38.9% reduction) with a smaller schedulability gap compared with single-cycle CQF. And we also improved the scheduleability based on MCCQF by utilizing internal offset. Yan Liu 0062, Shuangping Zhan, Yao Xin, Yi Wang 0004 |
Comput. Networks | 4 |
| 2024 | Recursive Multi-Tree Construction With Efficient Rule Sifting for Packet Classification on FPGAabstractAs a programmable accelerator, SmartNIC provides more opportunities for algorithmic packet classification. Our aim in this work is to achieve both line-speed rule search and efficient rule update, two highly desired metrics for SDN data plane. We leverage the parallelism offered by the FPGA in SmartNIC following an algorithm/hardware co-design paradigm. Particularly, we first design an algorithm that constructs multiple trees for the rule set with a recursive rule sifting process. Unlike traditional space-cutting-based multi-tree construction, our rule sifting mechanism breaks the space constraints of rule-to-tree mapping and enables bounded height on each tree, thus providing the potential of bounded worst-case and line-speed performance. We then design a flexible hardware architecture with multiple systolic arrays that can be implemented in parallel on FPGA. Each systolic array works as a coarse-grained pipeline, and the multiple trees constructed earlier will be mapped onto these pipeline stages. This hardware-software mapping enables bounded worst-case rule searching. Additionally, incremental rule update is achieved simply by traversing the pipeline in one pass, with little and bounded impact on rule searching. Experimental results show that our design achieves an average classification throughput of 600.8/147.5 MPPS and an update throughput of 8.2/5.9 MUPS for 10k/100k-scale 5-tuple and OpenFlow rule sets. Yao Xin, Wenjun Li 0004, Chengjun Jia, Yang Xu 0010, Bin Liu 0001, Zhihong Tian 0001, Weizhe Zhang |
IEEE/ACM Trans. Netw. | 1 |
| 2023 | Comparative analysis of neural networks techniques to forecast Airfare PricesabstractWith the growth of tourism industry, airplanes have became an affordable choice for medium- and long-distance travels. Accurate forecasting of flights tickets helps the aviation industry to match demand, supply flexibly and optimize aviation resources. Airline companies use dynamic pricing strategies to determine the price of airline tickets to maximize profits. Passengers want to purchase tickets at the lowest selling price for the flight of their choice. However, airline tickets are a special commodity that is time-sensitive and scarce, and the price of airline tickets is affected by various factors.Our research work provides a systematic comparison of various traditional machine learning methods (i.e., Ridge Regression, Lasso Regression, K-Nearest Neighbor, Decision Tree, XGBoost, Random Forest) and deep learning methods (e.g., Fully Connected Networks, Convolutional Neural Networks, Transformer) to address the problem of airfare prediction, by keeping the consumers’ needs. Moreover, we proposed innovative Bayesian neural networks, which represent the first exploitation attempt of Bayesian Inference for the airfare prediction task, to the best of our knowledge. Therefore, we evaluate the performance of our implemented and optimized models on an open dataset. The experimental results show that deep learning-based methods achieve better results on average than traditional ones, while Bayesian neural networks can achieve better performance among the other machine learning methods. However, taking into account both prediction performance and computational time, the Random Forest turns out to be the best choice to apply in this scenario. Alessandro Aliberti, Yao Xin, Alessio Viticchié, Enrico Macii, Edoardo Patti |
COMPSAC | 2 |
| 2023 | MCCQF: Low-Latency Transmission Based on IEEE 802.1 Qch For Hierarchical Networkingabstract5G and Industrial Internet are bringing a variety of applications with on-time and reliable demands. Cyclic queuing and forwarding (CQF), a well-known mechanism defined by IEEE 802.1 Qch in Time Sensitive Network (TSN), achieves deterministic end-to-end latency and jitter without complex gating calculations. However, most of the current work ignores the prevalence of hybrid networks with different link rates, resulting in low bandwidth utilization and high latency for single-cycle CQF. In this paper, we propose a multi-cycle CQF to address the transmission in multi-link-rate networking, reducing deterministic end-to-end latency and improving link bandwidth utilization. In addition, we formulate the scheduling constraints, being of guiding significance for designing the transmission of multi-link-rate networks, and we design an online scheduling algorithm based on it. We compare the proposed scheme with the single-cycle CQF online scheduling algorithm in hierarchical multi-link-rate networking scenarios, and the evaluation shows that our algorithm achieves better end-to-end ultra-low latency (38.9% reduction) with a smaller schedulability gap compared with single-cycle CQF. Yan Liu 0062, Dajun Zhou, Shuangping Zhan, Yao Xin, Jiashuo Lin, Xingbo Feng, Enze Shi, Ye Qi, Junqing Zheng, Yi Wang 0004 |
ICC | 4 |
| 2022 | HybridTSS: A Recursive Scheme Combining Coarse- and Fine- Grained Tuples for Packet ClassificationabstractThe popular OpenFlow virtual switch Open vSwitch (OVS) uses a variant of Tuple Space Search (TSS) for packet classification. Although it is easy for rule updates, the lookup performance is poor. By introducing partial trees into TSS, the recently proposed CutTSS improves the lookup performance of TSS. However, it is challenging to replace TSS in OVS for two reasons: (1) the hand-tuned partitioning heuristics are rule-set dependent; (2) the complex and irregular data structures make it difficult to be integrated and maintained in real systems. To address these issues, we propose HybridTSS, a recursive TSS scheme for fast packet classification in OVS, which exploits three novel ideas: (1) the recursive partitioning based on reinforcement learning balances global rule partitions with low training complexity; (2) a hybrid TSS scheme combining coarse-grained and fine-grained tuples suppresses tuple explosion in TSS; (3) a heterogeneous search algorithm consisting of TSS and linear search adapts to characteristics of rules at different scales for fast lookups. Using ClassBench, we show that, while immune from the main drawbacks of CutTSS, HybridTSS retains the update performance of TSS, and achieves almost an order of magnitude higher lookup performance than TSS, making it an ideal packet classification algorithm for OVS. Yuxi Liu 0017, Yao Xin, Wenjun Li 0004, Haoyu Song 0001, Ori Rottenstreich, Gaogang Xie, Weichao Li 0001, Yi Wang 0004 |
APNet | 2 |
| 2022 | Updatable Packet Classification on FPGA with Bounded Worst-Case PerformanceabstractFPGA has been recognized as an attractive acceler-ator for line-speed packet classification in SmartNIC due to its ability to reconfigure and provide massive parallelism. As a promising algorithmic approach that can fully exploit the FPGA characteristics, decision tree based packet classification on FPGA has been actively investigated in the past decade. However, most of them suffer from unbalanced tree structures with unpredictable depths under certain rule sets, so the potential of FPGA may not be brought into full play. Worse still, few of them can support efficient rule updates on-the-fly, which is highly required in virtualized data centers. To address these issues, we design and implement an efficient hardware ar-chitecture based on the recently proposed KickTree algorithm, which consists of multiple balanced trees with bounded depth. A strategy of multi-PE (processing element), parallel search, and serial update is adopted to decouple the search and update process. The parsing of multiple tree search results adopts a modular and hierarchical design, supporting architecture with various tree numbers. Additionally, incremental rule updates can be achieved simply by traversing all PEs in one pass, with little and bounded impact on rule searching. Experimental results on FPGA show that our design can achieve an average classification throughput of 182.6 MPPS and an average update throughput of 3.1 MUPS for various 100k-scale rule sets. Yao Xin, Wenjun Li 0004, Gaogang Xie, Yang Xu 0010, Yi Wang 0004 |
HOTI | 1 |
| 2022 | High Throughput Hardware/Software Heterogeneous System for RRPN-Based Scene Text DetectionabstractRotation Region Proposal Networks (RRPN) are used to generate rotated proposals with the information of text angle for arbitrary oriented scene text detection (STD). However, the computational complexity of RRPN inference is relatively high compared with other methods, which makes it difficult for massive deployment. In this paper, the first full-stack FPGA-CPU heterogeneous system design of RRPN-based STD algorithm is proposed. A hardware/software partition method is presented to analyze and split the tasks to enhance the computation efficiency of hardware. The fast 2D Winograd algorithm and block floating point are utilized to reduce computation complexity while maintaining a relatively high precision. The implementation results show that the peak performance of MAC arrays in the proposed architecture reaches 655.4 GOPS and the energy efficiency achieves 64.9 GOPS/W. By fully exploiting the parallel and pipelined merits in the algorithms, the first hardware architectures for skew non-maximum suppression (S-NMS) layer and rotation region-of-interest (RRoI) polling layer are proposed. The throughput of the proposed hardware/software heterogeneous system achieves 40 times and 1.4 times improvements compared with CPU and GPU, respectively. Moreover, the comprehensive operating expense ratio of pure CPU, GPU, and the proposed system is 80.7:2.5:1, which indicates that it is suitable for massive deployment. Yao Xin, Donald Donglong Chen, Chongyang Zeng, Yi Wang 0004, Ray C. C. Cheung |
IEEE Trans. Computers | 1 |
| 2022 | FPGA-Based Updatable Packet Classification Using TSS-Combined Bit-Selecting TreeabstractOpenFlow switches are being deployed in SDN to enable a wide spectrum of non-traditional applications. As a promising alternative to brutal force TCAMs, FPGA-based packet classification is being actively investigated. However, none of the existing FPGA designs can achieve high performance on both search and update for large-scale rule sets. To address this issue, we propose TcbTree, an FPGA-based algorithmic scheme for packet classification. Specifically, at the algorithmic side, i) a two-stage framework consisting of heterogeneous algorithms is proposed, where most rules can be mapped into several balanced trees without rule replications, ii) for the remaining few rules, a centralized TSS (Tuple Space Search) architecture together with a real-time feedback scheme is designed to enhance the efficiency of TSS search on FPGA, and iii) a tree dilution method is designed to equalize rule distribution in trees, so that the latency of tree search can be reduced. At the hardware side, i) an efficient data structure set is designed to convert tree traversal to addressing process, which breaks the constraints of limited tree depth and imbalanced node distribution, and ii) distinct from fully pipelined designs, multiple levels of parallelism are efficiently explored with multi-core, multi-search-engine and coarse-grained pipelines herein. Experimental results using ClassBench show that, with the implementation of TcbTree on FPGA, the average classification throughputs for 1k, 10k, 32k and 100k rule sets achieve 788.8 MPPS, 404.3 MPPS, 237 MPPS and 41.8 MPPS, respectively, and the update throughput for all benchmark rule sets is above 1 MUPS. Yao Xin, Wenjun Li 0004, Guoming Tang, Tong Yang 0003, Xiaohe Hu, Yi Wang 0004 |
IEEE/ACM Trans. Netw. | 1 |
| 2021 | KickTree: A Recursive Algorithmic Scheme for Packet Classification with Bounded Worst-Case PerformanceabstractAs a promising alternative to TCAM-based solutions for packet classification, FPGA has received increasing attention. Although extensive research has been conducted in this area, existing FPGA-based packet classifiers cannot satisfy the burgeoning needs from OpenFlow, which demands large-scale rule sets and frequent rule updates. As a recently proposed hardware-specific approach, TabTree avoids rule replication and supports dynamic rule update. However, it still faces problems of unbalanced rule subset partition, unevenly distributed subtrees and excessive TSS leaf nodes when implemented on FPGA. In this paper, we propose a hardware-friendly packet classification approach called KickTree, which is elaborated by considering hardware properties. To take advantage of intrinsic parallelism of FPGA, KickTree adopts multiple balanced decision trees which can run simultaneously. The bit selection is more flexible which breaks the restriction of rule subset. Moreover, each subset size is strictly limited, leading to bounded and evenly-distributed Yao Xin, Yuxi Liu 0017, Wenjun Li 0004, Ruyi Yao, Yang Xu 0010, Yi Wang 0004 |
ANCS | 1 |
| 2015 | An Application Specific Instruction Set Processor (ASIP) for Adaptive Filters in Neural ProstheticsabstractNeural coding is an essential process for neuroprosthetic design, in which adaptive filters have been widely utilized. In a practical application, it is needed to switch between different filters, which could be based on continuous observations or point process, when the neuron models, conditions, or system requirements have changed. As candidates of coding chip for neural prostheses, low-power general purpose processors are not computationally efficient especially for large scale neural population coding. Application specific integrated circuits (ASICs) do not have flexibility to switch between different adaptive filters while the cost for design and fabrication is formidable. In this research work, we explore an application specific instruction set processor (ASIP) for adaptive filters in neural decoding activity. The proposed architecture focuses on efficient computation for the most time-consuming matrix/vector operations among commonly used adaptive filters, being able to provide both flexibility and throughput. Evaluation and implementation results are provided to demonstrate that the proposed ASIP design is area-efficient while being competitive to commercial CPUs in computational performance. Yao Xin, Xiangyu Li 0005, Ray C. C. Cheung, Dong Song, Theodore W. Berger |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2014 | GPU-based biclustering for microarray data analysis in neurocomputing
Benben Liu, Yao Xin, Ray C. C. Cheung, Hong Yan 0001 |
Neurocomputing | 2 |