Bei Hua

dblp:74/1602 · DBLP profile ↗
← Back
42ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0001-7281-8977ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Databases, data management, data science and information retrieval · 9 · 4 since 2021Computer networks · 8 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Prototype-Calibrated Graph Prompting for Few-Shot Graph Adaptation
abstract
Graph Neural Networks (GNNs) increasingly follow the ''pre-training, adaptation'' paradigm, where a GNN is pre-trained on large-scale graphs and then adapted to downstream tasks. Graph prompting adapts to the frozen encoder by modifying the input graph structure, rather than fine-tuning the model parameters. However, existing graph prompting methods often rely on probabilistic rewiring and auxiliary regularizers to control sparsity, which makes the prompting process sensitive to hyperparameters and can introduce instability in few-shot settings. To address the issue, we propose ProtoCalib, a lightweight local graph prompt for few-shot adaptation of frozen GNNs. ProtoCalib uses prototypes from the support set to score candidate edges for each anchor node. It then calibrates these scores with a per-node edge budget, which keeps the prompted graph sparse and removes the need for extra sparsity or entropy losses. We further adopt a deterministic construction of the prompted adjacency to reduce sampling noise at inference time. Extensive experiments on five graph datasets under four pre-training strategies demonstrate that our proposed ProtoCalib outshines baselines on multiple node classification datasets.
Yan Yan 0030, Yingzi Shi, Bei Hua, Shuo Wen
SIGIR5
2025 SpatialSplat: Efficient Semantic 3D from Sparse Unposed Images
abstract
A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage costs of high-dimensional semantic features, existing methods extend this paradigm by associating each primitive with a compressed semantic feature vector. However, these methods have two major limitations: (a) the naively compressed feature compromises expressiveness, affecting the model's ability to capture fine-grained semantics, and (b) the pixel-wise primitive prediction introduces redundancy in overlapping areas, causing unnecessary memory overhead. To this end, we introduce \textbf{SpatialSplat}, a feedforward framework that produces redundancy-aware Gaussians and capitalizes on a dual-field semantic representation. Particularly, with the insight that primitives within the same instance exhibit high semantic consistency, we decompose the semantic representation into a coarse feature field that encodes uncompressed semantics with minimal primitives, and a fine-grained yet low-dimensional feature field that captures detailed inter-instance relationships. Moreover, we propose a selective Gaussian mechanism, which retains only essential Gaussians in the scene, effectively eliminating redundant primitives. Our proposed Spatialsplat learns accurate semantic information and detailed instances prior with more compact 3D Gaussians, making semantic 3D reconstruction more applicable. We conduct extensive experiments to evaluate our method, demonstrating a remarkable 60\% reduction in scene representation parameters while achieving superior performance over state-of-the-art methods. The code is available at https://github.com/shengyuuu/SpatialSplat.git
Yu Sheng, Jiajun Deng, Yu Zhang 0086, Bei Hua, Yanyong Zhang, Jianmin Ji
ICCV5
2025 AG-Mask: Augmented 3D Generative Masked Motion Model for Text-to-Motion
abstract
Generating natural 3D human motions that are consistent with textual descriptions is a key task in text-to-motion generation. The Transformer-based text-conditional mask motion generation model relies on the multi-head attention mechanism and mask training strategy, making significant progress in generating high-quality and high-fidelity 3D human motions. However, these models only use the multi-head attention mechanism to capture the long-distance dependence of features, lacking the modeling of spatio-temporal relationship of motion sequence, which affects the coherence of the generated motion. In addition, the multi-head attention mechanism has limited ability to capture the detailed information of motions, which affects the authenticity and accuracy of generated motions. Therefore, we propose an augmented 3D generative masked motion model (AG-Mask), which significantly enhances the model’s ability to capture spatiotemporal feature and detailed feature of motion sequence, and effectively generates high-quality motions consistent with text descriptions. Specifically, we design two augmented bidirectional transformer in AG-Mask: STM-Transformer and MDR-Transformer, which are used to process the basic and detailed information of motions respectively. STM-Transformer can boost the information extraction ability of the model in the motion channel and spatial dimension. MDR-Transformer can model the spatio-temporal relationship of motions and extract rich multidimensional features. Their combined processing promotes the generation of motion sequence from coarse to fine, optimizing the quality of generated motions. Experiments on the HumanML3D and KIT-ML datasets show that AG-Mask achieves state-of-the-art performance in generating high-quality motions. In addition, AG-Mask refines the motion editing function, making it more flexible for practical applications.
Zixin Su, Mengxiao Yin, Fancui Xie, Peihong Wu, Bei Hua, Feng Zhan
IJCNN5
2025 MD-Mono: Lightweight Self-Supervised Monocular Depth Estimation Based on Multi-Scale Adaptive Detail Enhancement
abstract
Self-supervised monocular depth estimation has garnered widespread attention because it does not require hard-to-obtain depth labels during training. Many existing studies have focused on the design of depth encoders, often neglecting the potential of decoders, which results in decoders that struggle to utilize the multi-scale features extracted by the encoder fully, lack the ability to capture the features comprehensively, and also fall short in recovering local details. To address these issues, this paper proposes a lightweight self-supervised monocular depth estimation architecture called MD-Mono. MD-Mono employs a hybrid depth encoder combining Convolutional Neural Networks (CNNs) and Transformers, aiming to capture both local features and global semantic information. In the depth decoder, we propose an Adaptive Depth Focus (ADF) module and an Implicit Detail Enhancement (IDE) module. The ADF module adaptively adjusts each stage of the decoding process according to the input features, effectively integrating and utilizing multi-scale features. The IDE module implicitly maps the input to a high-dimensional, nonlinear feature space, capturing more detailed feature information for recovering local details. The synergy of these two modules enables our architecture to achieve a semantically richer and spatially more accurate representation with fewer parameters. Experimental results show that MD-Mono significantly outperforms Monodepth2 in terms of accuracy and exhibits good generalization ability on the Make3D and DrivingStereo datasets.
Peihong Wu, Mengxiao Yin, Pengfei Lai, Zixin Su, Feng Zhan, Bei Hua
IJCNN6
2025 SMA-GNN: A Symbol-Aware Graph Neural Network for Signed Link Prediction in Recommender Systems
abstract
Recommender Systems (RS) play a critical role in enhancing user experiences across online platforms by modeling user-item interactions as bipartite graphs. Predicting signed links in such graphs remains challenging due to the sparsity and complexity of sign distributions and the limitations of traditional methods like matrix factorization and Graph Convolutional Networks (GCNs), which often fail to capture the intricate local topological and sign-based patterns essential for accurate predictions. To address these challenges, we propose SMA-GNN, a framework specifically designed for signed link prediction in bipartite graphs. SMA-GNN combines Local Subgraph Extraction, Two-Anchor Distance Labeling (TADL), and a Symbol-aware Multi-head Attention Mechanism to enhance predictive capability and interpretability. By extracting a closed local subgraph around the target link, our method captures relevant topological and sign contexts. TADL refines this by assigning unique structural labels to nodes based on their proximity to anchor nodes, encapsulating roles and relationships. The symbol-aware attention mechanism integrates edge sign information into the message-passing process, generating highly discriminative subgraph embeddings. Experiments on benchmark datasets show that SMA-GNN outperforms global embedding methods in prediction accuracy and provides deeper insights into user-item interactions, enabling more precise and personalized recommendations. Our code is avilable at https://github.com/xiaohuzidefeijian/SMAGNN/tree/master
Hongxiang Lin, Shuo Wen, Bei Hua
KDD (2)5
2024 NaviFormer: A Data-Driven Robot Navigation Approach via Sequence Modeling and Path Planning with Safety Verification
abstract
Reinforcement learning has shown great potential in improving the performance of robot navigation. In response to the increasing deployments of mobile robots within various scenarios, a data-driven paradigm of navigation approach with safety verification is preferred where one can train RL algorithms with large amounts of prior data, keep learning continuously, and ensure safe navigation in applications. Conventional end-to-end reinforcement learning navigation paradigms have encountered multiple challenges in meeting these demands. In this work, we introduce a novel robot navigation approach termed NaviFormer. This approach handles navigation tasks based on sequence modeling to obtain the data-driven ability. It also integrates rule-based verification for safety insurance. We conduct a series of experiments to validate the data-driven ability of our approach and to compare it with existing navigation methods. We also perform quantitative tests on a real-world robot platform, TurtleBot. The experimental results show our method’s outstanding data-driven ability and highlight its superior arrival rate and generalization compared to other state-of-the-art methods like the PPO-based navigation method.
Ziyang Feng, Quecheng Qiu, Yu'an Chen, Bei Hua, Jianmin Ji
ICRA5
2023 RoUD: Scalable RDMA over UD in Lossy Data Center Networks
abstract
Remote direct memory access (RDMA) has been widely deployed in data centers due to the lower latency and higher throughput of the kernel TCP/IP stack. However, RDMA still faces a scalability problem including connection scalability and network scalability issues. In this paper, we present RoUD, a userspace network stack that leverages the unreliable datagram (UD) transport mode of RDMA to improve connection scalability. RoUD also eliminates the dependency on PFC in data center networks, thereby enhancing network scalability. RoUdimplements three performance optimizations in the userspace network stack and introduces two types of flow control to avoid packet loss on the host from happening on the host for high performance. We built a prototype of RoUD based on the standard InfiniBand Verbs library. The evaluation results on a testbed with 100 Gbps RNICs show that in the case of large-scale connections its throughput is 1.4× better than the widely used reliable connection (RC) transport.
Bei Hua
CCGrid3
2023 A Fast, Reliable, Adaptive Multi-hop Broadcast Scheme for Vehicular Ad Hoc Networks
Ping Liu 0008, Xingfu Wang, Ammar Hawbani, Bei Hua, Liang Zhao 0004
ICA3PP (2)4
2023 Automatic Generation of Robot Facial Expressions with Preferences
abstract
The capability of humanoid robots to generate facial expressions is crucial for enhancing interactivity and emotional resonance in human-robot interaction. However, humanoid robots vary in mechanics, manufacturing, and ap-pearance. The lack of consistent processing techniques and the complexity of generating facial expressions pose significant challenges in the field. To acquire solutions with high confidence, it is necessary to enable robots to explore the solution space automatically based on performance feedback. To this end, we designed a physical robot with a human-like appearance and developed a general framework for automatic expression generation using the MAP-Elites algorithm. The main advan-tage of our framework is that it does not only generate facial expressions automatically but can also be customized according to user preferences. The experimental results demonstrate that our framework can efficiently generate realistic facial expressions without hard coding or prior knowledge of the robot kinematics. Moreover, it can guide the solution-generation process in accordance with user preferences, which is desirable in many real-world applications.
Bing Tang, Rongyun Cao, Rongya Chen, Bei Hua, Feng Wu 0001
ICRA5
2023 ST-MAN: Spatio-Temporal Multimodal Attention Network for Traffic Prediction
Ruozhou He, Bei Hua, Jianjun Tong
KSEM (2)3
2022 Reducing Write Amplification of LSM-Tree with Block-Grained Compaction
abstract
LSM-tree has been widely used as a write-optimized storage engine in many key-value stores, such as LevelDB and RocksDB. However, conventional compaction operations on the LSM-tree need to read, merge, and write many SSTables, which we call Table Compaction in this paper. Table Compaction will cause two major problems, namely write amplification and block-cache invalidation. They will lower both write and read performance of the LSM-tree. To address these issues, we propose a novel compaction scheme named Block Compaction that adopts a block-grained merging policy to perform compaction operations on the LSM-tree. Block Compaction identifies the boundaries of data blocks and tries to avoid reusing data blocks, which not only reduces the write amplification but also alleviates the block-cache invalidation. We present cost analysis to theoretically demonstrate that Block Compaction is more efficient than the existing Table Compaction. Furthermore, we analyze the side-effects of Block Compaction and present three optimizations: (1) Selective Compaction is to reduce the space amplification of Block Compaction by integrating Table Compaction with Block Compaction. (2) Parallel Merging divides a compaction task into several sub-tasks and uses multiple workers to accomplish sub-tasks in parallel. (3) Lazy Deletion mitigates the overhead caused by traversing files at the tail of compaction operations. We implement a new key-value store named BlockDB based on Block Compaction and its optimizations. Then, we compare BlockDB with LevelDB, RocksDB, and L2SM using the YCSB benchmark. The results show that BlockDB can reduce write amplification up to 32% and running time by up to 43.6%, compared to its competitors. In addition, it can maintain the high performance for point lookups and range scans.
Peiquan Jin, Bei Hua, Hai Long
ICDE3
2022 A Universal PINNs Method for Solving Partial Differential Equations with a Point Source
abstract
In recent years, deep learning technology has been used to solve partial differential equations (PDEs), among which the physics-informed neural networks (PINNs)method emerges to be a promising method for solving both forward and inverse PDE problems. PDEs with a point source that is expressed as a Dirac delta function in the governing equations are mathematical models of many physical processes. However, they cannot be solved directly by conventional PINNs method due to the singularity brought by the Dirac delta function. In this paper, we propose a universal solution to tackle this problem by proposing three novel techniques. Firstly the Dirac delta function is modeled as a continuous probability density function to eliminate the singularity at the point source; secondly a lower bound constrained uncertainty weighting algorithm is proposed to balance the physics-informed loss terms of point source area and the remaining areas; and thirdly a multi-scale deep neural network with periodic activation function is used to improve the accuracy and convergence speed. We evaluate the proposed method with three representative PDEs, and the experimental results show that our method outperforms existing deep learning based methods with respect to the accuracy, the efficiency and the versatility.
Hongsheng Liu 0002, Beiji Shi, Zidong Wang 0010, Yang Li 0106, Min Wang 0037, Haotian Chu, Fan Yu 0004, Bei Hua, Bin Dong 0001, Lei Chen 0002
IJCAI11
2022 DTS: A Dual Transport Switching Scheme for RDMA-based Applications
abstract
RDMA has been widely adopted by distributed applications for its low latency, high throughput, and near-zero CPU overhead. However, high availability is still lacked for RDMA-based applications, especially in the cases of network failure or NIC upgrade. In this paper, we propose DTS, a scheme for RDMA-based applications to provide high availability by adopting the active and standby transports and switching them transparently and fast when necessary. We implement the prototype of DTS based on the Libfabric RxD provider. Our evaluation demonstrates that with DTS existing RDMA-based applications can be deployed with zero downtime and the performance degradation is ignorable.
Junhong Ye, Bei Hua
IPCCC7
2022 Meta-Auto-Decoder for Solving Parametric Partial Differential Equations
abstract
Many important problems in science and engineering require solving the so-called parametric partial differential equations (PDEs), i.e., PDEs with different physical parameters, boundary conditions, shapes of computation domains, etc. Recently, building learning-based numerical solvers for parametric PDEs has become an emerging new field. One category of methods such as the Deep Galerkin Method (DGM) and Physics-Informed Neural Networks (PINNs) aim to approximate the solution of the PDEs. They are typically unsupervised and mesh-free, but require going through the time-consuming network training process from scratch for each set of parameters of the PDE. Another category of methods such as Fourier Neural Operator (FNO) and Deep Operator Network (DeepONet) try to approximate the solution mapping directly. Being fast with only one forward inference for each PDE parameter without retraining, they often require a large corpus of paired input-output observations drawn from numerical simulations, and most of them need a predefined mesh as well. In this paper, we propose Meta-Auto-Decoder (MAD), a mesh-free and unsupervised deep learning method that enables the pre-trained model to be quickly adapted to equation instances by implicitly encoding (possibly heterogenous) PDE parameters as latent vectors. The proposed method MAD can be interpreted by manifold learning in infinite-dimensional spaces, granting it a geometric insight. Extensive numerical experiments show that the MAD method exhibits faster convergence speed without losing accuracy than other deep learning-based methods.
Zhanhong Ye, Hongsheng Liu 0002, Beiji Shi, Zidong Wang 0010, Yang Li 0106, Min Wang 0037, Haotian Chu, Fan Yu 0004, Bei Hua, Lei Chen 0002, Bin Dong 0001
NeurIPS11
2022 BETA: Beacon-Based Traffic-Aware Routing in Vehicular Ad Hoc Networks
abstract
Data transmission in Vehicular Ad Hoc Networks (VANETs) often suffers from routing interruptions due to the unstable communication links between vehicles. Over the past decades, many traffic-aware routing protocols have been proposed to alleviate routing interruptions by sensing traffic conditions. However, in most traffic-aware routing protocols, vehicles must transmit a large number of control packets to accumulate traffic information, which may degrade network performance due to the resulting intense competition over the wireless medium. Instead of using control packets, we propose to leverage the beacon mechanism that has been widely used in VANETs to realize traffic awareness. Vehicles broadcast beacons to exchange necessary information with their neighbors periodically. We can leverage this information exchange process among vehicles to replace control packets. To realize this idea, first, a mathematical analysis is provided to demonstrate its feasibility. Then, we propose a concrete protocol to address the technical challenges of using beacons. Extensive simulation results show that our protocol performs better than the state-of-the-art counterparts regarding packet delivery ratio, average delivery time, and network overhead.
Ping Liu 0008, Xingfu Wang, Ammar Hawbani, Bei Hua, Liang Zhao 0004, Zhi Liu 0002
IEEE Trans. Intell. Transp. Syst.4
2021 RoBF: An Auto-Tuning Bloom Filter for Mixed Queries on LSM-Tree
abstract
Bloom filter is an efficient technique to improve query performance in LSM-tree-based databases, such as RocksDB, HBase, and Cassandra.However, the original Bloom filter uses a fixed false positive rate (FPR), which makes it inefficient for mixed queries that involve both point and range queries.To solve this problem, in this paper, we present an improved Bloom filter called RoBF (Range-Query-Oriented Bloom Filter), which uses a mixture of Bloom filters and can process mixed queries on LSM-tree efficiently.We design an efficient algorithm for generating the solution based on the query distribution.We compare our proposal with the trie-based filter and find out that each has its own advantages for various scenarios.Therefore, we propose to use different filters with varied sizes for different levels on LSM-tree.Following this idea, we present an algorithm to generate specific filters with a specific size for different levels on LSM-tree to optimize the performance of mixed queries under limited memory space.We conduct comparative experiments and compare the proposed RoBF with various competitors, and the results show that RoBF can improve the performance of evaluating mixed queries by up to 6x to 30x, compared to the original Bloom filter in RocksDB.
Ruicheng Liu, Peiquan Jin, Shouhong Wan, Bei Hua
SEKE4
2020 MasQ: RDMA for Virtual Private Cloud
abstract
RDMA communication in virtual private cloud (VPC) networks is still a challenging job due to the difficulty in fulfilling all virtualization requirements without sacrificing RDMA communication performance. To address this problem, this paper proposes a software-defined solution, namely, MasQ, which is short for "queue masquerade". The core insight of MasQ is that all RDMA communications should associate with at least one queue pair (QP). Thus, the requirements of virtualization, such as network isolation and the application of security rules, can be easily fulfilled if QP's behavior is properly defined. In particular, MasQ exploits the virtio-based paravirtualization technique to realize the control path. Moreover, to avoid performance overhead, MasQ leaves all data path operations, such as sending and receiving, to the hardware. We have implemented MasQ in the OpenFabrics Enterprise Distribution (OFED) framework and proved its scalability and performance efficiency by evaluating it against typical applications. The results demonstrate that MasQ achieves almost the same performance as bare-metal RDMA for data communication.
Binzhang Fu, Kun Tan 0002, Bei Hua, Zhi-Li Zhang, Kai Zheng 0003
SIGCOMM5
2020 Local Overlapping Community Detection
abstract
Local community detection refers to finding the community that contains the given node based on local information, which becomes very meaningful when global information about the network is unavailable or expensive to acquire. Most studies on local community detection focus on finding non-overlapping communities. However, many real-world networks contain overlapping communities like social networks. Given an overlapping node that belongs to multiple communities, the problem is to find communities to which it belongs according to local information. We propose a framework for local overlapping community detection. The framework has three steps. First, find nodes in multiple communities to which the given node belongs. Second, select representative nodes from nodes obtained above, which tends to be in different communities. Third, discover the communities to which these representative nodes belong. In addition, to demonstrate the effectiveness of the framework, we implement six versions of this framework. Experimental results demonstrate that the six implementation versions outperform the other algorithms.
Li Ni 0001, Wenjian Luo, Wenjie Zhu 0005, Bei Hua
ACM Trans. Knowl. Discov. Data4
2019 vSocket: virtual socket interface for RDMA in public clouds
abstract
RDMA has been widely adopted as a promising solution for high performance networks, but is still unavailable for a large number of socket-based applications running in public clouds due to the following reasons. There is no available virtualization technique of RDMA that can meet the cloud's requirements. Moreover, it is cost prohibitive to rewrite the socket-based applications with the Verbs API. To address the above problems, we present vSocket, a software-based RDMA virtualization framework for socket-based applications in public clouds. vSocket takes into account the demands of clouds such as security rules and network isolation, so it can be deployed in the current public clouds. Furthermore, vSocket provides native socket API so that socket-based applications can use it without any modifications. Finally, to validate the performance gains, we implemented a prototype and compared it with current virtual network solutions against 1) basic network benchmarks and 2) the Redis, a typical I/O intensive application. Experimental results show that the latency of basic benchmarks can be reduced by 88% and the throughput of Redis is improved by 4 times.
Binzhang Fu, Bei Hua
VEE5
2019 Increasing multicast transmission rate with localized multipath in software-defined networks
Bei Hua
Frontiers Comput. Sci.2
2018 G-NET: Effective GPU Sharing in NFV Systems
Kai Zhang 0006, Bingsheng He, Zeke Wang, Bei Hua, Jiayi Meng, Lishan Yang 0001
NSDI5
2017 DIDO: Dynamic Pipelines for In-Memory Key-Value Stores on Coupled CPU-GPU Architectures
abstract
As an emerging hardware, the coupled CPU-GPU architecture integrates a CPU and a GPU into a single chip, where the two processors share the same memory space. This special property opens up new opportunities for building in-memory keyvalue store systems, as it eliminates the data transfer costs on PCI-e bus, and enables fine-grained cooperation between the CPU and the GPU. In this paper, we propose DIDO, an in-memory key-value store system with dynamic pipeline executions on the coupled CPU-GPU architecture, to address the limitations and drawbacks of state-of-the-art system designs. DIDO is capable of adapting to different workloads through dynamically adjusting the pipeline with fine-grained task assignment to the CPU and the GPU at runtime. By exploiting the hardware features of coupled CPU-GPU architectures, DIDO achieves this goal with a set of techniques, including dynamic pipeline partitioning, flexible index operation assignment, and work stealing. We develop a cost model guided adaption mechanism to determine the optimal pipeline configuration. Our experiments have shown the effectiveness of DIDO in significantly enhancing the system throughput for diverse workloads.
Kai Zhang 0006, Bingsheng He, Bei Hua
ICDE4
2017 Zcopy-vhost: Eliminating Packet Copying in Virtual Network I/O
abstract
Virtualization has been widely used as a key technology in cloud computing. Although facilitating the deployment of applications, virtualization introduces huge processing overheads, among which is network I/O virtualization that has become a critical bottleneck of a virtual system. DPDK-vhost is currently the fastest para-virtualized network I/O backend, however it performs poorly when exchanging large packets between virtual machines. Its inefficiency comes from packet copying involved in packet transmission. A zero-copy solution was proposed in the literature to eliminate packet copying by use of shared memory, but it violates the isolation principle of virtual machines. This paper presents a zero-copy vhost design that eliminates packet copying by modifying the extended page tables (EPTs), and meanwhile keeps virtual machines isolated. A prototype adapted from DPDK-vhost is implemented in QEMU/KVM environment, and its performance is verified by experiments to be much higher than that of DPDK-vhost when large packets are transmitted.
Bei Hua, Heqing Zhu, Cunming Liang
LCN2
2017 A distributed in-memory key-value store system on heterogeneous CPU-GPU cluster
Kai Zhang 0006, Kaibo Wang, Yuan Yuan 0014, Lei Guo 0004, Rubao Li, Xiaodong Zhang 0001, Bingsheng He, Bei Hua
VLDB J.9
2015 A Heuristic Algorithm for Optimal Discrete Bandwidth Allocation in SDN Networks
abstract
Multicast inter-session fairness problem is about fairly sharing network resource among a set of multicast sessions. Besides resource allocation fairness, minimum resource adjustment and lowest computational complexity are another two important requirements in real networks. In this paper, we propose a heuristic algorithm to solve the problem in the context of SDN networks for delivering layered encoded video streaming. Compared with existing algorithms, the heuristic algorithm optimally balances resource allocation fairness and minimum resource adjustment, and moreover possesses the lowest computational complexity.
Bei Hua
GLOBECOM2
2015 A holistic approach to build real-time stream processing system with GPU
Kai Zhang 0006, Bei Hua
J. Parallel Distributed Comput.3
2014 NCoS: A framework for realizing network coding over software-defined network
abstract
Network coding is a transmission mechanism for improving the capacity of multicast applications. Practical problems arise when applying network coding in traditional wired networks: multipath multicast routing, backward compatibility with deployed base, and adding new functions in commercial routers. SDN network naturally solves these problems via logically centralized control plane and open APIs of Openflow switches. This paper proposes NCoS, a framework for realizing network coding over SDN networks. It gives a brief introduction on how to extend Openflow protocol to include new actions, and add network coding related functions in controller and switches.
Bei Hua
LCN2
2013 DHash: A cache-friendly TCP lookup algorithm for fast network processing
abstract
A typical hash based TCP lookup algorithm is hard to make a trade-off between speed and space. This paper presents DHash, a high-efficient TCP lookup algorithm that aims at supporting large number of sessions in high speed networks. DHash achieves this goal by designing a compact and cache-friendly lookup data structure that well fits the modern computer architectures. To show the power of DHash, we implement it in a user-space TCP/IP stack, and then parallelize the stack on the Intel multicore processors. Experiments show that DHash is able to achieve 16.3Mpps while handling one million concurrent sessions on our parallel platform.
Kai Zhang 0006, Junchang Wang, Bei Hua
LCN3
2012 User interest modeling and its application for question recommendation in user-interactive question answering systems
Xingliang Ni, Xiaojun Quan, Wenyin Liu, Bei Hua
Inf. Process. Manag.5
2012 Using surface variability characteristics for segmentation of deformable 3D objects with application to piecewise statistical deformable model
Horace Ho-Shing Ip, Bei Hua, Jun Feng 0003
Vis. Comput.3
2011 Building High-Performance Application Protocol Parsers on Multi-core Architectures
abstract
Parsing packet payloads according to the syntax and semantics of an application protocol is a key step in analyzing network traffic. However, it is still a challenge to fulfill this task with high speed(10Gbps+) because parsing packets through deep-content analysis to build a corresponding syntax tree requires tremendous computing resources. Multi-core architectures provide a viable solution for building high-performance parsers for application protocols. Existing sequential application protocol parsers are hard to be reused, and building a new protocol parser from scratch is error-prone and time-consuming. This paper proposes a general and efficient approach to building high-performance parallel application protocol parsers on multi-core platforms. First, the open-source lexical analyzer FLEX is used to describe a protocol and generate a sequential parser. Then a source-to-source translation is performed to transform the sequential parser into a parallel one. Finally, an efficient parallel run-time system is built by employing lock-free design principles from top to bottom to support multi-threaded execution on multi-core processors. Experimental results show that our parsers achieve nearly 20Gbps for average HTTP packets and 5Gbps for the challenging smaller FIX packets.
Kai Zhang 0006, Junchang Wang, Bei Hua, Xinan Tang
ICPADS3
2011 Short text clustering by finding core terms
Xingliang Ni, Xiaojun Quan, Wenyin Liu, Bei Hua
Knowl. Inf. Syst.5
2009 Segmenting deformable soft-body meshes based on statistical variation information for piecewise Active Shape Model
abstract
This paper proposes an algorithm for segmenting deforming soft-body meshes based on statistical variation information extracted from the deforming meshes. The variation information is extracted by performing a global principal component analysis (PCA) on the set of meshes. eigen-variation similarity (EVS) and eigen-variation magnitude (EVM) are then defined for the vertices and triangle faces of the meshes based on the extracted variation information. A multiple-source region growing algorithm is presented for segmenting a mesh that favors grouping faces with similar variations into a same component. We apply the proposed mesh segmentation algorithm to the construction of piecewise active shape model (ASM) and use such piecewise ASM to reconstruct unseen meshes. Experimental results show that our algorithm outperforms several state-of-the-art methods in terms of reconstruction accuracy.
Horace Ho-Shing Ip, Jun Feng 0003, Bei Hua
CAD/Graphics4
2009 Practice of parallelizing network applications on multi-core architectures
abstract
The industry wide shift to multi-core architectures arouses great interests in parallelizing sequential applications. However, it is very difficult to parallelize fine-grained applications for multi-core architectures due to insufficient hardware support of fast communication and synchronization. Fortunately, network applications can be decomposed into pipelined structures that are amenable to streaming based parallel processing. To realize the potential of pipelining on multi-core architectures, it requires reevaluating the basic tradeoffs in parallel processing, including the ones between load balance and data locality and between general lock mechanisms and special lock-free data structures. This paper presents the practice of building a high-performance multi-core based network processing platform in which connection-affinity and lock-free design principles are applied effectively for better data locality and faster core-to-core synchronization and communication.We parallelize a complete Layer 2 to Layer 7 (L2-L7) network processing system on an Intel Core 2 Quad processor, including a TCP/IP stack based on Libnids (L2-L4) and a port-independent protocol identification engine by deep packet inspection (L7+). Furthermore, we develop a compiling method to transform sequential network applications to parallel ones to enable those applications to run on multi-core architectures. Our experience suggests that (1) fine-grained pipelining can be a good software solution for parallelizing network applications on multi-core architectures if connection-affinity and lock-free are used as the first design principles; (2) a delicate partitioning scheme is required to map pipelined structures onto specific multi-core architecture; (3) an automatic parallelization approach can work if domain knowledge is considered in the parallelizing process. Our multi-core based network processing platform can deliver not only 6Gbps processing speed for large packet sizes but also more challenging 2Gbps speed for smaller packets.
Junchang Wang, Haipeng Cheng, Bei Hua, Xinan Tang
ICS3
2008 Scalable packet classification using interpreting: a cross-platform multi-core solution
abstract
Packet classification is an enabling technology to support advanced Internet services. It is still a challenge for a software solution to achieve 10Gbps (line-rate) classification speed. This paper presents a classification algorithm that can be efficiently implemented on a multi-core architecture with or without cache. The algorithm embraces the holistic notion of exploiting application characteristics, considering the capabilities of the CPU and the memory hierarchy, and performing appropriate data partitioning. The classification algorithm adopts two stages: searching on a reduction tree and searching on a list of ranges. This decision is made based on a classification heuristic: the size of the range list is limited after the first stage search. Optimizations are then designed to speed up the two-stage execution. To exploit the speed gap (1) between the CPU and external memory; (2) between internal memory (cache) and external memory, an interpreter is used to trade the CPU idle cycles with demanding memory access requirements. By applying the CISC style of instruction encoding to compress the range expressions, it not only significantly reduces the total memory requirement but also makes effective use of the internal memory (cache) bandwidth. We show that compressing data structures is an effective optimization across the multi-core architectures.
Haipeng Cheng, Bei Hua, Xinan Tang
PPoPP3
2008 A robust localization algorithm in wireless sensor networks
Bei Hua, Yi Shang
Frontiers Comput. Sci. China2
2008 High-performance packet classification algorithm for multithreaded IXP network processor
abstract
Packet classification is crucial for the Internet to provide more value-added services and guaranteed quality of service. Besides hardware-based solutions, many software-based classification algorithms have been proposed. However, classifying at 10 Gbps speed or higher is a challenging problem and it is still one of the performance bottlenecks in core routers. In general, classification algorithms face the same challenge of balancing between high classification speed and low memory requirements. This paper proposes a modified recursive flow classification (RFC) algorithm, Bitmap-RFC, which significantly reduces the memory requirements of RFC by applying a bitmap compression technique. To speed up classifying speed, we exploit the multithreaded architectural features in various algorithm development stages from algorithm design to algorithm implementation. As a result, Bitmap-RFC strikes a good balance between speed and space. It can significantly keep both high classification speed and reduce memory space consumption. This paper investigates the main NPU software design aspects that have dramatic performance impacts on any NPU-based implementations: memory space reduction , instruction selection , data allocation , task partitioning , and latency hiding . We experiment with an architecture-aware design principle to guarantee the high performance of the classification algorithm on an NPU implementation. The experimental results show that the Bitmap-RFC algorithm achieves 10 Gbps speed or higher and has a good scalability on Intel IXP2800 NPU.
Bei Hua, Nenghai Yu, Xinan Tang
ACM Trans. Embed. Comput. Syst.3
2007 Bilateration: An Attack-Resistant Localization Algorithm of Wireless Sensor Network
Bei Hua, Yi Shang, Lihua Yue
EUC2
2007 Study of a Cost-Effective Localization Algorithm in Wireless Sensor Networks
Bei Hua
MSN2
2007 Garbage Collector Verification for Proof-Carrying Code
Chunxiao Lin, Bei Hua
J. Comput. Sci. Technol.4
2006 High-performance packet classification algorithm for many-core and multithreaded network processor
abstract
Packet classification is crucial for the Internet to provide more value-added services and guaranteed quality of service. Besides hardware-based solutions, many software-based classification algorithms have been proposed. However, classifying at 10Gbps speed or higher is a challenging problem and it is still one of the performance bottlenecks in core routers. In general, classification algorithms face the same challenge of balancing between high classification speed and low memory requirements. This paper proposes a modified Recursive Flow Classification (RFC) algorithm, Bitmap-RFC, which significantly reduces the memory requirements of RFC by applying a bitmap compression technique. To speed up classifying speed, we experiment on exploiting the architectural features of a many-core and multithreaded architecture from algorithm design to algorithm implementation. As a result, Bitmap-RFC strikes a good balance between speed and space. It can not only keep high classification speed but also reduce memory space significantly.This paper investigates the main NPU software design aspects that have dramatic performance impacts on any NPU-based implementations: memory space reduction, instruction selection, data allocation, task partitioning, and latency hiding. We experiment with an architecture-aware design principle to guarantee the high performance of the classification algorithm on an NPU implementation. The experimental results show that the Bitmap-RFC algorithm achieves 10Gbps speed or higher and has a good scalability on Intel IXP2800 NP.
Bei Hua, Xianghui Hu, Xinan Tang
CASES2
2006 High-performance IPv6 forwarding algorithm for multi-core and multithreaded network processor
abstract
IP forwarding is one of the main bottlenecks in Internet backbone routers, as it requires performing the longest-prefix match at 10Gbps speed or higher. IPv6 forwarding further exacerbates the situation because its search space is quadrupled. We propose a high-performance IPv6 forwarding algorithm TrieC, and implement it efficiently on the Intel IXP2800 network processor (NPU). Programming the multi-core and multithreaded NPU is a daunting task. We study the interaction between the parallel algorithm design and the architecture mapping to facilitate efficient algorithm implementation. We experiment with an architecture-aware design principle to guarantee the high performance of the resulting algorithm.This paper investigates the main software design issues that have dramatic performance impacts on any NPU based implementation: memory space reduction, instruction selection, data allocation, task partitioning, latency hiding, and thread synchronization. In the paper, we provide insight on how to design an NPU-aware algorithm for high-performance networking applications. Based on the detailed performance analysis of the TrieC algorithm, we provide guidance on developing high-performance networking applications for the multi-core and multithreaded architecture.
Xianghui Hu, Xinan Tang, Bei Hua
PPoPP3