Zhengbin Pang

dblp:58/2743 · DBLP profile ↗
← Back
45ranked-venue papers
1as first author
24since 2021 · last 2026
0000-0003-2046-5043ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Energy-Efficient 0.56-pJ/cycle AVFS System Based on a Fast Transient Response Digital LDO and a Self-Calibrating Elastic Clock
Jiliang Liu, Zhengbin Pang, Fangxu Lv, Shijie Li 0002, Qiang Wang 0006, Lizhou Wu, Chengzhuo Zhao
ISCAS2
2026 HierMine: Accelerating Graph Pattern Mining via Hierarchical Sampling
abstract
Current Approximate Graph Pattern Mining (AGPM) systems rely on uniform and static sampler allocation strategies, which result in significant computational redundancy because most regions of the sampling space contribute little to the final estimate. However, existing systems overlook this critical characteristic and fail to account for the hierarchical memory architecture of modern servers, leading to slow convergence and substantial memory access overhead. To address these challenges, we propose HierMine, a novel AGPM system designed to overcome the aforementioned drawbacks. The core contributions of HierMine include: (i) a hierarchical sampling strategy that minimizes redundant computation and accelerates convergence; (ii) a dynamic sampling adjustment mechanism that partitions the sampling space online and reallocates samplers adaptively to enhance the hierarchical strategy; (iii) an online grouping-based convergence detection technique that enables the fine-grained dynamic sampling adjustment mechanism; and (iv) a hierarchical data layout optimized for memory access efficiency. Extensive experiments show that HierMine delivers an average speedup of up to 28.9× over state-of-the-art AGPM systems such as ScaleGPM, while maintaining strong theoretical guarantees on estimation quality. It also consistently outperforms exact graph pattern mining systems, highlighting the practical benefits of our approximate approach.
Xinbiao Gan, Songzhu Mei, Zhengbin Pang, Hongxu Jin
ACM Trans. Archit. Code Optim.4
2025 Learning Cognitive Agent-driven for Graph-aware Communication and Double Reward Path Finding Models
Jinsheng Deng, Zhengbin Pang
CogSci3
2025 Exploring Graph-aware Reasoning and Bidirectional Selection for Vision-Language Navigation
abstract
Inspired by structured state space models and graph neural network modeling, we proposes a novel graph-aware reasoning (GAR) model to effectively solve the problem between memory utilization efficiency and reasoning navigation. First, we introduce graph networks into the navigation framework to enhance the model modeling ability for long sequence dependencies. It is integration graph neural networks into the state space to ensures that the agent can accurately find topological paths based on instructions and environmental interactions. Then, we design a bidirectional selective state space model to enhance the agent-aware of spatial information. When without relying on the attention network, we capture the context information in the image through position embedding. Finally, we fuison the features in the bidirectional state space and effectively compresses and improves the fine-grained image features by a residual connection methods. Our experimental results on R2R and REVERIE datasets show that GAR reduces GPU memory usage by 4.45% compared with the DUET baseline model.
Dongming Zhou 0003, Jinsheng Deng, Zhengbin Pang
ICASSP3
2025 Dynamic Spatial-Aware Network for Occlusion Robust 3D Hand Mesh Reconstruction from RGB Images
Yibo Bai, Xiaoqing Yin, Zhengbin Pang, Jinsheng Deng
PRCV (10)3
2025 DHAA: Distributed heuristic action aware multi-agent path finding in high density scene
Dongming Zhou 0003, Zhengbin Pang
Multim. Tools Appl.2
2025 DICCR: Double-gated intervention and confounder causal reasoning for vision-language navigation
Dongming Zhou 0003, Jinsheng Deng, Zhengbin Pang
Neural Networks3
2025 nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs
abstract
Convolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures.
Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007
IEEE Trans. Computers5
2024 Learning Cross-modal Knowledge Reasoning and Heuristic-prompt for Visual-language Navigation
abstract
Visual language navigation is an exciting and challenging multi-modal task. Most existing research focuses on the fusion of visual features and semantic space, which ignoring the importance of local highlight features and semantic knowledge alignment in images for agent navigation. Therefore, this paper proposes a novel visual language model combining Knowledge-augmented Reasoning and Soft-Prompt (KRSP) learning. First, we perform fine-grained processing of local regions in the image and to map context image features and text knowledge to the same common sub-space. We focus on regional knowledge to increase the model reasoning ability. Next, soft-prompt learning aligns keywords and sub-visual information in instruction features to solve the path mismatch problem in coarse-grained instructions. We use a large-scale pre-training model CoCoOp to collect highly matched soft action prompts into a unified instruction set. Finally, we propose a general cross-modal feature alignment loss function. The potential semantic correlation between sub-visual information and instruction space is closer through the penalty mechanism of the alignment function. This paper verifies the method effectiveness on the R2R and REVERIE datasets, and the experimental results show that KRSP achieves state-of-the-art performance. Among them, the KRSP of SPL evaluation metric increased by 4.5% in unseen scenarios.
Dongming Zhou 0003, Zhengbin Pang
CIKM2
2024 FPGA Implementation of Sequence Detector for High-Speed PAM4 Wireline Transceiver
abstract
To solve the problem of high bit error rate (BER) due to high inter-symbol interference (ISI) in high-speed wireline transceivers, this paper proposes a low-complexity adaptive reduced-state sequence detector (ARSSD). The detector is based on the maximum likelihood sequence detection (MLSD) to reduce the detection bit error rate (BER), adopts the ISI parameter acquisition method based on the zero-forcing algorithm to achieve the adaptive detector parameters, and combines the viterbi algorithm and the set partitioning algorithm to reduce the complexity of operations. The behavioral simulation and the implementation of the hardware circuit are completed in this paper. The experimental results based on the analog front-end and the field programmable gate array (FPGA) show that when the pulse amplitude modulation 4 (PAM4) bit rate is 12 ∼ 56Gbps and the channel loss is -5dB ∼ -17dB@14GHz, the detection BERs of 32x4 parallel ARSSDs are reduced by two orders of magnitude compared to the conventional decision feedback equalization, which is consistent with the results of the behavioral simulation.
Chaolong Xu, Fangxu Lv, Zhengbin Pang, Liquan Xiao, Zhouhao Yang
ACM Great Lakes Symposium on VLSI3
2024 Heuristic Action-aware and Priority Communication for Multi-agent Path Finding
abstract
Large-scale multi-agent path finding (MAPF ) has the problem of shared features, which leads to a large amount of bandwidth and additional overhead. Therefore, this paper proposes a multi-agent path finding method that combines heuristic action-aware network and dual Q network(HADQ). First, we use a convolutional neural network to encoder observation input at field of view scope. On this basis, we combine self-attention network and graph neural network to query the priority of agent communication to reduce collisions. we improves the timeliness of agents timely communication by reducing the size of shared features. Then, we embedding the agent shortest path into the training process as a heuristic guide. Finally, this paper uses the deep Q network to map the observation values into the actions of the agent, thereby constructing a dynamic communication topology network. Compared with the state-of-the-art MAPF method, the method proposed in this paper has certain improvements in solution quality and success rate. Experimental results show that even in unseen data sets. HADQ can also show good generalization ability and robustness.
Dongming Zhou 0003, Zhengbin Pang
ICME2
2024 Enhanced Causal Reasoning and Graph Networks for Multi-agent Path Finding
abstract
Multi-Agent Path Finding refers to the problem of finding the optimal path set for multiple agents from the starting position to the target position without conflict. In multi-agent path planning, multiple factors need to be considered, such as the dynamic characteristics of the agent, environmental constraints, communication and collaboration mechanisms, etc. In order to solve such problems, researchers have proposed a variety of path planning algorithms, including centralized path planning algorithms, distributed path planning algorithms and collaborative path planning algorithms. In this paper, we propose a decentralized multi-agent path planning framework that combines causal relationship drive and graph neural network(CRGN). First, we propose an end-to-end framework to extract global features and local features in images. On this basis, we design the fusion of deep features and shallow features to increase visual fine-graining. Then, we propose to use graph neural networks to aggregate information communication between agents and build a dynamic communication topology network through agent communication within the local field of view. Finally, we propose a new reinforcement learning algorithm to train the agent’s generalization ability in unknown environments. We fit the state value function and action value function in reinforcement learning through deep neural networks. Thus, it can better solve the intelligent agent path planning under large-scale complex problems. We verified the effectiveness of the model on multiple different datasets. Experimental results show that our proposed approach is close to the performance of expert algorithms. At the same time, we show through extensive ablation experiments that our model can exhibit good performance and robustness even in unseen scenes.
Zhengbin Pang
IJCNN2
2024 Collective Communication Acceleration Architecture for Reduce Operation with Long Operands Based on Efficient Data Access and Reconfigurable On-Chip Buffer
abstract
The Reduce operation with long operands is widely used in high-performance computing and AI(Artificial Intelligence) computing. A hardware acceleration architecture for Reduce operation with long operands is designed and implemented in FPGA. Flow control of the accelerator is achieved by using a dedicated hardware trigger mechanism. Based on DMA(Direct Memory Access), an efficient data access method is proposed between host memory and the accelerator. A reconfigurable on-chip buffer structure is designed to achieve flexible and efficient buffering of a large number of operands. Calculations are directly performed on the accelerator with on-chip ALU(Arithmetic Logic Unit) arrays to reduce the communications between the host and NIC(Network Interface Card). Experimental results show that this work has a significant acceleration effect compared to the non-offloading method and the current offloading method in "Tianhe" interconnection.
Jinbo Xu, Zhengbin Pang
ISPA2
2024 Credit assignment for trained neural networks based on Koopman operator theory
Changyuan Zhao, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002, Zhengbin Pang
Frontiers Comput. Sci.6
2024 Qualitative and Quantitative Model Checking Against Recurrent Neural Networks
Wanwei Liu, Fu Song, Bai Xue 0001, Wenjing Yang 0002, Ji Wang 0001, Zhengbin Pang
J. Comput. Sci. Technol.7
2024 Joint Spatial-Spectral Optimization for the High-Magnification Fusion of Hyperspectral and Multispectral Images
abstract
The fusion of hyperspectral and multispectral images is an important strategy for enhancing the spatial resolution of hyperspectral images. With the rapid advancement of multispectral imaging technology, the disparity in spatial resolution between multispectral and hyperspectral images is increasing. In certain scenarios, termed high-magnification, this difference can exceed$32\times $. Previous methods do not perform well under high-magnification fusion, and naturally, a challenge arises in achieving effective high-magnification super-resolution fusion. In light of the above analysis, this article introduces a novel algorithm for high-magnification super-resolution fusion of hyperspectral and multispectral images based on the joint optimization of spatial and spectral information. Specifically, our algorithm consists of three stages: 1) a fast preliminary fusion stage based on the Moore-Penrose inverse and singular value correlation priors for the rapid acquisition of preliminary solutions; 2) a joint spatial-spectral optimization stage where a coupled optimization framework is constructed to achieve integrated optimization of spatial and spectral information; and 3) an error backpropagation optimization stage where an effective error optimization term is introduced to further refine the fusion performance. We conducted extensive experiments on widely employed publicly available simulated datasets and real datasets. The experimental results unequivocally indicate that our proposed methodology consistently exhibits superior fusion performance compared with state-of-the-art methods, even under the condition of${\geq }60\times $super-resolution.
Yibing Zhan, Zhengbin Pang, Tong Zhou 0008, Xueqiong Li, Long Lan, Yuanxi Peng
IEEE Trans. Geosci. Remote. Sens.3
2023 Predicting gene regulatory links from single-cell RNA-seq data using graph neural networks
abstract
Single-cell RNA-sequencing (scRNA-seq) has emerged as a powerful technique for studying gene expression patterns at the single-cell level. Inferring gene regulatory networks (GRNs) from scRNA-seq data provides insight into cellular phenotypes from the genomic level. However, the high sparsity, noise and dropout events inherent in scRNA-seq data present challenges for GRN inference. In recent years, the dramatic increase in data on experimentally validated transcription factors binding to DNA has made it possible to infer GRNs by supervised methods. In this study, we address the problem of GRN inference by framing it as a graph link prediction task. In this paper, we propose a novel framework called GNNLink, which leverages known GRNs to deduce the potential regulatory interdependencies between genes. First, we preprocess the raw scRNA-seq data. Then, we introduce a graph convolutional network-based interaction graph encoder to effectively refine gene features by capturing interdependencies between nodes in the network. Finally, the inference of GRN is obtained by performing matrix completion operation on node features. The features obtained from model training can be applied to downstream tasks such as measuring similarity and inferring causality between gene pairs. To evaluate the performance of GNNLink, we compare it with six existing GRN reconstruction methods using seven scRNA-seq datasets. These datasets encompass diverse ground truth networks, including functional interaction networks, Loss of Function/Gain of Function data, non-specific ChIP-seq data and cell-type-specific ChIP-seq data. Our experimental results demonstrate that GNNLink achieves comparable or superior performance across these datasets, showcasing its robustness and accuracy. Furthermore, we observe consistent performance across datasets of varying scales. For reproducibility, we provide the data and source code of GNNLink on our GitHub repository: https://github.com/sdesignates/GNNLink.
Guo Mao, Zhengbin Pang, Ke Zuo, Xiangdong Pei, Xinhai Chen 0001, Jie Liu 0002
Briefings Bioinform.2
2023 Towards robust neural networks via a global and monotonically decreasing robustness training strategy
abstract
Robustness of deep neural networks (DNNs) has caused great concerns in the academic and industrial communities, especially in safety-critical domains. Instead of verifying whether the robustness property holds or not in certain neural networks, this paper focuses on training robust neural networks with respect to given perturbations. State-of-the-art training methods, interval bound propagation (IBP) and CROWN-IBP, perform well with respect to small perturbations, but their performance declines significantly in large perturbation cases, which is termed “drawdown risk” in this paper. Specifically, drawdown risk refers to the phenomenon that IBP-family training methods cannot provide expected robust neural networks in larger perturbation cases, as in smaller perturbation cases. To alleviate the unexpected drawdown risk, we propose a global and monotonically decreasing robustness training strategy that takes multiple perturbations into account during each training epoch (global robustness training), and the corresponding robustness losses are combined with monotonically decreasing weights (monotonically decreasing robustness training). With experimental demonstrations, our presented strategy maintains performance on small perturbations and the drawdown risk on large perturbations is alleviated to a great extent. It is also noteworthy that our training method achieves higher model accuracy than the original training methods, which means that our presented training strategy gives more balanced consideration to robustness and accuracy.
Taoran Wu, Wanwei Liu, Bai Xue 0001, Wenjing Yang 0002, Ji Wang 0001, Zhengbin Pang
Frontiers Inf. Technol. Electron. Eng.7
2022 ERA: ECN-Ratio-Based Congestion Control in Datacenter Networks
abstract
The widespread deployment of Remote Direct Memory Access (RDMA) in datacenter networks increases the stringency for convergence speed when congestion occurs. Fast convergence significantly reduces buffer occupancy, which in turn lessens the probability of triggering Priority-based Flow Control (PFC). Besides, the propagation delay becomes shorter with rapidly growing link speed, which correspondingly makes the queueing delay a major part of end-to-end latency in datacenter networks. Fast convergence and low buffer occupancy become more essential for lowering queue delay and flow complete time. In this paper, we present ERA, an ecn-ratio-based congestion control scheme, which contributes to fast convergence for datacenter networks. ERA consists of two fundamental components: (i) an ECN-marking-ratio-based queue buffer occupancy estimating (QBOE) solution and (ii) a queue-building-rate driven rate adjustment (QDRA) mechanism to achieve fast convergence in several control periods. We conduct extensive experiments to evaluate the performance of ERA, and the results show that ERA greatly accelerates the convergence process compared to other solutions. ERA achieves low tail latency and low buffer occupancy simultaneously.
Dezun Dong, Zhengbin Pang, Junhong Ye
CCGRID3
2022 DNNEmu: A Lightweight Performance Emulator for Distributed DNN Training
Enda Yu, Dezun Dong, Zhengbin Pang
ICA3PP4
2022 STEGNN: Spatial-Temporal Embedding Graph Neural Networks for Road Network Forecasting
abstract
As intelligent transportation systems (ITS) are now being integrated into our everyday lives, it has been widely accepted that forecasting road networks is a promising killer engine for ITS with high social and economic benefits. However, current solutions ignore the heterogeneity of spatial-temporal traffic data and fail to capture hidden spatial-temporal correlations. This paper presents STEGNN: a novel spatial-temporal embedding graph neural network for road network forecasting. The key idea of STEGNN is utilizing Cosine Similarity to generate a high-quality temporal graph and thus fills the gap between the temporal-spatial correlations for traffic graph, which includes (i) a novel approach to construct temporal graph based on temporal-spatial similarity from traffic graphs, which is much more accurate on measured similarity of time series claimed by previous methods; (ii) an advanced spatial-temporal embedding model to exploit spatial-temporal dependencies by leveraging specific arrangements of temporal and spatial graphs; and (iii) an effective framework that gasps extensive spatial-temporal dependencies in the long-term by mixing multi-layer graph convolution with dilated convolution to understand wide-range spatial-temporal features. Extensive evaluations validate STEGNN by applying it to real-world traffic graphs and indicate that STEGNN outperforms state-of-the-art solutions with much more accurate forecasting of road networks.
Jiaqi Si, Xinbiao Gan, Tiaojie Xiao, Bo Yang 0023, Dezun Dong, Zhengbin Pang
ICPADS6
2022 Fast-Converging Congestion Control in Datacenter Networks
abstract
The widespread deployment of Remote Direct Memory Access (RDMA) in datacenter networks increases the stringency for convergence speed when congestion occurs. Fast convergence significantly reduces buffer occupancy, which in turn lessens the probability of triggering Priority-based Flow Control (PFC). Besides, the propagation delay becomes shorter with rapidly growing link speed, which correspondingly makes the queueing delay a major part of end-to-end latency. Fast convergence and low buffer occupancy become more essential for lowering queue delay and flow complete time. We present DQCC (Double-Q Congestion Control), a fast-converging congestion control scheme, which consists of two fundamental components: (i) an ECN-marking-ratio-based queue buffer occupancy estimating (QBOE) solution and (ii) a queue-building-rate driven rate adjustment (QDRA) mechanism to achieve fast convergence. We conduct extensive experiments to evaluate the performance of DQCC, and the results show that DQCC greatly accelerates the convergence process. DQCC achieves low tail latency and low buffer occupancy simultaneously.
Dezun Dong, Zhengbin Pang, Junhong Ye
ISCC3
2022 Reconstructing gene regulatory networks of biological function using differential equations of multilayer perceptrons
abstract
BACKGROUND: Building biological networks with a certain function is a challenge in systems biology. For the functionality of small (less than ten nodes) biological networks, most methods are implemented by exhausting all possible network topological spaces. This exhaustive approach is difficult to scale to large-scale biological networks. And regulatory relationships are complex and often nonlinear or non-monotonic, which makes inference using linear models challenging. RESULTS: In this paper, we propose a multi-layer perceptron-based differential equation method, which operates by training a fully connected neural network (NN) to simulate the transcription rate of genes in traditional differential equations. We verify whether the regulatory network constructed by the NN method can continue to achieve the expected biological function by verifying the degree of overlap between the regulatory network discovered by NN and the regulatory network constructed by the Hill function. And we validate our approach by adapting to noise signals, regulator knockout, and constructing large-scale gene regulatory networks using link-knockout techniques. We apply a real dataset (the mesoderm inducer Xenopus Brachyury expression) to construct the core topology of the gene regulatory network and find that Xbra is only strongly expressed at moderate levels of activin signaling. CONCLUSION: We have demonstrated from the results that this method has the ability to identify the underlying network topology and functional mechanisms, and can also be applied to larger and more complex gene network topologies.
Guo Mao, Ruigeng Zeng, Jintao Peng, Ke Zuo, Zhengbin Pang, Jie Liu 0002
BMC Bioinform.5
2021 NEPG: Partitioning Large-Scale Power-Law Graphs
Jiaqi Si, Xinbiao Gan, Dezun Dong, Zhengbin Pang
ICA3PP (3)5
2019 Efficient Management and Intelligent Fault Tolerance for HPC Interconnect Networks
abstract
Interconnect Network is the key component in high performance computing system. With the incoming era of Exa-Scale (1018FLOPS) computing, designing large scale interconnect networks is facing with serious challenges in network management and fault tolerance. To construct higher performance and more reliable interconnect networks, we propose an efficient and intelligent network management architecture for indirect interconnect networks. This paper emphatically introduces the network management architecture, the in-band management channels in the Network Interface Chips (NIC) and Network Routing Chips (NRC) respectively, efficient centralized network management approach, distributed intelligent faulttolerant routing management, and so on. Based on the prototype system, the Control and Status Registers (CSR) accessing latency performance of in-band network management is evaluated. Also based on a customized simulation model for indirect interconnect networks, the typical distributed fault-tolerant scenes are tested. The experiment results show the in-band network management can averagely achieves 716 times improvements than the out of-band network management. Moreover, by running heuristic algorithms to automatically reconstruct routing tables for failure links or routers, the intellectual network management Engine can achieve fault-tolerant routing to maximizing network performance.
Jijun Cao, Zhang Luo, Zhengbin Pang
ICPADS5
2018 RSON: An inter/intra-chip silicon photonic network for rack-scale computing systems
abstract
The increasing demand for more computational power from scientific computing, big data processing, and machine learning is pushing the development of HPC (high-performance computing) systems. As the basic HPC building blocks, modularized server racks with a large number of multicore nodes are facing performance and energy efficiency challenges. This paper proposes RSON, an optical network for rack-scale computing systems. RSON connects processor cores, caches, local memories, and remote memories through a novel inter/intra-chip silicon photonic network architecture. We develop a low-latency scalable channel partition and low-power dynamic path priority control scheme for RSON. Experimental results show that RSON can help rack-scale computing systems achieve up to 6.8X higher performance under the same energy consumption than state-of-the-art systems under the latest APEX (application performance at extreme scale) benchmarks.
Peng Yang 0003, Zhengbin Pang, Zhehui Wang, Xuanqi Chen, Luan H. K. Duong, Jiang Xu 0001
DATE2
2018 Integrated High-Speed Optical SerDes over 100GBd Based on Optical Time Division Multiplexing
abstract
An on-chip optical transceiver for transmission system over 100GBd is proposed based on optical time division multiplexing (OTDM) technology, and the performances, such as the insertion loss, the inter-symbol interference (ISI) crosstalk, and the potential symbol rate, are analyzed in detail. Co-designed with the double rail driver, on-chip Mach-Zehnder interferometer switch repeatedly generates extremely narrow sampling pulses of only 12ps full width at half maximum. Based on such narrow optical sampling pulse train, a four-stage cascaded optical switch divides the 25GHz clock cycle into four recurrent 9.5ps time slots and one blank time slot of 2ps. Thus, a 100GBd optical transmission channel is realized based on 4-bit 25Gbps bit-streams at the electrical interface. The ISI extinction ratio at the worst channel is 1.9dB with 10dB depth modulator, and the insertion loss caused by the OTDM mechanism is about 16dB. Further, taking advantages of dark modulation, an OTDM system with 5-bit 25Gbps bit-streams at the electrical interface is proposed to generate a 125GBd transmission utilizing the same optical sampling pulse. The ISI performance is much better and the extinction ratio at the worst channel is enhanced to 3.99dB.
Zhang Luo, Zhengbin Pang, Renfa Li
ACM J. Emerg. Technol. Comput. Syst.4
2017 MOCA: an Inter/Intra-Chip Optical Network for Memory
abstract
The memory wall problem is due to the imbalanced developments and separation of processors and memories. It is becoming acute as more and more processor cores are integrated into a single chip and demand higher memory bandwidth through limited chip pins. Optical memory interconnection network (OMIN) promises high bandwidth, bandwidth density, and energy efficiency, and can potentially alleviate the memory wall problem. In this paper, we propose an optical inter/intra-chip processor-memory communication architecture, called MOCA. Experimental results and analysis show that MOCA can significantly improve system performance and energy efficiency. For example, comparing to Hybrid Memory Cube (HMC), MOCA can speedup application execution time by 2.6x, reduce communication latency by 75%, and improve energy efficiency by 3.4x for 256-core processors in 7 nm technology.
Zhehui Wang, Zhengbin Pang, Peng Yang 0003, Jiang Xu 0001, Xuanqi Chen, Rafael Kioji Vivas Maeda, Luan H. K. Duong, Haoran Li 0002, Zhe Wang 0003
DAC2
2016 The Efficient In-band Management for Interconnect Network in Tianhe-2 System
abstract
Interconnect network plays an important role in high performance computing systems. And its manageability directly affects the RAS (i.e., Reliability, Availability, and Serviceability) of the whole system. The Tianhe-2 system located in NSCC-gz (i.e., National Supercomputing Center of China in Guangzhou) uses proprietary interconnect network, which includes 5,856 high-radix network router chips (i.e., NRC) and 18,304 network interface chips (i.e., NIC). For such a very large-scale interconnect network, it is a great challenge to manage (such as configure, monitor, and debug) the numerous network chips and its network ports in an efficient way. By implementing the in-band management with very few hardware resources, the interconnect network in Tianhe-2 system achieves a highly efficient network management. In this paper, we introduce the design and implementation of the in-band management for interconnect network in Tianhe-2 system, especially emphasizing on several key features, including the set of achieved management functionalities, the architecture of network management, the format of management packets, the data flow and processing of management packets, etc. In this paper, we also evaluate the performance of in-band management by mainly comparing with out-band management scheme. The preliminary results demonstrate the efficiency of the in-band management for interconnect network in Tianhe-2 system.
Jijun Cao, Liquan Xiao, Zhengbin Pang, Kefei Wang
PDP3
2015 A low-latency fine-grained dynamic shared cache management scheme for chip multi-processor
abstract
In order to utilize the shared last-level cache (LLC) in chip multi-processors (CMP) more efficiently, the partitioning of LLC resources among all cores should have the characteristics of low-latency for access, fine granularity for migration and simple hardware complexity for implementation. This paper proposes a dynamic LLC management scheme to achieve these goals. The proposed scheme migrates cache resources among different cores at the granularity of cache blocks, instead of ways. The quantity of victim cache blocks that each victim core can migrate to other target cores are related to an eviction probability, which are calculated according to the performance goal. Then the victim cache blocks for a target core is chosen from the nearest victim core who has non-zero eviction probability by introducing innovate E-Table structure in CMP. The eviction probabilities are updated periodically. With the help of E-Tables, the proposal achieves low-latency accesses by always keeping the required cache blocks near to the target cores. And fine granularity is guaranteed by maintaining an eviction probability for each core. In addition, only little additional hardware changes to traditional cache structure is required. Simulation results suggest significant performance improvements from 6.8% to 22.7% over related works.
Jinbo Xu, Zhengbin Pang
IPCCC3
2015 High Performance Interconnect Network for Tianhe System
Xiangke Liao, Zhengbin Pang, Kefei Wang, Yutong Lu, Dezun Dong, Guang Suo
J. Comput. Sci. Technol.2
2014 Fast NIC based RDMA implementation for adaptive unreliable networks
abstract
Remote Direct Memory Access (RDMA) is one of the basic communication fabrics for parallel computers. Its performance is crucial for most parallel workloads. For next generation super computers, the system usually extends to extremely large scale with millions of computation cores connected through the interconnection network. Most interconnection network features out of order packets delivery and low system-wide network reliability. Building high performance and reliable point-to-point communication over next generation interconnection network is a challenging task. Current system usually implements RDMA through two approaches: 1) the receiver side counters approach; 2) sender side sliding window approach. We see that current approach works better on reliable interconnection networks, but has obvious performance degradation over the un-reliable network. This paper proposes a new hardware approach for fast RDMA transmission, which can provide scalable performance on un-reliable network. Our approach uses a novel receiver side sliding window to support out-of-order packets delivery. Our approach proposes a new implementation of the traditional sliding window approach, which uses receiver side sliding window to efficiently retransmit partial RDMA data when part of the packets fail to reach the receiver. The experiments show that for low reliability interconnection network, our approach has scalable performance benefit over current RDMA implementations.
Zhengbin Pang
AICCSA4
2014 A Low Overhead Last-Write-Touch Prediction Scheme
abstract
Last-write-touch prediction can reduce cache-to-cache transfer latency by converting 3-hop misses into 2-hop misses in directory-based shared-memory multiprocessors. By predicting a last-write-touch and self-downgrading a cache block in advance, a processor can get the data from the memory directly and the coherence overhead is significantly reduced. In this paper, we propose a new low overhead last-write-touch prediction scheme that exploits the inherent write burst characteristics of programs. The scheme uses write burst numbers to compute history traces and generate signatures. Compared with the existing instruction-based prediction technique, much storage overhead can be reduced. The experimental results show that our last-write-touch prediction scheme can achieve almost the same prediction accuracy as the instruction-based prediction scheme with the storage overheads of the history table reduced by 69% and the storage overheads of the signature table reduced by 36%.
Zhengbin Pang, Junsheng Chang
DASC3
2014 Low-latency last-level cache structure based on grouped cores in Chip Multi-Processor
abstract
Last-Level Cache (LLC) plays an important role in Chip Multi-Processor (CMP). The objective of this work is to optimize the structure and management strategy of LLC. Based on 8-core CMP, a LLC structure based on grouped cores is proposed, where 8 cores are divided into 4 groups. All LLC resources are classified into three types, which are fixed private cache, dynamic private cache and dynamic shared cache. The layout of the LLC structure and the corresponding dynamic partitioning strategy are designed to achieve low access latency and high efficiency. Experimental results on full-system simulator suggest that the proposed structure and method are able to reduce the access latency by 2% to 12% compared with previous works, such as tiled structure, cache-centered structure and core-centered structure. Consequently, performance measured by IPC is improved up to 7%. The contribution of this paper is useful for CMP performance, and applies to not only 8-core CMP but also all small-scale CMPs.
Jinbo Xu, Kefei Wang, Zhengbin Pang
IPCCC4
2014 Selective Extension of Routing Algorithms Based on Turn Model
abstract
Turn-model is a classical method for designing partially adaptive routing algorithms without virtual channels, and can also be the basis of fully adaptive routing algorithms. We propose a novel scheme, Selective Extension of Routing Algorithms based on turn model (SERA), which alleviates restrictions on turn and path selections if possible without adding any new buffers or virtual channels. SERA can improve adaptivity of the original routing algorithms, and maintain the deadlock-free property. Thus, SERA is an important extension of the previous turn model theory. To present the effectiveness of SERA in adaptive algorithms, we redesign two existing routing algorithms, Odd-Even and LEAR. Simulation results show that the SERA scheme achieves an average delay reduction of 6% compared to the original routing algorithms.
Liquan Xiao, Sheng Ma, Zhengbin Pang, Kefei Wang
PDP4
2014 The TH Express high performance interconnect networks
Zhengbin Pang, Guibin Wang, Dezun Dong, Guang Suo
Frontiers Comput. Sci.1
2014 An incentive compatible reputation mechanism for P2P systems
Junsheng Chang, Zhengbin Pang, Huaimin Wang 0001, Gang Yin
J. Supercomput.2
2013 Scalable NIC Architecture to Support Offloading of Large Scale MPI Barrier
Weixia Xu 0001, Zhengbin Pang, Pingjing Lu
APPT4
2013 Fine-Grained Location-Free Planarization in Wireless Sensor Networks
abstract
Extracting planar graph from network topologies is of great importance for efficient protocol design in wireless ad hoc and sensor networks. Previous techniques of planar topology extraction are often based on ideal assumptions, such as UDG communication model and accurate node location measurements. To make these protocols work effectively in practice, we need extract a planar topology in a location-free and distributed manner with small stretch factors. The planar topologies constructed by current location-free methods often have large stretch factors. In this paper, we present a fine-grained and location-free network planarization method under ρ-quasi-UDG communication model with ρ≥1/√2. Compared with existing location-free planarization approaches, our method can extract a provably connected planar graph, called topological planar simplification (TPS), from the connectivity graph in a fine-grained manner using local connectivity information. We evaluate our design through extensive simulations and compare with the state-of-the-art approaches. The simulation results show that our method produces high-quality planar graphs with a small stretch factor in practical large-scale networks.
Dezun Dong, Xiangke Liao, Yunhao Liu 0001, Xiang-Yang Li 0001, Zhengbin Pang
IEEE Trans. Mob. Comput.5
2012 Inferring Assertion for Complementary Synthesis
abstract
Complementary synthesis can automatically synthesize the decoder circuit of an encoder. However, its user needs to manually specify an assertion on some configuration pins to prevent the encoder from reaching the nonworking states. To avoid this tedious task, this paper propose an automatic approach to infer this assertion, by iteratively discovering and removing cases without decoders. To discover all decoders that may exist simultaneously under this assertion, a second algorithm based on functional dependency is proposed to decompose , the Boolean relation that uniquely determines the encoder's input, into all possible decoders. To help the user select the correct decoder, a third algorithm is proposed to infer each decoder's precondition formula, which represents those cases that lead to this decoder's existence. Experimental results on several complex encoders indicate that our algorithm can always infer assertions and generate decoders for them. Moreover, when multiple decoders exist simultaneously, the user can easily select the correct one by inspecting their precondition formulas.
ShengYu Shen, Kefei Wang, Zhengbin Pang, Sikun Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2011 A Parallel Processing Scheme for Large-Size Sliding-Window Applications
abstract
There exists large gap between the data input speed and processing speed in large-size sliding-window applications. To shorten this gap, a parallel processing scheme is proposed, which achieves high data reusability and parallelism with memory resources as few as possible and memory access control logics as simple as possible. This scheme combines the advantages of parallelism among different sliding-windows and parallelism among different data in a single window. For different windows, they are divided into groups and mapped into multiple processing elements. And for data in a single window, multi-module memory structure is introduced to buffer them, where module assignment and addressing scheme is designed for conflict-free parallel access. Experimental results on FPGA show that this work can improve the processing speed significantly without incurring too many memory resources and too complicated memory access control logics.
Jinbo Xu, Zhengbin Pang
HPCC3
2011 Finding First-Order Minimal Unsatisfiable Cores with a Heuristic Depth-First-Search Algorithm
ShengYu Shen, Zhengbin Pang, Sikun Li
IDEAL5
2009 DTM: Decoupled Hardware Transactional Memory to Support Unbounded Transaction and Operating System
abstract
Supporting unbounded transactions and operating system (OS) are two notable challenges that need to be efficiently resolved by practically accepted hardware transaction memory (HTM) systems. Current proposals that heavily rely on traditional cache system to handle version management or conflict detection support poorly to resolve these two challenges. Traditional approach brings either unavoidable design complexity or large performance costs. We propose DTM (decoupled transactional memory), a comprehensive hardware-based solution which follows a new design approach that fully decouples transaction processing from traditional cache system to resolve these two challenges.
Zhengbin Pang, Qiang Dou
ICPP3
2008 Software Assisted Transact Cache to Support Efficient Unbounded Transactional Memory
abstract
Transactional memory (TM) provides efficient, easy, deadlock-free parallel programming model for today's multicore-ubiquitous hardware platform. Implementation of TM needs to guarantee that the transaction is executed atomically and in isolation. Our paper proposes an efficient and unbounded hybrid-mode TM system with strong isolation guarantee, called HybridTCache. HybridTCache optimizes the common case by executing small transactions completely by hardware, and triggers operating system (OS) support with low overhead for the uncommon case when transaction size exceeds the hardware capacity. HybridTCache adds a new L1 cache, named TCache, to buffer transactional data for the active transaction executed by the processor. Compared with traditional log based approach, TCache provides fast bookkeeping which eliminates software logging overhead for the un-overflowed blocks, thus making both transaction commit and abort fast. A key design point of hardware TM is to support unbounded transactions. HybridTCache achieves this by introducing TCache overflow exceptions and resorting to OS to handle the overflowed blocks.
Zhengbin Pang
HPCC3
2007 Exploring Data Reusing of Failed Transaction
Zhengbin Pang
APPT4