Qinfen Hao

dblp:67/1667 · DBLP profile ↗
← Back
17ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-1545-1600ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CAMI: A Context-Aware Isolation Architecture for GPU Memories
abstract
The widespread use of GPUs in cloud and high-performance computing makes memory isolation a critical security requirement. While the programming model assumes that each thread local memory is private, the underlying hardware does not always enforce this guarantee. Weaknesses in address translation can allow one thread to access another local memory, creating a semantic gap that enables cross-thread corruption and exploitation. To address these challenges, we propose CAMI, a hardware-level framework that integrates fine-grained execution context into the memory translation pipeline. CAMI enforces a binding between the execution context of each memory access and the ownership of its target memory page, ensuring that even subtle inconsistencies in translation cannot be exploited. By introducing an efficient hardware enforcement unit within the MMU and extending page table entries with ownership metadata, CAMI achieves strong, fine-grained isolation while maintaining low performance overhead. We implement CAMI in a cycle-accurate GPU simulator and conduct comprehensive evaluations. Results show that CAMI effectively eliminates cross-thread memory access vulnerabilities with minimal runtime cost, offering a practical path toward secure and high-performance GPU architectures.
Wei Yan 0005, Qinfen Hao, Xiaochun Ye, Yier Jin, Ninghui Sun
DATE3
2025 Accelerating Oblivious Transfer with a Pipelined Architecture
abstract
With the rapid development of machine learning and big data technologies, ensuring user privacy has become a pressing challenge. Secure multi-party computation offers a solution to this challenge by enabling privacy-preserving computations, but it also incurs significant performance overhead, thus limiting its further application. Our analysis reveals that the oblivious transfer protocol accounts for up to 96.64% of execution time. To address these challenges, we propose POTA, a high-performance pipelined OT hardware acceleration architecture supporting the silent OT protocol. Finally, we implement a POTA prototype on Xilinx VCU129 FPGAs. Experimental results demonstrate that under various network settings, POTA achieves significant speedups, with maximum improvements of 22.67x for OT efficiency and 192.57x for basic operations in MPC applications.
Wei Yan 0005, Qinfen Hao, Ninghui Sun
DATE5
2025 SSMDVFS: Microsecond-Scale DVFS on GPGPUs with Supervised and Self-Calibrated ML
abstract
Over the past decade, as GPUs have evolved to achieve higher computational performance, their power density has also accelerated. Consequently, improving energy efficiency and reducing power consumption has become critically important. Dynamic voltage and frequency scaling (DVFS) is an effective technique for enhancing energy efficiency. With the advent of integrated voltage regulators, DVFS can now operate on microsecond$(\boldsymbol{\mu}\mathbf{s})$timescales. However, developing a practical and effective strategy to guide rapid DVFS remains a significant challenge. This paper proposes a supervised and self-calibrated machine learning framework (SSMDVFS) to guide microsecond-scale GPU voltage and frequency scaling. This framework features an end-to-end design that encompasses data generation, neural network model design, training, compression, and final runtime calibration. Unlike analytical models, which struggle to accurately represent GPU architectures, and reinforcement learning approaches, which can be challenging to converge during runtime, the SSMDVFS offers a practical solution for guiding microsecond-scale voltage and frequency scaling. Experimental results demonstrate that the proposed framework improves energy-delay product (EDP) by 11.09% and outperforms analytical models and reinforcement learning approaches by 13.17% and 36.80 %, respectively.
Minqing Sun, Yingtao Shen, Wei Yan 0005, Qinfen Hao, An Zou
DATE5
2025 ParTEE: A Framework for Secure Parallel Computing of RISC-V Trusted Execution Environments
Ziang Zhou, Wei Yan 0005, Qinfen Hao, Xiaochun Ye, Ninghui Sun
Euro-Par (2)5
2025 An Efficient Paillier Homomorphic Encryption Circuit With Optional CRT Acceleration for IoT
abstract
The Paillier scheme, widely recognized as the most prevalent additive homomorphic encryption paradigm, faces significant challenges in Internet of Things (IoT) applications due to latency, power, and hardware overhead. This paper proposes an efficient Paillier homomorphic encryption circuit for IoT, integrating Chinese Remainder Theorem (CRT) acceleration. First, we propose an algorithm framework tailored for hardware reuse that supports multiple functionalities of the Paillier scheme. It introduces an Montgomery modular multiplication (MMM) algorithm with superior Area-Time Product (ATP) to implement core computations, and reduces hardware cost by reusing MMM to replace other computational units. Then, a computational unit reuse architecture based on the algorithmic framework is designed to reduce resource overhead. Moreover, a split-coupled MMM circuit design is proposed to counteract computational resource expansion induced by CRT operations. The hardware design is synthesized under SMIC 40 nm CMOS technology. The evaluation shows that the proposed scheme provides a high-performance Paillier circuit design with less area and lower power, offering an effective solution for data security processing in IoT.
Jundong Feng, Zeljko Zilic, Qinfen Hao
IEEE Internet Things J.4
2024 MPC-PAT: A Pipeline Architecture for Beaver Triple Generation in Secure Multi-party Computation
abstract
Secure Multi-Party Computation (MPC) is proposed to protect the data privacy from a group of parties, enabling collaborative computation of correct results for target functions. SPDZ, a set of mature MPC protocols widely used in machine learning and other scenarios, requires a significant number of Beaver triples for secure multiplications among parties. Given no Trusted Third Party (TTP) participated, the generation time constitutes over 92% of the total running time. This paper introduces MPC-PAT, a high-performance pipeline architecture designed for efficient Beaver triple generation. MPC-PAT accelerates random number generation, hash function, and modular multiplication(MM) in two finite fields. The evaluation results from its FPGA implementation demonstrate 99× speed-up for basic operations and 136× speed-up for various convolutional networks compared to the existing SPDZ works.
Wei Yan 0005, Qinfen Hao, Ninghui Sun
ITC-Asia5
2023 CXL over Ethernet: A Novel FPGA-based Memory Disaggregation Design in Data Centers
abstract
Memory resources in data centers generally suffer from low utilization and lack of dynamics. Memory disaggregation solves these problems by decoupling CPU and memory, which currently includes approaches based on RDMA or interconnection protocols such as Compute Express Link (CXL). However, the RDMA-based approach involves code refactoring and higher latency. The CXL-based approach supports native memory semantics and overcomes the shortcomings of RDMA, but is limited within rack level. In addition, memory pooling and sharing based on CXL products are currently in the process of early exploration and still take time to be available in the future. In this paper, we propose the CXL over Ethernet approach that the host processor can access the remote memory with memory semantics through Ethernet. Our approach can support native memory load/store access and extends the physical range to cross server and rack levels by taking advantage of CXL and RDMA technologies. We prototype our approach with one server and two FPGA boards with 100 Gbps network and measure the memory access latency. Furthermore, we optimize the memory access path by using data cache and congestion control algorithm in the critical path to further lower access latency. The evaluation results show that the average latency for the server to access remote memory is$1.97\ \mu\mathrm{s}$, which is about 37% lower than the baseline latency in the industry. The latency can be reduced to 415 ns when memory accesses hit cache on FPGA.
Chenjiu Wang, Ruiqi Fan, Qinfen Hao
FCCM6
2022 HIRE: Distilling high-order relational knowledge from heterogeneous graph neural networks
Jing Liu 0080, Tongya Zheng, Qinfen Hao
Neurocomputing3
2016 A Highly Scalable Optical Network-on-Chip With Small Network Diameter and Deadlock Freedom
abstract
To increase the performance of chip multiprocessors, optical network-on-chip (ONoC) becomes promising because of its high bandwidth and low energy consumption. In this paper, we propose an architecture called RPNoC (Ring-based Packet-switched NoC), which uses few optical devices. Specifically, Single-waveguide RPNoC employs only one waveguide. Multiwaveguide RPNoC introduces space division multiplexing to make the architecture highly scalable. A novel wavelength assignment method and a deadlock-free deterministic routing algorithm are jointly designed, which make the network diameter quite small. This design also guarantees deadlock freedom, a little resource use, and low complexity at the same time. Evaluation is carried out for the 64-node RPNoC under different synthetic and realistic traffic patterns. The simulation result shows that it yields high throughput and low latency. Comparison with other packet-switched ONoCs shows that RPNoC has the lowest energy consumption.
Huaxi Gu, Yintang Yang, Kun Wang 0001, Qinfen Hao
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Alleviate chip I/O pin constraints for multicore processors through optical interconnects
abstract
Chip I/O pins are an increasingly limited resource and significantly affect the performance, power and cost of multicore processors. Optical interconnects promise low power and high bandwidth, and are potential alternatives to electrical interconnects. This work systematically developed a set of analytical models for electrical and optical interconnects to study their structures, receiver sensitivities, crosstalk noises, and attenuations. We verified the models by published implementation results. The analytical models quantitatively identified the advantages of optical interconnects in terms of bandwidth, energy consumption, and transmission distance. We showed that optical interconnects can significantly reduce chip pin counts. For example, compared to electrical interconnects, optical interconnects can save at least 92% signal pins when connecting chips more than 25 cm (10 inches) apart.
Zhehui Wang, Jiang Xu 0001, Peng Yang 0003, Xuan Wang 0001, Zhe Wang 0003, Luan H. K. Duong, Haoran Li 0002, Rafael Kioji Vivas Maeda, Xiaowen Wu, Yaoyao Ye, Qinfen Hao
ASP-DAC12
2015 Crosstalk Noise in WDM-Based Optical Networks-on-Chip: A Formal Study and Comparison
abstract
Optical networks-on-chip (ONoCs) using wavelength-division multiplexing (WDM) technology have progressively attracted more and more attention for their use in tackling the high-power consumption and low bandwidth issues in growing metallic interconnection networks in multiprocessor systems-on-chip. However, the basic optical devices employed to construct WDM-based ONoCs are imperfect and suffer from inevitable power loss and crosstalk noise. Furthermore, when employing WDM, optical signals of various wavelengths can interfere with each other through different optical switching elements within the network, creating crosstalk noise. As a result, the crosstalk noise in large-scale WDM-based ONoCs accumulates and causes severe performance degradation, restricts the network scalability, and considerably attenuates the signal-to-noise ratio (SNR). In this paper, we systematically study and compare the worst case as well as the average crosstalk noise and SNR in three well-known optical interconnect architectures, mesh-based, folded-torus-based, and fat-tree-based ONoCs using WDM. The analytical models for the worst case and the average crosstalk noise and SNR in the different architectures are presented. Furthermore, the proposed analytical models are integrated into a newly developed crosstalk noise and loss analysis platform (CLAP) to analyze the crosstalk noise and SNR in WDM-based ONoCs of any network size using an arbitrary optical router. Utilizing CLAP, we compare the worst case as well as the average crosstalk noise and SNR in different WDM-based ONoC architectures. Furthermore, we indicate how the SNR changes in respect to variations in the number of optical wavelengths in use, the free-spectral range, and the microresonators$\boldsymbol {Q}$factor. The analyses’ results demonstrate that the crosstalk noise is of critical concern to WDM-based ONoCs: in the worst case, the crosstalk noise power exceeds the signal power in all three WDM-based ONoC architectures, even when the number of processor cores is small, e.g., 64.
Mahdi Nikdast, Jiang Xu 0001, Luan H. K. Duong, Xiaowen Wu, Xuan Wang 0001, Zhehui Wang, Zhe Wang 0003, Peng Yang 0003, Yaoyao Ye, Qinfen Hao
IEEE Trans. Very Large Scale Integr. Syst.10
2014 Paraio: A scalable network I/O framework for many-core systems
abstract
Many high-performance networked applications are designed using the event-driven paradigm. In many-core era, hundreds or even thousands of processor cores can be utilized to serve more clients. However, data race and load imbalance in current event-driven hybrid models will be a key bottleneck which challenges developers to fully exploit many-core resources to develop high-performance networked applications. In this paper, we extend the symmetric multi-thread event-driven model and present Paraio, a scalable network I/O framework to improve performance of networked applications such as web servers and software-defined network (SDN) controllers. In order to maximize the degree of parallelism in the event-based application execution, Paraio features the shared-data marking method that divides the event-processing logic and marks event handlers from the essential shared data perspective. In Paraio runtime, workloads are balanced among threads by an efficient work stealing, and new connection is allocated according to threads' load to obtain a fast response. Evaluation on web server and SDN controller, shows that Paraio applications with work stealing achieve better performance and scalability.
Yi Liu 0013, Depei Qian 0001, Qinfen Hao
ICPADS5
2014 Keynote: "High throughput computing data center"
abstract
Over the last few decades, data center (DC) technology has evolved from DC 1.0 (tightly-coupled silos) to DC 2.0 (computer virtualization) in order to enhance data processing capability. In the era of big data, highly diversified analytics applications stress data centers. The mounting requirements on throughput, resource utilization, manageability and energy efficiency demand seamless integration of heterogeneous system resources to adapt to varied big data applications, for which DC 2.0 does not sufficient. By rethinking the challenges of big data applications, Huawei proposes High Throughput Computing Data Center architecture (HTC-DC) toward the design of DC 3.0. HTC-DC features resource disaggregation via unified interconnection. It offers PB-level data processing capability, intelligent manageability, high scalability and high energy efficiency, hence a promising candidate for DC 3.0.
Qinfen Hao
RTCSA1
2013 Shedder: A Metadata Sharing Management Method across Multi-clusters
Qinfen Hao, Qianqian Zhong, Zhenzhong Zhang
ICA3PP (1)1
2011 An Integrated Approach to Automatic Management of Virtualized Resources in Cloud Environments
abstract
Cloud computing, as a newly emergent computing environment, promises dynamic flexible infrastructures required to host Internet applications and application service level objects (SLOs) guaranteed services in a pay-as-you-go manner to the public. However, an important problem that remains to be effectively addressed is how to offer a cloud resource management solution that saves hardware and operations and management costs while meeting various SLOs. It faces the following challenges: complex dynamic relationships between application workload and SLOs and resource utilization, and the virtual machine (VM) placement problem in cloud environments. In this paper, we present an integrated approach that employs three-layered resource controllers using different analytic techniques, including the feedback control theory, statistical machine learning and system identification etc. Compared with Xen, KVM is chosen as the VM monitor to implement the proposed approach. Our experimental results show that the integration of layered controllers can reasonably allocate multiple resources to applications which execute on different VMs in cloud environments to achieve application SLOs under fluctuating time-varying workloads and unpredictable variations of system situations. In addition, it provides application SLO differentiation.
Qinfen Hao, Zhoujun Li 0001
Comput. J.2
2010 Formal Discussion on Relationship between Virtualization and Cloud Computing
abstract
As a most prevalent topic in recent few years, the public has shown great interest on cloud computing - the brand new concept about service pattern for IT industry. However, the public view is now mainly focused on the technical development of such a framework that is practically available, yet little research has been taken from the aspect of the academic. This article is written with the intention for some discussion and exploration on relationships between virtualization and Cloud Computing. The article provides one way it provides an attempt to give out a formal definition of cloud computing from the viewpoint of virtualization and the set theory upon which the concept of virtualization is based.
Hanfei Dong, Qinfen Hao, Tiegang Zhang
PDCAT2
2010 DeepComp: towards a balanced system design for high performance computer systems
Mingfa Zhu, Qinfen Hao
Frontiers Comput. Sci. China4