Penghao Zhang

dblp:249/2206 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
abstract
Although Vision Language Models (VLMs) have shown generalization in medical imaging, pathology presents unique challenges due to ultra-high resolution, complex tissue structures, and nuanced semantics. These factors make pathology VLMs prone to hallucinations, i.e., generating outputs inconsistent with visual evidence, which undermines clinical trust. Existing RAG approaches in this domain largely depend on text-based knowledge bases, limiting their ability to leverage diagnostic visual cues. To address this, we propose Patho-AgenticRAG, a multimodal RAG framework with a database built on page-level embeddings from authoritative pathology textbooks. Unlike traditional text-only retrieval systems, it supports joint text–image search, enabling retrieval of textbook pages that contain both the queried text and relevant visual cues, thus avoiding the loss of critical image-based information. Patho-AgenticRAG also supports reasoning, task decomposition, and multi-turn search interactions, improving accuracy in complex diagnostic scenarios. Experiments show that Patho-AgenticRAG significantly outperforms existing multimodal models in complex pathology tasks like multiple-choice diagnosis and visual question answering.
Wenchuan Zhang, Jingru Guo, Hengzhe Zhang, Penghao Zhang, Shuwan Zhang, Yuhao Yi, Hong Bu
AAAI4
2026 Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
abstract
Recent advances in vision-language models (VLMs) have enabled broad progress in the general medical field. However, pathology still remains a more challenging sub-domain, with current pathology-specific VLMs exhibiting limitations in both diagnostic accuracy and reasoning plausibility. Such shortcomings are largely attributable to the nature of current pathology datasets, which are primarily composed of image–description pairs that lack the depth and structured diagnostic paradigms employed by real-world pathologists. In this study, we leverage pathology textbooks and real-world pathology experts to construct high-quality, reasoning-oriented datasets. Building on this, we introduce Patho-R1, a multimodal RL-based pathology Reasoner, trained through a three-stage pipeline: (1) continued pretraining on 3.5 million image-text pairs for knowledge infusion; (2) supervised fine-tuning on 500k high-quality Chain-of-Thought samples for reasoning incentivizing; (3) reinforcement learning using Group Relative Policy Optimization and Decoupled Clip and Dynamic sAmpling Policy Optimization strategies for multimodal reasoning quality refinement. To further assess the alignment quality of our dataset, we propose Patho-CLIP, trained on the same figure-caption corpus used for continued pretraining. Comprehensive experimental results demonstrate that both Patho-CLIP and Patho-R1 achieve robust performance across a wide range of pathology-related tasks, including zero-shot classification, cross-modal retrieval, Visual Question Answering, and Multiple Choice Question.
Wenchuan Zhang, Penghao Zhang, Jingru Guo, Tao Cheng 0006, Shuwan Zhang, Yuhao Yi, Hong Bu
AAAI2
2025 Cuckoo: Deadline-Aware Job Packing on Heterogeneous GPUs for DL Model Training
abstract
The growing scale and heterogeneity of GPU clusters pose new challenges to deep learning (DL) job scheduling. While existing schedulers primarily focus on GPU utilization, they often ignore multi-dimensional resource demands of DL workloads and lack precise execution time estimation for co-located jobs. While Muri pioneered the use of interleaving execution to improve resource efficiency, it simplified interference when jobs using one resource simultaneously and is agnostic to the deadline constraints. The job grouping also comes to suboptimal when heterogeneous GPU devices are taken into account. In this paper, we propose Cuckoo, a scheduling system that packs deep learning jobs with stringent deadline requirements over a set of heterogeneous GPU devices where multi-dimensional resources are interleaved and shared by a group of jobs. Specifically, the interleaving execution of simultaneous jobs is characterized and modeled through stage-grained execution time estimation considering the runtime performance interference and the impact of GPU heterogeneity on the job performance. The job packing is formulated as a multi-objective optimization problem which is then solved by the maximum weight matching algorithm. Cuckoo then allocates heterogeneous resources to the packed job groups through a graph-based maximum flow and minimum cut algorithm. Experiments show that Cuckoo improves deadline satisfaction rate by 2.38x and reduces average job completion time (JCT) by 1.81x compared with the state-of-the-art approaches. Cuckoo is implemented based on Kubernetes and has been deployed in Kuaishou to serve thousands of model training jobs that can be interleaved on shared heterogeneous GPU clusters
Yuzheng Zhang, Renyu Yang, Weihan Jiang, Tianyu Ye, Yiqiao Liao, Penghao Zhang, Tiezi Zhang, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang
SoCC7
2025 KAIOPS: A Platform Solution of End-to-End Multi-Modal AIOps for AI Training at Scale
abstract
The resilience of large-scale AI training platforms are fundamental to enabling contemporary AI innovation and business development. However, with the rapid increase in the scale and complexity of AI model training tasks, anomalies become the norm rather than the exception at scale. Failing to handle them properly may lead to enormous resource waste and prolonged development cycles. Traditional anomaly detection methods struggle to tackle the complex temporal characteristics and extreme class imbalance inherently manifesting in training tasks, and fall short in automated solution to root cause analysis and the follow-up remediation. This paper proposes KAIOPS, an end-to-end automated platform solution for handling anomalies and engineering experience of daily operational maintenance for large-scale AI training clusters at Kuaishou. KAIOPS employs a Temporal Context Encoding mechanism to precisely capture and encode long-term trends and critical temporal context information within fault evolution. The detection model elaborates a dynamic class-weighted loss function for enhancing the detection performance. To deliver a complete end-to-end intelligent processing pipeline, KAIOPS further leverages knowledge graph and LLMs for automated root cause analysis and actionable solution generation. Extensive experiments, on the basis of data collected from Kuaishou’s production-grade training clusters, show the superior performance of our proposed approach. KAIOPS has been deployed in Kuaishou, in both testbed and production grade environments, consisting of with over 10,000 GPUs, and accelerate the reliability assurance for industry-scale model training and serving.
Zeying Wang, Penghao Zhang, Xu Wang 0007, Tianyu Wo, Chunming Hu, Chengru Song, Jin Ouyang, Renyu Yang
ASE3
2025 Brief Announcement: Accelerating Distributed Search System with In-network Computation
abstract
Locality Sensitive Hashing (LSH) has proven to be an effective technique for indexing similar data in high-dimensional spaces, enabling efficient approximate nearest neighbor searches. However, as data sets continue to grow in size and complexity, traditional LSH-based distributed search systems face challenges in terms of network bottlenecks. Specifically, the transmission of large packets containing candidate answers to agent servers can significantly affect response time and throughput.
Penghao Zhang
SPAA1
2024 Zebra: Accelerating Distributed Sparse Deep Training With in-Network Gradient Aggregation for Hot Parameters
abstract
Distributed sparse deep learning has been widely used in many Internet-scale applications. Network communication is one of the major hurdles for training performance. In-network gradient aggregation on programmable switches is a promising solution for speeding up the performance. Nevertheless, existing in-network aggregation solutions are designed for the dense deep training, and fall short when used for the sparse training. To address this gap, we present Zebra based on our key observation on the extremely biased update frequency of parameters in distributed sparse deep training. Specifically, Zebra offloads only the aggregation for “hot” parameters that are updated frequently onto programmable switches. To enable this offloading and achieve high aggregation throughput, we propose solutions to address the challenges related to hot parameter identification, parameter orchestration and gradient aggregation as well as system reliability. We implemented Zebra on Intel Tofino switches and integrated it with PS-lite. Finally, we evaluate Zebra's performance through extensive experiments and show that it can speed up the gradient aggregation by$1.5 \sim 4 \times$and the end-to-end performance by$1.4 \sim 2.6 \times$.
Penglai Cui, Zhenyu Li 0001, Ru Jia, Penghao Zhang, Mathy Lauren, Gaogang Xie
ICNP5
2024 NetDS: Distributed Search Framework with Hybrid Acceleration Methods
abstract
Approximate neighbour nearest search has achieved great success for indexing similar high-dimensional data in distributed search systems. As the scale of data vectors grows, distributed search require large storage, low latency, and high throughput on processing vectors. To achieve this, researchers tend to load balance data with more machines and implement efficient distributed frameworks, but they need to pay huge storage overhead, which leads to inefficient network transmission.To address this gap, we propose NetDS, which exploits the computational capacity of in-network computation and the storage capacity of solid-state drives. NetDS utilizes a multi-level constrained balanced tree to process data vectors and construct multi-level tables. Then, NetDS proposes a heuristic neighbour graph to solve the boundary data problem. NetDS also offloads central tables and graphs into switches to accelerate vector classification. Finally, NetDS designs hybrid storage and query pre-match methods to accelerate the ANNS distributed system. We deploy NetDS on a programmable switch and evaluate it. NetDS completes the data and query processing in a shorter time than other typical distributed frameworks.
Penghao Zhang, Zhiguo Hu
ISPA1
2024 A novel attribute reduction method with constraints on empirical risk and decision rule length
Penghao Zhang, Yanjun Liu 0008, Guoyin Wang 0001
Inf. Sci.2
2023 An Efficient Network Flow Rules Compression Method Based On In-Network Computation On Cloud Computing
abstract
The development of cloud computing has led to the explosion of network traffic. The switches of cloud computing make it hard to process large-scale network traffic. Prior approaches proposed flow rules compression methods with splitting matching fields, expanding switch storage space, or reconstructing network architecture. The methods incurred little performance improvement in the network. The technology of innetwork computation(INC) has been widely used to accelerate the network of cloud computing and data-intensive distributed applications. The computing tasks performed on servers are offloaded to the network through INC. We implement a flow rules compression method(NetFR) based on INC to accelerate the network performance of cloud computing. NetFR adopts heuristic methods which include rule-deduplication and optimal coverage to compress flow rules, and offloads compressed rules into switches to accelerate the performance of the network on cloud computing. Finally, we evaluate NetFR and show that the performance of NetFR is 1.3x better than other methods of network packet matching. Furthermore, NetFR can achieve 98% compression ratio on average and 99.99% on highest.
Penghao Zhang
ICPADS1
2023 Misconfiguration-Free Compositional SDN for Cloud Networks
abstract
Cloud computing provides a new paradigm to offer flexible IT infrastructures. In IaaS clouds, tenants deploy software-defined networking (SDN) policies to simplify network management and customize network behaviors. However, programming SDN networks is error-prone no matter using low-level APIs or high-level programming languages. Specifically, SDN policies may contain misconfigurations that do not break the pre-defined network invariants (e.g., black holes), but either degrade the deployment efficiency or mistakenly translate tenants intents. Prior studies for checking either traditional access control policies or network-wide invariants, are thus fail to detect these misconfigurations. To address this gap, this paper presents PMM, a misconfiguration checking tool for compositional SDN that works at the data plane of cloud networks. We first propose a new data structure, minimal interval set, to represent the match patterns of rulesets. This representation serves the basis for composition algebra construction and misconfiguration checking. We then propose the principles, algorithms and also optimisations for fast and accurate checking. We finally implement PMM in Covisor. Experiments with both real-world rulesets and synthetic rulesets show that PMM can detect misconfigurations of SDN policies in cloud networks within hundreds of milliseconds.
Zhenyu Li 0001, Penghao Zhang, Penglai Cui, Kavé Salamatian, Gaogang Xie
IEEE Trans. Dependable Secur. Comput.3
2022 A Multicriteria Ranking Approach for Evaluating Best Cities for International Students
abstract
The dramatic increase in the number of students enrolled in higher education programs outside their country of citizenship during the last half-century has created a huge demand for study abroad-related information circulation. Nonetheless, media nowadays attempts to break information barriers by gathering and processing data from multiple sources, mainly focusing on university academic competency. Although important, academic competency cannot represent international students’ overall quality of life when spending their time in unfamiliar foreign cities. This research is designed to provide solutions to the problem from another perspective Utilizing the Multiple Criteria Decision Making (MCDM) method to evaluate international study destinations through various dimensions comprehensively. The research integrates various aspects such as economic development, culture inclusiveness, and personal safety into account. It then adopts the Technique for Order Preference by Similarity to Ideal Solution (TOPSIS) to obtain a detailed index for each destination in the rank. Our work-Global Ranking of Study Destination (GRSD) by cities, is accomplished to contribute to the international student community. With a ranking system that considers vital facets of life in certain cities, prospective international students can make assessments and decisions better for their future.
Penghao Zhang, Zeyu Hou, Junyi Chai 0001
SMC1
2022 Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading
abstract
Programmable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance.
Penglai Cui, Zhenyu Li 0001, Penghao Zhang, Tianhao Miao, Jianer Zhou, Hongtao Guan, Gaogang Xie
IEEE Trans. Parallel Distributed Syst.4
2022 NetSHa: In-Network Acceleration of LSH-Based Distributed Search
abstract
Locality Sensitive Hashing (LSH) is widely adopted to index similar data in high-dimensional space for approximate nearest neighbor search. Demanding applications (e.g. web search) mean that LSH must exhibit low response times and high throughput. To achieve this, they tend to load balance between multiple machines. However, as the scale of concurrent queries and the volume of data grow, large numbers of index messages are required. Hence, the network is a key bottleneck. To address this gap, we propose NetSHa, which exploits the computational capacity of programmable switches. Specifically, we introduce a heuristic sort-reduce approach to drop potentially poor candidate answers while preserving search quality. Then, NetSHa aggregates good candidate answers from different index messages when transmitting them. Through this, it reduces the network communication cost. Furthermore, we introduce a best-effort replacement mechanism to improve its concurrency. We implement NetSHa on a Barefoot Tofino programmable switch and evaluate it using 7 real-world datasets. The experimental results show that NetSHa reduces the packet volume by$4\sim 10$times and improves the search efficiency by least 3× in comparison with typical LSH-based distributed search frameworks.
Penghao Zhang, Zhenyu Li 0001, Penglai Cui, Ru Jia, Peng He 0003, Gareth Tyson, Gaogang Xie
IEEE Trans. Parallel Distributed Syst.1
2021 Accelerating LSH-based Distributed Search with In-network Computation
abstract
Locality Sensitive Hashing (LSH) is widely adopted to index similar data in high-dimensional space for approximate nearest neighbor search. With the rapid increase of datasets, recent interests in LSH have moved to the implementation of distributed search systems with low response time and high throughput. However, as the scale of the concurrent queries and the volume of available data grow, large amounts of index messages still need to be transmitted to centralized servers for the candidate answer reducing and resorting. Hence, the network remains the bottleneck in distributed search systems.To address this gap, we turn our efforts to the network itself and propose NetSHa. NetSHa exploits the in-network computational capacity provided by programmable switches. Specially, NetSHa designs a sort-reduce approach to drop the potential poor candidate answers and aggregates the good candidate answers on programmable switches, while preserving the search quality. We implement NetSHa on Barefoot Tofino switches and evaluate it using 3 datasets (i.e., Random, Wiki and Image). The experimental results show that NetSHa reduces the packet volume by 10 times at most and improves the search efficiency by 3x at least, in comparison with typical LSH-based distributed search frameworks.
Penghao Zhang, Zhenyu Li 0001, Peng He 0003, Gareth Tyson, Gaogang Xie
INFOCOM1
2021 Fast Online Packet Classification With Convolutional Neural Network
abstract
Packet classification is a critical component in network appliances. Software Defined Networking and cloud computing update the rulesets frequently for flexible policy configuration. Tuple Space Search (TSS), implemented in Open vSwitch (OVS), achieves fast rule updating at the sacrifice of the classification rate. In TSS, each tuple is managed by a hash table and classifying a packet needs to go through all hash tables. Merging tuples can reduce the number of hash tables, but inevitably increases the hash conflicts that may even worsen the classification performance in some cases. No existing algorithm meets the need of both fast packet classification and online rule updating. In this paper, we propose Convolutional Neural Network (CNN)-based Range Partition (CRP) to achieve fast packet classification and online update simultaneously. CRP exploits CNN-based image recognition to quickly partition tuples into range spaces upon the change of ruleset distribution, which reduces hash operations while avoiding rule overlapping caused by hashing many rules to the same location of the hash table. Experimental results demonstrate that CRP achieves$3.2\times $classification speed and$4.2\times $update speed on average compared with state-of-the-art algorithms. We also implement CRP in OVS. The throughput of CRP-OVS is$10\times $that of native OVS.
Xinyi Zhang 0004, Gaogang Xie, Xin Wang 0001, Penghao Zhang, Yanbiao Li 0001, Kavé Salamatian
IEEE/ACM Trans. Netw.4
2020 Misconfiguration Checking for SDN: Data Structure, Theory and Algorithms
abstract
Software-Defined Networking (SDN) facilitates net-work innovations with programmability. However, programming the network is error-prone no matter using low-level APIs or high-level programming languages. That said, SDN policies deployed in networks may contain misconfigurations. Prior studies focus on either traditional access control policies or network-wide states, and thus are unable to effectively detect potential misconfigurations in SDN policies with bitmask patterns and complex action behaviorsTo address this gap, this paper first presents a new data structure, minimal interval set, to represent the match patterns of rulesets. This representation serves the basis for composition algebra construction and fast misconfiguration checking. We then propose the principles and algorithms for fast and accurate con-figuration verification. We finally implement a misconfiguration checking tool in Covisor with optimisations to further reduce the overhead. Experiments with synthetic and random rulesets show its fitness for purpose.
Zhenyu Li 0001, Penghao Zhang, Kavé Salamatian, Gaogang Xie
ICNP3