Wenhao Gu

dblp:221/1031 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Enhanced Reasoning for Biomedical Document-Level Relation Extraction via a Novel Cascade Language Model Framework
abstract
Haohua Song, Wenhao Gu, Zhijing Li, Yunwenyu, Tiantian Zhu, Xiao Yang, Zexuan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Haohua Song, Wenhao Gu, Zhijing Li 0007, Yunwen Yu, Xiao Yang 0019, Zexuan Zhu 0001
ACL (1)2
2026 DPIO: A Unified I/O Architecture for Heterogeneous CPU and DPU NVMeoF
Wenhao Gu, Xuchao Xie, Yujuan Tan, Dezun Dong
HPDC1
2025 Bridging Metadata Service and CXL: A Metadata-Grained and Directory-Aware Storage Engine for Distributed Storage Systems
abstract
The AI training and inference workloads are particularly metadata-intensive and drive an urgent need for distributed file systems (DFS) with high IOPS metadata service. Meanwhile, the emerging Compute Express Link (CXL) protocol introduces memory semantics to the PCIe-attached storage devices and is compelling for building high-performance metadata storage. However, a fundamental mismatch exists between the metadata access granularity and the internal storage granularity of typical CXL-enabled storage devices. Besides, existing metadata storage engines lack the perception of the DFS directory structures and metadata semantics in distributed storage systems. In this paper, we investigate the way to employ CXL-enabled devices as the storage backend for DFS metadata and propose MDSec, a metadata-grained and directory-aware metadata storage engine that bridges the semantic gaps between the DFS metadata service and CXL-enabled storage devices. MDSec unifies the granularity of all kinds of metadata for better metadata placement across CXL-enabled persistent storage, designs a directory-aware metadata grouping and placement strategy to improve the spatial locality of metadata access, and employs fully parallel metadata handlers to enhance metadata processing parallelism for parallel DFS client accesses. We evaluate MDSec on a cluster with 25 nodes. MDSec improves the throughput of Ext4 and NOVA by 258% and 53%, while reducing their latency by 46% and 11%, respectively. These results indicate that MDSec efficiently integrates CXL storage with DFS metadata service.
Xuchao Xie, Xinghan Qiao, Qiulin Wu, Wenhao Gu, Liquan Xiao
CLUSTER6
2025 MetaWriter: Personalized Handwritten Text Recognition Using Meta-Learned Prompt Tuning
abstract
Recent advancements in handwritten text recognition (HTR) have enabled the effective conversion of handwritten text to digital formats. However, achieving robust recognition across diverse writing styles remains challenging. Traditional HTR methods lack writer-specific personalization at test time due to limitations in model architecture and training strategies. Existing attempts to bridge this gap, through gradient-based meta-learning, still require labeled examples and suffer from parameter-inefficient fine-tuning, leading to substantial computational and memory overhead. To overcome these challenges, we propose an efficient framework that formulates personalization as prompt tuning, incorporating an auxiliary image reconstruction task with a self-supervised loss to guide prompt adaptation with unlabeled test-time examples. To ensure self-supervised loss effectively minimizes text recognition error, we leverage meta-learning to learn the optimal initialization of the prompts. As a result, our method allows the model to efficiently capture unique writing styles by updating less than 1% of its parameters and eliminating the need for time-intensive annotation processes. We validate our approach on the RIMES and IAM Handwriting Database benchmarks, where it consistently outperforms previous state-ofthe-art methods while using 20x fewer parameters. We believe this represents a significant advancement in personalized handwritten text recognition, paving the way for more reliable and practical deployment in resource-constrained scenarios.
Wenhao Gu, Li Gu, Chingyee Yee Suen, Yang Wang 0003
CVPR1
2025 DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning
abstract
Despite recent significant advancements in Handwritten Document Recognition (HDR), the efficient and accurate recognition of text against complex backgrounds, diverse handwriting styles, and varying document layouts remains a practical challenge. Moreover, this issue is seldom addressed in academic research, particularly in scenarios with minimal annotated data available. In this paper, we introduce the DocTTT framework to address these challenges. The key innovation of our approach is that it uses test-time training to adapt the model to each specific input during testing. We propose a novel Meta-Auxiliary learning approach that combines Meta-learning and self-supervised Masked Autoencoder (MAE). During testing, we adapt the visual representation parameters using a self-supervised MAE loss. During training, we learn the model parameters using a meta-learning framework, so that the model parameters are learned to adapt to a new input effectively. Experimental results show that our proposed method significantly outperforms existing state-of-the-art approaches on benchmark datasets.
Wenhao Gu, Li Gu, Ziqiang Wang 0003, Ching Y. Suen, Yang Wang 0003
WACV1
2022 LTNoT: Realizing the Trade-Offs Between Latency and Throughput in NVMe over TCP
Wenhao Gu, Xuchao Xie, Dezun Dong
ICA3PP1
2022 A Transformable NVMeoF Queue Design for Better Differentiating Read and Write Request Processing
abstract
NVMeoF is the latest extension of NVMe for remote storage access which allows remote access to NVMe controllers through high-speed RDMA, FC, and TCP networks. NVMe over TCP (NoT) can build on the basis of large-scale common network infrastructure in datacenters and standard TCP/IP software protocol stack, enabling a wide availability compared with RDMA-enabled specific network infrastructure for NVMe-overRDMA. However, the processing of read/write I/O at the host and target prominently shows significantly different characteristics and requirements, where one side sends the NVMeoF instruction of the request, while the other side sends the requested data. The existing NoT implementation can not meet the different characteristics of requests in the datacenter, which eventually results in the I/O performance being limited by the common processing pipeline and sending strategy. In this paper, we propose RNoT, a transformable queue that can meet the differentiated processing scheme of read or write request characteristics respectively in NoT implementation. Specifically, RNoT defines a switchable working attribute and separates resources for read and write I/O to achieve intra-queue long-term exclusivity, delivers read and write requests into other RNoT queue pairs to achieve inter-queue I/O scheduling, and transfers request command and data with targeted approaches to achieve short and long flow optimization. We implemented RNoT in Linux Kernel and evaluated it using realistic benchmarks and applications. Our experimental results demonstrate that RNoT can achieve 30.39% and 29.27% lower latency than i10 and NoT respectively, increase IOPS by up to 41.34% than NoT on average, thus RNoT can effectively optimize the read and write I/O performance in NoT with dedicated processing scheme.
Wenhao Gu, Xuchao Xie, Dezun Dong
ICPADS1
2022 Alleviating Performance Interference Through Intra-Queue I/O Isolation for NVMe-over-Fabrics
Wenhao Gu, Xuchao Xie, Dezun Dong
NPC1
2022 Prediction of biomarker-disease associations based on graph attention network and text representation
abstract
MOTIVATION: The associations between biomarkers and human diseases play a key role in understanding complex pathology and developing targeted therapies. Wet lab experiments for biomarker discovery are costly, laborious and time-consuming. Computational prediction methods can be used to greatly expedite the identification of candidate biomarkers. RESULTS: Here, we present a novel computational model named GTGenie for predicting the biomarker-disease associations based on graph and text features. In GTGenie, a graph attention network is utilized to characterize diverse similarities of biomarkers and diseases from heterogeneous information resources. Meanwhile, a pretrained BERT-based model is applied to learn the text-based representation of biomarker-disease relation from biomedical literature. The captured graph and text features are then integrated in a bimodal fusion network to model the hybrid entity representation. Finally, inductive matrix completion is adopted to infer the missing entries for reconstructing relation matrix, with which the unknown biomarker-disease associations are predicted. Experimental results on HMDD, HMDAD and LncRNADisease data sets showed that GTGenie can obtain competitive prediction performance with other state-of-the-art methods. AVAILABILITY: The source code of GTGenie and the test data are available at: https://github.com/Wolverinerine/GTGenie.
Zhi-an Huang, Wenhao Gu, Wenying Pan, Xiao Yang 0019, Zexuan Zhu 0001
Briefings Bioinform.3
2020 Generalizing Spatial Transformers to Projective Geometry with Applications to 2D/3D Registration
Cong Gao 0003, Xingtong Liu, Wenhao Gu, Benjamin Killeen, Mehran Armand, Russell H. Taylor, Mathias Unberath
MICCAI (3)3