VLDB 2026 Research / reviewers in the wild / expert
Yufeng Gu
dblp:253/2398
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FCL: frequency-based contrastive learning for generalizable face forgery detection
Yu Zhu 0005, Shengze Wang 0008, Yufeng Gu, Nan Wang 0003 |
Multim. Syst. | 3 |
| 2025 | PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model InferenceabstractLarge Language Model (LLM) inference uses an autoregressive manner to generate one token at a time, which exhibits notably lower operational intensity compared to earlier Machine Learning (ML) models such as encoder-only transformers and Convolutional Neural Networks. At the same time, LLMs possess large parameter sizes and use key-value caches to store context information. Modern LLMs support context windows with up to 1 million tokens to generate versatile text, audio, and video content. A large key-value cache unique to each prompt requires a large memory capacity, limiting the inference batch size. Both low operational intensity and limited batch size necessitate a high memory bandwidth. However, contemporary hardware systems for ML model deployment, such as GPUs and TPUs, are primarily optimized for compute throughput. This mismatch challenges the efficient deployment of advanced LLMs and makes users to pay for expensive compute resources that are poorly utilized for the memory-bound LLM inference tasks. Yufeng Gu, Alireza Khadem, Sumanth Umesh, Xavier Servot, Onur Mutlu, Ravi R. Iyer 0001, Reetuparna Das |
ASPLOS (2) | 1 |
| 2025 | Multi-Dimensional Vector ISA Extension for Mobile In-Cache Computing
Alireza Khadem, Daichi Fujiki, Hilbert Chen, Yufeng Gu, Nishil Talati, Scott A. Mahlke, Reetuparna Das |
HPCA | 4 |
| 2025 | DX100: Programmable Data Access Accelerator for IndirectionabstractIndirect memory accesses frequently appear in applications where memory bandwidth is a critical bottleneck.Prior indirect memory access proposals, such as indirect prefetchers, runahead execution, fetchers, and decoupled access/execute architectures, primarily focus on improving memory access latency by loading data ahead of computation but still rely on the DRAM controllers to reorder memory requests and enhance memory bandwidth utilization.DRAM controllers have limited visibility to future memory accesses due to the small capacity of request buffers and the restricted memorylevel parallelism of conventional core and memory systems.We introduce DX100, a programmable data access accelerator for indirect memory accesses.DX100 is shared across cores to offload bulk indirect memory accesses and associated address calculation operations.DX100 reorders, interleaves, and coalesces memory requests to improve DRAM row-buffer hit rate and memory bandwidth utilization.DX100 provides a general-purpose ISA to support diverse access types, loop patterns, conditional accesses Alireza Khadem, Kamalakkannan Kamalavasan, Zhenyan Zhu, Akash Poptani, Yufeng Gu, Jered Dominguez-Trujillo, Nishil Talati, Daichi Fujiki, Scott A. Mahlke, Galen M. Shipman, Reetuparna Das |
ISCA | 5 |
| 2025 | LWD-IUM: A Lightweight Detector for Advancing Robotic Grasp in VR-Based Industrial and Underwater MetaverseabstractIn the burgeoning field of virtual reality (VR) metaverse, the sophistication of interactions between robotic agents and their environment has become a critical concern. In this work, we present LWD-IUM, a novel light-weight detector designed to enhance robotic grasp capabilities in the VR metaverse. LWD-IUM applies deep learning techniques to discern and navigate the complex VR metaverse environment, aiding robotic agents in the identification and grasping of objects with high precision and efficiency. The algorithm is constructed with an advanced lightweight neural network structure based on self-attention mechanism that ensures optimal balance between computational cost and performance, making it highly suitable for real-time applications in VR. Evaluation on the KITTI 3D dataset demonstrated real-time detection capabilities (24-30 fps) of LWD-IUM, with its mean average precision (mAP) remaining 80% above standard 3D detectors, even with a 50% parameter reduction. In addition, we show that LWD-IUM outperforms existing models for object detection and grasping tasks through the real environment testing on a Baxter dual-arm collaborative robot. By pioneering advancements in robotic grasp in the VR metaverse, LWD-IUM promotes more immersive and realistic interactions, pushing the boundaries of what’s possible in virtual experiences. Liangfan Shi, Yufeng Gu, Yuchao Zheng 0001, Shintaro Kameda, Huimin Lu 0001 |
IWCMC | 2 |
| 2023 | GenDP: A Framework of Dynamic Programming Acceleration for Genome Sequencing AnalysisabstractGenomics is playing an important role in transforming healthcare. Genetic data, however, is being produced at a rate that far outpaces Moore's Law. Many efforts have been made to accelerate genomics kernels on modern commodity hardware such as CPUs and GPUs, as well as custom accelerators (ASICs) for specific genomics kernels. While ASICs provide higher performance and energy efficiency than general-purpose hardware, they incur a high hardware design cost. Moreover, in order to extract the best performance, ASICs tend to have significantly different architectures for different kernels. The divergence of ASIC designs makes it difficult to run commonly used modern sequencing analysis pipelines due to software integration and programming challenges. Yufeng Gu, Arun Subramaniyan 0001, Timothy Dunn, Alireza Khadem, Kuan-Yu Chen 0001, Somnath Paul, Md. Vasimuddin, Sanchit Misra, David T. Blaauw, Satish Narayanasamy, Reetuparna Das |
ISCA | 1 |
| 2021 | GenomicsBench: A Benchmark Suite for GenomicsabstractOver the last decade, advances in high-throughput sequencing and the availability of portable sequencers have enabled fast and cheap access to genetic data. For a given sample, sequencers typically output fragments of the DNA in the sample. Depending on the sequencing technology, the fragments range from a length of 150-250 at high accuracy to lengths in few tens of thousands but at much lower accuracy. Sequencing data is now being produced at a rate that far outpaces Moore's law and poses significant computational challenges on commodity hardware. To meet this demand, software tools have been extensively redesigned and new algorithms and custom hardware have been developed to deal with the diversity in sequencing data. However, a standard set of benchmarks that captures the diverse behaviors of these recent algorithms and can facilitate future architectural exploration is lacking. To that end, we present the GenomicsBench benchmark suite which contains 12 computationally intensive data-parallel kernels drawn from popular bioinformatics software tools. It covers the major steps in short and long-read genome sequence analysis pipelines such as basecalling, sequence mapping, de-novo assembly, variant calling and polishing. We observe that while these genomics kernels have abundant data level parallelism, it is often hard to exploit on commodity processors because of input-dependent irregularities. We also perform a detailed microarchitectural characterization of these kernels and identify their bottlenecks. GenomicsBench includes parallel versions of the source code with CPU and GPU implementations as applicable along with representative input datasets of two sizes - small and large. Arun Subramaniyan 0001, Yufeng Gu, Timothy Dunn, Somnath Paul, Md. Vasimuddin, Sanchit Misra, David T. Blaauw, Satish Narayanasamy, Reetuparna Das |
ISPASS | 2 |
| 2020 | Efficient Shapley Explanation for Features Importance Estimation Under Uncertainty
Xiaoxiao Li 0001, Yuan Zhou 0004, Nicha C. Dvornek, Yufeng Gu, Pamela Ventola, James S. Duncan |
MICCAI (1) | 4 |
| 2020 | Multi-site fMRI analysis using privacy-preserving federated learning and domain adaptation: ABIDE resultsabstractDeep learning models have shown their advantage in many different tasks, including neuroimage analysis. However, to effectively train a high-quality deep learning model, the aggregation of a significant amount of patient information is required. The time and cost for acquisition and annotation in assembling, for example, large fMRI datasets make it difficult to acquire large numbers at a single site. However, due to the need to protect the privacy of patient data, it is hard to assemble a central database from multiple institutions. Federated learning allows for population-level models to be trained without centralizing entities' data by transmitting the global model to local entities, training the model locally, and then averaging the gradients or weights in the global model. However, some studies suggest that private information can be recovered from the model gradients or weights. In this work, we address the problem of multi-site fMRI classification with a privacy-preserving strategy. To solve the problem, we propose a federated learning approach, where a decentralized iterative optimization algorithm is implemented and shared local model weights are altered by a randomization mechanism. Considering the systemic differences of fMRI distributions from different sites, we further propose two domain adaptation methods in this federated learning formulation. We investigate various practical aspects of federated model optimization and compare federated learning with alternative training strategies. Overall, our results demonstrate that it is promising to utilize multi-site data without data sharing to boost neuroimage analysis performance and find reliable disease-related biomarkers. Our proposed pipeline can be generalized to other privacy-sensitive medical data analysis problems. Our code is publicly available at: https://github.com/xxlya/Fed_ABIDE/. Xiaoxiao Li 0001, Yufeng Gu, Nicha C. Dvornek, Lawrence H. Staib, Pamela Ventola, James S. Duncan |
Medical Image Anal. | 2 |