EDBT 2026 Demo / reviewers in the wild / expert
Bin Gao 0013
dblp:181/2330-13
· DBLP profile ↗
9ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0001-5009-3514ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CARBS: Compiler Autotuning via Randomized Biased SearchabstractPerformance tuning is a persistent challenge in high-performance computing (HPC), where large codebases, architectural diversity, and limited evaluation budgets make manual optimization impractical. Modern compilers expose hundreds of optimization flags, yielding an enormous configuration space that existing autotuners explore inefficiently, often requiring thousands of costly trials. Bin Gao 0013, Weng-Fai Wong |
HPDC | 2 |
| 2026 | DANMP: Accelerating Multi-Scale Deformable Attention Using Near-Memory-Processing ArchitectureabstractMulti-Scale Deformable Attention (MSDAttn) has become a fundamental component in various vision tasks due to its effective multi-scale grid sampling (MSGS). However, its reliance on random sampling results in highly irregular memory access patterns, making it a memory-intensive operation inefficient for GPUs. Near-memory processing (NMP) offers a promising solution for accelerating memory-bound kernels, yet existing NMP-based attention accelerators remain suboptimal for MSDAttn due to incompatible load balancing and data reuse strategies. Specifically, current NMP solutions uniformly distribute processing elements (PEs) across all banks, leading to significant PE underutilization and excessive cross-bank data transfers. Moreover, most rely on locality-based reuse, which fails under MSDAttn’s unpredictable sampling patterns. Huize Li, Qinggang Wang, Bin Gao 0013, Dan Chen 0006, Yu Huang 0013, Xin Xin 0008 |
ICS | 3 |
| 2026 | Texplorer: Efficient Tensor Program Optimization for GPUs Using a Highly Constrained Search SpaceabstractOptimizing tensor programs for deep learning inference is costly, often requiring hours of search over massive search spaces ($\gt 10^{10}$). Prior machine learning-guided approaches are stochastic and inefficient because they treat both the search and programs as black boxes, overlooking two key properties: (1) the search space is highly structured, with large regions of invalid or low-performing configurations, and (2) tensor programs exhibit analyzable patterns in memory access and pipelined execution. We present Texplorer, a deterministic tensor program optimizer that exploits these structures for robust and efficient search. Texplorer introduces a three-stage pruning algorithm to eliminate invalid, redundant, and inefficient candidates, shrinking the search space from billions to thousands. It replaces costly learned models with a lightweight analytical cost model that captures GPU pipeline and memory behavior, enabling accurate, training-free performance prediction and strong cross-operator generalization. Finally, Texplorer integrates pruning and cost modeling into a one-pass analytical evaluation pipeline that ranks all remaining candidates, enabling rapid convergence without iterative retraining or stochastic sampling. Across 121 operators and 11 DNNs on NVIDIA V100, A100, and H100 GPUs, Texplorer matches state-of-the-art performance while reducing search time by up to$300\times$. An enhanced variant,Texplorer$^+$, further improves performance by up to 3.1% with up to$72\times$faster tuning. Bin Gao 0013, Weng-Fai Wong |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2025 | Online Context Caching for Distributed Large Language Models Serving
Bin Gao 0013, Zhuomin He, Yizhen Yao, Zhanzhi Lou Lou, Zhi Zhou 0006, Weng-Fai Wong |
INFOCOM | 1 |
| 2024 | IMI: In-memory Multi-job Inference Acceleration for Large Language ModelsabstractLarge Language Models (LLMs) are increasingly used in various applications but are computationally complex and energy-consuming due to the high volume of off-chip memory accesses. Processing-in-Memory (PIM) has emerged as a potential solution for efficient inference. However, existing PIM accelerators designed for deep neural networks (DNNs) aren’t suitable for LLMs because of differences in operations, input sizes, and job completion times. This leads to performance issues like head-of-line blocking where an earlier job can monopolize resources at the expense of later jobs, and low resource utilization. To improve efficiency, a time-multiplex solution and job colocation accelerator could be beneficial. However, facilitating multi-job execution with in-memory acceleration is challenging due to limitations in memristor architecture, inefficiency of non-stationary weight programming, difficulty in dynamic partitioning of hardware resources, and complexity in dynamic job scheduling. This work proposes a PIM-based LLM accelerator to enable the concurrent LLM inference job execution. The experiment shows that IMI can significantly improve resource utilization as well as the rate of satisfying the service level requirements of jobs. Bin Gao 0013, Zhehui Wang, Zhuomin He, Tao Luo 0014, Weng-Fai Wong, Zhi Zhou 0006 |
ICPP | 1 |
| 2023 | DAG-Aware Optimization for Geo-Distributed Data AnalyticsabstractGeo-distributed data analytics has been proposed to analyze geographically distributed data. Existing studies have achieved significant reductions in execution time and data transfer cost ($) of data analytics jobs by optimizing task placement. Given a directed acyclic graph (DAG)-style job, however, they mainly optimize each stage independently, and they tend to distribute tasks and intermediate data across all locations, potentially inflating execution time and data transfer cost of descendent stages and the whole job. Qingyuan Wang 0005, Bin Gao 0013, Zhi Zhou 0006, Fei Xu 0009, Chenghao Ouyang |
ICPP | 2 |
| 2022 | EAGAN: Efficient Two-Stage Evolutionary Architecture Search for GANs
Guohao Ying, Xin He 0019, Bin Gao 0013, Bo Han 0003, Xiaowen Chu 0001 |
ECCV (16) | 3 |
| 2022 | An Online Framework for Joint Network Selection and Service Placement in Mobile Edge ComputingabstractWith the rapid development and deployment of 5G wireless technology, mobile edge computing (MEC) has emerged as a new computing paradigm to facilitate a large variety of infrastructures at the network edge to reduce user-perceived communication delay. One of the fundamental problems in this new paradigm is to preserve satisfactory quality-of-service (QoS) for mobile users in light of densely dispersed wireless communication environment and often capacity-constrained MEC nodes. Such user-perceived QoS, typically in terms of the end-to-end delay, is highly vulnerable to both access network bottleneck and communication delay. Previous works have primarily focused on optimizing the communication delay through dynamic service placement, while ignoring the critical effect of access network selection on the access delay. In this work, we study the problem of jointly optimizing the access network selection and service placement for MEC, with the objective of improving the QoS in a cost-efficient manner by judiciously balancing the access delay, communication delay, and service switching cost. Specifically, we propose an efficient online framework to decompose a long-term time-varying optimization problem into a series of one-shot subproblems. To address the NP-hardness of the one-shot problem, we design a computationally-efficient two-phase algorithm based on matching and game theory, which achieves a near-optimal solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations are conducted to validate the efficacy of our proposed solution. Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009, Bo Li 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2019 | Winning at the Starting Line: Joint Network Selection and Service Placement for Mobile Edge ComputingabstractMobile Edge Computing (MEC) is an emerging computing paradigm in which computational capabilities are pushed from the central cloud to the network edges. However, preserving the satisfactory quality-of-service (QoS) for user applications is non-trivial among multiple densely dispersed yet capacity constrained MEC nodes. This is mainly because both the access network and edge nodes are vulnerable to network congestion. Previous works are mostly limited to optimizing the QoS through dynamic service placement, while ignoring the critical effects of access network selection on the network congestion. In this paper, we study the problem of jointly optimizing the access network selection and service placement for MEC, towards the goal of improving the QoS by balancing the access, switching and communication delay. Specifically, we first design an efficient online framework to decompose the long-term optimization problem into a series of one-shot problems. To address the NP-hardness of the one-shot problem, we further propose an iteration-based algorithm to derive a computation efficient solution. Both rigorous theoretical analysis on the optimality gap and extensive trace-driven simulations validate the efficacy of our proposed solution. Bin Gao 0013, Zhi Zhou 0006, Fangming Liu, Fei Xu 0009 |
INFOCOM | 1 |