Tianao Ge

dblp:322/4187 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0003-0605-0888ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Interleaved Bitstream Execution for Multi-Pattern Regex Matching on GPUs
abstract
Pattern matching is a key operation in unstructured data analytics, commonly supported by regular expression (regex) engines.Bitparallel regex engines compile regexes into bitstream programs, which expose fine-grained parallelism and are well-suited for GPU execution.A straightforward strategy executes each bitstream instruction sequentially, processing all data blocks in a loop.However, this execution suffers from poor data reuse and high memory consumption, limiting throughput.Our key insight is to adopt an interleaved execution model, where all bitstream instructions are fused into a single loop and executed block-wise.While interleaved execution could improve data reuse, enabling it on GPUs is non-trivial due to cross-block data dependencies.To address this, we introduce 1) Dependency-Aware Thread-Data Mapping, which resolves cross-block dependencies via selective recomputation.We further improve interleaved execution performance with two additional optimizations: 2) Shift Rebalancing, which balances dependency chains to reduce synchronization barriers; and 3) Zero Block Skipping, which exploits bitstream sparsity to skip computation on zero blocks.Together, these techniques make interleaved execution practical and efficient.Experiments on real-world regex benchmarks demonstrate a 19.5× geometric mean speedup over the state-of-theart GPU regex engine.
Tianao Ge, Xiaowen Chu 0001, Hongyuan Liu 0002
MICRO1
2024 ngAP: Non-blocking Large-scale Automata Processing on GPUs
abstract
Finite automata serve as compute kernels for various applications that require high throughput. However, despite the increasing compute power of GPUs, their potential in processing automata remains underutilized. In this work, we identify three major challenges that limit GPU throughput. 1) The available parallelism is insufficient, resulting in underutilized GPU threads. 2) Automata workloads involve significant redundant computations since a portion of states matches with repeated symbols. 3) The mapping between threads and states is switched dynamically, leading to poor data locality. Our key insight is that processing automata "one-symbol-at-a-time" serializes the execution, and thus needs to be revamped. To address these challenges, we propose Non-blocking Automata Processing, which allows parallel processing of different symbols in the input stream and also enables further optimizations: 1) We prefetch a portion of computations to increase the chances of processing multiple symbols simultaneously, thereby utilizing GPU threads better. 2) To reduce redundant computations, we store repeated computations in a memoization table, enabling us to substitute them with table lookups. 3) We privatize some computations to preserve the mapping between threads and states, thus improving data locality. Experimental results demonstrate that our approach outperforms the state-of-the-art GPU automata processing engine by an average of 7.9× and up to 901× across 20 applications.
Tianao Ge, Hongyuan Liu 0002
ASPLOS (1)1
2022 RAISE: Efficient GPU Resource Management via Hybrid Scheduling
abstract
As the de facto high-throughput accelerators, graphics processing units (G PU s) are now used in a wide spec-trum of fields, including artificial intelligence, high performance computing and finance. While with excessive computing and memory resources, G PU s are facing significant challenges to reach high utilization by a monolithic task. Multiple tasks are thus concurrently running to share the GPUs, but they may adversely affect each other, causing performance degradation. As a result, it is extremely critical to manage resources in a reasonable way to strike a balance between utilization and performance. Targeting the issue, this paper proposes an effective resource management design via hybrid task scheduling. Our design continuously tracks the G PU executions and collects the usage statistics, which are then used to direct the task selection and dispatch, including the type, starting time and kernel dimensions. A prototype is developed on off-the-shelf GPUs by moderately refactoring the CUDA source codes. Experimental results show that the design can achieve up to 1.96x performance improvement (1.51x on average), meanwhile effectively boosting resource utilization.
Yue Weng, Tianao Ge, Xianwei Zhang 0001, Yutong Lu
CCGRID2
2022 RollBin: reducing code-size via loop rerolling at binary level
abstract
Code size is an increasing concern on resource constrained systems, ranging from embedded devices to cloud servers. To address the issue, lowering memory occupancy has become a priority in developing and deploying applications, and accordingly compiler-based optimizations have been proposed to reduce program footprint. However, prior arts are generally dealing with source codes or intermediate representations, and thus are very limited in scope in real scenarios where only binary files are commonly provided. To fill the gap, this paper presents a novel code-size optimization RollBin to reroll loops at binary level. RollBin first locates the unrolled loops in binary files, and then probes to decide the unrolling factor by identifying regular memory address patterns. To reconstruct the iterations, we propose a customized data dependency analysis that tackles the challenges brought by shuffled instructions and loop-carry dependencies. Next, the recognized iterations are rolled up through instruction removal and update, which are generally reverting the normal unrolling procedure. The evaluations on standard SPEC2006/2017 and MiBench demonstrate that RollBin effectively shrinks code size by 1.7% and 2.2% on average (up to 7.8%), which respectively outperforms the state-of-the-arts by 31% and 38%. In addition, the use cases of representative realistic applications manifest that RollBin can be applicable in practices.
Tianao Ge, Zewei Mo, Xianwei Zhang 0001, Yutong Lu
LCTES1