Gansen Hu

dblp:259/0380 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0006-0468-3744ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Indexing and storage engines · 55% Query processing and optimization · 46%
Software engineering, system software, and programming languages
3 papers
Program verification · 52% Program synthesis and code generation · 34% Compilers and program optimization · 14%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 12 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
model compression
1.012026
LLMFolder: Revisiting Constant Folding in Large Language Models · EuroSys 2026
Query processing and optimization › query execution
batch query processing
0.812024
WeBridge: Synthesizing Stored Procedures for Large-Scale Real-World Web Applications · Proc. ACM Manag. Data 2024
Query processing and optimization
query rewriting
0.612022
WeTune: Automatic Discovery and Verification of Query Rewrite Rules · SIGMOD Conference 2022
Program verification
equivalence checking
0.612022
WeTune: Automatic Discovery and Verification of Query Rewrite Rules · SIGMOD Conference 2022
Program verification
SMT-based verification
0.612022
WeTune: Automatic Discovery and Verification of Query Rewrite Rules · SIGMOD Conference 2022
Indexing and storage engines
concurrent index
0.412020
XIndex: a scalable learned index for multicore data storage · PPoPP 2020
Indexing and storage engines › learned index
concurrent learned index
0.412020
XIndex: a scalable learned index for multicore data storage · PPoPP 2020
Indexing and storage engines
learned index
0.412020
XIndex: a scalable learned index for multicore data storage · PPoPP 2020
Indexing and storage engines › access methods
ordered index
0.412020
XIndex: a scalable learned index for multicore data storage · PPoPP 2020
Compilers and program optimization › compiler optimization › local optimization
constant folding
0.312026
LLMFolder: Revisiting Constant Folding in Large Language Models · EuroSys 2026
Query processing and optimization
database access optimization
0.212024
WeBridge: Synthesizing Stored Procedures for Large-Scale Real-World Web Applications · Proc. ACM Manag. Data 2024
Indexing and storage engines
key-value store
0.112020
XIndex: a scalable learned index for multicore data storage · PPoPP 2020

Methods — techniques the papers use, named apart from their topics

activation function linearization · 2.0speculative execution · 1.5program analysis · 1.5concolic execution · 1.5superoptimization · 1.1constraint enumeration · 1.1SMT solving · 1.1two-phase compaction · 0.4learned model · 0.4fine-grained synchronization · 0.4
YearPublicationVenuePosition
2026 LLMFolder: Revisiting Constant Folding in Large Language Models
abstract
Large language models (LLMs) demonstrate remarkable capabilities but face deployment challenges due to their massive parameter counts. While pruning can reduce model size, it leads to significant accuracy degradation under high compression ratios. We present a novel perspective inspired by constant folding in compiler optimization. Our approach enables parameter reduction by treating activation functions in LLMs as linear functions.
Gansen Hu, Jinglin Wei, Haibo Chen 0001
EuroSys1
2026 A study and formal framework of the composability of LLM compression techniques
Gansen Hu
Frontiers Comput. Sci.1
2024 WeBridge: Synthesizing Stored Procedures for Large-Scale Real-World Web Applications
abstract
Modern web applications use databases to store their data. When processing user requests, these applications retrieve and store data in the database server, which incurs network round trips. These round trips significantly increase the application's latency. Previous approaches have attempted to reduce these round trips by prefetching query results or batching database accesses. However, neither method can efficiently reduce the latency when some queries depend on previous queries' results. In real-world applications, nearly 50% of the queries depend on the result of other queries. This paper presents WeBridge, the first system capable of synthesizing stored procedures for large-scale real-world web applications. First, WeBridge employs concolic execution technique to analyze the applications and generate stored procedures for hot program paths. Then, it seamlessly integrates the stored procedures into the application by extending the database access library. Finally, it improves the efficiency of the stored procedures with speculative execution. Evaluation using real-world web applications and workloads show that WeBridge achieves up to 79.8% median latency reduction and up to 2× peak throughput.
Gansen Hu, Chuzhe Tang, Jiahuan Shen, Zhiyuan Dong, Sheng Yao 0006, Haibo Chen 0001
Proc. ACM Manag. Data1
2022 WeTune: Automatic Discovery and Verification of Query Rewrite Rules
abstract
Query rewriting transforms a relational database query into an equivalent but more efficient one, which is crucial for the performance of database-backed applications. Such rewriting relies on pre-specified rewrite rules. In existing systems, these rewrite rules are discovered through manual insights and accumulate slowly over the years. In this paper, we present WeTune, a rule generator that automatically discovers new rewrite rules. Inspired by compiler superoptimization, WeTune enumerates all valid logical query plans up to a certain size and tries to discover equivalent plans that could potentially lead to more efficient rewrites. The core challenge is to determine which set of conditions (aka constraints) allows one to prove the equivalence between a pair of query plans. We address this challenge by enumerating combinations of "interesting" constraints that relate tables and their attributes between each pair of queries. We also propose a new SMT-based verifier to verify the equivalence of a query pair under different enumerated constraints. To evaluate the usefulness of rewrite rules discovered by WeTune, we apply them on the SQL queries collected from the 20 most popular open-source web applications on GitHub. WeTune successfully optimizes 247 queries that existing databases cannot optimize, resulting in substantial performance improvements.
Yicun Yang, Gansen Hu, Chuzhe Tang, Haibo Chen 0001, Jinyang Li 0001
SIGMOD Conference5
2020 XIndex: a scalable learned index for multicore data storage
abstract
We present XIndex, a concurrent ordered index designed for fast queries. Similar to a recent proposal of the learned index, XIndex uses learned models to optimize index efficiency. Comparing with the learned index, XIndex is able to effectively handle concurrent writes without affecting the query performance by leveraging fine-grained synchronization and a new compaction scheme, Two-Phase Compaction. Furthermore, XIndex adapts its structure according to run-time workload characteristics to support dynamic workload. We demonstrate the advantages of XIndex with both YCSB and TPC-C (KV), a TPC-C variant for key-value stores. XIndex achieves up to 3.2X and 4.4X performance improvement comparing with Masstree and Wormhole, respectively, on a 24-core machine, and it is open-sourced1.
Chuzhe Tang, Youyun Wang, Zhiyuan Dong, Gansen Hu, Haibo Chen 0001
PPoPP4