Zhihan Guo

dblp:241/5293 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MisEdu-RAG: A Misconception-Aware Dual-Hypergraph RAG for Novice Math Teachers
Zhihan Guo, Rundong Xue, Jionghao Lin
AIED1
2026 Simulating Novice Students Using Machine Unlearning and Relearning in Large Language Models
Zhihan Guo, Jionghao Lin
AIED (1)2
2026 Exploring Creator-Centric Methods for LLM-Assisted Interactive Storytelling
Yuelu Li, Lujin Zhang, Zhihan Guo, Wenchuan Lu, David Kei-Man Yip
CHI4
2025 From General Reward to Targeted Reward: Improving Open-ended Long-context Generation Models
abstract
Current research on long-form context in Large Language Models (LLMs) primarily focuses on the understanding of long-contexts, the Openended Long Text Generation (Open-LTG) remains insufficiently explored.Training a longcontext generation model requires curation of gold-standard reference data, which is typically nonexistent for informative Open-LTG tasks.However, previous methods only utilize general assessments as reward signals, which limits accuracy.To bridge this gap, we introduce ProxyReward, an innovative reinforcement learning (RL) based framework, which includes a dataset and a reward signal computation method.Firstly, ProxyReward Dataset generation is accomplished through simple prompts that enables the model to create automatically, obviating extensive labeled data or significant manual effort.Secondly, ProxyReward Signal offers a targeted evaluation of information comprehensiveness and accuracy for specific questions.The experimental results indicate that our method Prox-yReward surpasses even GPT-4-Turbo.It can significantly enhance performance by 20% on the Open-LTG task when training widely used open-source models, while also surpassing the LLM-as-a-Judge approach.Our work presents effective methods to enhance the ability of LLMs to address complex open-ended questions posed by humans.
Zhihan Guo, Jiele Wu, Wenqian Cui, Minda Hu, Yufei Wang 0005, Irwin King
EMNLP1
2022 How Good is My HTAP System?
abstract
Hybrid Transactional and Analytical Processing (HTAP) systems have recently gained popularity as they combine OLAP and OLTP processing to reduce administrative and synchronization costs between dedicated systems. However, there is no precise characterization of the features that distinguish a good HTAP system from a poor one. In this paper, we seek to solve this problem from the perspectives of both performance and freshness. To simultaneously capture the performance of both transactional and analytical processing, we introduce a new concept called throughput frontier, which visualizes both transactional and analytical throughput in a single 2D graph. The throughput frontier can capture information regarding the performance of each engine, the interference between the two engines, and various system design decisions. To capture how well an HTAP system supports real-time analytics, we define a freshness metric which quantifies how recent is the snapshot of the data seen by each analytical query. We also develop a practical way to measure freshness in a real system. We design a new hybrid benchmark called HATtrick which incorporates both throughput frontier and freshness as metrics. Using the benchmark, we evaluate three representative HTAP systems under various data size and system configurations and demonstrate how the metrics reveal important system characteristics and performance information.
Elena Milkai, Yannis Chronis, Kevin P. Gaffney, Zhihan Guo, Jignesh M. Patel, Xiangyao Yu
SIGMOD Conference4
2022 Cornus: Atomic Commit for a Cloud DBMS with Storage Disaggregation
abstract
Two-phase commit (2PC) is widely used in distributed databases to ensure atomicity of distributed transactions. Conventional 2PC was originally designed for the shared-nothing architecture and has two limitations: long latency due to two eager log writes on the critical path, and blocking of progress when a coordinator fails. Modern cloud-native databases are moving to a storage disaggregation architecture where storage is a shared highly-available service. Our key observation is that disaggregated storage enables protocol innovations that can address both the long-latency and blocking problems. We develop Cornus, an optimized 2PC protocol to achieve this goal. The only extra functionality Cornus requires is an atomic compare-and-swap capability in the storage layer, which many existing storage services already support. We present Cornus in detail and show how it addresses the two limitations. We also deploy it on real storage services including Azure Blob Storage and Redis. Empirical evaluations show that Cornus can achieve up to 1.9X latency reduction over conventional 2PC.
Zhihan Guo, Xinyu Zeng, Wuh-Chwen Hwang, Ziwei Ren, Xiangyao Yu, Mahesh Balakrishnan 0001, Philip A. Bernstein
Proc. VLDB Endow.1
2021 The Storage Hierarchy is Not a Hierarchy: Optimizing Caching on Modern Storage Devices with Orthus
Zhihan Guo, Guanzhou Hu, Kaiwei Tu, Ramnatthan Alagappan, Rathijit Sen, Kwanghyun Park 0001, Andrea C. Arpaci-Dusseau, Remzi H. Arpaci-Dusseau
FAST2
2021 Releasing Locks As Early As You Can: Reducing Contention of Hotspots by Violating Two-Phase Locking
abstract
Hotspots, a small set of tuples frequently read/written by a large number of transactions, cause contention in a concurrency control protocol. While a hotspot may comprise only a small fraction of a transaction's execution time, conventional strict two-phase locking allows a transaction to release lock only after the transaction completes, which leaves significant parallelism unexploited. Ideally, a concurrency control protocol serializes transactions only for the duration of the hotspots, rather than the duration of transactions.
Zhihan Guo, Cong Yan, Xiangyao Yu
SIGMOD Conference1
2020 A Statistical Perspective on Discovering Functional Dependencies in Noisy Data
abstract
We study the problem of discovering functional dependencies (FD) from a noisy data set. We adopt a statistical perspective and draw connections between FD discovery and structure learning in probabilistic graphical models. We show that discovering FDs from a noisy data set is equivalent to learning the structure of a model over binary random variables, where each random variable corresponds to a functional of the data set attributes. We build upon this observation to introduce FDX a conceptually simple framework in which learning functional dependencies corresponds to solving a sparse regression problem. We show that FDX can recover true functional dependencies across a diverse array of real-world and synthetic data sets, even in the presence of noisy or missing data. We find that FDX scales to large data instances with millions of tuples and hundreds of attributes while it yields an average F1 improvement of 2x against state-of-the-art FD discovery methods.
Zhihan Guo, Theodoros Rekatsinas
SIGMOD Conference2