VLDB 2026 Research / reviewers in the wild / expert
Chunhua Liao
dblp:44/6478
· DBLP profile ↗
22ranked-venue papers
5as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 6 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Code Pairs: Dialogue-Based Data Generation for LLM Code TranslationabstractLe Chen, Nuo Xu, Winson Chen, Bin Lei, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nuo Xu 0013, Winson Chen, Pei-Hung Lin, Dunzhi Zhou, Rajeev Thakur, Caiwen Ding, Ali Jannesari, Chunhua Liao |
ACL (1) | 10 |
| 2025 | REALM: Recursive Relevance Modeling for LLM-based Document Re-RankingabstractLarge Language Models (LLMs) have shown strong capabilities in document re-ranking, a key component in modern Information Retrieval (IR) systems.However, existing LLMbased approaches face notable limitations, including ranking uncertainty, unstable top-k recovery, and high token cost due to tokenintensive prompting.To effectively address these limitations, we propose REALM, an uncertainty-aware re-ranking framework that models LLM-derived relevance as Gaussian distributions and refines them through recursive Bayesian updates.By explicitly capturing uncertainty and minimizing redundant queries, REALM achieves better rankings more efficiently.Experimental results demonstrate that our REALM surpasses state-of-the-art rerankers while significantly reducing token usage and latency, improving NDCG@10 by 0.7 -11.9 and simultaneously reducing the number of LLM inferences by 23.4 -84.4%, promoting it as the next-generation re-ranker for modern IR systems. Pinhuan Wang, Chunhua Liao, Feiyi Wang, Hang Liu 0001 |
EMNLP | 3 |
| 2025 | Reductive Analysis with Compiler-Guided Large Language Models for Input-Centric Code OptimizationsabstractInput-centric program optimization aims to optimize code by considering the relations between program inputs and program behaviors. Despite its promise, a long-standing barrier for its adoption is the difficulty of automatically identifying critical features of complex inputs. This paper introduces a novel technique, reductive analysis through compiler-guided Large Language Models (LLMs) , to solve the problem through a synergy between compilers and LLMs. It uses a reductive approach to overcome the scalability and other limitations of LLMs in program code analysis. The solution, for the first time, automates the identification of critical input features without heavy instrumentation or profiling, cutting the time needed for input identification by 44× (or 450× for local LLMs), reduced from 9.6 hours to 13 minutes (with remote LLMs) or 77 seconds (with local LLMs) on average, making input characterization possible to be integrated into the workflow of program compilations. Optimizations on those identified input features show similar or even better results than those identified by previous profiling-based methods, leading to optimizations that yield 92.6% accuracy in selecting the appropriate adaptive OpenMP parallelization decisions, and 20-30% performance improvement of serverless computing while reducing resource usage by 50-60%. Xinning Hui, Chunhua Liao, Xipeng Shen |
Proc. ACM Program. Lang. | 3 |
| 2024 | Enhancing Performance Through Control-Flow Unmerging and Loop Unrolling on GPUsabstractCompilers use a wide range of advanced optimizations to improve the quality of the machine code they generate. In most cases, compiler optimizations rely on precise analyses to be able to perform the optimizations. However, whenever a control-flow merge is performed information is lost as it is not possible to precisely reason about the program anymore. One existing solution to this issue is code duplication, which involves duplicating instructions from merge blocks to their predecessors. This paper introduces a novel and more aggressive approach to code duplication, grounded in loop unrolling and control-flow unmerging that enables subsequent optimizations that cannot be enabled by applying only one of these transformations. We implemented our approach inside LLVM, and evaluated its performance on a collection of GPU benchmarks in CUDA. Our results demonstrate that, even when faced with branch divergence, which complicates code duplication across multiple branches and increases the associated cost, our optimization technique achieves performance improvements of up to 81%. Alnis Murtovi, Giorgis Georgakoudis, Konstantinos Parasyris, Chunhua Liao, Ignacio Laguna, Bernhard Steffen |
CGO | 4 |
| 2024 | Bilateral transformer 3D planar recoveryabstractIn recent years, deep learning based methods for single image 3D planar recovery have made significant progress, but most of the research has focused on overall plane segmentation performance rather than the accuracy of small scale plane segmentation. In order to solve the problem of feature loss in the feature extraction process of small target object features, a three dimensional planar recovery method based on bilateral transformer was proposed. The two sided network branches capture rich small object target features through different scale sampling, and are used for detecting planar and non-planar regions respectively. In addition, the loss of variational information is used to share the parameters of the bilateral network, which achieves the output consistency of the bilateral network and alleviates the problem of feature loss of small target objects. The method is verified on Scannet and Nyu V2 datasets, and a variety of evaluation indexes are superior to the current popular algorithms, proving the effectiveness of the method in three dimensional planar recovery. Chunhua Liao, Zhina Xie |
Graph. Model. | 2 |
| 2022 | Automatic inspection of program state in an uncooperative environmentabstractAbstract The program state is formed by the values that the program manipulates. These values are stored in the stack, in the heap, or in static memory. The ability to inspect the program state is useful as a debugging or as a verification aid. Yet, there exists no general technique to insert inspection points in type‐unsafe languages such as C or C++. The difficulty comes from the need to traverse the memory graph in a so‐called uncooperative environment. In this article, we propose an automatic technique to deal with this problem. We introduce a static code transformation approach that inserts in a program the instrumentation necessary to report its internal state. Our technique has been implemented in LLVM. It is possible to adjust the granularity of inspection points trading precision for performance. In this article, we demonstrate how to use inspection points to debug compiler optimizations; to augment benchmarks with verification code; and to visualize data structures. José Wesley de S. Magalhães, Chunhua Liao, Fernando Magno Quintão Pereira |
Softw. Pract. Exp. | 2 |
| 2021 | Deep NLP-based co-evolvement for synthesizing code analysis from natural languageabstractThis paper presents Deepsy, a Natural Language-based synthesizer to assist source code analysis. It takes English descriptions of to-be-found code patterns as its inputs, and automatically produces ASTMatcher expressions that are directly usable by LLVM/Clang to materialize intended code analysis. The code analysis domain features profuse complexities in data types and operations, which make it elusive for prior rule-based synthesizers to tackle. On the other hand, machine learning-based solutions are neither applicable due to the scarcity of well labeled examples. This paper presents how Deepsy addresses the challenges by leveraging deep Natural Language Processing (NLP) and creating a new technique named dependency tree-based co-evolvement.Deepsy features an effective design that seamlessly integrates Natural Language dependency analysis into code analysis and meanwhile synergizes it with type-based narrowing and domain-specific guidance. Deepsy achieves over 70.0% expression-level accuracy and 85.1% individual API-level accuracy, significantly outperforming previous solutions. Zifan Nan, Hui Guan 0001, Xipeng Shen, Chunhua Liao |
CC | 4 |
| 2020 | AutoParBench: a unified test framework for OpenMP-based parallelizersabstractThis paper describes AutoParBench, a framework to test OpenMP-based automatic parallelization tools. The core idea of this framework is a common representation, called a "JSON snapshot", that normalizes the output produced by auto-parallelizers. By converting---automatically---this output to the common representation, AutoPar-Bench lets us compare auto-parallelizers among themselves, and compare them semantically against a reference collection. Currently, this reference collection consists of 99 programs with 1,579 loops. AutoParBench produces graphic or quantitative reports that lead to fast bug discovery. By investigating differences in snapshots produced by separate sources, i.e., tool-vs-tool or tool-vs-reference, we have discovered 3 unique bugs in ICC, 2 in DawnCC, 4 in AutoPar and 2 in Cetus. These bugs have been acknowledged, and at least one of them was repaired as direct consequence of this work. Gleison Souza Diniz Mendonca, Chunhua Liao, Fernando Magno Quintão Pereira |
ICS | 2 |
| 2020 | Estimating Chlorophyll Content of Rice Based on UAV-Based Hyperspectral Imagery and Continuous Wavelet TransformabstractChlorophyll is an essential pigment for photosynthesis of crops, which indicates the growth status of crops. Accurate and robust information on the spatial dynamics of chlorophyll content is of critical importance for crop growth status assessment and corresponding response activities. However, previous studies mainly focus on the methodologies based on in situ hyperspectral data, which are not applicative for regional chlorophyll content mapping. In this content, Unmanned Aerial Vehicle (UAV) based hyperspectral imagery, with high spatial and spectral resolution, may provide the spatial distribution of chlorophyll content accurately over crop fields. In this study, an empirical model between wavelet features, derived from UAV based hyperspectral imagery by continuous wavelet transform (CWT), and chlorophyll content (measured by a portable soil-plant analysis development meter) is proposed by support vector regression (SVR). The results suggest that the UAV based hyperspectral imagery combined with CWT demonstrates good performance in rice chlorophyll content estimation with R and RMSE of 0.81 and 3.51, respectively. Moreover, the wavelet coefficients corresponding to two bands at red (640nm, 628nm) and one band at green (548nm) are the most effective wavelet features to estimate chlorophyll content of rice. Gangqiang An, Minfeng Xing, Chunhua Liao, Binbin He |
IGARSS | 3 |
| 2020 | XPlacer: Automatic Analysis of Data Access Patterns on Heterogeneous CPU/GPU SystemsabstractThis paper presents XPlacer, a framework to automatically analyze problematic data access patterns in C++ and CUDA code. XPlacer records heap memory operations in both host and device code for later analysis. To this end, XPlacer instruments read and write operations, function calls, and kernel launches. Programmers mark points in the program execution where the recorded data is analyzed and anomalies diagnosed. XPlacer reports data access anti-patterns, including alternating CPU/GPU accesses to the same memory, memory with low access density, and unnecessary data transfers. The diagnostic also produces summative information about the recorded accesses, which aids users in identifying code that could degrade performance. The paper evaluates XPlacer using LULESH, a Lawrence Livermore proxy application, Rodina benchmarks, and an implementation of the Smith-Waterman algorithm. XPlacer diagnosed several performance issues in these codes. The elimination of a performance problem in LULESH resulted in a 3x speedup on a heterogeneous platform combining Intel CPUs and Nvidia GPUs. Peter Pirkelbauer, Pei-Hung Lin, Tristan Vanderbruggen, Chunhua Liao |
IPDPS | 4 |
| 2019 | Corn Biomass Estimation Using Sentinel-2 and VENµS Data Based on A Simple Light Use Efficiency MethodabstractIn this study, Sentinel-2 and VENμS remote sensing data were combined to estimate corn biomass based on a simple light use efficiency method. The RMSE between the measured dry aboveground biomass and estimated dry aboveground biomass is 112.36 g/m2. The estimates show good agreements with the measured biomass, indicating that the effective light use efficiency (ELUE) can be effectively estimated using maximal fAPAR throughout the growing season. Chunhua Liao, Jinfei Wang, Bo Shan |
IGARSS | 1 |
| 2018 | Runtime and Memory Evaluation of Data Race Detection Tools
Pei-Hung Lin, Chunhua Liao, Markus Schordan, Ian Karlin |
ISoLA (2) | 2 |
| 2018 | Bridging the gap between deep learning and sparse matrix format selectionabstractThis work presents a systematic exploration on the promise and special challenges of deep learning for sparse matrix format selection---a problem of determining the best storage format for a matrix to maximize the performance of Sparse Matrix Vector Multiplication (SpMV). It describes how to effectively bridge the gap between deep learning and the special needs of the pillar HPC problem through a set of techniques on matrix representations, deep learning structure, and cross-architecture model migrations. The new solution cuts format selection errors by two thirds, and improves SpMV performance by 1.73X on average over the state of the art. Yue Zhao 0011, Jiajia Li 0001, Chunhua Liao, Xipeng Shen |
PPoPP | 3 |
| 2017 | POSTER: Bridging the Gap Between Deep Learning and Sparse Matrix Format SelectionabstractIn this work, we conduct a systematic exploration on the promise and challenges of deep learning for the sparse matrix format selection. We propose a set of novel techniques to solve special challenges to deep learning, including input matrix representations, a late-merging deep neural network structure design, and the use of transfer learning to alleviate cross-architecture portability issues. Yue Zhao 0011, Jiajia Li 0001, Chunhua Liao, Xipeng Shen |
PACT | 3 |
| 2017 | POSTER: An Infrastructure for HPC Knowledge Sharing and ReuseabstractThis paper presents a prototype infrastructure for addressing the barriers for effective accumulation, sharing, and reuse of the various types of knowledge for high performance parallel computing. Yue Zhao 0011, Chunhua Liao, Xipeng Shen |
PPoPP | 2 |
| 2017 | DataRaceBench: a benchmark suite for systematic evaluation of data race detection toolsabstractData races in multi-threaded parallel applications are notoriously damaging while extremely difficult to detect. Many tools have been developed to help programmers find data races. However, there is no dedicated OpenMP benchmark suite to systematically evaluate data race detection tools for their strengths and limitations. Chunhua Liao, Pei-Hung Lin, Joshua Asplund, Markus Schordan, Ian Karlin |
SC | 1 |
| 2016 | Towards Ontology-Based Program AnalysisabstractProgram analysis is fundamental for program optimizations, debugging, and many other tasks. But developing program analyses has been a challenging and error-prone process for general users. Declarative program analysis has shown the promise to dramatically improve the productivity in the development of program analyses. Current declarative program analysis is however subject to some major limitations in supporting cooperations among analysis tools, guiding program optimizations, and often requires much effort for repeated program preprocessing. In this work, we advocate the integration of ontology into declarative program analysis. As a way to standardize the definitions of concepts in a domain and the representation of the knowledge in the domain, ontology offers a promising way to address the limitations of current declarative program analysis. We develop a prototype framework named PATO for conducting program analysis upon ontology-based program representation. Experiments on six program analyses confirm the potential of ontology for complementing existing declarative program analysis. It supports multiple analyses without separate program preprocessing, promotes cooperative Liveness analysis between two compilers, and effectively guides a data placement optimization for Graphic Processing Units (GPU). Yue Zhao 0011, Guoyang Chen, Chunhua Liao, Xipeng Shen |
ECOOP | 3 |
| 2016 | Evaluation of spatio-temporal data fusion methods for generating NDVI time series in cropland areasabstractSpatio-temporal data fusion model is a feasible way to obtain high spatial resolution and high temporal resolution images in crop monitoring. As vegetation indices such as Normalized Difference Vegetation Index (NDVI) are generally used directly to monitor the vegetation growth, in this study, two recently proposed spatio-temporal data fusion methods (FSDAF and DPM-STVIFM) were evaluated for generating NDVI time series in cropland areas. It is found that both methods have limitations and the performances of the two methods vary with the dates of available fine-resolution images and the degree of land cover changes between the available fine-resolution images and the synthetic fine-resolution images. Chunhua Liao, Jinfei Wang |
IGARSS | 1 |
| 2013 | Symbolic Analysis of Concurrency Errors in OpenMP ProgramsabstractIn this paper we present the OpenMP Analysis Toolkit (OAT), which uses Satisfiability Modulo Theories (SMT) solver based symbolic analysis to detect data races and deadlocks in OpenMP codes. Our approach approximately simulates real executions of an OpenMP program through schedule permutation. We conducted experiments on real-world OpenMP benchmarks and student homework assignments by comparing our OAT tool with two commercial dynamic analysis tools: Intel Thread Checker and Sun Thread Analyzer, and one commercial static analysis tool: Viva64 PVS Studio. The experiments show that our symbolic analysis approach is more accurate than static analysis and more efficient and scalable than dynamic analysis tools with less false positives and negatives. Hongyi Ma, Steve Diersen, Chunhua Liao, Daniel J. Quinlan, Zijiang Yang 0006 |
ICPP | 4 |
| 2013 | An IconMap-based exploratory analytical approach for multivariate geospatial data
Xianfeng Zhang, Chunhua Liao, Jonathan Li 0001 |
Sci. China Inf. Sci. | 2 |
| 2007 | Invited Paper: A Compile-time Cost Model for OpenMPabstractOpenMP has gained wide popularity as an API for parallel programming on shared memory and distributed shared memory platforms. It is also a promising candidate to exploit the emerging multicore, multithreaded processors. In addition, there is an increasing trend to combine OpenMP with MPI to take full advantage of mainstream supercomputers consisting of clustered SMPs. All of these require that attention be paid to the quality of the compiler's translation of OpenMP and the flexibility of runtime support. Many compilers and runtime libraries have an internal cost model that helps evaluate compiler transformations, guides adaptive runtime systems, and helps achieve load balancing. But existing models are not sufficient to support OpenMP, especially on new platforms. In this paper we present our experience adapting the cost models in OpenUH, a branch of Open64, to estimate the execution cycles of parallel OpenMP regions using knowledge of both software and hardware. Our OpenMP cost model reuses major components from Open64, along with extensions to consider more OpenMP details. Preliminary evaluations of the model are presented using kernel benchmarks. The challenges and possible extensions for modeling OpenMP on multicore platforms are also discussed. Chunhua Liao, Barbara M. Chapman |
IPDPS | 1 |
| 2007 | OpenUH: an optimizing, portable OpenMP compilerabstractAbstract OpenMP has gained wide popularity as an API for parallel programming on shared memory and distributed shared memory platforms. Despite its broad availability, there remains a need for a portable, robust, open source, optimizing OpenMP compiler for C/C++/Fortran 90, especially for teaching and research, for example into its use on new target architectures, such as SMPs with chip multi‐threading, as well as learning how to translate for clusters of SMPs. In this paper, we present our efforts to design and implement such an OpenMP compiler on top of Open64, an open source compiler framework, by extending its existing analysis and optimization and adopting a source‐to‐source translator approach where a native back end is not available. The compilation strategy we have adopted and the corresponding runtime support are described. The OpenMP validation suite is used to determine the correctness of the translation. The compiler's behavior is evaluated using benchmark tests from the EPCC microbenchmarks and the NAS parallel benchmark. Copyright © 2007 John Wiley & Sons, Ltd. Chunhua Liao, Oscar R. Hernandez, Barbara M. Chapman |
Concurr. Comput. Pract. Exp. | 1 |