VLDB 2026 Research / reviewers in the wild / expert
Luanzheng Guo
dblp:200/8189 · also Lenny Guo
· DBLP profile ↗
26ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0001-8266-0923ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Order-Preserving Dimension Reduction for Multimodal Semantic EmbeddingabstractSearching for the k-nearest neighbors in multimodal data retrieval is computationally expensive, particularly due to the inherent difficulty in comparing similarity measures across different modalities. Recent advances in multimodal machine learning address this issue by mapping data into a shared embedding space; however, the high dimensionality of these embeddings (hundreds to thousands of dimensions) presents a challenge for time-sensitive vision applications. This work proposes Order-Preserving Dimension Reduction (OPDR), aiming to reduce the dimensionality of embeddings while preserving the ranking of KNN in the lower-dimensional space. One notable component of OPDR is a new measure function to quantify KNN quality as a global metric, based on which we derive a closed-form map between target dimensionality and key contextual parameters. We have integrated OPDR with multiple state-of-the-art dimension-reduction techniques, distance functions, and embedding models; experiments on a variety of multimodal datasets demonstrate that OPDR effectively retains recall high accuracy while significantly reducing computational costs. Chengyu Gong, Gefei Shen, Luanzheng Guo, Nathan R. Tallent, Dongfang Zhao 0001 |
AAAI | 3 |
| 2026 | ADVICE: Automatic Identification of Variables to Checkpoint Through Compiler Augmentation
Luanzheng Guo, Nathan R. Tallent, Kento Sato |
CCGrid | 2 |
| 2026 | PowerMorph: Shaping LLM Training for Data Center Demand Response
Boqiang Li, Luanzheng Guo, Buxin She, Nathan R. Tallent, Veronica Adetola, Rong Ge 0002 |
IPDPS | 2 |
| 2026 | QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
Md. Hasanur Rashid, Jesun Sahariar Firoz, Nathan R. Tallent, Luanzheng Guo, Dong Dai 0001 |
IPDPS | 4 |
| 2026 | Characterizing Dataflow for I/O-Aware Scheduling in HPC Workflows
Luanzheng Guo, Antonios Kougkas, Xian-He Sun, Nathan R. Tallent |
IPDPS | 2 |
| 2026 | Accelerating AI Compression through Lightweight Lossless Encoding and Pipelined Workflows
Boyuan Zhang 0002, Luanzheng Guo, Jiannan Tian, Jinyang Liu 0003, Daoce Wang, Chengming Zhang 0006, Bo Fang 0002, Fengguang Song, Jan Strube 0001, Nathan R. Tallent, Dingwen Tao |
IPDPS | 2 |
| 2025 | PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML TrainingabstractThe exponential growth of large-scale AI models has led to computational and power demands that can exceed the capacity of a single data center. This is due to the limited power supplied by regional grids that leads to limited regional computational power. Consequently, distributing training workloads across geographically distributed sites has become essential. However, this approach introduces a significant challenge in the form of communication overhead, creating a fundamental trade-off between the performance gains from accessing greater aggregate power and the performance losses from increased network latency. Although prior work has focused on reducing communication volume or using heuristics for distribution, these methods assume constant homogeneous power supplies and ignore the challenge of heterogeneous power availability between sites. Talha Mehboob, Luanzheng Guo, Nathan R. Tallent, Michael Zink, David Irwin 0001 |
SoCC | 2 |
| 2025 | ProHD: Projection-Based Hausdorff Distance ApproximationabstractThe Hausdorff distance (HD) is a robust measure of set dissimilarity, but computing it exactly on large, high-dimensional datasets is prohibitively expensive. We propose ProHD, a projection-guided approximation algorithm that dramatically accelerates HD computation while maintaining high accuracy. ProHD identifies a small subset of candidate “extreme” points by projecting the data onto a few informative directions (such as the centroid axis and top principal components) and computing the HD on this subset. This approach guarantees an underestimate of the true HD with a bounded additive error and typically achieves results within a few percent of the exact value. In extensive experiments on image, physics, and synthetic datasets (up to two million points in D = 256), ProHD runs 10-100× faster than exact algorithms while attaining 5-20× lower error than random sampling-based approximations. Our method enables practical HD calculations in scenarios like large vector databases and streaming data, where quick and reliable set distance estimation is needed. Jiuzhou Fu, Luanzheng Guo, Nathan R. Tallent, Dongfang Zhao 0001 |
ICDM | 2 |
| 2025 | BMQSim: Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
Boyuan Zhang 0002, Bo Fang 0002, Fanjiang Ye, Luanzheng Guo, Fengguang Song, Nathan R. Tallent, Dingwen Tao |
ICS | 4 |
| 2025 | FlowForecaster: Automatically Inferring Detailed & Interpretable Workflow Scaling Models for ForecastsabstractDistributed scientific workflows underpin many areas of scientific exploration. To enable good scheduling decisions, we introduce a novel method for predicting their expected task dependences and data flow when scaling data sizes and task parallelism. Most workflows, following the 80-20% rule, execute in predictable patterns relative to concurrency and input data sizes. We develop FlowForecaster, an efficient method for automatically inferring detailed and interpretable workflow scaling models from a few empirical task property graphs (3–5). Our model is an abstract directed acyclic graph (DAG) with analytical expressions to describe how the DAG scales and how data flows along edges. Importantly, our expression language and rules can explain data dependent structure and flow. Our model inference finds repeated substructure, infers analytical rules to explain substructure scaling (edge branching and joining), and predicts edge properties such as data accesses, access size, and data volume. From the model, we can predict entire DAG substructures. We validate FlowForecaster on several workflows and find that we can use interpretable rules to explain 97% of observed results on task and data scaling. Hyungro Lee, Jesun Sahariar Firoz, Nathan R. Tallent, Luanzheng Guo, Mahantesh Halappanavar |
IPDPS | 4 |
| 2025 | High-performance Visual Semantics Compression for AI-Driven ScienceabstractScientific images play a crucial role in many experimental sciences; however, the large volumes of data generated present significant challenges. Effective image compression must be fast, achieve high compression ratios, and preserve critical domain-specific features. Existing compressors, such as JPEG and SZ, often distort important textures when operating at high compression ratios. Conversely, AI-based compressors offer superior image quality and higher compression ratios but are significantly slower than traditional methods. To address this trade-off, we developed ViSemZ, a high-performance AI-based compressor specifically designed to preserve visual semantics. Our approach enhances AI compression by integrating sparse encoding with variable-length integer truncation, optimized lossless encoding using bitshuffle and a decoupled lookback prefix-sum, and pipelining techniques to enable efficient data streaming and asynchronous processing. Evaluations on scientific datasets demonstrate that, at comparable compression ratios, ViSemZ achieves performance almost on par with existing AI-based compressors while delivering a 9.6× overall compression speedup. These results effectively bridge the performance gap between traditional and AI-based compression methods. Boyuan Zhang 0002, Luanzheng Guo, Jiannan Tian, Jinyang Liu 0003, Daoce Wang, Fanjiang Ye, Chengming Zhang 0006, Jan Strube 0001, Nathan R. Tallent, Dingwen Tao |
PPoPP | 2 |
| 2025 | FastFlow: Rapid Workflow Response By Prioritizing Critical Data Flows and their Interactions
Jesun Sahariar Firoz, Hyungro Lee, Luanzheng Guo, Nathan R. Tallent |
SSDBM | 3 |
| 2024 | Identifying Outliers in AI-based Image CompressionabstractImage compression using artificial intelligence (AI) is gaining importance in scientific research, where instruments and simulations can produce hundreds of images per second. Effective compression with high ratios is essential for facilitating discoveries. A key challenge is the automatic detection of outliers—cases where compression fails or significant phenomena are present. To address this, we developed a consensus-driven methodology using unsupervised machine learning techniques for identifying outlier compressed images. We evaluated our approach on unlabeled datasets, including microscopy and X-ray images, successfully identifying multiple outliers using metrics such as peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), structural texture similarity index measure (STSIM) and deep image and structural texture similarity index (DISTS). Rizwan A. Ashraf, Luanzheng Guo, Hyungro Lee, Nathan R. Tallent |
IEEE Big Data | 2 |
| 2024 | Improving I/O-aware Workflow Scheduling via Data Flow Characterization and trade-off AnalysisabstractThe scientific computing paradigm has transitioned from compute-intensive to I/O-intensive and memory-intensive in the past decade, especially when data-driven science has become common practice. Numerous empirical I/O-aware scheduling optimizations have been developed by incorporating I/O capacity and bandwidth as constraints into scheduling. Unfortunately, there is a lack of data flow (I/O) characterization tools and an understanding of trade-offs between concurrency, locality, and I/O bandwidth. To bridge the gap, this work 1) presents a set of descriptors to characterize, organize, and visualize I/O profiles, including flow size, I/O bandwidth, and operation count, which group data flows by I/O types, tasks, and files; 2) proposes an I/O Roofline model-based trade-off analysis to find the optimal trade-off between flow operational intensity, concurrency, and flow performance. The I/O descriptors generate useful insights into complicated I/O behaviors, suggesting distinct concurrency, storage, and scheduling to be used by types, tasks, and files. The proposed trade-off analysis guides scheduling decisions that generate resource assignment with the best flow parallelism. We evaluate our I/O-aware scheduling methodology on a highly I/O-intensive workflow–1000 Genomes. The experimental results demonstrate speedups of up to 2.4× compared to the state-of-the-art methods. Luanzheng Guo, Hyungro Lee, Jesun Sahariar Firoz, Nathan R. Tallent |
IEEE Big Data | 1 |
| 2024 | Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi - Threaded ProgramsabstractAfter all these years and all these other shared memory programming frameworks, OpenMP is still the most popular one. However, its greater levels of non-deterministic execution makes debugging and testing more challenging. The ability to record and deterministically replay the program execution is key to address this challenge. However, scalably replaying OpenMP programs is still an unresolved problem. In this paper, we propose two novel techniques that use Distributed Clock (DC) and Distributed Epoch (DE) recording schemes to eliminate excessive thread synchronization for OpenMP record and replay. Our evaluation on representative HPC applications with ReOMP, which we used to realize DC and DE recording, shows that our approach is 2-5x more efficient than traditional approaches that synchronize on every shared-memory access. Furthermore, we demonstrate that our approach can be easily combined with MPI-Ievel replay tools to replay non-trivial MPI+OpenMP applications. We achieve this by integrating ReOMP into ReMPI, an existing scalable MPI record-and-replay tool, with only a small MPI-scale-independent runtime overhead. Shiman Meng, Luanzheng Guo, Kento Sato, Dong H. Ahn, Ignacio Laguna, Gregory L. Lee, Martin Schulz 0001 |
CLUSTER | 4 |
| 2024 | DaYu: Optimizing Distributed Scientific Workflows by Decoding Dataflow Semantics and DynamicsabstractThe combination of ever-growing scientific datasets and distributed workflow complexity creates I/O performance bottlenecks due to data volume, velocity, and variety. Although the increasing use of descriptive data formats (e.g., HDF5, netCDF) helps organize these datasets, it also introduces obscure bottlenecks due to the need to translate high-level operations into file addresses and then into low-level I/O operations. To address this challenge, we introduce DaYu, a method and toolset for analyzing (a) semantic relationships between logical datasets and file addresses, (b) how dataset operations translate into I/O, and (c) the combination across entire workflows. DaYu's analysis and visualization enable the identification of critical bottlenecks and the reasoning about remediation. We describe our methodology and propose optimization guidelines. Evaluation on scientific workflows demonstrates up to a 3.7x performance improvement in I/O time for obscure bottlenecks. The time and storage overhead for DaYu's time-ordered data are typically under 0.2% of runtime and 0.25% of data volume, respectively. Jaime Cernuda, Luanzheng Guo, Nathan R. Tallent, Antonios Kougkas, Xian-He Sun |
CLUSTER | 4 |
| 2024 | AutoCheck: Automatically Identifying Variables for Checkpointing by Data Dependency AnalysisabstractCheckpoint/Restart (C/R) has been widely deployed in numerous HPC systems, Clouds, and industrial data centers, which are typically operated by system engineers. Nevertheless, there is no existing approach that helps system engineers without domain expertise, and domain scientists without system fault tolerance knowledge identify those critical variables accounted for correct application execution restoration in a failure for C/R. To address this problem, we propose an analytical model and a tool (AutoCheck) that can automatically identify critical variables to checkpoint for C/R. AutoCheck relies on first, analytically tracking and optimizing data dependency between variables and other application execution state, and second, a set of heuristics that identify critical variables for checkpointing from the refined data dependency graph (DDG). AutoCheck allows programmers to pinpoint critical variables to checkpoint quickly within a few minutes. We evaluate AutoCheck on 14 representative HPC benchmarks, demonstrating that AutoCheck can efficiently identify correct critical variables to checkpoint. Shiman Meng, Wubiao Xu, Luanzheng Guo, Kento Sato |
SC | 6 |
| 2024 | A Visual Comparison of Silent Error PropagationabstractHigh-performance computing (HPC) systems play a critical role in facilitating scientific discoveries. Their scale and complexity (e.g., the number of computational units and software stack) continue to grow as new systems are expected to process increasingly more data and reduce computing time. However, with more processing elements, the probability that these systems will experience a random bit-flip error that corrupts a program's output also increases, which is often recognized as silent data corruption. Analyzing the resiliency of HPC applications in extreme-scale computing to silent data corruption is crucial but difficult. An HPC application often contains a large number of computation units that need to be tested, and error propagation caused by error corruption is complex and difficult to interpret. To accommodate this challenge, we propose an interactive visualization system that helps HPC researchers understand the resiliency of HPC applications and compare their error propagation. Our system models an application's error propagation to study a program's resiliency by constructing and visualizing its fault tolerance boundary. Coordinating with multiple interactive designs, our system enables domain experts to efficiently explore the complicated spatial and temporal correlation between error propagations. At the end, the system integrated a nonmonotonic error propagation analysis with an adjustable graph propagation visualization to help domain experts examine the details of error propagation and answer such questions as why an error is mitigated or amplified by program execution. Harshitha Menon, Kathryn Mohror, Shusen Liu 0001, Luanzheng Guo, Peer-Timo Bremer, Valerio Pascucci |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Automatic Code Generation for High-Performance Graph AlgorithmsabstractGraph problems are common across many fields, from scientific computing to social sciences. Despite their importance and the attention received, implementing graph algorithms effectively on modern computing systems remains a challenging task that requires significant programming effort and generally results in customized implementations. Current computing and memory hierarchies are not architected for irregular computations, resulting performance that is far from the theoretical architectural peak. In this paper, we propose a compiler framework to simplify the development of graph algorihtm implementations that can achieve high performance on modern computing systems. We provide a high-level domain specific language (DSL) to represent graph algorithms through sparse linear algebra expressions and graph primitives including semiring and masking. The compiler leverages the semantics information expressed through the DSL during the optimization and code transformation passes, resulting in more efficient IR passed to the compiler backend. In particular, we introduce an Index Tree Dialect that preserves the semantic information of the graph algorithm to perform high-level, domain-specific optimizations, including workspace transformation, two-phase computation, and automatic parallelization. We demonstrate that this work outperforms state-of-the-art graph libraries LAGraph by up to 3.7 × speedup in semiring operations, 2.19 ×speedup in an important sparse computational kernel, and 9.05 × speedup in graph processing algorithms. Rizwan A. Ashraf, Luanzheng Guo, Ruiqin Tian, Gokcen Kestor |
PACT | 3 |
| 2023 | Im2win: An Efficient Convolution Paradigm on GPU
Luanzheng Guo, Xu T. Liu |
Euro-Par | 3 |
| 2023 | Data Flow Lifecycles for Optimizing Workflow CoordinationabstractA critical performance challenge in distributed scientific workflows is coordinating tasks and data flows on distributed resources. To guide these decisions, this paper introduces data flow lifecycle analysis. Workflows are commonly represented using directed acyclic graphs (DAGs). Data flow lifecycles (DFL) enrich task DAGs with data objects and properties that describe data flow and how tasks interact with that flow. Lifecycles enable analysis from several important perspectives: task, data, and data flow. We describe representation, measurement, analysis, visualization, and opportunity identification for DFLs. Our measurement is both distributed and scalable, using space that is constant per data file. We use lifecycles and opportunity analysis to reason about improved task placement and reduced data movement for five scientific workflows with different characteristics. Case studies show improvements of 15×, 1.9×, and 10--30×. Our work is implemented in the DataLife tool. Hyungro Lee, Luanzheng Guo, Jesun Sahariar Firoz, Nathan R. Tallent, Antonios Kougkas, Xian-He Sun |
SC | 2 |
| 2022 | Towards Supporting Semiring in MLIR-Based COMET CompilerabstractSemirings are widely used in large-scale scientific applications of high-dimensional data and graph analytics for linear algebra computations. In this work, we propose a semiring compiler for today's high-performance computing (HPC) systems, often armed with heterogeneous devices, as an alternative to library-based approaches. In particular, we extend a domain-specific language (DSL) and compiler framework to automatically generate kernel code for semiring operations within the COMpiler for Extreme Targets (COMET) based on the Multi-Level Intermediate Representation (MLIR) framework. We provide a high-level programming abstraction representing various semiring operations with the familiar Einstein notation. We also build a semiring dialect and efficient code generation based on MLIR's extensible framework that can process a variety of semiring operators. By leveraging high-level semantics information and progressive lowering in code generation, we achieved better performance with up to 3.8x speedup compared with operations in the LAGraph library. Luanzheng Guo, Rizwan A. Ashraf, Ryan D. Friese, Gokcen Kestor |
PACT | 1 |
| 2021 | PARIS: Predicting application resilience using machine learning
Luanzheng Guo, Dong Li 0001, Ignacio Laguna |
J. Parallel Distributed Comput. | 1 |
| 2019 | MOARD: Modeling Application Resilience to Transient Faults on Data ObjectsabstractUnderstanding application resilience (or error tolerance) in the presence of hardware transient faults on data objects is critical to ensure computing integrity and enable efficient application-level fault tolerance mechanisms. However, we lack a method and a tool to quantify application resilience to transient faults on data objects. The traditional method, random fault injection, cannot help, because of losing data semantics and insufficient information on how and where errors are tolerated. In this paper, we introduce a method and a tool (called “MOARD”) to model and quantify application resilience to transient faults on data objects. Our method is based on systematically quantifying error masking events caused by application-inherent semantics and program constructs. We use MOARD to study how and why errors in data objects can be tolerated by the application. We demonstrate tangible benefits of using MOARD to direct a fault tolerance mechanism to protect data objects. Luanzheng Guo, Dong Li 0001 |
IPDPS | 1 |
| 2018 | FlipTracker: understanding natural error resilience in HPC applications
Luanzheng Guo, Dong Li 0001, Ignacio Laguna, Martin Schulz 0001 |
SC | 1 |
| 2013 | Indoor frame recovering via line segments refinement and votingabstractFrame structure estimation from line segments is an important yet challenging problem in understanding indoor scenes. In practice, line segment extraction can be affected by occlusions, illumination variations, and weak object boundaries. To address this problem, an approach for frame structure recovery based on line segment refinement and voting is proposed. We refined line segments by the revising, connecting, and adding operations. We then propose an iterative voting mechanism for selecting refined line segments, where a cross ratio constraint is enforced to build crab-like models. Our algorithm outperforms state-of-the-art approaches, especially when considering complex indoor scenes. Luanzheng Guo, Lingfeng Wang 0002, Chunhong Pan, Shiming Xiang |
ICASSP | 2 |