EDBT 2026 Demo / reviewers in the wild / expert
Zhe Wang 0059
dblp:75/3158-59
· DBLP profile ↗
15ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0003-1123-9925ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Gaussian Mixture Model-Based Splatting for Rapid Rendering and Time Series Analysis of Large-Scale Particle DataabstractDriven by advances in supercomputing, the scale of scientific simulation data has grown dramatically. In fields such as cosmology, particle data have become a common representation, with state-of-the-art simulations now exceeding the trillion-particle mark. Consequently, the challenge of visually analyzing such massive datasets has become increasingly urgent. The traditional visual analysis workflow typically follows a "compression $\rightarrow$→ storage $\rightarrow$→ reconstruction $\rightarrow$→ visualization" pipeline. However, this process is hampered by an extremely time-consuming reconstruction stage, which severely impedes real-time interactive visualization. Moreover, in multi-time-step analyses, the enormous volume of reconstructed data creates significant I/O bottlenecks. In this work, we draw inspiration from 3D Gaussian splatting and compress the simulation data using Gaussian Mixture Models (GMMs), treating the resulting Gaussian kernels as fundamental rendering primitives. Our method renders billion-scale particles for each timestep in approximately 32 ms, requiring only 645 MB of GPU memory per timestep - nearly 20× smaller than the original 12 GB raw data. This eliminates costly reconstruction, accelerates the visual analysis pipeline, and overcomes I/O bottlenecks in multi-time-step analysis. Extensive experiments and comparisons across multiple datasets validate the effectiveness of our method. Ruixiao Peng, Guan Li 0002, Zhe Wang 0059, Yu Dong 0001, Tianchi Zhang 0003, Xuyi Lu, Yifei Jia, Guihua Shan, Dong Tian |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Enhancing INR-Based Super-Resolution Performance in Scientific Visualization via a Priori and a Posteriori Constraints
Yang Liu 0469, Guan Li 0002, Weiqun Cao, Guihua Shan, Dong Tian, Zhe Wang 0059 |
CGI (3) | 6 |
| 2025 | In Situ Workload Estimation for Block Assignment and Duplication in Parallelization-Over-Data Particle AdvectionabstractAbstract Particle advection is a foundational algorithm for analyzing a flow field. The commonly used Parallelization‐Over‐Data (POD) strategy for particle advection can become slow and inefficient when there are unbalanced workloads, which are particularly prevalent in in situ workflows. In this work, we present an in situ workflow containing workload estimation for block assignment and duplication in a parallelization‐over‐data algorithm. With tightly coupled workload estimation and load‐balanced block assignment strategy, our workflow offers a considerable improvement over the traditional round‐robin block assignment strategy. Our experiments demonstrate that particle advection is up to 3X faster and associated workflow saves approximately 30% of execution time after adopting strategies presented in this work. Zhe Wang 0059, Kenneth Moreland, Matthew Larsen, James Kress, Hank Childs, Guan Li 0002, Guihua Shan, David Pugmire |
Comput. Graph. Forum | 1 |
| 2025 | Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic ModelsabstractThis paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. In this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets. Tushar M. Athawale, Zhe Wang 0059, David Pugmire, Kenneth Moreland, Qian Gong, Scott Klasky, Chris R. Johnson 0001, Paul Rosen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | Parallelize Over Data Particle Advection: Participation, Ping Pong Particles, and OverheadabstractParticle advection is one of the foundational algorithms for visualization and analysis and is central to understanding vector fields common to scientific simulations. Achieving efficient performance with large data in a distributed memory setting is notoriously difficult. Because of its simplicity and minimized movement of large vector field data, the Parallelize over Data (POD) algorithm has become a de facto standard. Despite its simplicity and ubiquitous usage, the scaling issues with the POD algorithm are known and have been described throughout the literature. In this paper, we describe a set of in-depth analyses of the POD algorithm that shed new light on the underlying causes for the poor performance of this algorithm. We designed a series of representative workloads to study the performance of the POD algorithm and executed them on a supercomputer while collecting timing and statistical data for analysis. we then performed two different types of analysis. In the first analysis, we introduce two novel metrics for measuring algorithmic efficiency over the course of a workload run. The second analysis was from the perspective of the particles being advected. Using particle-centric analysis, we identify that the overheads associated with particle movement between processes (not the communication itself) have a dramatic impact on the overall execution time. These overheads become particularly costly when flow features span multiple blocks, resulting in repeated particle circulation (which we term "ping pong particles") between blocks. Our findings shed important light on the underlying causes of poor performance and offer directions for future research to address these limitations. Zhe Wang 0059, Kenneth Moreland, Matthew Larsen, James Kress, Hank Childs, David Pugmire |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2024 | A General Framework for Error-controlled Unstructured Scientific Data CompressionabstractData compression plays a key role in reducing storage and I/O costs. Traditional lossy methods primarily target data on rectilinear grids and cannot leverage the spatial coherence in unstructured mesh data, leading to suboptimal compression ratios. We present a multi-component, error-bounded compression framework designed to enhance the compression of floating-point unstructured mesh data, which is common in scientific applications. Our approach involves interpolating mesh data onto a rectilinear grid and then separately compressing the grid interpolation and the interpolation residuals. This method is general, independent of mesh types and typologies, and can be seamlessly integrated with existing lossy compressors for improved performance. We evaluated our framework across twelve variables from two synthetic datasets and two real-world simulation datasets. The results indicate that the multi-component framework consistently outperforms state-of-the-art lossy compressors on unstructured data, achieving, on average, a 2.3 − 3.5× improvement in compression ratios, with error bounds ranging from 1 × 10 the−6to 1×10−2. We further investigate impact of hyperparameters, such as grid spacing and error allocation, to deliver optimal compression ratios in diverse datasets. Qian Gong, Zhe Wang 0059, Viktor Reshniak, Xin Liang 0001, Jieyang Chen, Qing Liu 0002, Tushar M. Athawale, Yi Ju, Anand Rangarajan 0001, Sanjay Ranka, Norbert Podhorszki, Rick Archibald, Scott Klasky |
e-Science | 2 |
| 2024 | A High-Quality Workflow for Multi-Resolution Scientific Data Reduction and VisualizationabstractMulti-resolution methods such as Adaptive Mesh Refinement (AMR) can enhance storage efficiency for HPC applications generating vast volumes of data. However, their applicability is limited and cannot be universally deployed across all applications. Furthermore, integrating lossy compression with multi-resolution techniques to further boost storage efficiency encounters significant barriers. To this end, we introduce an innovative workflow that facilitates high-quality multi-resolution data compression for both uniform and AMR simulations. Initially, to extend the usability of multi-resolution techniques, our workflow employs a compression-oriented Region of Interest (ROI) extraction method, transforming uniform data into a multi-resolution format. Subsequently, to bridge the gap between multi-resolution techniques and lossy compressors, we optimize three distinct compressors, ensuring their optimal performance on multi-resolution data. These optimizations can improve the compression ratio of SOTA approaches by up to $3.3 \times$ under the same data quality loss. Lastly, we incorporate an advanced uncertainty visualization method into our workflow to understand the potential impacts of lossy compression. Experimental evaluation demonstrates that our workflow achieves significant compression quality improvements. Daoce Wang, Pascal Grosset, Jesus Pulido, Tushar M. Athawale, Jiannan Tian, Kai Zhao 0008, Zarija Lukic, Axel Huebl, Zhe Wang 0059, James P. Ahrens, Dingwen Tao |
SC | 9 |
| 2023 | Adaptive elasticity policies for staging-based in situ visualization
Zhe Wang 0059, Matthieu Dorier, Pradeep Subedi, Philip E. Davis, Manish Parashar |
Future Gener. Comput. Syst. | 1 |
| 2023 | Towards elastic in situ analysis for high-performance computing simulations
Matthieu Dorier, Zhe Wang 0059, Srinivasan Ramesh, Utkarsh Ayachit, Shane Snyder, Robert B. Ross, Manish Parashar |
J. Parallel Distributed Comput. | 2 |
| 2022 | Colza: Enabling Elastic In Situ Visualization for High-performance Computing SimulationsabstractIn situ analysis and visualization have grown increasingly popular for enabling direct access to data from high-performance computing (HPC) simulations. As a simulation progresses and interesting physical phenomena emerge, however, the data produced may become increasingly complex, and users may need to dynamically change the type and scale of in situ analysis tasks being carried out and consequently adapt the amount of resources allocated to such tasks. To date, none of the production in situ analysis frameworks offer such an elasticity feature, and for good reason: the assumption that the number of processes could vary during run time would force developers to rethink software and algorithms at every level of the in situ analysis stack. In this paper we present Colza, a data staging service with elastic in situ visualization capabilities. Colza relies on the widely used ParaView Catalyst in situ visualization framework and enables elasticity by replacing MPI with a custom collective communication library based on the Mochi suite of libraries. To the best of our knowledge, this work is the first to enable elastic in situ visualization capabilities for HPC applications on top of existing production analysis tools. Matthieu Dorier, Zhe Wang 0059, Utkarsh Ayachit, Shane Snyder, Robert B. Ross, Manish Parashar |
IPDPS | 2 |
| 2021 | Adaptive Placement of Data Analysis Tasks For Staging Based In-Situ ProcessingabstractIn-situ processing addresses the gap between speeds of computing and I/O capabilities by processing data close to the data source, i.e., on the same system as the data source (e.g., a simulation). However, the effective implementation of in-situ processing workflows requires the optimization of several design parameters such as where on the system workflow data analysis/visualization (ana/vis) as placed and how execution as well as the interaction and data exchanges between ana/vis are coordinated. For example, in the case of hybrid in-situ processing, interacting ana/vis may be tightly or loosely coupled depending on their placement, and this can lead to very different performance and scalability. A key challenge is deciding the most appropriate ana/vis placement, which depends on dynamic applications, workflow, and system characteristics that might change at runtime. In this paper, we present a framework to support online adaptive data analysis placement during the execution of an in-situ workflow. Specifically, the paper presents a model and architecture, and explores several data analysis placement strategies. Evaluation results show that dynamically choosing appropriate data analysis placement strategies can balance the benefits and overhead of different data analysis placement patterns to reduce in-situ processing time. Zhe Wang 0059, Pradeep Subedi, Matthieu Dorier, Philip E. Davis, Manish Parashar |
HiPC | 1 |
| 2020 | Staging Based Task Execution for Data-driven, In-Situ Scientific WorkflowsabstractAs scientific workflows increasingly use extreme-scale resources, the imbalance between higher computational capabilities, generated data volumes, and available I/O bandwidth is limiting the ability to translate these scales into insights. In-situ workflows (and the in-situ approach) are leveraging storage levels close to the computation in novel ways in order to reduce the required I/O. However, to be effective, it is important that the mapping and execution of such in-situ workflows adopts a data-driven approach, enabling in-situ tasks to be executed flexibly based upon data content. This paper first explores the design space for data-driven in-situ workflows. Specifically, it presents a model that captures different factors that influence the mapping, execution, and performance of data-driven in-situ workflows and experimentally studies the impact of different mapping decisions and execution patterns. The paper then presents the design, implementation, and experimental evaluation of a data-driven in-situ workflow execution framework that leverages in-memory distributed data management and user-defined task-triggers to enable efficient and scalable in-situ workflow execution. Zhe Wang 0059, Pradeep Subedi, Matthieu Dorier, Philip E. Davis, Manish Parashar |
CLUSTER | 1 |
| 2014 | Task scheduling for energy optimization and temperature improvementsabstractIn this paper, we develop several algorithms for simultaneously optimizing for energy and temperature for executing a set of independent tasks on multicore machines. Simulation results demonstrate that our algorithms are effectively in achieving these goals. Zhe Wang 0059, Sanjay Ranka |
ISCC | 2 |
| 2010 | A simple thermal model for multi-core processors and its application to slack allocationabstractPower density and heat density of multicore processor system are increasing exponentially with Moore's Law. High temperature on chip greatly affects its reliability, and the cost of packaging and cooling system increases exponentially with power consumption. For a multicore processor, the peak temperature of a block depends on its own power density as well as power density of other blocks on chip. In this paper, we have developed a simple thermal model, called Matrix Model (MM), that can be used to derive temperature profiles for all the cores of a multicore processor. We theoretically demonstrate the correctness and efficiency of MM. Our simulation results show that the model is comparable to the HotSpot Model for predicting the peak temperature. Besides having lower computational cost, the MM is succinct (a single matrix) and can be used to derive algorithms for a variety of scenarios. We use this model to develop a novel slack allocation algorithm for a workflow represented by Directed Acyclic Graph on a multicore processor. Zhe Wang 0059, Sanjay Ranka |
IPDPS | 1 |
| 2009 | Slotted Wavelength Scheduling for Bulk Transfers in Research NetworksabstractThe advancement of optical network technologies has enabled data-intensive e-science collaborations, which often require the transfer of large files with predictable performance. To support such applications, we design and evaluate two algorithms for scheduling time-constrained bulk transfers on wavelength-based optical research networks. The first one seeks to maximize the network throughput while maintaining a level of fairness among the jobs. The second algorithm works in an overloaded network and serves as an alternative to the first algorithm. It seeks to extend the end times by the smallest possible proportion and complete all the jobs by the extended end times. The main challenge is that the underlying problems are integer optimization problems for wavelength assignment, which have no known fast optimal solutions. We present a heuristic sub-algorithm called LPDAR, which converts fractional solutions from linear programming into integer solutions. LPDAR is the key component used in both aforementioned algorithms. Evaluation shows that LPDAR leads to very good algorithms with a performance level and speed both comparable to those of the LP fractional solutions. Zhe Wang 0059, Sanjay Ranka, Ye Xia 0001 |
ICPP | 1 |