Jian Zhang 0070

dblp:07/314-70 · DBLP profile ↗
← Back
44ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0003-1348-8124ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 11 since 2021Systems, architecture and hardware · 16 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 WindStencil: Unleashing GPU Potential for High-Order Stencil Computation in High-Performance Inviscid CFD Simulations
Xiazhen Liu, Runfeng Jin, Jian Zhang 0070, Wu Yuan 0002, Shan Liang 0005, Zhonghua Lu
ICS6
2026 From Optimal Solutions to Reusable Rules: A Knowledge-Driven Stencil Scheduling Framework
Jian Zhang 0070, Xiazhen Liu, Wu Yuan 0002, Shan Liang 0005
KSEM (6)3
2025 An Integrated Topology-Aware Mapping Scheme for Large-Scale CFD Applications on the ORISE Supercomputer
Lin Gao 0004, Wu Yuan 0002, Xiazhen Liu, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu
ICA3PP (6)5
2025 FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
abstract
Efficiently solving large-scale linear systems is a critical challenge in electromagnetic simulations, particularly when using the Crank-Nicolson Finite-Difference Time-Domain method. Existing iterative solvers are commonly employed to handle the resulting sparse systems but suffer from slow convergence due to the ill-conditioned nature of the double-curl operator. Approximate preconditioners, like SOR and Incomplete LU decomposition (ILU) provide insufficient convergence, while direct solvers are impractical due to excessive memory requirements. To address this, we propose FlashMP, a novel preconditioning system that designs a subdomain exact solver based on discrete transforms. FlashMP provides an efficient GPU implementation that achieves multi-GPU scalability through domain decomposition. Evaluations on AMD MI60 GPU clusters (up to$\mathbf{1 0 0 0 ~ G P U s}$) show that FlashMP reduces iteration counts by up to$16 \times$and achieves speedups of$2.5 \times$to$4.9 \times$compared to baseline implementations in state-of-the-art libraries Hypre. Weak scalability tests show parallel efficiencies up to 84.1 %.
Yaqian Gao, Runfeng Jin, Yidong Chen 0014, Wu Yuan 0002, Wenpeng Ma, Shan Liang 0005, Jian Zhang 0070, Zhonghua Lu
ICCD11
2024 MIST: Efficient Mixed-Precision Preconditioning Through Iterative Sparse- Triangular Solver Design
abstract
Exact sparse-triangular solvers are highly sequential and difficult to implement efficiently on GPUs with ILU preconditioning. Lower precision is crucial for reducing data movement and storage demands in memory-bound problems. However, current mixed-precision systems struggle to achieve performance gains for ILU preconditioning on multi-GPU platforms due to challenges in (1) utilizing two levels of parallelism and (2) minimizing off-chip memory bandwidth while maintaining accuracy. Additionally, these systems focus on scalar operations and lack support for point-block matrices, which arise naturally in multiphysics problems and require tailored algorithm designs. To address these challenges, we propose MIST, a novel Mixed-precision Iterative Sparse-Triangular solver optimized for GPUs to accelerate preconditioning in Krylov methods. We (1) implement an efficient mixed-precision Jacobi iterative local solver to harness single-GPU parallelism and scale it to multi- GPU via domain decomposition, and (2) design a BSpMVA kernel to reduce bandwidth while achieving high double-precision accuracy. Integrated into a widely-used numerical library, MIST offers end-to-end support for solving sparse linear systems, balancing efficiency and convergence. Experimental results show that MIST provides a 3.38× average speedup over cuSPARSE�s exact sparse-triangular solver, with an additional 1.37× speedup when using low-precision, while maintaining robustness
Yidong Chen 0014, Wenpeng Ma, Wu Yuan 0002, Jian Zhang 0070, Zhonghua Lu
ICCD5
2024 Large-scale Phase-Field Simulations for Solid-Solid Phase Transformations involving Elastic Energy
abstract
Phase-field models have been used extensively in studying microstructure evolution in alloys and have the superiority of comprehending, predicting, and optimizing microstructure-sensitive macroscopic material properties. The elastic strain energy is a vital factor in modeling crystal structure formation in solid-solid phase transformations. Conventionally, it is computed in the reciprocal space according to the famous Khachaturyan-Shatalov theory. In large-scale simulations, the full-space Fourier transform becomes extremely time-consuming.
Yaqian Gao, Jian Zhang 0070, Huang Ye, Xuebin Chi
ICPP2
2024 High-Performance 3D convolution on the Latest Generation Sunway Processor
abstract
The emergence of High-Performance Computing (HPC) and Artificial Intelligence (AI) has significantly expanded the applications of three-dimensional convolutional neural networks (3D CNNs). At the same time, the next-generation Sunway supercomputer has evidenced its superior computational capabilities in the HPC+AI domain. However, complex 3D convolution remains a primary performance limitation in many applications. The optimization of tensor-like operators on the Sunway processor is usually implemented via a multi-level blocking approach, adapting to its architecture. Although it can effectively mitigate the differences in memory access latency among different memory hierarchies, the performance of 3D convolutions is still frequently limited by the transfer bandwidth.
Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070
ICPP7
2024 A Feature Extraction Framework for 3D Scientific Voxel Object Using SVD and Neural Network
abstract
Latest advances in computational methods and high-performance computing have enabled large-scale scientific simulations to become feasible. However, efficiently analyzing the resulting large datasets remains challenging. Typically, 3D voxel data occupies huge storage space which causes serious storage pressure. Performing real-time feature extraction during simulations could mitigate this demand. In this paper, We propose a three-decker structure framework for locating and extracting voxel object features such as pose, class, and size. The framework first utilizes a 3D CNN to localize the object’s center and size along each axis, enabling adaptive cropping of the target from the original voxel space. Then the singular value decomposition (SVD) will applied to the extracted object to preliminarily extract and refine its posture. Finally, a neural network will refine the SVD-extracted pose and predict the category and size. Finally, a neural network is utilized to infer residual pose corrections as well as category and size information for the object after preliminary pose alignment via SVD decomposition. Our method achieves accurate feature extraction for voxel objects with complex poses. More importantly, we achieve pose estimation without an initial position. By storing extracted features instead of full data, we significantly reduce storage needs. We validate our approach using the phase field simulations dataset and ModelNet40 dataset. At last, based on the proposed framework we develop an in-situ feature extraction library for running with large-scale scientific computing programs and performing real-time feature extraction on the large amount of computation data it generates. As a concrete running example, we selected the microstructure evolution program governed by the phase-field method for a common run and achieved high-accuracy feature extraction.
Zhichen Feng, Yaqian Gao, Huang Ye, Jian Zhang 0070
IJCNN4
2024 A General Parallel Framework for Material Point Method Based on the p4est Library
abstract
Material Point Method(MPM) is widely used to simulate large deformation processes such as material fracture, collision, and fluid structure interaction. In general, an MPM simulation evolves a dynamic process in which a large number of Lagrangian particles move on top of Eulerian grids. During the process, physical quantities such as momentum and force are interpolated frequently between the particles and their corresponding grids. High Performance Computing(HPC) has become an indispensable tool for large-scale MPM simulations, while dynamic task decomposition and load balancing are critical for efficiency.In this paper, we present a general parallel framework integrating MPM and a well-established oct-tree mesh management library, p4est, aiming to improve the efficiency of large-scale MPM simulations. Through careful design of data structure and interfaces to connect the original Grid class of MPM with the p4est mesh, a highly modular structure is realized for the framework. This design allows the users to easily add new features or optimize existing ones, thereby enhancing its flexibility and reusability. On top of this, dynamic load-balancing strategies for MPM can be realized without much effort. A recommended strategy that considers both particle and grid workload is presented. In addition, the advantage of dynamic load balancing is demonstrated and analyzed through practical simulations. The versatility and efficiency of the proposed framework are also demonstrated through concrete real-world applications including penetration and building implosion simulations with up to 1.5 billion degrees of freedom. The code achieves 87% overall parallel efficiency scaling up to 2048 processors.
Shaobo Tian, Huang Ye, Jian Zhang 0070
ISPA5
2024 POSTER: Enabling Extreme-Scale Phase Field Simulation with In-situ Feature Extraction
abstract
In this paper, we present an integrated framework composed of a highly efficient phase field simulator and an in-situ feature extraction library. This novel framework enables us to conduct extreme-scale micro-structure evolution simulations while the characteristic features of each individual grain are extracted on the fly. After systematic design and optimization on the new generation Sunway supercomputer, the code scales up to 39 million cores and achieves 582 PFlops in double precision and 637 POps in mixed precision.
Zhichen Feng, Yaqian Gao, Shaobo Tian, Huang Ye, Jian Zhang 0070
PPoPP6
2024 Mixed-precision block incomplete sparse approximate preconditioner on Tensor core
Wenpeng Ma, Wu Yuan 0002, Jian Zhang 0070, Zhonghua Lu
CCF Trans. High Perform. Comput.4
2024 A novel transformer-based graph generation model for vectorized road design
abstract
Abstract Road network design, as an important part of landscape modeling, shows a great significance in automatic driving, video game development, and disaster simulation. To date, this task remains labor‐intensive, tedious and time‐consuming. Many improved techniques have been proposed during the last two decades. Nevertheless, most of the state‐of‐the‐art methods still encounter problems of intuitiveness, usefulness and/or interactivity. As a rapid deviation from the conventional road design, this paper advocates an improved road modeling framework for automatic and interactive road production driven by geographical maps (including elevation, water, vegetation maps). Our method integrates the capability of flexible image generation models with powerful transformer architecture to afford a vectorized road network. We firstly construct a dataset that includes road graphs, density map and their corresponding geographical maps. Secondly, we develop a density map generation network based on image translation model with an attention mechanism to predict a road density map. The usage of density map facilitates faster convergence and better performance, which also serves as the input for road graph generation. Thirdly, we employ the transformer architecture to evolve density maps to road graphs. Our comprehensive experimental results have verified the efficiency, robustness and applicability of our newly‐proposed framework for road design.
Peichi Zhou, Chen Li 0035, Jian Zhang 0070, Changbo Wang, Hong Qin 0001
Comput. Animat. Virtual Worlds3
2024 Force-Directed Graph Layouts Revisited: A New Force Based on the T-Distribution
abstract
In this article, we propose the t-FDP model, a force-directed placement method based on a novel bounded short-range force (t-force) defined by Student's t-distribution. Our formulation is flexible, exerts limited repulsive forces for nearby nodes and can be adapted separately in its short- and long-range effects. Using such forces in force-directed graph layouts yields better neighborhood preservation than current methods, while maintaining low stress errors. Our efficient implementation using a Fast Fourier Transform is one order of magnitude faster than state-of-the-art methods and two orders faster on the GPU, enabling us to perform parameter tuning by globally and locally adjusting the t-force in real-time for complex graphs. We demonstrate the quality of our approach by numerical evaluation against state-of-the-art approaches and extensions for interactive exploration.
Fahai Zhong, Mingliang Xue, Jian Zhang 0070, Fan Zhang 0045, Rui Ban, Oliver Deussen, Yunhai Wang
IEEE Trans. Vis. Comput. Graph.3
2023 udPINNs: An Enhanced PDE Solving Algorithm Incorporating Domain of Dependence Knowledge
Nanxi Chen, Jiyan Qiu, Wu Yuan 0002, Jian Zhang 0070
KSEM (4)5
2023 Target Netgrams: An Annulus-Constrained Stress Model for Radial Graph Visualization
abstract
We present Target Netgrams as a visualization technique for radial layouts of graphs. Inspired by manually created target sociograms, we propose an annulus-constrained stress model that aims to position nodes onto the annuli between adjacent circles for indicating their radial hierarchy, while maintaining the network structure (clusters and neighborhoods) and improving readability as much as possible. This is achieved by having more space on the annuli than traditional layout techniques. By adapting stress majorization to this model, the layout is computed as a constrained least square optimization problem. Additional constraints (e.g., parent-child preservation, attribute-based clusters and structure-aware radii) are provided for exploring nodes, edges, and levels of interest. We demonstrate the effectiveness of our method through a comprehensive evaluation, a user study, and a case study.
Mingliang Xue, Yunhai Wang, Chang Han, Jian Zhang 0070, Kaiyi Zhang 0003, Christophe Hurter, Jian Zhao 0010, Oliver Deussen
IEEE Trans. Vis. Comput. Graph.4
2022 A Fine-grained Prefetching Scheme for DGEMM Kernels on GPU with Auto-tuning Compatibility
abstract
General Matrix Multiplication (GEMM) is one of the fundamental kernels for scientific and high-performance computing. When optimizing the performance of GEMM on GPU, the matrix is usually partitioned into a hierarchy of tiles to fit the thread hierarchy. In practice, the thread-level parallelism is affected not only by the tiling scheme but also by the resources that each tile consumes, such as registers and local data share memory. This paper presents a fine-grained prefetching scheme that improves the thread-level parallelism by balancing the usage of such resources. The gain and loss on instruction and thread level parallelism are analyzed and a mathematical model is developed to estimate the overall performance gain. Moreover, the proposed scheme is integrated into the open-source tool Tensile to automatically generate assembly and tune a collection of kernels to maximize the performance of DGEMM for a family of problem sizes. Experiments show about 1.10X performance speedup on a wide range of matrix sizes for both single and batched matrix-matrix multiplication.
Huang Ye, Shaobo Tian, Jian Zhang 0070
IPDPS5
2022 Unsupervised Textured Terrain Generation via Differentiable Rendering
abstract
Constructing large-scale realistic terrains using modern modeling tools is an extremely challenging task even for professional users, undermining the effectiveness of video games, virtual reality, and other applications. In this paper, we present a step towards unsupervised and realistic modeling of textured terrains from DEM and satellite imagery, built upon two-stage illumination and texture optimization via differentiable rendering. First, a differentiable renderer for satellite imagery is established based on the Lambert diffuse model that allows inverse optimization of material and lighting parameters towards specific objective. Second, the original illumination direction of satellite imagery is recovered by reducing the difference between the shadow distribution generated by the renderer and that of the satellite image in YCrCb colour space, leveraging the abundant geometric information of DEM. Third, we propose to generate the original texture of the shadowed region by introducing visual consistency and smoothness constraints via differentiable rendering to arrive at an end-to-end unsupervised architecture. Comprehensive experiments demonstrate the effectiveness and efficiency of our proposed method as a potential tool to achieve virtual terrain modeling for widespread graphics applications.
Peichi Zhou, Dingbo Lu, Chen Li 0035, Jian Zhang 0070, Changbo Wang
ACM Multimedia4
2022 Sparse Reconstruction Method for Flow Fields Based on Mode Decomposition Autoencoder
Jiyan Qiu, Wu Yuan 0002, Jian Zhang 0070, Xuebin Chi
PRICAI (1)4
2022 User-level parallel file system: Case studies and performance optimizations
abstract
Abstract User‐level file systems are usually adopted to bridge the gap between efficacy and efficiency of file system developments for new applications' I/O demands. And the widely known user‐space file system framework, FUSE, is commonly utilized to deployed user‐level file systems. This article first uses a popular stack‐able file system as a case study to exam how FUSE affects I/O performance. Based on the testing and analytical results, this article then presents SHC, an implementation method to implement a user‐level file system without FUSE intervention. Experimental results indicate that SHC improves write bandwidth by up to 5.6x compared with that of FUSE and present leading superiority on read cases.
Yanliang Zou, Chen Chen 0124, Tongliang Deng, Jian Zhang 0070, Xiaomin Zhu 0001, Si Chen 0009, Shu Yin 0001
Concurr. Comput. Pract. Exp.4
2022 Authoring multi-style terrain with global-to-local control
Jian Zhang 0070, Chen Li 0035, Peichi Zhou, Changbo Wang, Gaoqi He, Hong Qin 0001
Graph. Model.1
2022 Pyramid-based Scatterplots Sampling for Progressive and Streaming Data Visualization
abstract
We present a pyramid-based scatterplot sampling technique to avoid overplotting and enable progressive and streaming visualization of large data. Our technique is based on a multiresolution pyramid-based decomposition of the underlying density map and makes use of the density values in the pyramid to guide the sampling at each scale for preserving the relative data densities and outliers. We show that our technique is competitive in quality with state-of-the-art methods and runs faster by about an order of magnitude. Also, we have adapted it to deliver progressive and streaming data visualization by processing the data in chunks and updating the scatterplot areas with visible changes in the density map. A quantitative evaluation shows that our approach generates stable and faithful progressive samples that are comparable to the state-of-the-art method in preserving relative densities and superior to it in keeping outliers and stability when switching frames. We present two case studies that demonstrate the effectiveness of our approach for exploring large data.
Xin Chen 0075, Jian Zhang 0070, Chi-Wing Fu, Jean-Daniel Fekete, Yunhai Wang
IEEE Trans. Vis. Comput. Graph.2
2022 F2-Bubbles: Faithful Bubble Set Construction and Flexible Editing
abstract
In this paper, we propose F2-Bubbles, a set overlay visualization technique that addresses overlapping artifacts and supports interactive editing with intelligent suggestions. The core of our method is a new, efficient set overlay construction algorithm that approximates the optimal set overlay by considering set elements and their non-set neighbors. Thanks to the efficiency of the algorithm, interactive editing is achieved, and with intelligent suggestions, users can easily and flexibly edit visualizations through direct manipulations with local adaptations. A quantitative comparison with state-of-the-art set visualization techniques and case studies demonstrate the effectiveness of our method and suggests that F2-Bubbles is a helpful technique for set visualization.
Yunhai Wang, Da Cheng, Jian Zhang 0070, Liang Zhou 0001, Gaoqi He, Oliver Deussen
IEEE Trans. Vis. Comput. Graph.4
2022 Data-Driven Colormap Adjustment for Exploring Spatial Variations in Scalar Fields
abstract
Colormapping is an effective and popular visualization technique for analyzing patterns in scalar fields. Scientists usually adjust a default colormap to show hidden patterns by shifting the colors in a trial-and-error process. To improve efficiency, efforts have been made to automate the colormap adjustment process based on data properties (e.g., statistical data value or histogram distribution). However, as the data properties have no direct correlation to the spatial variations, previous methods may be insufficient to reveal the dynamic range of spatial variations hidden in the data. To address the above issues, we conduct a pilot analysis with domain experts and summarize three requirements for the colormap adjustment process. Based on the requirements, we formulate colormap adjustment as an objective function, composed of a boundary term and a fidelity term, which is flexible enough to support interactive functionalities. We compare our approach with alternative methods under a quantitative measure and a qualitative user study (25 participants), based on a set of data with broad distribution diversity. We further evaluate our approach via three case studies with six domain experts. Our method is not necessarily more optimal than alternative methods of revealing patterns, but rather is an additional color adjustment option for exploring data with a dynamic range of spatial variations.
Qiong Zeng, Yongwei Zhao 0002, Yinqiao Wang, Jian Zhang 0070, Yi Cao 0005, Changhe Tu, Ivan Viola, Yunhai Wang
IEEE Trans. Vis. Comput. Graph.4
2022 KD-Box: Line-segment-based KD-tree for Interactive Exploration of Large-scale Time-Series Data
abstract
Time-series data-usually presented in the form of lines-plays an important role in many domains such as finance, meteorology, health, and urban informatics. Yet, little has been done to support interactive exploration of large-scale time-series data, which requires a clutter-free visual representation with low-latency interactions. In this paper, we contribute a novel line-segment-based KD-tree method to enable interactive analysis of many time series. Our method enables not only fast queries over time series in selected regions of interest but also a line splatting method for efficient computation of the density field and selection of representative lines. Further, we develop KD-Box, an interactive system that provides rich interactions, e.g., timebox, attribute filtering, and coordinated multiple views. We demonstrate the effectiveness of KD-Box in supporting efficient line query and density field computation through a quantitative comparison and show its usefulness for interactive visual analysis on several real-world datasets.
Yue Zhao 0033, Yunhai Wang, Jian Zhang 0070, Chi-Wing Fu, Mingliang Xu 0001, Dominik Moritz
IEEE Trans. Vis. Comput. Graph.3
2021 Redesigning Peridigm on SIMT Accelerators for High-performance Peridynamics Simulations
abstract
Peridigm is one of the most frequently utilized Peridynamics (PD) simulation software for problems involving discontinuity, such as cracks and fragmentation. However, performing long-term and large-scale simulations is very time-consuming for Peridigm. To enhance the performance and scalability of Peridigm, we port and optimize Peridigm on the SIMT accelerators. Challenges are imposed on efficient Peridigm on the SIMT architecture by the complex calculations and massive memory access of PD simulations. In this study, a series of strategies and techniques are proposed to optimize the performance of Peridigm. We first adjust the algorithms of bond-based calculations to eliminate the data conflicts with minimized overhead in order to achieve parallel Peridigm on accelerators. Furthermore, we propose thread grouping and collaborative memory access strategies to decrease the overhead of data fetch from device memory. To improve the efficiency of calculations, we also refine the calculation instructions. Finally, we offer a transmission-computation overlapping strategy for reducing the overhead brought by the data transmissions and improving the scalability. The optimized Peridigm on 4 Nvidia Tesla V100 GPUs accelerates the basic parallel Peridigm on 4 V100 GPUs 10.24 times. Compared to the original Peridigm run on 8 Intel Xeon Gold 6248 CPUs (160 cores, 320 threads) and the optimized PD application run on 4 SW26010 processors (1,040 cores), our work on 4 V100 GPUs accelerates the simulation 9 times and 4 times respectively. As for large-scale simulations, because we don't have enough V100 GPUs, we run our work on noncommercial SIMT accelerators which have similar performance to the V100 of the PCIe version, with the example scales from 282,000 points to 36,096,000 points and the number of accelerators scales from 4 to 512, near-linear scalability is observed and the performance ultimately reaching 825.72 TFLOPS with 98.81% parallel efficiency
Huang Ye, Jian Zhang 0070
IPDPS3
2021 Implicit Multidimensional Projection of Local Subspaces
abstract
We propose a visualization method to understand the effect of multidimensional projection on local subspaces, using implicit function differentiation. Here, we understand the local subspace as the multidimensional local neighborhood of data points. Existing methods focus on the projection of multidimensional data points, and the neighborhood information is ignored. Our method is able to analyze the shape and directional information of the local subspace to gain more insights into the global structure of the data through the perception of local structures. Local subspaces are fitted by multidimensional ellipses that are spanned by basis vectors. An accurate and efficient vector transformation method is proposed based on analytical differentiation of multidimensional projections formulated as implicit functions. The results are visualized as glyphs and analyzed using a full set of specifically-designed interactions supported in our efficient web-based visualization tool. The usefulness of our method is demonstrated using various multi- and high-dimensional benchmark datasets. Our implicit differentiation vector transformation is evaluated through numerical comparisons; the overall method is evaluated through exploration examples and use cases.
Rongzheng Bian, Yumeng Xue, Liang Zhou 0001, Jian Zhang 0070, Baoquan Chen, Daniel Weiskopf, Yunhai Wang
IEEE Trans. Vis. Comput. Graph.4
2021 SineStream: Improving the Readability of Streamgraphs by Minimizing Sine Illusion Effects
abstract
In this paper, we propose SineStream, a new variant of streamgraphs that improves their readability by minimizing sine illusion effects. Such effects reflect the tendency of humans to take the orthogonal rather than the vertical distance between two curves as their distance. In SineStream, we connect the readability of streamgraphs with minimizing sine illusions and by doing so provide a perceptual foundation for their design. As the geometry of a streamgraph is controlled by its baseline (the bottom-most curve) and the ordering of the layers, we re-interpret baseline computation and layer ordering algorithms in terms of reducing sine illusion effects. For baseline computation, we improve previous methods by introducing a Gaussian weight to penalize layers with large thickness changes. For layer ordering, three design requirements are proposed and implemented through a hierarchical clustering algorithm. Quantitative experiments and user studies demonstrate that SineStream improves the readability and aesthetics of streamgraphs compared to state-of-the-art methods.
Chuan Bu, Quanjie Zhang, Qianwen Wang 0001, Jian Zhang 0070, Michael Sedlmair, Oliver Deussen, Yunhai Wang
IEEE Trans. Vis. Comput. Graph.4
2020 Novel Sketch-Based 3D Model Retrieval via Cross-domain Feature Clustering and Matching
Jian Zhang 0070, Chen Li 0035, Changbo Wang, Gaoqi He, Hong Qin 0001
ICANN (1)2
2020 FILT: Optimizing KV-Embedded File Systems through Flat Indexing
abstract
The effectiveness of applying key-value store mechanisms to manage metadata of file systems has been demonstrated recently. However, traditional indirect metadata indexing schemes are not in concert with modern key-value data structures, which could degrade the performance of a KV-embedded file system due to the overhead of hierarchical path queries. In this paper, we propose FILT, a proof-of-concept file system middleware that can solve this problem by employing flat indexing. FILT exploits the benefits of both flat indexing and LSM-tree structure to eliminate redundant path lookups. Our extensive performance evaluation studies show that FILT can offer up to 5.8x performance gain compared with sophisticated local file systems.
Chen Chen 0124, Tongliang Deng, Jian Zhang 0070, Yanliang Zou, Xiaomin Zhu 0001, Shu Yin 0001
ICDCS3
2020 Large-scale Simulations of Peridynamics on Sunway Taihulight Supercomputer
abstract
Peridynamics (PD) methods are good at describing solid mechanical behaviours and have the superiority on simulating the discontinuous problems. They can be applied to many fields, such as materials science, human health, and industrial manufacturing, etc., which motivates us to provide their efficient numerical simulations on the Sunway TaihuLight supercomputer. However, massive and complex calculations of PD simulations and the characteristics of Sunway TaihuLight bring challenges to efficient parallel PD simulations. In this paper, we present a series of performance optimization techniques to perform a large-scale parallel PD simulation application on Sunway TaihuLight. We first design the data grouping and SPM-based caching to increase the bandwidth of data transmission and reduce the time of the main memory access. Further, we design and implement vectorization and instruction-level optimization for PD applications to improve computational performance. Finally, we offer the overlapping strategies of data transmission and computation so that data transmission can be covered by computation. Our work in a core group improves the performance of the serial version on the SW26010 processor by 181 times. Compared to the serial and single-CPU Peridigm-based simulations on Intel Xeon E5-2680 V3, our work gets a speedup of 60 times and 6 times, respectively. Near linear scalability is also obtained. When testing the weak scaling, the simulation of a 296,222,720-point example achieves 1.14 PFLOPS with 8192 (532,480 cores) processes. When testing the strong scaling, 90% parallel efficiency is observed as the number of processes increases 64 times to 4096 processes.
Huang Ye, Jian Zhang 0070
ICPP3
2020 BORA: a bag optimizer for robotic analysis
abstract
We present BORA (Bag Optimizer for Robotic Analysis), a file system middleware that optimizes the acquisition of bags, which are specially formatted files used to store timestamped ROS (robot operating system) messages. BORA sits between ROS and an existing file system to conduct semantic-aware data pre-processing. In particular, it categorizes ROS bag data into multiple groups with each having a distinct label. BORA predigests data index constructions and reduces file open time via a hash-based label management scheme. It is also capable of providing ROS analytic applications with only data needed without a sequence of data searching and locating operations. We implement a BORA prototype, which is then integrated into three computing platforms: a single-node server, a four-node PVFS storage cluster, and a Tianhe-1A Supercomputer storage subsystem. Next, we evaluate the BORA prototype on the three platforms using four real-world ROS applications. Our experimental results show that compared to a traditional bag management scheme BORA improves data acquisition performance by up to 11x. In addition, it offers up to 10x data acquisition performance improvement and 3,100x bags open improvement under a swarm robotics data analysis scenario where data is retrieved across multiple bags simultaneously.
Jian Zhang 0070, Tao Xie 0004, Yuzhuo Jing, Guanzhou Hu, Si Chen 0009, Shu Yin 0001
SC1
2020 A Recursive Subdivision Technique for Sampling Multi-class Scatterplots
abstract
We present a non-uniform recursive sampling technique for multi-class scatterplots, with the specific goal of faithfully presenting relative data and class densities, while preserving major outliers in the plots. Our technique is based on a customized binary kd-tree, in which leaf nodes are created by recursively subdividing the underlying multi-class density map. By backtracking, we merge leaf nodes until they encompass points of all classes for our subsequently applied outlier-aware multi-class sampling strategy. A quantitative evaluation shows that our approach can better preserve outliers and at the same time relative densities in multi-class scatterplots compared to the previous approaches, several case studies demonstrate the effectiveness of our approach in exploring complex and real world data.
Xin Chen 0075, Tong Ge, Jian Zhang 0070, Baoquan Chen, Chi-Wing Fu, Oliver Deussen, Yunhai Wang
IEEE Trans. Vis. Comput. Graph.3
2020 ShapeWordle: Tailoring Wordles using Shape-aware Archimedean Spirals
abstract
We present a new technique to enable the creation of shape-bounded Wordles, we call ShapeWordle, in which we fit words to form a given shape. To guide word placement within a shape, we extend the traditional Archimedean spirals to be shape-aware by formulating the spirals in a differential form using the distance field of the shape. To handle non-convex shapes, we introduce a multi-centric Wordle layout method that segments the shape into parts for our shape-aware spirals to adaptively fill the space and generate word placements. In addition, we offer a set of editing interactions to facilitate the creation of semantically-meaningful Wordles. Lastly, we present three evaluations: a comprehensive comparison of our results against the state-of-the-art technique (WordArt), case studies with 14 users, and a gallery to showcase the coverage of our technique.
Yunhai Wang, Kaiyi Zhang 0003, Chen Bao, Jian Zhang 0070, Chi-Wing Fu, Christophe Hurter, Bongshin Lee, Oliver Deussen
IEEE Trans. Vis. Comput. Graph.6
2019 Example-based rapid generation of vegetation on terrain via CNN-based distribution learning
Jian Zhang 0070, Changbo Wang, Chen Li 0035, Hong Qin 0001
Vis. Comput.1
2019 Procedural modeling of rivers from single image toward natural scene production
Jian Zhang 0070, Changbo Wang, Hong Qin 0001, Yan Gao 0004
Vis. Comput.1
2018 A Perception-Driven Approach to Supervised Dimensionality Reduction for Visualization
abstract
Dimensionality reduction (DR) is a common strategy for visual analysis of labeled high-dimensional data. Low-dimensional representations of the data help, for instance, to explore the class separability and the spatial distribution of the data. Widely-used unsupervised DR methods like PCA do not aim to maximize the class separation, while supervised DR methods like LDA often assume certain spatial distributions and do not take perceptual capabilities of humans into account. These issues make them ineffective for complicated class structures. Towards filling this gap, we present a perception-driven linear dimensionality reduction approach that maximizes the perceived class separation in projections. Our approach builds on recent developments in perception-based separation measures that have achieved good results in imitating human perception. We extend these measures to be density-aware and incorporate them into a customized simulated annealing algorithm, which can rapidly generate a near optimal DR projection. We demonstrate the effectiveness of our approach by comparing it to state-of-the-art DR methods on 93 datasets, using both quantitative measure and human judgments. We also provide case studies with class-imbalanced and unlabeled data.
Yunhai Wang, Kang Feng, Jian Zhang 0070, Chi-Wing Fu, Michael Sedlmair, Xiaohui Yu 0001, Baoquan Chen
IEEE Trans. Vis. Comput. Graph.4
2018 Is There a Robust Technique for Selecting Aspect Ratios in Line Charts?
abstract
The aspect ratio of a line chart heavily influences the perception of the underlying data. Different methods explore different criteria in choosing aspect ratios, but so far, it was still unclear how to select aspect ratios appropriately for any given data. This paper provides a guideline for the user to choose aspect ratios for any input 1D curves by conducting an in-depth analysis of aspect ratio selection methods both theoretically and experimentally. By formulating several existing methods as line integrals, we explain their parameterization invariance. Moreover, we derive a new and improved aspect ratio selection method, namely the -LOR (local orientation resolution), with a certain degree of parameterization invariance. Furthermore, we connect different methods, including AL (arc length based method), the banking to 45 principle, RV (resultant vector) and AS (average absolute slope), as well as -LOR and AO (average absolute orientation). We verify these connections by a comparative evaluation involving various data sets, and show that the selections by RV and -LOR are complementary to each other for most data. Accordingly, we propose the dual-scale banking technique that combines the strengths of RV and -LOR, and demonstrate its practicability using multiple real-world data sets.
Yunhai Wang, Zeyu Wang 0005, Lifeng Zhu, Jian Zhang 0070, Chi-Wing Fu, Zhanglin Cheng, Changhe Tu, Baoquan Chen
IEEE Trans. Vis. Comput. Graph.4
2016 Mathematical foundations of arc length-based aspect ratio selection
abstract
The aspect ratio of a plot can strongly influence the perception of trends in the data. Arc length based aspect ratio selection (AL) has demonstrated many empirical advantages over previous methods. However, it is still not clear why and when this method works. In this paper, we attempt to unravel its mystery by exploring its mathematical foundation. First, we explain the rationale why this method is parameterization invariant and follow the same rationale to extend previous methods which are not parameterization invariant. As such, we propose maximizing weighted local curvature (MLC), a parameterization invariant form of local orientation resolution (LOR) and reveal the theoretical connection between average slope (AS) and resultant vector (RV). Furthermore, we establish a mathematical connection between AL and banking to 45 degrees and derive the upper and lower bounds of its average absolute slopes. Finally, we conduct a quantitative comparison that revises the understanding of aspect ratio selection methods in three aspects: (1) showing that AL, AWO and RV always perform very similarly while MS is not; (2) demonstrating the advantages in the robustness of RV over AL; (3) providing a counterexample where all previous methods produce poor results while MLC works well.
Fubo Han, Yunhai Wang, Jian Zhang 0070, Oliver Deussen, Baoquan Chen
PacificVis3
2016 Extreme-scale phase field simulations of coarsening dynamics on the sunway taihulight supercomputer
abstract
Many important properties of materials such as strength, ductility, hardness and conductivity are determined by the microstructures of the material. During the formation of these microstructures, grain coarsening plays an important role. The Cahn-Hilliard equation has been applied extensively to simulate the coarsening kinetics of a two-phase microstructure. It is well accepted that the limited capabilities in conducting large scale, long time simulations constitute bottlenecks in predicting microstructure evolution based on the phase field approach. We present here a scalable time integration algorithm with large stepsizes and its efficient implementation on the Sunway TaihuLight supercomputer. The highly nonlinear and severely stiff Cahn-Hilliard equations with degenerate mobility for microstructure evolution are solved at extreme scale, demonstrating that the latest advent of high performance computing platform and the new advances in algorithm design are now offering us the possibility to simulate the coarsening dynamics accurately at unprecedented spatial and time scales.
Jian Zhang 0070, Chunbao Zhou, Yangang Wang 0002, Lili Ju, Qiang Du 0001, Xuebin Chi, Dexun Chen
SC1
2016 The Sunway TaihuLight supercomputer: system and applications
Haohuan Fu, Junfeng Liao, Jinzhe Yang, Lanning Wang, Zhenya Song, Xiaomeng Huang, Chao Yang 0002, Wei Xue 0003, Fangfang Liu 0004, Fangli Qiao, Xunqiang Yin, Chaofeng Hou, Jian Zhang 0070, Yangang Wang 0002, Chunbo Zhou, Guangwen Yang 0002
Sci. China Inf. Sci.16
2015 Forecast Verification and Visualization based on Gaussian Mixture Model Co-estimation
abstract
Abstract Precipitation forecast verification is essential to the quality of a forecast. The Gaussian mixture model (GMM) can be used to approximate the precipitation of several rain bands and provide a concise view of the data, which is especially useful for comparing forecast and observation data. The robustness of such comparison mainly depends on the consistency of and the correspondence between the extracted rain bands in the forecast and observation data. We propose a novel co‐estimation approach based on GMM in which forecast and observation data are analysed simultaneously. This approach naturally increases the consistency of and correspondence between the extracted rain bands by exploiting the similarity between both forecast and observation data. Moreover, a novel visualization and exploration framework is implemented to help the meteorologists gain insight from the forecast. The proposed approach was applied to the forecast and observation data provided by the China Meteorological Administration. The results are evaluated by meteorologists and novel insight has been gained.
Yunhai Wang, Chaoran Fan, Jian Zhang 0070, Tao Niu, Song Zhang 0004, Jinrong Jiang
Comput. Graph. Forum3
2012 Automating Transfer Function Design with Valley Cell-Based Clustering of 2D Density Plots
abstract
Abstract Two‐dimensional transfer functions are an effective and well‐accepted tool in volume classification. The design of them mostly depends on the user's experience and thus remains a challenge. Therefore, we present an approach in this paper to automate the transfer function design based on 2D density plots. By exploiting their smoothness, we adopted the Morse theory to automatically decompose the feature space into a set of valley cells. We design a simplification process based on cell separability to eliminate cells which are mainly caused by noise in the original volume data. Boundary persistence is first introduced to measure the separability between adjacent cells and to suitably merge them. Afterward, a reasonable classification result is achieved where each cell represents a potential feature in the volume data. This classification procedure is automatic and facilitates an arbitrary number and shape of features in the feature space. The opacity of each feature is determined by its persistence and size. To further incorporate the user's prior knowledge, a hierarchical feature representation is created by successive merging of the cells. With this representation, the user is allowed to merge or split features of interest and set opacity and color freely. Experiments on various volumetric data sets demonstrate the effectiveness and usefulness of our approach in transfer function generation.
Yunhai Wang, Jian Zhang 0070, Dirk J. Lehmann, Holger Theisel, Xuebin Chi
Comput. Graph. Forum2
2011 Efficient opacity specification based on feature visibilities in direct volume rendering
abstract
Abstract Due to 3D occlusion, the specification of proper opacities in direct volume rendering is a time‐consuming and unintuitive process. The visibility histograms introduced by Correa and Ma reflect the effect of occlusion by measuring the influence of each sample in the histogram to the rendered image. However, the visibility is defined on individual samples, while volume exploration focuses on conveying the spatial relationships between features. Moreover, the high computational cost and large memory requirement limits its application in multi‐dimensional transfer function design. In this paper, we extend visibility histograms to feature visibility, which measures the contribution of each feature in the rendered image. Compared to visibility histograms, it has two distinctive advantages for opacity specification. First, the user can directly specify the visibilities for features and the opacities are automatically generated using an optimization algorithm. Second, its calculation requires only one rendering pass with no additional memory requirement. This feature visibility based opacity specification is fast and compatible with all types of transfer function design. Furthermore, we introduce a two‐step volume exploration scheme, in which an automatic optimization is first performed to provide a clear illustration of the spatial relationship and then the user adjusts the visibilities directly to achieve the desired feature enhancement. The effectiveness of this scheme is demonstrated by experimental results on several volumetric datasets.
Yunhai Wang, Jian Zhang 0070, Wei Chen 0001, Huai Zhang, Xuebin Chi
Comput. Graph. Forum2
2011 Efficient Volume Exploration Using the Gaussian Mixture Model
abstract
The multidimensional transfer function is a flexible and effective tool for exploring volume data. However, designing an appropriate transfer function is a trial-and-error process and remains a challenge. In this paper, we propose a novel volume exploration scheme that explores volumetric structures in the feature space by modeling the space using the Gaussian mixture model (GMM). Our new approach has three distinctive advantages. First, an initial feature separation can be automatically achieved through GMM estimation. Second, the calculated Gaussians can be directly mapped to a set of elliptical transfer functions (ETFs), facilitating a fast pre-integrated volume rendering process. Third, an inexperienced user can flexibly manipulate the ETFs with the assistance of a suite of simple widgets, and discover potential features with several interactions. We further extend the GMM-based exploration scheme to time-varying data sets using an incremental GMM estimation algorithm. The algorithm estimates the GMM for one time step by using itself and the GMM generated from its previous steps. Sequentially applying the incremental algorithm to all time steps in a selected time interval yields a preliminary classification for each time step. In addition, the computed ETFs can be freely adjusted. The adjustments are then automatically propagated to other time steps. In this way, coherent user-guided exploration of a given time interval is achieved. Our GPU implementation demonstrates interactive performance and good scalability. The effectiveness of our approach is verified on several data sets.
Yunhai Wang, Wei Chen 0001, Jian Zhang 0070, Tingxin Dong, Guihua Shan, Xuebin Chi
IEEE Trans. Vis. Comput. Graph.3