Jun Han 0010

dblp:02/3721-10 · DBLP profile ↗
← Back
30ranked-venue papers
14as first author
22since 2021 · last 2026
0000-0002-7286-062XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 14 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2Computer networks · 2Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Towards Understanding Time-Varying Spatial 3D Data Analysis with Animation and Small Multiples in Virtual Reality and Desktop
abstract
The growing availability of time-varying spatial 3D (S4D) data, such as ocean and atmospheric datasets, has created opportunities for studying dynamic phenomena across time and 3D space. However, designing effective visualizations for S4D data remains challenging due to the high cognitive demands and complexity of these datasets. While techniques like animation and small multiples have been applied in Virtual Reality (VR) and desktop environments, the lack of understanding of analysts’ tasks and challenges limits the development of better visualization techniques. To fill this gap, we conducted an empirical study with domain experts across various fields, comparing four visualization techniques: VR animation, VR small multiples, desktop animation, and desktop small multiples. We identified the strengths and weaknesses of the four techniques, as well as key analytical tasks, current practices, and challenges in S4D data analysis. Finally, we outlined future research opportunities for advancing S4D visualization techniques.
Linping Yuan, Le Lin, Yuquan Lin, Jun Han 0010, Zikun Deng, Weicong Cheng, Huamin Qu
VR4
2026 CD-TVD: Contrastive Diffusion for 3D Super-Resolution with Scarce High-Resolution Time-Varying Data
abstract
Large-scale scientific simulations require significant resources to generate high-resolution time-varying data (TVD). While super-resolution is an efficient post-processing strategy to reduce costs, existing methods rely on a large amount of HR training data, limiting their applicability to diverse simulation scenarios. To address this constraint, we proposed CD-TVD, a novel framework that combines contrastive learning and an improved diffusion-based super-resolution model to achieve accurate 3D super-resolution from limited time-step high-resolution data. During pre-training on historical simulation data, the contrastive encoder and diffusion super-resolution modules learn degradation patterns and detailed features of high-resolution and low-resolution samples. In the training phase, the improved diffusion model with a local attention mechanism is fine-tuned using only one newly generated high-resolution timestep, leveraging the degradation knowledge learned by the encoder. This design minimizes the reliance on large-scale high-resolution datasets while maintaining the capability to recover fine-grained details. Experimental results on fluid and atmospheric simulation datasets confirm that CD-TVD delivers accurate and resource-efficient 3D super-resolution, marking a significant advancement in data augmentation for large-scale scientific simulations. The code is available at https://github.com/Xin-Gao-private/CD-TVD.
Chongke Bi, Jiakang Deng, Guan Li 0002, Jun Han 0010
IEEE Trans. Vis. Comput. Graph.5
2026 A Few-Shot Learning Framework for Time-Varying Scientific Data Generation via Conditional Diffusion Model
abstract
A key factor in successfully achieving remarkable performance for deep learning models is the availability of large amounts of data. However, in scientific visualization, providing such extensive volumetric data is often infeasible due to the high computational cost of simulations and the challenges of data storage. To address this data sparsity issue in model training, we propose a few-shot learning framework that leverages only few training samples (e.g., 1, 3, or 5) to ensure both generalization capability and performance through a conditional diffusion model. Our approach consists of two stages: the forward process and the reverse process. In the forward process, we inject noise at various levels into the few samples. In the reverse process, we design a time-aware UNet that iteratively learns to denoise the noisy data. Additionally, we introduce a noise-aware loss function that dynamically adjusts optimization weights based on the noise levels in the training data. Our method demonstrates consistent and robust performance, regardless of the selection method for the few training samples. Furthermore, it achieves superior results in both quantitative and qualitative evaluations compared to state-of-the-art solutions across three scientific visualization tasks: spatial super-resolution, temporal super-resolution, and variable translation.
Jun Han 0010
IEEE Trans. Vis. Comput. Graph.1
2026 MoE-INR: Implicit Neural Representation with Mixture-of-Experts for Time-Varying Volumetric Data Compression
abstract
Implicit neural representations (INRs) have emerged as a transformative paradigm for time-varying volumetric data compression and representation, owing to their ability to model high-dimensional signals effectively. INRs represent scalar fields based on sampled coordinates, typically using either a single network for the entire field or multiple networks across different spatial domains. However, these approaches often face challenges in modeling complex patterns and introducing boundary artifacts. To address these limitations, we propose MoE-INR, an INR architecture based on a mixture-of-experts (MoE) framework. MoE-INR automates irregular subdivisions of spatiotemporal fields and dynamically assigns them to different expert networks. The architecture comprises three key components: a policy network, a shared encoder, and multiple expert decoders. The policy network subdivides the field and determines which expert decoder is responsible for a given input coordinate. The shared encoder extracts hidden representations from the input coordinates, and the expert decoders transform these high-dimensional features into scalar values. This design results in a unified framework accommodating diverse INR types, including conventional, grid-based, and ensemble. We evaluate the effectiveness of MoE-INR on multiple time-varying datasets with varying characteristics. Experimental results demonstrate that MoE-INR significantly outperforms existing non-MoE and MoE-based INRs and traditional lossy compression methods across quantitative and qualitative metrics under various compression ratios.
Jun Han 0010, Kaiyuan Tang 0001, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2026 TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting
abstract
Advancements in volume visualization (VolVis) focus on extracting insights from 3D volumetric data by generating visually compelling renderings that reveal complex internal structures. Existing VolVis approaches have explored non-photorealistic rendering techniques to enhance the clarity, expressiveness, and informativeness of visual communication. While effective, these methods often rely on complex predefined rules and are limited to transferring a single style, restricting their flexibility. To overcome these limitations, we advocate the representation of VolVis scenes using differentiable Gaussian primitives combined with pretrained large models to enable arbitrary style transfer and real-time rendering. However, conventional 3D Gaussian primitives tightly couple geometry and appearance, leading to suboptimal stylization results. To address this, we introduce TexGS-VolVis, a textured Gaussian splatting framework for VolVis. TexGS-VolVis employs 2D Gaussian primitives, extending each Gaussian with additional texture and shading attributes, resulting in higher-quality, geometry-consistent stylization and enhanced lighting control during inference. Despite these improvements, achieving flexible and controllable scene editing remains challenging. To further enhance stylization, we develop image-and text-driven non-photorealistic scene editing tailored for TexGS-VolVis and 2D-lift-3D segmentation to enable partial editing with fine-grained control. We evaluate TexGS-VolVis both qualitatively and quantitatively across various volume rendering scenes, demonstrating its superiority over existing methods in terms of efficiency, visual quality, and editing flexibility.
Kaiyuan Tang 0001, Kuangshi Ai, Jun Han 0010, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.3
2025 ST2VR: An Interactive Authoring System for SpatioTemporal STorytelling in Virtual Reality with Hierarchical Narrative Structure
abstract
The increasing popularity of Virtual Reality (VR) has provided a new medium for narrating spatiotemporal data stories. Compared to traditional 2D environments, telling spatiotemporal stories in VR holds the promise of offering a more immersive and engaging experience for audiences. However, when creating VR spatiotemporal data stories, story creators may feel overwhelmed by the numerous narrative elements involved and there is a disconnect between the creation and viewing environments. To address these challenges, we designed a hierarchical narrative structure with three levels: Story Line, Story Piece, and Viewpoint. We then introduce ST2VR, an interactive authoring system that supports the creation of spatiotemporal data stories within an immersive environment. The system features two types of interactive interfaces to support the organization of the overall narrative as well as the immersive design of story details. A use case demonstrates the process of authoring VR spatiotemporal data stories using ST2VR, and a user study evaluates the system’s efficiency and effectiveness.
Ziyue Lin, Linping Yuan, Jun Han 0010, Yalong Yang 0001, Siming Chen 0001
PacificVis5
2025 DCINR: A Divide-and-Conquer Implicit Neural Representation for Compressing Time-Varying Volumetric Data in Hours
abstract
Implicit neural representation (INR) has been a powerful paradigm for effectively compressing time-varying volumetric data. However, the optimization process can span days or even weeks due to its reliance on coordinate-based inputs and outputs for modeling volumetric data. To address this issue, we introduce a divide-and-conquer INR (DCINR), significantly accelerating the compressing process of time-varying volumetric data in hours. Our approach starts by dividing the data set into a set of non-overlapping blocks. Then, we apply a block selection strategy to weed out redundant blocks to reduce the computation cost without sacrificing performance. In parallel, each selected block is modeled by a tiny INR, with the size of the INR being adapted to match the information richness in the block. The block size is determined by maximizing the average network capacity. After optimization, the optimized INRs are utilized to decompress the data set. By evaluating our approach across various time-varying volumetric data sets, DCINR surpasses learning-based and lossy compression approaches in compression ratio, visual fidelity, and various performance metrics. Additionally, this method operates within a comparable compression time to that of lossy compressors, achieves extreme compression ratios ranging from thousands to tens of thousands, and preserves features with high quality.
Jun Han 0010
IEEE Trans. Vis. Comput. Graph.1
2025 A Study of Data Augmentation for Learning-Driven Scientific Visualization
abstract
The success of deep learning heavily relies on the large amount of training samples. However, in scientific visualization, due to the high computational cost, only few data are available during training, which limits the performance of deep learning. A common technique to address the data sparsity issue is data augmentation. In this paper, we present a comprehensive study on nine data augmentation techniques (i.e., noise injection, interpolation, scale, flip, rotation, variational auto-encoder, generative adversarial network, diffusion model, and implicit neural representation) for understanding their effectiveness on two scientific visualization tasks, i.e., spatial super-resolution and ambient occlusion prediction. We compare the data quality, rendering fidelity, optimization time, and memory consumption of these data augmentation techniques using several scientific datasets with various characteristics. We investigate the effects of data augmentation on the method, quantity, and diversity for these tasks with various deep learning models. Our study shows that increasing the quantity and single-domain diversity of augmented data can boost model performance, while the method and cross-domain diversity of the augmented data do not have the same impact. Based on our findings, we discuss the opportunities and future directions for scientific data augmentation.
Jun Han 0010, Hao Zheng 0006, Jun Tao 0002
IEEE Trans. Vis. Comput. Graph.1
2025 DTBIA: An Immersive Visual Analytics System for Brain-Inspired Research
abstract
The Digital Twin Brain (DTB) is an advanced artificial intelligence framework that integrates spiking neurons to simulate complex cognitive functions and collaborative behaviors. For domain experts, visualizing the DTB's simulation outcomes is essential to understanding complex cognitive activities. However, this task poses significant challenges due to DTB data's inherent characteristics, including its high-dimensionality, temporal dynamics, and spatial complexity. To address these challenges, we developed DTBIA, an Immersive Visual Analytics System for Brain-Inspired Research. In collaboration with domain experts, we identified key requirements for effectively visualizing spatiotemporal and topological patterns at multiple levels of detail. DTBIA incorporates a hierarchical workflow - ranging from brain regions to voxels and slice sections - along with immersive navigation and a 3D edge bundling algorithm to enhance clarity and provide deeper insights into both functional (BOLD) and structural (DTI) brain data. The utility and effectiveness of DTBIA are validated through two case studies involving with brain research experts. The results underscore the system's role in enhancing the comprehension of complex neural behaviors and interactions.
Jun-Hsiang Yao, Mingzheng Li, Yuxiao Li 0002, Jielin Feng, Jun Han 0010, Qibao Zheng, Jianfeng Feng, Siming Chen 0001
IEEE Trans. Vis. Comput. Graph.6
2024 KD-INR: Time-Varying Volumetric Data Compression via Knowledge Distillation-Based Implicit Neural Representation
abstract
Traditional deep learning algorithms assume that all data is available during training, which presents challenges when handling large-scale time-varying data. To address this issue, we propose a data reduction pipeline called knowledge distillation-based implicit neural representation (KD-INR) for compressing large-scale time-varying data. The approach consists of two stages: spatial compression and model aggregation. In the first stage, each time step is compressed using an implicit neural representation with bottleneck layers and features of interest preservation-based sampling. In the second stage, we utilize an offline knowledge distillation algorithm to extract knowledge from the trained models and aggregate it into a single model. We evaluated our approach on a variety of time-varying volumetric data sets. Both quantitative and qualitative results, such as PSNR, LPIPS, and rendered images, demonstrate that KD-INR surpasses the state-of-the-art approaches, including learning-based (i.e., CoordNet, NeurComp, and SIREN) and lossy compression (i.e., SZ3, ZFP, and TTHRESH) methods, at various compression ratios ranging from hundreds to ten thousand.
Jun Han 0010, Hao Zheng 0006, Chongke Bi
IEEE Trans. Vis. Comput. Graph.1
2023 GMT: A deep learning approach to generalized multivariate translation for scientific data analysis and visualization
Siyuan Yao, Jun Han 0010, Chaoli Wang 0001
Comput. Graph.2
2023 CoordNet: Data Generation and Visualization Generation for Time-Varying Volumes via a Coordinate-Based Neural Network
abstract
Although deep learning has demonstrated its capability in solving diverse scientific visualization problems, it still lacks generalization power across different tasks. To address this challenge, we propose CoordNet, a single coordinate-based framework that tackles various tasks relevant to time-varying volumetric data visualization without modifying the network architecture. The core idea of our approach is to decompose diverse task inputs and outputs into a unified representation (i.e., coordinates and values) and learn a function from coordinates to their corresponding values. We achieve this goal using a residual block-based implicit neural representation architecture with periodic activation functions. We evaluate CoordNet on data generation (i.e., temporal super-resolution and spatial super-resolution) and visualization generation (i.e., view synthesis and ambient occlusion prediction) tasks using time-varying volumetric data sets of various characteristics. The experimental results indicate that CoordNet achieves better quantitative and qualitative results than the state-of-the-art approaches across all the evaluated tasks. Source code and pre-trained models are available at https://github.com/stevenhan1991/CoordNet.
Jun Han 0010, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2023 DL4SciVis: A State-of-the-Art Survey on Deep Learning for Scientific Visualization
abstract
Since 2016, we have witnessed the tremendous growth of artificial intelligence+visualization (AI+VIS) research. However, existing survey articles on AI+VIS focus on visual analytics and information visualization, not scientific visualization (SciVis). In this article, we survey related deep learning (DL) works in SciVis, specifically in the direction of DL4SciVis: designing DL solutions for solving SciVis problems. To stay focused, we primarily consider works that handle scalar and vector field data but exclude mesh data. We classify and discuss these works along six dimensions: domain setting, research task, learning type, network architecture, loss function, and evaluation metric. The article concludes with a discussion of the remaining gaps to fill along the discussed dimensions and the grand challenges we need to tackle as a community. This state-of-the-art survey guides SciVis researchers in gaining an overview of this emerging topic and points out future directions to grow this research.
Chaoli Wang 0001, Jun Han 0010
IEEE Trans. Vis. Comput. Graph.2
2022 Scalar2Vec: Translating Scalar Fields to Vector Fields via Deep Learning
abstract
We introduce Scalar2Vec, a new deep learning solution that translates scalar fields to velocity vector fields for scientific visualization. Given multivariate or ensemble scalar field volumes and their velocity vector field counterparts, Scalar2Vec first identifies suitable variables for scalar-to-vector translation. It then leverages a k-complete bipartite translation network (kCBT-Net) to complete the translation task. kCBT-Net takes a set of sampled scalar volumes of the same variable as input, extracts their multi -scale information, and learns to synthesize the corresponding vector volumes. Ground-truth vector fields and their derived quantities are utilized for loss computation and network training. After training, Scalar2Vec can infer unseen velocity vector fields of the same data set directly from their scalar field counterparts. We demonstrate the effectiveness of Scalar2Vec with quantitative and qualitative results on multiple data sets and compare it with three other state-of-the-art deep learning methods.
Pengfei Gu, Jun Han 0010, Danny Ziyi Chen, Chaoli Wang 0001
PacificVis2
2022 AQX: Explaining Air Quality Forecast for Verifying Domain Knowledge using Feature Importance Visualization
abstract
Air pollution forecast has become critical because of its direct impact on human health and its increased production caused by rapid industrialization. Machine learning (ML) solutions are being drastically explored in this domain because they can potentially produce highly accurate results with access to historical data. However, experts in the environmental area are skeptical about adopting ML solutions in real-world applications and policy making due to their black-box nature. In contrast, despite having low accuracy sometimes, the existing traditional simulation model (e.g., CMAQ) are widely used and follows well-defined and transparent equations. Therefore, presenting the knowledge learned by the ML model can make it transparent as well as comprehensible. In addition, validating the ML model’s learning with the existing domain knowledge might aid in addressing their skepticism, building appropriate trust, and better utilizing ML models. In collaboration with three experts with an average of five years of research experience in the air pollution domain, we identified that feature (meteorological feature like wind) contribution, towards the final forecast as the major information to be verified with domain knowledge. In addition, the accuracy of ML models compared with traditional simulation models and raw wind trajectories are essential for domain experts to validate the feature contribution. Based on the identified information, we designed and developed AQX, a visual analytics system to help experts validate and verify the ML model’s learning with their domain knowledge. The system includes multiple coordinated views to present the contributions of input features at different levels of aggregation in both temporal and spatial dimensions. It also provides a performance comparison of ML and traditional models in terms of accuracy and spatial map, along with the animation of raw wind trajectories for the input period. We further demonstrated two case studies and conducted expert interviews with two domain experts to show the effectiveness and usefulness of AQX.
Reshika Palaniyappan Velumani, Meng Xia 0002, Jun Han 0010, Chaoli Wang 0001, Alexis Kai-Hon Lau, Huamin Qu
IUI3
2022 TSR-VFD: Generating temporal super-resolution for unsteady vector field data
Jun Han 0010, Chaoli Wang 0001
Comput. Graph.1
2022 SurfNet: Learning Surface Representations via Graph Convolutional Network
abstract
Abstract For scientific visualization applications, understanding the structure of a single surface (e.g., stream surface, isosurface) and selecting representative surfaces play a crucial role. In response, we propose SurfNet, a graph‐based deep learning approach for representing a surface locally at the node level and globally at the surface level. By treating surfaces as graphs, we leverage a graph convolutional network to learn node embedding on a surface. To make the learned embedding effective, we consider various pieces of information (e.g., position, normal, velocity) for network input and investigate multiple losses. Furthermore, we apply dimensionality reduction to transform the learned embeddings into 2D space for understanding and exploration. To demonstrate the effectiveness of SurfNet, we evaluate the embeddings in node clustering (node‐level) and surface selection (surface‐level) tasks. We compare SurfNet against state‐of‐the‐art node embedding approaches and surface selection methods. We also demonstrate the superiority of SurfNet by comparing it against a spectral‐based mesh segmentation approach. The results show that SurfNet can learn better representations at the node and surface levels with less training time and fewer training samples while generating comparable or better clustering and selection results.
Jun Han 0010, Chaoli Wang 0001
Comput. Graph. Forum1
2022 SSR-TVD: Spatial Super-Resolution for Time-Varying Data Analysis and Visualization
abstract
We present SSR-TVD, a novel deep learning framework that produces coherent spatial super-resolution (SSR) of time-varying data (TVD) using adversarial learning. In scientific visualization, SSR-TVD is the first work that applies the generative adversarial network (GAN) to generate high-resolution volumes for three-dimensional time-varying data sets. The design of SSR-TVD includes a generator and two discriminators (spatial and temporal discriminators). The generator takes a low-resolution volume as input and outputs a synthesized high-resolution volume. To capture spatial and temporal coherence in the volume sequence, the two discriminators take the synthesized high-resolution volume(s) as input and produce a score indicating the realness of the volume(s). Our method can work in the in situ visualization setting by downscaling volumetric data from selected time steps as the simulation runs and upscaling downsampled volumes to their original resolution during postprocessing. To demonstrate the effectiveness of SSR-TVD, we show quantitative and qualitative results with several time-varying data sets of different characteristics and compare our method against volume upscaling using bicubic interpolation and a solution solely based on CNN.
Jun Han 0010, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2022 STNet: An End-to-End Generative Framework for Synthesizing Spatiotemporal Super-Resolution Volumes
abstract
We present STNet, an end-to-end generative framework that synthesizes spatiotemporal super-resolution volumes with high fidelity for time-varying data. STNet includes two modules: a generator and a spatiotemporal discriminator. The input to the generator is two low-resolution volumes at both ends, and the output is the intermediate and the two-ending spatiotemporal super-resolution volumes. The spatiotemporal discriminator, leveraging convolutional long short-term memory, accepts a spatiotemporal super-resolution sequence as input and predicts a conditional score for each volume based on its spatial (the volume itself) and temporal (the previous volumes) information. We propose an unsupervised pre-training stage using cycle loss to improve the generalization of STNet. Once trained, STNet can generate spatiotemporal super-resolution volumes from low-resolution ones, offering scientists an option to save data storage (i.e., sparsely sampling the simulation output in both spatial and temporal dimensions). We compare STNet with the baseline bicubic+linear interpolation, two deep learning solutions ( SSR+TSF, STD), and a state-of-the-art tensor compression solution (TTHRESH) to show the effectiveness of STNet.
Jun Han 0010, Hao Zheng 0006, Danny Ziyi Chen, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2022 VCNet: A generative model for volume completion
abstract
We present VCNet, a new deep learning approach for volume completion by synthesizing missing subvolumes. Our solution leverages a generative adversarial network (GAN) that learns to complete volumes using the adversarial and volumetric losses. The core design of VCNet features a dilated residual block and long-term connection. During training, VCNet first randomly masks basic subvolumes (e.g., cuboids, slices) from complete volumes and learns to recover them. Moreover, we design a two-stage algorithm for stabilizing and accelerating network optimization. Once trained, VCNet takes an incomplete volume as input and automatically identifies and fills in the missing subvolumes with high quality. We quantitatively and qualitatively test VCNet with volumetric data sets of various characteristics to demonstrate its effectiveness. We also compare VCNet against a diffusion-based solution and two GAN-based solutions.
Jun Han 0010, Chaoli Wang 0001
Vis. Informatics1
2021 Hierarchical Self-supervised Learning for Medical Image Segmentation Based on Multi-domain Data Aggregation
Hao Zheng 0006, Jun Han 0010, Lin Yang 0003, Zhuo Zhao, Chaoli Wang 0001, Danny Ziyi Chen
MICCAI (1)2
2021 V2V: A Deep Learning Approach to Variable-to-Variable Selection and Translation for Multivariate Time-Varying Data
abstract
We present V2V, a novel deep learning framework, as a general-purpose solution to the variable-to-variable (V2V) selection and translation problem for multivariate time-varying data (MTVD) analysis and visualization. V2V leverages a representation learning algorithm to identify transferable variables and utilizes Kullback-Leibler divergence to determine the source and target variables. It then uses a generative adversarial network (GAN) to learn the mapping from the source variable to the target variable via the adversarial, volumetric, and feature losses. V2V takes the pairs of time steps of the source and target variable as input for training, Once trained, it can infer unseen time steps of the target variable given the corresponding time steps of the source variable. Several multivariate time-varying data sets of different characteristics are used to demonstrate the effectiveness of V2V, both quantitatively and qualitatively. We compare V2V against histogram matching and two other deep learning solutions (Pix2Pix and CycleGAN).
Jun Han 0010, Hao Zheng 0006, Yunhao Xing, Danny Ziyi Chen, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2020 SSR-VFD: Spatial Super-Resolution for Vector Field Data Analysis and Visualization
abstract
We present SSR-VFD, a novel deep learning framework that produces coherent spatial super-resolution (SSR) of three-dimensional vector field data (VFD). SSR-VFD is the first work that advocates a machine learning approach to generate high-resolution vector fields from low-resolution ones. The core of SSR-VFD lies in the use of three separate neural nets that take the three components of a low-resolution vector field as input and jointly output a synthesized high-resolution vector field. To capture spatial coherence, we take into account magnitude and angle losses in network optimization. Our method can work in the in situ scenario where VFD are down-sampled at simulation time for storage saving and these reduced VFD are upsampled back to their original resolution during postprocessing. To demonstrate the effectiveness of SSR-VFD, we show quantitative and qualitative results with several vector field data sets of different characteristics and compare our method against volume upscaling using bicubic interpolation, and two solutions based on CNN and GAN, respectively.
Shaojie Ye, Jun Han 0010, Hao Zheng 0006, Han Gao 0005, Danny Ziyi Chen, Jian-Xun Wang 0001, Chaoli Wang 0001
PacificVis3
2020 PQA-CNN: Towards Perceptual Quality Assured Single-Image Super-Resolution in Remote Sensing
abstract
Recent advances in remote sensing open up unprecedented opportunities to obtain a rich set of visual features of objects on the earth's surface. In this paper, we focus on a single-image super-resolution (SISR) problem in remote sensing, where the objective is to generate a reconstructed satellite image of high quality (i.e., a high spatial resolution) from a satellite image of relatively low quality. This problem is motivated by the lack of high quality satellite images in many remote sensing applications (e.g., due to the cost of high resolution sensors, communication bandwidth constraints, and historic hardware limitations). Two important challenges exist in solving our problem: i) it is not a trivial task to reconstruct a satellite image of high quality that meets the human perceptual requirement from a single low quality image; ii) it is challenging to rigorously quantify the uncertainty of the results of an SISR scheme in the absence of ground truth data. To address the above challenges, we develop PQA-CNN, a perceptual quality-assured conventional neural network framework, to reconstruct a high quality satellite image from a low quality one by designing novel uncertainty-driven neural network architectures and integrating an uncertainty quantification model with the framework. We evaluate PQA-CNN on a real-world remote sensing application on land usage classifications. The results show that PQA-CNN significantly outperforms the state-of-the-art super-resolution baselines in terms of accurately reconstructing high-resolution satellite images under various evaluation scenarios.
Yang Zhang 0031, Xiangyu Dong 0004, Md Tahmid Rashid, Lanyu Shang, Jun Han 0010, Daniel Yue Zhang, Dong Wang 0002
IWQoS5
2020 TransRes: A Deep Transfer Learning Approach to Migratable Image Super-Resolution in Remote Urban Sensing
abstract
Recent advances in remote sensing provide a powerful and scalable sensing paradigm to capture abundant visual information about the urban environments. We refer to such a sensing paradigm as remote urban sensing. In this paper, we focus on a migratable satellite image super-resolution problem in remote urban sensing applications. Our goal is to reconstruct satellite images of a high resolution in a target area where the high-resolution training data is not available by transferring a super-resolution model learned in a source area where such data is available. This problem is motivated by the limitation of current solutions that primarily rely on a rich set of high-resolution satellite images in the studied area that are not always available. Two important challenges exist in solving our problem: i) the target and source areas often have very different urban characteristics that prevent the direct application of a super-resolution model learned from the source area to the target area; ii) it is not a trivial task to ensure effective model migration with desirable quality without sufficient high quality training data. To address the above challenges, we develop TransRes, a deep adversarial transfer learning framework, to effectively reconstruct high-resolution satellite images without requiring any ground-truth training data from the studied area. We evaluate the TransRes framework using the real-world satellite imagery data collected from three different cities in Europe. The results show that TransRes consistently outperforms the state-of-the-art baselines by achieving the lowest perception errors under various application scenarios.
Yang Zhang 0031, Ruohan Zong, Jun Han 0010, Daniel Yue Zhang, Md Tahmid Rashid, Dong Wang 0002
SECON3
2020 FlowNet: A Deep Learning Framework for Clustering and Selection of Streamlines and Stream Surfaces
abstract
For effective flow visualization, identifying representative flow lines or surfaces is an important problem which has been studied. However, no work can solve the problem for both lines and surfaces. In this paper, we present FlowNet, a single deep learning framework for clustering and selection of streamlines and stream surfaces. Given a collection of streamlines or stream surfaces generated from a flow field data set, our approach converts them into binary volumes and then employs an autoencoder to learn their respective latent feature descriptors. These descriptors are used to reconstruct binary volumes for error estimation and network training. Once converged, the feature descriptors can well represent flow lines or surfaces in the latent space. We perform dimensionality reduction of these feature descriptors and cluster the projection results accordingly. This leads to a visual interface for exploring the collection of flow lines or surfaces via clustering, filtering, and selection of representatives. Intuitive user interactions are provided for visual reasoning of the collection with ease. We validate and explain our deep learning framework from multiple perspectives, demonstrate the effectiveness of FlowNet using several flow field data sets of different characteristics, and compare our approach against state-of-the-art streamline and stream surface selection algorithms.
Jun Han 0010, Jun Tao 0002, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2020 TSR-TVD: Temporal Super-Resolution for Time-Varying Data Analysis and Visualization
abstract
We present SSR-TVD, a novel deep learning framework that produces coherent spatial super-resolution (SSR) of time-varying data (TVD) using adversarial learning. In scientific visualization, SSR-TVD is the first work that applies the generative adversarial network (GAN) to generate high-resolution volumes for three-dimensional time-varying data sets. The design of SSR-TVD includes a generator and two discriminators (spatial and temporal discriminators). The generator takes a low-resolution volume as input and outputs a synthesized high-resolution volume. To capture spatial and temporal coherence in the volume sequence, the two discriminators take the synthesized high-resolution volume(s) as input and produce a score indicating the realness of the volume(s). Our method can work in the in situ visualization setting by downscaling volumetric data from selected time steps as the simulation runs and upscaling downsampled volumes to their original resolution during postprocessing. To demonstrate the effectiveness of SSR-TVD, we show quantitative and qualitative results with several time-varying data sets of different characteristics and compare our method against volume upscaling using bicubic interpolation and a solution solely based on CNN.
Jun Han 0010, Chaoli Wang 0001
IEEE Trans. Vis. Comput. Graph.1
2019 Biomedical Image Segmentation via Representative Annotation
abstract
Deep learning has been applied successfully to many biomedical image segmentation tasks. However, due to the diversity and complexity of biomedical image data, manual annotation for training common deep learning models is very timeconsuming and labor-intensive, especially because normally only biomedical experts can annotate image data well. Human experts are often involved in a long and iterative process of annotation, as in active learning type annotation schemes. In this paper, we propose representative annotation (RA), a new deep learning framework for reducing annotation effort in biomedical image segmentation. RA uses unsupervised networks for feature extraction and selects representative image patches for annotation in the latent space of learned feature descriptors, which implicitly characterizes the underlying data while minimizing redundancy. A fully convolutional network (FCN) is then trained using the annotated selected image patches for image segmentation. Our RA scheme offers three compelling advantages: (1) It leverages the ability of deep neural networks to learn better representations of image data; (2) it performs one-shot selection for manual annotation and frees annotators from the iterative process of common active learning based annotation schemes; (3) it can be deployed to 3D images with simple extensions. We evaluate our RA approach using three datasets (two 2D and one 3D) and show our framework yields competitive segmentation results comparing with state-of-the-art methods.
Hao Zheng 0006, Lin Yang 0003, Jianxu Chen 0001, Jun Han 0010, Yizhe Zhang 0001, Peixian Liang, Zhuo Zhao, Chaoli Wang 0001, Danny Ziyi Chen
AAAI4
2019 TransLand: An Adversarial Transfer Learning Approach for Migratable Urban Land Usage Classification using Remote Sensing
abstract
Urban land usage classification is a critical task in big data based smart city applications that aim to understand the social-economic land functions and physical land attributes in urban environments. This paper focuses on a migratable urban land usage classification problem using remote sensing data (i.e., satellite images). Our goal is to accurately classify the land usage of locations in a target city where the ground truth land usage data is not available by leveraging a classification model from a source city where such data is available. This problem is motivated by the limitation of current solutions that primarily rely on a rich set of ground-truth data for accurate model training, which encounters high annotation costs. Two important challenges exist in solving our problem: i) the target and source cities often have different urban characteristics that prevent the direct application of a model learned from the source city to the target city; ii) the complex visual features in satellite images make it non-trivial to “translate” the images from the target city to the source city for an accurate classification. To address the above challenges, we develop TransLand, an adversarial transfer learning framework to translate the satellite images from the target city to the source city for accurate land usage classification. We evaluate our scheme on the real-world satellite imagery and land usage datasets collected from live different cities in Europe. The results show that TransLand significantly outperforms the state-of-the-art land usage classification baselines in classifying the land usage of locations in a city.
Yang Zhang 0031, Ruohan Zong, Jun Han 0010, Hao Zheng 0006, Qiuwen Lou, Daniel Yue Zhang, Dong Wang 0002
IEEE BigData3
2019 HFA-Net: 3D Cardiovascular Image Segmentation with Asymmetrical Pooling and Content-Aware Fusion
Hao Zheng 0006, Lin Yang 0003, Jun Han 0010, Yizhe Zhang 0001, Peixian Liang, Zhuo Zhao, Chaoli Wang 0001, Danny Ziyi Chen
MICCAI (2)3