Xingdi Zhang

dblp:232/1995 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2026
0000-0002-9426-0721ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 Exploring 3D Unsteady Flow using 6D Observer Space Interactions
abstract
Visualizing and analyzing 3D unsteady flow fields is a very challenging task. We approach this problem by leveraging the mathematical foundations of 3D observer fields to explore and analyze 3D flows in reference frames that are more suitable to visual analysis than the input reference frame. We design novel interactive tools for determining, filtering, and combining reference frames for observer-aware 3D unsteady flow visualization. We represent the space of reference frame motions in a 3D spatial domain via a 6D parameter space, in which every observer is a time-dependent curve. Our framework supports operations in this 6D observer space by separately focusing on two 3D subspaces, for 3D translations, and 3D rotations, respectively. We show that this approach facilitates a variety of interactions with 3D flow fields. Building on the interactive selection of observers, we furthermore introduce novel techniques such as observer-aware streamline- and pathline-filtering as well as observer-aware isosurface animations of scalar fluid properties for the enhanced visualization and analysis of 3D unsteady flows. We discuss the theoretical underpinnings as well as practical implementation considerations of our approach, and demonstrate the benefits of its 6+1D observer-based methodology on several 3D unsteady flow datasets.
Xingdi Zhang, Amani Ageeli, Thomas Theußl, Markus Hadwiger, Peter Rautek
IEEE Trans. Vis. Comput. Graph.1
2025 VortexTransformer: End-to-End Objective Vortex Detection in 2D Unsteady Flow Using Transformers
abstract
Abstract Vortex structures play a pivotal role in understanding complex fluid dynamics, yet defining them rigorously remains challenging. One hard criterion is that a vortex detector must be objective, i.e., it needs to be indifferent to reference frame transformations. We propose VortexTransformer, a novel deep learning approach using point transformer architectures to directly extract vortex structures from pathlines. Unlike traditional methods that rely on grid‐based velocity fields in the Eulerian frame, our approach operates entirely on a Lagrangian representation of the flow field (i.e., pathlines), enabling objective identification of both strong and weak vortex structures. To train VortexTransformer, we generate a large synthetic dataset using parametric flow models to simulate diverse vortex configurations, ensuring a robust ground truth. We compare our method against CNN and U‐Net architectures, applying the trained models to real‐world flow datasets. VortexTransformer is an end‐to‐end detector, which means that reference frame transformations as well as vortex detection are handled implicitly by the network, demonstrating the ability to extract vortex boundaries without the need for parameters such as arbitrary thresholds, or an explicit definition of a vortex. Our method offers a new approach to determining objective vortex labels by using the objective pairwise distances of material points for vortex detection and is adaptable to various flow conditions.
Xingdi Zhang, Peter Rautek, Markus Hadwiger
Comput. Graph. Forum1
2025 Enhancing Material Boundary Visualizations in 2D Unsteady Flow through Local Reference Frame Transformations
abstract
Abstract We present a novel technique for the extraction, visualization, and analysis of material boundaries and Lagrangian coherent structures (LCS) in 2D unsteady flow fields relative to local reference frame transformations. In addition to the input flow field, we leverage existing methods for computing reference frames adapted to local fluid features, in particular those that minimize the observed time derivative. Although, by definition, transforming objective tensor fields between reference frames does not change the tensor field, we show that transforming objective tensors, such as the finite‐time Lyapunov exponent (FTLE) or Lagrangian‐averaged vorticity deviation (LAVD), or the second‐order rate‐of‐strain tensor, into local reference frames that are naturally adapted to coherent fluid structures has several advantages: (1) The transformed fields enable analyzing LCS in space‐time visualizations that are adapted to each structure; (2) They facilitate extracting geometric features, such as iso‐surfaces and ridge lines, in a straightforward manner with high accuracy. The resulting visualizations are characterized by lower geometric complexity and enhanced topological fidelity. To demonstrate the effectiveness of our technique, we measure geometric complexity and compare it with iso‐surfaces extracted in the conventional reference frame. We show that the decreased geometric complexity of the iso‐surfaces in the local reference frame, not only leads to improved geometric and topological results, but also to a decrease in computation time.
Xingdi Zhang, Peter Rautek, Thomas Theußl, Markus Hadwiger
Comput. Graph. Forum1
2024 Pix4Point: Image Pretrained Standard Transformers for 3D Point Cloud Understanding
abstract
While Transformers have achieved impressive success in natural language processing and computer vision, their performance on 3D point clouds is relatively poor. This is mainly due to the limitation of Transformers: a demanding need for extensive training data. Unfortunately, in the realm of 3D point clouds, the availability of large datasets is a challenge, exacerbating the issue of training Transformers for 3D tasks. In this work, we solve the data issue of point cloud Transformers from two perspectives: (i) introducing more inductive bias to reduce the dependency of Transformers on data, and (ii) relying on cross-modality pretraining. More specifically, we first present Progressive Point Patch Embedding and present a new point cloud Transformer model namely PViT. PViT shares the same backbone as Transformer but is shown to be less hungry for data, enabling Transformer to achieve performance comparable to the state-of-the-art. Second, we formulate a simple yet effective pipeline dubbed Pix4Point that allows harnessing Transformers pretrained in the image domain to enhance downstream point cloud understanding. This is achieved through a modality-agnostic Transformer backbone with the help of a tokenizer and decoder specialized in the different domains. Pretrained on a large number of widely available images, significant gains of PViT are observed in the tasks of 3D point cloud classification, part segmentation, and semantic segmentation on ScanObjectNN, ShapeNetPart, and S3DIS, respectively. Our code and models are available at https: //github.com/guochengqian/Pix4Point.
Guocheng Qian, Abdullah Hamdi, Xingdi Zhang, Bernard Ghanem
3DV3
2024 Vortex Lens: Interactive Vortex Core Line Extraction using Observed Line Integral Convolution
abstract
This paper describes a novel method for detecting and visualizing vortex structures in unsteady 2D fluid flows. The method is based on an interactive local reference frame estimation that minimizes the observed time derivative of the input flow field v(x,t). A locally optimal reference frame w($x,t$) assists the user in the identification of physically observable vortex structures inObserved Line Integral Convolution(LIC) visualizations. The observed LIC visualizations are interactively computed and displayed in a user-steered vortex lens region, embedded in the context of a conventional LIC visualization outside the lens. The locally optimal reference frame is then used to detect observed critical points, where v = w, which are used to seed vortex core lines. Each vortex core line is computed as a solution of the ordinary differential equation (ODE) ˙$w(t)=w(w(t),t)$, with an observed critical point as initial condition$(w(t_0),t_0)$. During integration, we enforce a strict error bound on the difference between the extracted core line and the integration of a path line of the input vector field, i.e., a solution to the ODE ˙$v(t)=v(v(t),t)$. We experimentally verify that this error depends on the step size of the core line integration. This ensures that our method extracts Lagrangian vortex core lines that are the simultaneous solution of both ODEs with a numerical error that is controllable by the integration step size. We show the usability of our method in the context of an interactive system using a lens metaphor, and evaluate the results in comparison to state-of-the-art vortex core line extraction methods
Peter Rautek, Xingdi Zhang, Bernhard Woschizka, Thomas Theußl, Markus Hadwiger
IEEE Trans. Vis. Comput. Graph.2
2022 Interactive Exploration of Physically-Observable Objective Vortices in Unsteady 2D Flow
abstract
State-of-the-art computation and visualization of vortices in unsteady fluid flow employ objective vortex criteria, which makes them independent of reference frames or observers. However, objectivity by itself, although crucial, is not sufficient to guarantee that one can identify physically-realizable observers that would perceive or detect the same vortices. Moreover, a significant challenge is that a single reference frame is often not sufficient to accurately observe multiple vortices that follow different motions. This paper presents a novel framework for the exploration and use of an interactively-chosen set of observers, of the resulting relative velocity fields, and of objective vortex structures. We show that our approach facilitates the objective detection and visualization of vortices relative to well-adapted reference frame motions, while at the same time guaranteeing that these observers are in fact physically realizable. In order to represent and manipulate observers efficiently, we make use of the low-dimensional vector space structure of the Lie algebra of physically-realizable observer motions. We illustrate that our framework facilitates the efficient choice and guided exploration of objective vortices in unsteady 2D flow, on planar as well as on spherical domains, using well-adapted reference frames.
Xingdi Zhang, Markus Hadwiger, Thomas Theußl, Peter Rautek
IEEE Trans. Vis. Comput. Graph.1
2019 DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color Image
abstract
In this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and can be trained end-to-end. With a modified encoder-decoder structure, our network effectively fuses the dense color image and the sparse LiDAR depth. To address outdoor specific challenges, our network predicts a confidence mask to handle mixed LiDAR signals near foreground boundaries due to occlusion, and combines estimates from the color image and surface normals with learned attention maps to improve the depth accuracy especially for distant areas. Extensive experiments demonstrate that our model improves upon the state-of-the-art performance on KITTI depth completion benchmark. Ablation study shows the positive impact of each model components to the final performance, and comprehensive analysis shows that our model generalizes well to the input with higher sparsity or from indoor scenes.
Jiaxiong Qiu, Zhaopeng Cui, Yinda Zhang 0001, Xingdi Zhang, Shuaicheng Liu, Bing Zeng 0001, Marc Pollefeys
CVPR4
2018 Multi-exposure Fusion With JPEG Compression Guidance
abstract
Construct a High Dynamic Range (HDR) image is the primary method to solve the information loss caused by insufficient dynamic range of cameras. We propose a technique for fusing a bracketed low dynamic range (LDR) image sequence of varying exposures into an HDR image, skipping the physically-based HDR assembly step. Traditionally approaches often rely on complicated algorithms to select good regions from the input LDR images for the fusion. However, we found that the selection strategy can purely base on JPEG compression bits, bypassing the calculations of image low-level features, such as image gradients, local saturations, over/under exposure evaluations, as long as the input LDR image is compressed by the JPEG formats. In this way, lots of computations can be saved. In particular, we extract the coding bits from the intermediate product of the JPEG. The coding bits of blocks can be modified as the weights for the exposure fusion. Well-exposed regions often require higher bits for the compression while overexposure or saturated regions often correspond to lower bits. The objective and subjective evaluations demonstrate the effectiveness of our method.
Xingdi Zhang, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
VCIP1