Ziyun Wang 0001

dblp:60/7180-1 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-9803-7949ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Continuous-Time Human Motion Field from Event Cameras
Ziyun Wang 0001, Ruijun Zhang, Yufu Wang, Kostas Daniilidis
ICCV1
2025 EqNIO: Subequivariant Neural Inertial Odometry
abstract
Neural network-based odometry using accelerometer and gyroscope readings from a single IMU can achieve robust, and low-drift localization capabilities, through the use of _neural displacement priors (NDPs)_. These priors learn to produce denoised displacement measurements but need to ignore data variations due to specific IMU mount orientation and motion directions, hindering generalization. This work introduces EqNIO, which addresses this challenge with _canonical displacement priors_, i.e., priors that are invariant to the orientation of the gravity-aligned frame in which the IMU data is expressed. We train such priors on IMU measurements, that are mapped into a learnable canonical frame, which is uniquely defined via three axes: the first is gravity, making the frame gravity aligned, while the second and third are predicted from IMU data. The outputs (displacement and covariance) are mapped back to the original gravity-aligned frame. To maximize generalization, we find that these learnable frames must transform equivariantly with global gravity-preserving roto-reflections from the subgroup $O_g(3)\subset O(3)$, acting on the trajectory, rendering the NDP $O(3)$-_subequivariant_. We tailor specific linear, convolutional, and non-linear layers that commute with the actions of the group. Moreover, we introduce a bijective decomposition of angular rates into vectors that transform similarly to accelerations, allowing us to leverage both measurement types. Natively, angular rates would need to be inverted upon reflection, unlike acceleration, which hinders their joint processing. We highlight EqNIO's flexibility and generalization capabilities by applying it to both filter-based (TLIO), and end-to-end (RONIN) architectures, and outperforming existing methods that use _soft equivariance from auxiliary losses or data augmentation on various datasets. We believe this work paves the way for low-drift and generalizable neural inertial odometry on edge devices. The project details and code can be found at [https://github.com/RoyinaJayanth/EqNIO](https://github.com/RoyinaJayanth/EqNIO).
Royina Karegoudra Jayanth, Yinshuang Xu, Ziyun Wang 0001, Evangelos Chatzipantazis, Kostas Daniilidis, Daniel Gehrig
ICLR3
2024 Motion-Prior Contrast Maximization for Dense Continuous-Time Motion Estimation
Friedhelm Hamann, Ziyun Wang 0001, Ioannis Asmanis, Kenneth Chaney, Guillermo Gallego 0002, Kostas Daniilidis
ECCV (3)2
2024 Track Everything Everywhere Fast and Robustly
Yunzhou Song, Jiahui Lei, Ziyun Wang 0001, Lingjie Liu, Kostas Daniilidis
ECCV (3)3
2024 Un-EVIMO: Unsupervised Event-Based Independent Motion Segmentation
Ziyun Wang 0001, Kostas Daniilidis
ECCV (16)1
2024 TRAM: Global Trajectory and Motion of 3D Humans from in-the-Wild Videos
Yufu Wang, Ziyun Wang 0001, Lingjie Liu, Kostas Daniilidis
ECCV (11)2
2022 EvAC3D: From Event-Based Apparent Contours to 3D Models via Continuous Visual Hulls
Ziyun Wang 0001, Kenneth Chaney, Kostas Daniilidis
ECCV (7)1
2021 Geodesic-HOF: 3D Reconstruction Without Cutting Corners
Ziyun Wang 0001, Eric Mitchell, Volkan Isler, Daniel D. Lee
AAAI1
2021 EventGAN: Leveraging Large Scale Image Datasets for Event Cameras
abstract
Event cameras provide a number of benefits over traditional cameras, such as the ability to track incredibly fast motions, high dynamic range, and low power consumption. However, their application into computer vision problems, many of which are primarily dominated by deep learning solutions, has been limited by the lack of labeled training data for events. In this work, we propose a method which leverages the existing labeled data for images by simulating events from a pair of temporal image frames, using a convolutional neural network. We train this network on pairs of images and events, using an adversarial discriminator loss and a pair of cycle consistency losses. The cycle consistency losses utilize a pair of pre-trained self-supervised networks which perform optical flow estimation and image reconstruction from events, and constrain our network to generate events which result in accurate outputs from both of these networks. Trained fully end to end, our network learns a generative model for events from images without the need for accurate modeling of the motion in the scene, exhibited by modeling based methods, while also implicitly modeling event noise. Using this simulator, we train a pair of downstream networks on object detection and 2D human pose estimation from events, using simulated data from large scale image datasets, and demonstrate the networks' abilities to generalize to datasets with real events. The code and dataset in this paper are available here: https://github.com/alexzzhu/EventGAN.
Alex Zihao Zhu, Ziyun Wang 0001, Kaung Khant, Kostas Daniilidis
ICCP2
2020 Surface Hof: Surface Reconstruction From A Single Image Using Higher Order Function Networks
abstract
We address the problem of reconstructing a high-resolution surface representing an object from a single image. We present Surface HOF, which takes an image of an object as input and generates a mapping function for surface generation. The mapping function takes samples from a canonical domain and maps each sample to a local tangent plane on the 3D reconstruction of the object. By efficiently learning a continuous mapping function, the surface can be generated at arbitrary resolution in contrast to other methods which generate fixed resolution outputs. Experiments show that Surface HOF is more accurate and uses more efficient representations than other state of the art methods for surface reconstruction. Surface HOF is also easier to train: it requires minimal input pre-processing and output post-processing and generates surface representations that are more parameter efficient. Its accuracy and convenience make Surface HOF an appealing method for single image reconstruction.
Ziyun Wang 0001, Volkan Isler, Daniel D. Lee
ICIP1