Zhiyu Yao

dblp:230/4609 · DBLP profile ↗
← Back
11ranked-venue papers
9as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Deep-Learning-Based Zero-Sample Gradient Guidance Spatial Resolution Enhancement for Microwave Radiometer in Fengyun-3D
abstract
For satellite brightness temperature images, researchers are constantly pursuing higher resolutions to obtain more detailed meteorological information. In this paper, a novel deep-learning-based modelling approach, named Zero-Sample Gradient Guidance Spatial Resolution Enhancement (ZSGRE), is developed explicitly for microwave radiometers. The detailed model, including mathematical derivation and key parameters, is presented. Subsequently, the proposed approach is applied in four scenarios: synthetic scene, simulated geographical brightness temperature, practical measurement of microwave radiometer in Fengyun-3D (FY-3D), and a cyclone analysis on the Atlantic. Compared with other methods, the proposed ZSGRE method improves 2.51% of SSIM (structural similarity), enhances 2.3 dB of PSNR (Peak Signal-to-Noise Ratio), and decreases 15.8% of IFOV (Instantaneous Field of View). Such applications demonstrate ZSGRE’s significant performance: zero-sample preparation and spatial resolution enhancement.
Minghao Feng, Weidong Hu, Yuming Bai, Zhiyu Yao, Vahid Rastinasab, Jian Shang
IEEE Trans. Geosci. Remote. Sens.4
2025 Physics-Constrained Automated Well-to-Seismic Tie Based on Time-Frequency Key Feature Point Matching
abstract
High-resolution characterization of subsurface hydrocarbon reservoirs critically depends on the integration of seismic and well log data. However, the precision of well-to-seismic tie procedures is often compromised by inaccuracies in time-depth conversion curves, stemming from imprecise migration velocity errors, scale discrepancy and the inherent domain disparity between seismic data (time-domain) and well log data (depth-domain). Traditional methods rely on iterative wavelet estimation and manual adjustments, leading to wavelet-depth curve coupling problems, reduced accuracy, and significant time consumption. To address these limitations, we propose a novel automated time-depth conversion curve correction method. This approach leverages key feature matching in the time-frequency domain of well log reflection coefficients and borehole-side seismic traces, incorporating dual constraints: stratigraphic constraints in the time domain and spectral notch points in frequency domain. By directly aligning with well log reflection coefficients, this time-frequency joint analysis eliminates the need for wavelet estimation as well as iterative and interactive processes. Experiments performed using both synthetic and field data demonstrate the validity and effectiveness of the proposed method compared to traditional iterative matching method and interactive commercial software.
Zhiyu Yao, Wenkai Lu, Weiheng Geng, Jialin Wang 0003
IEEE Trans. Geosci. Remote. Sens.1
2024 Mobile Attention: Mobile-Friendly Linear-Attention for Vision Transformers
abstract
Vision Transformers (ViTs) excel in computer vision tasks due to their ability to capture global context among tokens. However, their quadratic complexity $\mathcal{O}(N^2D)$ in terms of token number $N$ and feature dimension $D$ limits practical use on mobile devices, necessitating more mobile-friendly ViTs with reduced latency. Multi-head linear-attention is emerging as a promising alternative with linear complexity $\mathcal{O}(NDd)$, where $d$ is the per-head dimension. Still, more compute is needed as $d$ gets large for model accuracy. Reducing $d$ improves mobile friendliness at the expense of excessive small heads weak at learning valuable subspaces, ultimately impeding model capability. To overcome this efficiency-capability dilemma, we propose a novel Mobile-Attention design with a head-competition mechanism empowered by information flow, which prevents overemphasis on less important subspaces upon trivial heads while preserving essential subspaces to ensure Transformer's capability. It enables linear-time complexity on mobile devices by supporting a small per-head dimension $d$ for mobile efficiency. By replacing the standard attention of ViTs with Mobile-Attention, our optimized ViTs achieved enhanced model capacity and competitive performance in a range of computer vision tasks. Specifically, we have achieved remarkable reductions in latency on the iPhone 12. Code is available at https://github.com/thuml/MobileAttention.
Zhiyu Yao, Jian Wang 0066, Haixu Wu, Jingdong Wang 0001, Mingsheng Long
ICML1
2024 Recommender Transformers with Behavior Pathways
abstract
Sequential recommendation requires the recommender to capture the evolving behavior characteristics from logged user behavior data for accurate recommendations. Nevertheless, user behavior sequences are viewed as a script with multiple ongoing threads intertwined. We find that only a small set of pivotal behaviors can be evolved into the user's future action. As a result, the future behavior of the user is hard to predict. We conclude this characteristic for sequential behaviors of each user as thebehavior pathway. Different users have their unique behavior pathways. Among existing sequential models, transformers have shown great capacity in capturing global-dependent characteristics. However, these models mainly provide a dense distribution over all previous behaviors using the self-attention mechanism, making the final predictions overwhelmed by the trivial behaviors not adjusted to each user. In this paper, we build the Recommender Transformer (RETR) with a novel Pathway Attention mechanism. RETR can dynamically plan the behavior pathway specified for each user, and sparingly activate the network through this behavior pathway to effectively capture evolving patterns useful for recommendation. The key design is a learned binary route to prevent the behavior pathway from being overwhelmed by trivial behaviors. Pathway attention is model-agnostic and can be applied to a series of transformer-based models for sequential recommendation. We empirically evaluate RETR on seven intra-domain benchmarks and RETR yields state-of-the-art performance. On another five cross-domain benchmarks, RETR can capture more domain-invariant representations for sequential recommendation.
Zhiyu Yao, Xinyang Chen 0001, Qinyan Dai, Tanchao Zhu, Mingsheng Long
WWW1
2024 SeisLFMFlow: Seismic Common Image Gathers Enhancement Using Self-Supervised Optical Flow Estimation Based on Local Feature Matching
abstract
Seismic imaging technology, which analyzes seismic wave propagation and reflection to gather data on underground geological structures, is vital for geological exploration. Due to factors such as the anisotropy of subsurface media, migration velocity errors, and drift of seismic streamers in marine environments, observation points at the same position exhibit horizontal and vertical displacements in different common offset gathers (COGs), thereby diminishing stacking coherence and compromising imaging quality. Consequently, the nonflattened seismic events in common image gathers (CIGs) extracted from COGs can lead to false amplitude variations with offset. Traditional CIG enhancement methods like cross-correlation matching encounter challenges such as slow inference speed, limited accuracy, the capability to predict only a single directional displacement, and difficulty in parameter tuning. Therefore, based on optimizing local normalized cross-correlation matching, an interpretable deep learning method to enhance CIGs using a self-supervised optical flow estimation network is proposed. Experiments performed using both synthetic and field data demonstrate the validity and effectiveness of the method.
Zhiyu Yao, Weiheng Geng, Wenkai Lu
IEEE Trans. Geosci. Remote. Sens.1
2023 ModeRNN: Harnessing Spatiotemporal Mode Collapse in Unsupervised Predictive Learning
abstract
Learning predictive models for unlabeled spatiotemporal data is challenging in part because visual dynamics can be highly entangled, especially in real scenes. In this paper, we refer to the multi-modal output distribution of predictive learning as spatiotemporal modes. We find an experimental phenomenon named spatiotemporal mode collapse (STMC) on most existing video prediction models, that is, features collapse into invalid representation subspaces due to the ambiguous understanding of mixed physical processes. We propose to quantify STMC and explore its solution for the first time in the context of unsupervised predictive learning. To this end, we present ModeRNN, a decoupling-aggregation framework that has a strong inductive bias of discovering the compositional structures of spatiotemporal modes between recurrent states. We first leverage a set of dynamic slots with independent parameters to extract individual building components of spatiotemporal modes. We then perform a weighted fusion of slot features to adaptively aggregate them into a unified hidden representation for recurrent updates. Through a series of experiments, we show high correlation between STMC and the fuzzy prediction results of future video frames. Besides, ModeRNN is shown to better mitigate STMC and achieve the state of the art on five video prediction datasets.
Zhiyu Yao, Yunbo Wang, Haixu Wu, Jianmin Wang 0001, Mingsheng Long
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Spatial Resolution Matching of Microwave Radiometer Measurements Using Iterative Deconvolution With Close Loop Priors (ICLP)
abstract
Passive multi-frequency microwave sensors frequently struggle with difficulties of non-uniform spatial resolution among multiple channels. The raw measurements in the land-sea transition zone are seriously contaminated. Conventional analytical deconvolution techniques suffer from the trade-off between spatial resolution enhancement and noise amplification, leading to low data integrity in the practical spatial resolution matching application. In order to provide multi-channel microwave radiometer data with matching levels of spatial resolution, a method based on iterative deconvolution with close loop priors(ICLP) is proposed. Specifically, a destriping module is first utilized as pre-processing step to maintain high data integrity. Then, the close loop mechanism using sparse adaptive priors is proposed to balance the spatial resolution and data integrity enhancement. Also, progressively iterative deconvolution is introduced to realize controllable levels of spatial resolution enhancement(spatial resolution matching) for multi-channel data to reach a consistent level. Experiments performed using both simulated and actual Microwave Radiation Imager(MWRI) data demonstrate the validity and effectiveness of the method.
Zhiyu Yao, Weidong Hu, Zhiyan Feng, Yang Liu 0215, Leo P. Ligthart
IEEE Trans. Geosci. Remote. Sens.1
2022 VideoDG: Generalizing Temporal Relations in Videos to Novel Domains
abstract
This paper introduces video domain generalization where most video classification networks degenerate due to the lack of exposure to the target domains of divergent distributions. We observe that the global temporal features are less generalizable, due to the temporal domain shift that videos from other unseen domains may have an unexpected absence or misalignment of the temporal relations. This finding has motivated us to solve video domain generalization by effectively learning the local-relation features of different timescales that are more generalizable, and exploiting them along with the global-relation features to maintain the discriminability. This paper presents the VideoDG framework with two technical contributions. The first is a new deep architecture named the Adversarial Pyramid Network, which improves the generalizability of video features by capturing the local-relation, global-relation, and cross-relation features progressively. On the basis of pyramid features, the second contribution is a new and robust approach of adversarial data augmentation that can bridge different video domains by improving the diversity and quality of augmented data. We construct three video domain generalization benchmarks in which domains are divided according to different datasets, different consequences of actions, or different camera views, respectively. VideoDG consistently outperforms the combinations of previous video classification models and existing domain generalization methods on all benchmarks.
Zhiyu Yao, Yunbo Wang, Jianmin Wang 0001, Philip S. Yu, Mingsheng Long
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 MotionRNN: A Flexible Model for Video Prediction With Spacetime-Varying Motions
abstract
This paper tackles video prediction from a new dimension of predicting spacetime-varying motions that are incessantly changing across both space and time. Prior methods mainly capture the temporal state transitions but overlook the complex spatiotemporal variations of the motion itself, making them difficult to adapt to ever-changing motions. We observe that physical world motions can be decomposed into transient variation and motion trend, while the latter can be regarded as the accumulation of previous motions. Thus, simultaneously capturing the transient variation and the motion trend is the key to make spacetime-varying motions more predictable. Based on these observations, we propose the MotionRNN framework, which can capture the complex variations within motions and adapt to spacetime-varying scenarios. MotionRNN has two main contributions. The first is that we design the MotionGRU unit, which can model the transient variation and motion trend in a unified way. The second is that we apply the MotionGRU to RNN-based predictive models and indicate a new flexible video prediction architecture with a Motion Highway, which can significantly improve the ability to predict changeable motions and avoid motion vanishing for stacked multiple-layer predictive models. With high flexibility, this framework can adapt to a series of models for deterministic spatiotemporal prediction. Our MotionRNN can yield significant improvements on three challenging benchmarks for video prediction with spacetime-varying motions.
Haixu Wu, Zhiyu Yao, Jianmin Wang 0001, Mingsheng Long
CVPR2
2020 Multi-Task Learning of Generalizable Representations for Video Action Recognition
abstract
In classic video action recognition, labels may not contain enough information about the diverse video appearance and dynamics, thus, existing models that are trained under the standard supervised learning paradigm may extract less generalizable features. We evaluate these models under a cross-dataset experiment setting, as the above label bias problem in video analysis is even more prominent across different data sources. We find that using the optical flows as model inputs harms the generalization ability of most video recognition models.Based on these findings, we present a multi-task learning paradigm for video classification. Our key idea is to avoid label bias and improve the generalization ability by taking data as its own supervision or supervising constraints on the data. First, we take the optical flows and the RGB frames by taking them as auxiliary supervisions, and thus naming our model as Reversed Two-Stream Networks (Rev2Net). Further, we collaborate the auxiliary flow prediction task and the frame reconstruction task by introducing a new training objective to Rev2Net, named Decoding Discrepancy Penalty (DDP), which constraints the discrepancy of the multi-task features in a self-supervised manner. Rev2Net is shown to be effective on the classic action recognition task. It specifically shows a strong generalization ability in the cross-dataset experiments.
Zhiyu Yao, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001, Philip S. Yu, Jia-Guang Sun 0001
ICME1
2020 Unsupervised Transfer Learning for Spatiotemporal Predictive Networks
abstract
This paper explores a new research problem of unsupervised transfer learning across multiple spatiotemporal prediction tasks. Unlike most existing transfer learning methods that focus on fixing the discrepancy between supervised tasks, we study how to transfer knowledge from a zoo of unsupervisedly learned models towards another predictive network. Our motivation is that models from different sources are expected to understand the complex spatiotemporal dynamics from different perspectives, thereby effectively supplementing the new task, even if the task has sufficient training samples. Technically, we propose a differentiable framework named transferable memory. It adaptively distills knowledge from a bank of memory states of multiple pretrained RNNs, and applies it to the target network via a novel recurrent structure called the Transferable Memory Unit (TMU). Compared with finetuning, our approach yields significant improvements on three benchmarks for spatiotemporal prediction, and benefits the target task even from less relevant pretext ones.
Zhiyu Yao, Yunbo Wang, Mingsheng Long, Jianmin Wang 0001
ICML1