Liangyu Xu

dblp:18/10367 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0009-0001-0847-6594ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FIRE: Flexible Integration of Data Quality Ratings for Effective Pretraining
abstract
Selecting high-quality data can improve the pretraining efficiency of large language models (LLMs).Existing methods generally rely on heuristic techniques or single quality signals, limiting their ability to evaluate data quality comprehensively.In this work, we propose FIRE, a flexible and scalable framework for integrating multiple data quality raters, which allows for a comprehensive assessment of data quality across various dimensions.FIRE aligns multiple quality signals into a unified space, and integrates diverse data quality raters to provide a comprehensive quality signal for each data point.Further, we introduce a progressive data selection scheme based on FIRE that iteratively refines the selection of high-quality data points.Extensive experiments show that FIRE outperforms other data selection methods and significantly boosts pretrained model performance across a wide range of downstream tasks, while requiring less than 37.5% tokens needed by the Random baseline to reach the target performance.
Liangyu Xu, Xuemiao Zhang, Feiyu Duan, Rongxiang Weng, Jingang Wang
EMNLP1
2025 Hypergraph-Guided Multimodal Prototype for Remote Sensing Scene Understanding
abstract
Noticeable achievements have been made in entity-level perception tasks (e.g., object detection) in remote sensing (RS) image interpretation. But for RS images carrying rich content, individual perception cannot well obtain the interaction patterns between entities. The recognition of relationships between entities is the key to deeply understanding RS scenes. In this article, we propose a hypergraph-guided multimodal prototype network (HMPNet), which performs relation recognition by matching relation representations with multimodal predicate prototypes. To overcome the imbalance of modal information in the matching process, a multimodal calibration strategy is devised, taking into account the image subprototype and text subprototype, which makes prediction results more reliable. Meanwhile, to align image and text subprototypes and explore relevant semantic patterns, the multimodal hypergraph is constructed to efficiently capture the associations between heterogeneous prototypes. Experimental results show that the performance of our model can reach the state-of-the-art (SOTA) level on the RS scene graph generation (SGG) task.
Chubo Deng, Qiwei Yan, Liangyu Xu, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Ringmo-SenseV2: Remote Sensing Foundation Model for Spatiotemporal Prediction Based on Multisource Heterogeneous Time-Series Data
abstract
The rapid development of Remote Sensing (RS) technology has generated a vast amount of heterogeneous time series data from various sources, including drone videos, satellite time-series images, and multi-object trajectories. Effectively processing and analyzing this multi-source heterogeneous data for accurate spatiotemporal prediction is crucial in fields such as environmental protection and disaster response. In this paper, we propose a universal predictive foundation model named Ringmo-SenseV2 to learn the general evolutionary patterns of RS elements from massive heterogeneous data. Ringmo- SenseV2 features a Mixture-of-Heterogeneous-Experts (MoHE) Transformer, which unifies the modeling of multi-source heterogeneous time-series data. Additionally, to better capture the complex dependencies across different spatiotemporal locations, we introduce a hypergraph translator, treating embeddings of different spatiotemporal locations as nodes and employing hypergraph convolution for information propagation. Furthermore, to enhance the model’s adaptability to different evolution speeds during pre-training, we implement the Adaptive tube Masking (AM) strategy, which controls prediction difficulty by adaptively setting mask proportions for sequences with varying evolution speeds. Extensive experiments demonstrate that Ringmo-SenseV2 exhibits outstanding performance across various RS prediction tasks. Further tests on scene graph generation for RS images showcase the model’s ability to extract image features, thereby enhancing image perception tasks.
Liangyu Xu, Wanxuan Lu, Leiyi Hu, Heming Yang 0003, Chubo Deng, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Detecting Objects as Cascade Corners
abstract
The corner-based detection paradigm enjoys the potential to produce high-quality boxes. But the development is constrained by three factors: 1) Hard to match corners. Heuristic corner matching algorithms can lead to incorrect boxes, especially when similar-looking objects co-occur. 2) Poor instance context. Two separate corners preserve few instance semantics, so it is difficult to guarantee getting both two class-specific corners on the same heatmap channel. 3) Unfriendly backbone. The training cost of the hourglass network is high. Accordingly, we build a novel corner-based framework, named Corner2Net. To achieve the corner-matching-free manner, we devise the cascade corner pipeline which progressively predicts the associated corner pair in two steps instead of synchronously searching two independent corners via parallel heads. Corner2Net decouples corner localization and object classification. Both two corners are class-agnostic and the instance-specific bottom-right corner further simplifies its search space. Meanwhile, RoI features with rich semantics are extracted for classification. Popular backbones (e.g., ResNeXt) can be easily connected to Corner2Net. Experimental results on COCO show Corner2Net surpasses all existing corner-based detectors by a large margin in accuracy and speed.
Haorao Wei, Jinze Yang, Liangyu Xu, Lu Fang 0001
ECAI5
2024 SDL-MVS: View Space and Depth Deformable Learning Paradigm for Multiview Stereo Reconstruction in Remote Sensing
abstract
Research on multiview stereo (MVS) based on remote sensing images has promoted the development of large-scale urban 3-D reconstruction. However, remote sensing multiview image data suffer from the problems of occlusion and uneven brightness between views during acquisition, which leads to the problem of blurred details in depth estimation. To solve the above problem, we reexamine the deformable learning method in the MVS task and propose a novel paradigm based on view space and depth deformable learning (SDL-MVS), aiming to learn deformable interactions of features in different view spaces and deformably model the depth ranges and intervals to enable high accurate depth estimation. Specifically, to solve the problem of view noise caused by occlusion and uneven brightness, we propose a progressive space deformable sampling (PSS) mechanism, which performs deformable learning of sampling points in the 3-D frustum space and the 2-D image space in a progressive manner to embed source features to the reference feature adaptively. To further optimize the depth, we introduce depth hypothesis deformable discretization (DHD), which achieves precise positioning of the depth prior by adaptively adjusting the depth range hypothesis and performing deformable discretization of the depth interval hypothesis. Finally, our SDL-MVS achieves explicit modeling of occlusion and uneven brightness faced in MVS through the deformable learning paradigm of view space and depth, achieving accurate multiview depth estimation. Extensive experiments on LuoJia-MVS and WHU datasets show that our SDL-MVS reaches state-of-the-art performance. It is worth noting that our SDL-MVS achieves a mean absolute error (MAE) error of 0.086 and an accuracy of 98.9% for Acc$_{\lt 0.6\,\text {m}}$and 98.9% for Acc$_{\lt 3-\text {interval}}$on the LuoJia-MVS dataset under the premise of three views as input.
Yongqiang Mao, Hanbo Bi, Liangyu Xu, Kaiqiang Chen, Zhirui Wang 0003, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 TAFormer: A Unified Target-Aware Transformer for Video and Motion Joint Prediction in Aerial Scenes
abstract
As drone technology advances, using unmanned aerial vehicles for aerial surveys has become the dominant trend in modern low-altitude remote sensing. The surge in aerial video data necessitates accurate prediction for future scenarios and motion states of the interested target, particularly in applications like traffic management and disaster response. Existing video prediction methods focus solely on predicting future scenes (video frames), suffering from the neglect of explicitly modeling target’s motion states, which is crucial for aerial video interpretation. To address this issue, we introduce a novel task called Target-Aware Aerial Video Prediction, aiming to simultaneously predict future scenes and motion states of the target. Further, we design a model specifically for this task, named TAFormer, which provides a unified modeling approach for both video and target motion states. Specifically, we introduce Spatiotemporal Attention (STA), which decouples the learning of video dynamics into spatial static attention and temporal dynamic attention, effectively modeling the scene appearance and motion. Additionally, we design an Information Sharing Mechanism (ISM), which elegantly unifies the modeling of video and target motion by facilitating information interaction through two sets of messenger tokens. Moreover, to alleviate the difficulty of distinguishing targets in blurry predictions, we introduce Target-Sensitive Gaussian Loss (TSGL), enhancing the model’s sensitivity to both target’s position and content. Extensive experiments on UAV123VP and VisDroneVP (derived from single-object tracking datasets) demonstrate the exceptional performance of TAFormer in target-aware video prediction, showcasing its adaptability to the additional requirements of aerial video interpretation for target awareness.
Liangyu Xu, Wanxuan Lu, Yongqiang Mao, Hanbo Bi, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
abstract
Extrapolating future weather radar echoes from past observations is a complex task vital for precipitation nowcasting. The spatial morphology and temporal evolution of radar echoes exhibit a certain degree of correlation, yet they also possess independent characteristics. Existing methods learn unified spatial and temporal representations in a highly coupled feature space, emphasizing the correlation between spatial and temporal features but neglecting the explicit modeling of their independent characteristics, which may result in mutual interference between them. To effectively model the spatiotemporal dynamics of radar echoes, we propose a spatial-frequency-temporal correlation-decoupling transformer (SFTformer). The model leverages stacked multiple SFT-Blocks to not only mine the correlation of the spatiotemporal dynamics of echo cells but also avoid the mutual interference between the temporal modeling and the spatial morphology refinement by decoupling them. Furthermore, inspired by the practice that weather forecast experts effectively review historical echo evolution to make accurate predictions, SFTfomer incorporates a joint training paradigm for historical echo sequence reconstruction and future echo sequence prediction. Experimental results on the HKO-7 dataset and ChinaNorth-2021 dataset demonstrate the superior performance of SFTfomer in short-term (1 h), mid-term (2 h), and long-term (3 h) precipitation nowcasting.
Liangyu Xu, Wanxuan Lu, Fanglong Yao, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 RingMo-Sense: Remote Sensing Foundation Model for Spatiotemporal Prediction via Spatiotemporal Evolution Disentangling
abstract
Remote sensing spatiotemporal prediction aims to infer future trends from historical spatiotemporal data, e.g., videos and time series images, has a broad application prospect in many fields. The foundation model is a promising research direction for spatiotemporal information mining because of its robust feature extraction capability, and has made rapid progress in natural scenes. Nevertheless, due to the spatially multi-scale and temporally multi-scale properties in remote sensing data, these methods still encounter bottlenecks when applied to remote sensing. Therefore, we propose a foundation model for remote sensing spatiotemporal prediction via spatiotemporal evolution decoupling, abbreviated as RingMo-Sense. Considering spatial affinity, temporal continuity, and spatiotemporal interaction, we construct spatial, temporal, and spatiotemporal triple-branch prediction networks. Specifically, we use parameter-sharing and progressive joint training strategies to achieve stable long-range prediction and parameter reduction simultaneously. In addition, we build a remote sensing spatiotemporal dataset by collecting various remote sensing videos and time series images. The experimental results on six downstream spatiotemporal tasks demonstrate that the proposed model yields competitive performance.
Fanglong Yao, Wanxuan Lu, Heming Yang 0003, Liangyu Xu, Leiyi Hu, Nayu Liu, Chubo Deng, Deke Tang, Changshuo Chen, Xian Sun 0001, Kun Fu 0001
IEEE Trans. Geosci. Remote. Sens.4