EDBT 2026 Demo / reviewers in the wild / expert
Lixiang Ru
dblp:232/2035
· DBLP profile ↗
16ranked-venue papers
6as first author
14since 2021 · last 2025
0000-0002-9129-2453ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography EstimationabstractFeature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority of existing methods focus on improving coarse feature representation rather than the fine-matching module. Prior fine-matching techniques, which rely on point-to-patch matching probability expectation or direct regression, often lack precision and do not guarantee the continuity of feature points across sequential images. To address this limitation, this paper concentrates on enhancing the fine-matching module in the semi-dense matching framework. We employ a lightweight and efficient homography estimation network to generate the perspective mapping between patches obtained from coarse matching. This patch-to-patch approach achieves the overall alignment of two patches, resulting in a higher sub-pixel accuracy by incorporating additional constraints. By leveraging the homography estimation between patches, we can achieve a dense matching result with low computational cost. Extensive experiments demonstrate that our method achieves higher accuracy compared to previous semi-dense matchers. Meanwhile, our dense matching results exhibit similar end-point-error accuracy compared to previous dense matchers while maintaining semi-dense efficiency. Xiaolong Wang 0013, Lei Yu 0005, Jiangwei Lao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Yu Zhang 0018, Ming Yang 0007 |
AAAI | 5 |
| 2025 | SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingabstractOpen-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two challenges. 1) Existing RS semantic categories are limited, particularly for pixel-level interpretation datasets. 2) Distinguishing among diverse RS spatial regions solely by language space is challenging due to the dense and intricate spatial distribution in open-world RS imagery. To address the first issue, we develop a fine-grained RS interpretation dataset, Sky-SA, which contains 183,375 high-quality local image-text pairs with full-pixel manual annotations, covering 1,763 category labels, exhibiting richer semantics and higher density than previous datasets. Afterwards, to solve the second issue, we introduce the vision-centric principle for vision-language modeling. Specifically, in the pre-training stage, the visual self-supervised paradigm is incorporated into image-text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across open-category texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarthOV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https://github.com/zqcrafts/SkySense-O. Qi Zhu 0010, Jiangwei Lao, Deyi Ji, Lixiang Ru, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Dong Liu 0002, Feng Zhao 0004 |
CVPR | 7 |
| 2025 | CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidanceabstractSemi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we propose a novel pipeline, CasP, which leverages cascaded correspondence priors for guidance. Specifically, the matching stage is decomposed into two progressive phases, bridged by a region-based selective cross-attention mechanism designed to enhance feature discriminability. In the second phase, one-to-one matches are determined by restricting the search range to the one-to-many prior areas identified in the first phase. Additionally, this pipeline benefits from incorporating high-level features, which helps reduce the computational costs of low-level feature extraction. The acceleration gains of CasP increase with higher resolution, and our lite model achieves a speedup of $\sim2.2\times$ at a resolution of 1152 compared to the most efficient method, ELoFTR. Furthermore, extensive experiments demonstrate its superiority in geometric estimation, particularly with impressive cross-domain generalization. These advantages highlight its potential for latency-sensitive and high-robustness applications, such as SLAM and UAV systems. Code is available at https://github.com/pq-chen/CasP. Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yingying Pei, Xinyi Liu 0002, Yongxiang Yao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002 |
ICCV | 8 |
| 2025 | SkySense V2: A Unified Foundation Model for Multi-Modal Remote SensingabstractThe multi-modal remote sensing foundation model (MM-RSFM) has significantly advanced various Earth observation tasks, such as urban planning, environmental monitoring, and natural disaster management. However, most existing approaches generally require the training of separate backbone networks for each data modality, leading to redundancy and inefficient parameter utilization. Moreover, prevalent pre-training methods typically apply self-supervised learning (SSL) techniques from natural images without adequately accommodating the characteristics of remote sensing (RS) images, such as the complicated semantic distribution within a single RS image. In this work, we present SkySense V2, a unified MM-RSFM that employs a single transformer backbone to handle multiple modalities. This backbone is pre-trained with a novel SSL strategy tailored to the distinct traits of RS data. In particular, SkySense V2 incorporates an innovative adaptive patch merging module and learnable modality prompt tokens to address challenges related to varying resolutions and limited feature diversity across modalities. In additional, we incorporate the mixture of experts (MoE) module to further enhance the performance of the foundation model. SkySense V2 demonstrates impressive generalization abilities through an extensive evaluation involving 16 datasets over 7 tasks, outperforming SkySense by an average of 1.8 points. Lixiang Ru, Lei Yu 0005, Yansheng Li 0001, Jingdong Chen |
ICCV | 2 |
| 2025 | ARGenSeg: Image Segmentation with Autoregressive Image Generation ModelabstractWe propose a novel AutoRegressive Generation-based paradigm for image Segmentation (ARGenSeg), achieving multimodal understanding and pixel-level perception within a unified framework.
Prior works integrating image segmentation into multimodal large language models (MLLMs) typically employ either boundary points representation or dedicated segmentation heads.
These methods rely on discrete representations or semantic prompts fed into task-specific decoders, which limits the ability of the MLLM to capture fine-grained visual details.
To address these challenges,
we introduce a segmentation framework for MLLM based on image generation, which naturally produces dense masks for target objects.
We leverage MLLM to output visual tokens and detokenize them into images using an universal VQ-VAE,
making the segmentation fully dependent on the pixel-level understanding of the MLLM.
To reduce inference latency,
we employ a next-scale-prediction strategy to generate required visual tokens in parallel.
Extensive experiments demonstrate that our method surpasses prior state-of-the-art approaches on multiple segmentation datasets with a remarkable boost in inference speed, while maintaining strong understanding capabilities. Xiaolong Wang 0013, Lixiang Ru, Kaixiang Ji, Jingdong Chen, Jun Zhou 0011 |
NeurIPS | 2 |
| 2025 | TransWCD: Scene-Adaptive Joint Constrained Framework for Weakly Supervised Change DetectionabstractChange detection (CD) based on deep learning typically requires costly pixel-level change labels. Recently, weakly supervised CD (WSCD) has emerged as a more label-efficient approach, using scene-level (i.e., image-level) labels to identify pixel-level changes in bitemporal images. With only scene-level labels, existing WSCD methods are typically trained as scene-level change classification models. However, these methods often suffer from label-prediction inconsistency, with false changes frequently predicted in unchanged scenes. To address this issue, we propose TransWCD-SA, an end-to-end classifier-predictor framework. TransWCD-SA consists of a hierarchical transformer-based TransWCD classifier and a scene-adaptive (SA) predictor. This classifier-predictor framework is trained with two-stage joint constraints in an end-to-end learning manner. Specifically, the TransWCD classifier integrates hierarchical transformer blocks and multiscale class activation maps (CAMs), capturing pixel-level changes across various scales under weak supervision. The SA predictor dynamically introduces different pixel-level information for scenes labeled as changed and unchanged. Furthermore, a scene gated constraint is proposed as a penalty for label-prediction inconsistency, which is activated by the Dirac delta function and rectify features of mispredicted pixels in the embedding space. We validate the effectiveness of TransWCD-SA on three datasets: Wuhan University building CD (WHU-CD), learning, vision, and remote sensing CD (LEVIR-CD), and DSIFN-CD, demonstrating significant improvement. The code is available athttps://github.com/zhenghuizhao/TransWCD. Zhenghui Zhao, Lixiang Ru, Chen Wu 0003, Di Wang 0023 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryabstractPrior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primar-ily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pretrained on a curated multimodal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multimodal spatiotemporal encoder taking temporal sequences of opti-cal and Synthetic Aperture Radar (SAR) data as input. This encoder is pretrained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multimodal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thor-ough evaluation encompassing 16 datasets over 7 tasks, from single- to multimodal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pretrained weights to facilitate future research and Earth Observation applications. Xin Guo 0010, Jiangwei Lao, Bo Dang 0002, Lei Yu 0005, Lixiang Ru, Liheng Zhong, Dingxiang Hu, Huimei He, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002, Yansheng Li 0001 |
CVPR | 6 |
| 2024 | POA: Pre-training Once for Models of All Sizes
Xin Guo 0010, Jiangwei Lao, Lei Yu 0005, Lixiang Ru, Jian Wang 0108, Guo Ye, Huimei He, Jingdong Chen, Ming Yang 0007 |
ECCV (3) | 5 |
| 2024 | Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual RecognitionabstractLong-tailed recognition (LTR) aims to learn balanced models from extremely unbalanced training data. Fine-tuning pretrained foundation models has recently emerged as a promising research direction for LTR. However, we observe that the fine-tuning process tends to degrade the intrinsic representation capability of pretrained models and lead to model bias towards certain classes, thereby hindering the overall recognition performance. To unleash the intrinsic representation capability of pretrained foundation models, in this work, we propose a new Parameter-Efficient Complementary Expert Learning (PECEL) for LTR. Specifically, PECEL consists of multiple experts, where individual experts are trained via Parameter-Efficient Fine-Tuning (PEFT) and encouraged to learn different expertise on complementary sub-categories via the proposed sample-aware logit adjustment loss. By aggregating the predictions of different experts, PECEL effectively achieves a balanced performance on long-tailed classes. Nevertheless, learning multiple experts generally introduces extra trainable parameters. To ensure parameter efficiency, we further propose a parameter sharing strategy which decomposes and shares the parameters in each expert. Extensive experiments on 4 LTR benchmarks show that the proposed PECEL can effectively learn multiple complementary experts without increasing the trainable parameters and achieve new state-of-the-art performance. Lixiang Ru, Xin Guo 0010, Lei Yu 0005, Jiangwei Lao, Jian Wang 0108, Jingdong Chen, Yansheng Li 0001, Ming Yang 0007 |
ACM Multimedia | 1 |
| 2023 | Token Contrast for Weakly-Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) using image-level labels typically utilizes Class Activation Map (CAM) to generate the pseudo labels. Limited by the local structure perception of CNN, CAM usually cannot identify the integral object regions. Though the recent Vision Transformer (ViT) can remedy this flaw, we observe it also brings the over-smoothing issue, i.e., the final patch tokens incline to be uniform. In this work, we propose Token Contrast (ToCo) to address this issue and further explore the virtue of ViT for WSSS. Firstly, motivated by the observation that intermediate layers in ViT can still retain semantic diversity, we designed a Patch Token Contrast module (PTC). PTC supervises the final patch tokens with the pseudo token relations derived from intermediate layers, allowing them to align the semantic regions and thus yield more accurate CAM. Secondly, to further differentiate the low-confidence regions in CAM, we devised a Class Token Contrast module (CTC) inspired by the fact that class tokens in ViT can capture high-level semantics. CTC facilitates the representation consistency between uncertain local regions and global objects by contrasting their class tokens. Experiments on the PASCAL VOC and MS COCO datasets show the proposed ToCo can remarkably surpass other single-stage competitors and achieve comparable performance with state-of-the-art multi-stage methods. Code is available at https://github.com/rulixiang/ToCo. Lixiang Ru, Heliang Zheng, Yibing Zhan, Bo Du 0001 |
CVPR | 1 |
| 2022 | Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersabstractWeakly-supervised semantic segmentation (WSSS) with image-level labels is an important and challenging task. Due to the high training efficiency, end-to-end solutions for WSSS have received increasing attention from the community. However, current methods are mainly based on convolutional neural networks and fail to explore the global information properly, thus usually resulting in incomplete object regions. In this paper, to address the aforementioned problem, we introduce Transformers, which naturally integrate global information, to generate more integral initial pseudo labels for end-to-end WSSS. Motivated by the inherent consistency between the self-attention in Transformers and the semantic affinity, we propose an Affinity from Attention (AFA) module to learn semantic affinity from the multi-head self-attention (MHSA) in Transformers. The learned affinity is then leveraged to refine the initial pseudo labels for segmentation. In addition, to efficiently derive reliable affinity labels for supervising AFA and ensure the local consistency of pseudo labels, we devise a Pixel-Adaptive Refinement module that incorporates low-level image appearance information to refine the pseudo labels. We perform extensive experiments and our method achieves 66.0% and 38.9% mIoU on the PASCAL VOC 2012 and MS COCO 2014 datasets, respectively, significantly outperforming recent end-to-end methods and several multi-stage competitors. Code is available at https://github.com/rulixiang/afa. Lixiang Ru, Yibing Zhan, Baosheng Yu, Bo Du 0001 |
CVPR | 1 |
| 2022 | Weakly-Supervised Semantic Segmentation with Visual Words Learning and Hybrid Pooling
Lixiang Ru, Bo Du 0001, Yibing Zhan, Chen Wu 0003 |
Int. J. Comput. Vis. | 1 |
| 2021 | Learning Visual Words for Weakly-Supervised Semantic SegmentationabstractCurrent weakly-supervised semantic segmentation (WSSS) methods with image-level labels mainly adopt class activation maps (CAM) to generate the initial pseudo labels. However, CAM usually only identifies the most discriminative object extents, which is attributed to the fact that the network doesn't need to discover the integral object to recognize image-level labels. In this work, to tackle this problem, we proposed to simultaneously learn the image-level labels and local visual word labels. Specifically, in each forward propagation, the feature maps of the input image will be encoded to visual words with a learnable codebook. By enforcing the network to classify the encoded fine-grained visual words, the generated CAM could cover more semantic regions. Besides, we also proposed a hybrid spatial pyramid pooling module that could preserve local maximum and global average values of feature maps, so that more object details and less background were considered. Based on the proposed methods, we conducted experiments on the PASCAL VOC 2012 dataset. Our proposed method achieved 67.2% mIoU on the val set and 67.3% mIoU on the test set, which outperformed recent state-of-the-art methods. Lixiang Ru, Bo Du 0001, Chen Wu 0003 |
IJCAI | 1 |
| 2021 | Multi-Temporal Scene Classification and Scene Change Detection With Correlation Based FusionabstractClassifying multi-temporal scene land-use categories and detecting their semantic scene-level changes for remote sensing imagery covering urban regions could straightly reflect the land-use transitions. Existing methods for scene change detection rarely focus on the temporal correlation of bi-temporal features, and are mainly evaluated on small scale scene change detection datasets. In this work, we proposed a CorrFusion module that fuses the highly correlated components in bi-temporal feature embeddings. We first extract the deep representations of the bi-temporal inputs with deep convolutional networks. Then the extracted features will be projected into a lower-dimensional space to extract the most correlated components and compute the instance-level correlation. The cross-temporal fusion will be performed based on the computed correlation in CorrFusion module. The final scene classification results are obtained with softmax layers. In the objective function, we introduced a new formulation to calculate the temporal correlation more efficiently and stably. The detailed derivation of backpropagation gradients for the proposed module is also given. Besides, we presented a much larger scale scene change detection dataset with more semantic categories and conducted extensive experiments on this dataset. The experimental results demonstrated that our proposed CorrFusion module could remarkably improve the multi-temporal scene classification and scene change detection results. Lixiang Ru, Bo Du 0001, Chen Wu 0003 |
IEEE Trans. Image Process. | 1 |
| 2019 | Scene Change Detection VIA Deep Convolution Canonical Correlation Analysis Neural NetworkabstractScene change detection is the process of identifying the differences between the multi-temporal image scenes at the semantic level, which has significant potential in the application of urban development and land management. In this paper, we propose a novel deep convolution canonical correlation analysis neural network (DCCANet) architecture, which could consider the spectral-spatial-temporal correlation feature for scene change detection in remote sensing images. For this purpose, we put together the convolutional neural networks (CNNs) and the deep canonical correlation analysis (DCCA) into the end-to-end network. The CNN could get the spectral-spatial feature information for scene representation, while the later could enhance the temporal correlation by nonlinear high- dimensional transformations between the multi-temporal image scenes for scene change detection. Experiments with high-resolution remote sensing image scene datasets demonstrated that our proposed approach can get a better performance in scene classification and change detection. Bo Du 0001, Lixiang Ru, Chen Wu 0003, Hui Luo 0012 |
IGARSS | 3 |
| 2019 | Unsupervised Deep Slow Feature Analysis for Change Detection in Multi-Temporal Remote Sensing ImagesabstractChange detection has been a hotspot in the remote sensing technology for a long time. With the increasing availability of multi-temporal remote sensing images, numerous change detection algorithms have been proposed. Among these methods, image transformation methods with feature extraction and mapping could effectively highlight the changed information and thus has a better change detection performance. However, the changes of multi-temporal images are usually complex, and the existing methods are not effective enough. In recent years, the deep network has shown its brilliant performance in many fields, including feature extraction and projection. Therefore, in this paper, based on the deep network and slow feature analysis (SFA) theory, we proposed a new change detection algorithm for multi-temporal remotes sensing images called deep SFA (DSFA). In the DSFA model, two symmetric deep networks are utilized for projecting the input data of bi-temporal imagery. Then, the SFA module is deployed to suppress the unchanged components and highlight the changed components of the transformed features. The change vector analysis pre-detection is employed to find unchanged pixels with high confidence as training samples. Finally, the change intensity is calculated with chi-square distance and the changes are determined by threshold algorithms. The experiments are performed on two real-world data sets and a public hyperspectral data set. The visual comparison and the quantitative evaluation have shown that DSFA could outperform the other state-of-the-art algorithms, including other SFA-based and deep learning methods. Bo Du 0001, Lixiang Ru, Chen Wu 0003, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |