Mengke Yuan

dblp:163/0469 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
16since 2021 · last 2025
0000-0001-9277-2654ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 DTESR: Remote Sensing Imagery Super-Resolution With Dynamic Reference Textures Exploitation
abstract
Reference-based remote sensing super-resolution (RefRS-SR) method shows great potential for improving both spatial resolution and coverage area of remote sensing images, by which high-resolution (HR) reference images can supplement fine details for low-resolution (LR) but wide coverage images. However, most RefRS-SR methods treat the reference as a static template and unidirectionally transfer the high-frequency information to the LR input. To address the issue of inefficient and inaccurate guided super-resolving, we propose a new RefRS-SR method with dynamic reference textures exploitation dubbed DTESR. The key referenced restoration (Ref Restoration) module consists of three components: correlation generation, texture enhancement and refinement (TER), and adaptive similarity-based fusion to progressively reconstruct high correlation and delicate textures for the LR input. Specifically, both the LR input and reference features are utilized for precise correlation generation. Next, both features are enhanced and refined with the most suitable reference under the guidance of the correlation map. Moreover, a learnable fusion method is designed to maintain the consistency of adjacent pixels. These operations will be iteratively applied to the three reconstruction scales to promote the exploitation of the Ref features. Through comprehensive quantitative and qualitative evaluations, our experimental results demonstrate that DTESR surpasses the current state-of-the-art RefRS-SR methods.
Jingliang Guo, Mengke Yuan, Tong Wang 0013, Zhifeng Li 0001, Xiaohong Jia 0001, Dong-Ming Yan 0001
IEEE Geosci. Remote. Sens. Lett.2
2025 Pattern Integration and Enhancement Vision Transformer for Self-Supervised Learning in Remote Sensing
abstract
Recent self-supervised learning (SSL) methods have demonstrated impressive results in learning visual representations from unlabeled remote sensing (RS) images. However, most RS images predominantly consist of scenographic scenes containing multiple ground objects without explicit foreground targets, which limits the performance of existing SSL methods that focus on foreground targets. This raises the question: Is there a method that can automatically aggregate similar objects within scenographic RS images, thereby enabling models to differentiate knowledge embedded in various geospatial patterns for improved feature representation? In this work, we present the pattern integration and enhancement vision transformer (PIEViT), a novel SSL framework designed specifically for RS imagery. PIEViT utilizes a teacher-student architecture to address both image-level and patch-level tasks. It employs a proposed, geospatial pattern cohesion (GPC) module to explore the natural clustering of patches, enhancing the differentiation of individual features. A feature integration projection (FIP) module is employed to further refine masked token reconstruction using geospatially clustered patches. We validated PIEViT across multiple downstream tasks, including object detection, semantic segmentation, and change detection. Experiments demonstrated that PIEViT enhances the representation of internal patch features, providing significant improvements over existing self-supervised baselines. It achieves excellent results in object detection, land cover classification, and change detection, underscoring its robustness, generalization, and transferability for RS image interpretation tasks.
Kaixuan Lu, Ruiqian Zhang, Xiao Huang 0003, Yuxing Xie, Xiaogang Ning, Hanchao Zhang, Mengke Yuan, Pan Zhang 0001, Tao Wang 0119, Tongkui Liao
IEEE Trans. Geosci. Remote. Sens.7
2025 ResGEM: Multi-Scale Graph Embedding Network for Residual Mesh Denoising
abstract
Mesh denoising is a crucial technology that aims to recover a high-fidelity 3D mesh from a noise-corrupted one. Deep learning methods, particularly graph convolutional networks (GCNs) based mesh denoisers, have demonstrated their effectiveness in removing various complex real-world noises while preserving authentic geometry. However, it is still a quite challenging work to faithfully regress uncontaminated normals and vertices on meshes with irregular topology. In this article, we propose a novel pipeline that incorporates two parallel normal-aware and vertex-aware branches to achieve a balance between smoothness and geometric details while maintaining the flexibility of surface topology. We introduce ResGEM, a new GCN, with multi-scale embedding modules and residual decoding structures to facilitate normal regression and vertex modification for mesh denoising. To effectively extract multi-scale surface features while avoiding the loss of topological information caused by graph pooling or coarsening operations, we encode the noisy normal and vertex graphs using four edge-conditioned embedding modules (EEMs) at different scales. This allows us to obtain favorable feature representations with multiple receptive field sizes. Formulating the denoising problem into a residual learning problem, the decoder incorporates residual blocks to accurately predict true normals and vertex offsets from the embedded feature space. Moreover, we propose novel regularization terms in the loss function that enhance the smoothing and generalization ability of our network by imposing constraints on normal fidelity and consistency. Comprehensive experiments have been conducted to demonstrate the superiority of our method over the state-of-the-art on both synthetic and real-scanned datasets.
Mengke Yuan, Mingyang Zhao 0001, Jianwei Guo 0003, Dong-Ming Yan 0001
IEEE Trans. Vis. Comput. Graph.2
2024 Dynamic Loss Decay Based Robust Oriented Object Detection on Remote Sensing Images with Noisy Labels
Guozhang Liu, Ting Liu 0010, Mengke Yuan, Guangxing Yang, Tao Wang 0119, Tongkui Liao
ICPR (29)3
2024 Water Body Extraction from SAR and Multi-Source Data Using Siamese Network-Based Segmentation
abstract
Extracting water bodies using C-band synthetic aperture radar (SAR) images has high reliability. However, the limited backscatter sensitivity of SAR data makes it difficult to differentiate between water and other features such as roads, building shadows, and more. To tackle these issues, we have implemented several strategies: Firstly, we have augmented our input data by incorporating polarized VV and VH data, as well as additional modalities such as DTM, DSM, and land use information. Through rigorous ablation studies, we have validated the efficacy of these data modalities, even in the presence of potential inaccuracies. Secondly, we have constructed a Siamese network for feature extraction and fusion across multiple modalities. Experimental results reveal that our proposed framework surpasses traditional feature extraction methods, yielding a 0.8% increase in F1 score. Finally, we have employed techniques such as stochastic multi-scale training, two-stage training, and model ensemble to further enhance performance. Notably, our method secured the second position in the 2024 IEEE GRSS Data Fusion Contest (DFC) Track 1 test phase, achieving an F1 score of 79.170%.
Ting Liu 0010, Mengke Yuan, Chaoran Lu, Kaixuan Lu, Baochai Peng, Heyang Duan, Pan Zhang 0001, Tao Wang 0119, Tongkui Liao
IGARSS2
2024 HeightFormer: Single-Imagery Height Estimation Transformer With Bilateral Feature Pyramid Fusion
abstract
Despite their ill-posedness and inherent ambiguity, recent deep learning approaches have demonstrated promising capability to estimate plausible height information from single spaceborne and airborne imagery. However, accurately predicting the height and preserving the rich geometric detailing of aerial images with limited resolution and complex structural variations remains a challenge. To address these issues, we introduce a novel transformer-based architecture for single-imagery height estimation (SIHE) dubbed as HeightFormer. Specifically, the building-block multiscale vision transformer (MViT) constitutes the encoder and decoder of HeightFormer to facilitate the capturing of long-range dependencies across a feature pyramid. Furthermore, we propose the bilateral feature pyramid fusion scheme, which consists of step-by-step and one-stop decoder feature map augmentation, to enhance global and local information reconstruction. The stepwise fusion module (SFM) iteratively fuses encoder and decoder features, while the multiscale fusion module (MFM) combines the final decoder feature with multiscale encoder features. In the end, the Heightbins module is designed to generate the attention map and the adaptive bin width. Then, the bin centers at each pixel are linearly combined as the final estimated height. Extensive experiments validate the effectiveness of HeightFormer on the Vaihingen dataset, the Potsdam dataset, and the DFC2019 dataset. Compared with the state-of-the-art, our method improves accuracy metrics and provides the ability to preserve structure and details. Building height estimation, transformer, attention, progressive refinement.
Jiangyan Wu, Mengke Yuan, Tong Wang 0013, Xiaohong Jia 0001, Dong-Ming Yan 0001
IEEE Geosci. Remote. Sens. Lett.2
2023 Fine-Grained Building Roof Instance Segmentation Based on Domain Adapted Pretraining and Composite Dual-Backbone
abstract
The diversity of building architecture styles of global cities situated on various landforms, the degraded optical imagery affected by clouds and shadows, and the significant inter-class imbalance of roof types pose challenges for designing a robust and accurate building roof instance segmentor. To address these issues, we propose an effective framework to fulfill semantic interpretation of individual buildings with high-resolution optical satellite imagery. Specifically, the leveraged domain adapted pretraining strategy and composite dual-backbone greatly facilitates the discriminative feature learning. Moreover, new data augmentation pipeline, stochastic weight averaging (SWA) training and instance segmentation based model ensemble in testing are utilized to acquire additional performance boost. Experiment results show that our approach ranks in the first place of the 2023 IEEE GRSS Data Fusion Contest (DFC) Track 1 test phase (mAP50:50.6%). Note-worthily, we have also explored the potential of multimodal data fusion with both optical satellite imagery and SAR data.
Guozhang Liu, Baochai Peng, Ting Liu 0010, Pan Zhang 0001, Mengke Yuan, Chaoran Lu, Ningning Cao, Simin Huang, Tao Wang 0119
IGARSS5
2023 Hgdnet: A Height-Hierarchy Guided Dual-Decoder Network for Single View Building Extraction and Height Estimation
abstract
Unifying the correlative single-view satellite image building extraction and height estimation tasks indicates a promising way to share representations and acquire generalist model for large-scale urban 3D reconstruction. However, the common spatial misalignment between building footprints and stereo-reconstructed nDSM height labels incurs degraded performance on both tasks. To address this issue, we propose a Height-hierarchy Guided Dual-decoder Network (HGDNet) to estimate building height. Under the guidance of synthesized discrete height-hierarchy nDSM, auxiliary height-hierarchical building extraction branch enhance the height estimation branch with implicit constraints, yielding an accuracy improvement of more than 6% on the DFC 2023 track2 dataset. Additional two-stage cascade architecture is adopted to achieve more accurate building extraction. Experiments on the DFC 2023 Track 2 dataset shows the superiority of the proposed method in building height estimation (δ1:0.8012), instance extraction (AP50:0.7730), and the final average score 0.7871 ranks in the first place in test phase.
Chaoran Lu, Ningning Cao, Pan Zhang 0001, Ting Liu 0010, Baochai Peng, Guozhang Liu, Mengke Yuan, Simin Huang, Tao Wang 0119
IGARSS7
2023 Deep unfolding multi-scale regularizer network for image denoising
abstract
Existing deep unfolding methods unroll an optimization algorithm with a fixed number of steps, and utilize convolutional neural networks (CNNs) to learn data-driven priors. However, their performance is limited for two main reasons. Firstly, priors learned in deep feature space need to be converted to the image space at each iteration step, which limits the depth of CNNs and prevents CNNs from exploiting contextual information. Secondly, existing methods only learn deep priors at the single full-resolution scale, so ignore the benefits of multi-scale context in dealing with high level noise. To address these issues, we explicitly consider the image denoising process in the deep feature space and propose the deep unfolding multi-scale regularizer network (DUMRN) for image denoising. The core of DUMRN is the feature-based denoising module (FDM) that directly removes noise in the deep feature space. In each FDM, we construct a multi-scale regularizer block to learn deep prior information from multi-resolution features. We build the DUMRN by stacking a sequence of FDMs and train it in an end-to-end manner. Experimental results on synthetic and real-world benchmarks demonstrate that DUMRN performs favorably compared to state-of-the-art methods.
Jingzhao Xu, Mengke Yuan, Dong-Ming Yan 0001, Tieru Wu
Comput. Vis. Media2
2023 Illumination Guided Attentive Wavelet Network for Low-Light Image Enhancement
abstract
Deep convolutional neural networks have recently been applied to improve the quality of low-light images and have achieved promising results. However, most existing methods cannot suppress noise during the enhancement process effectively, resulting in unknown artifacts and color distortions. In addition, these methods do not fully utilize illumination information and perform poorly under extremely low-light condition. To alleviate these problems, we propose theillumination guided attentive wavelet network(IGAWN) for low-light image enhancement (LLIE). Considering that the wavelet transform can separate high-frequency noise and desired low-frequency content effectively, we enhance low-light images in the frequency domain. By integrating attention mechanisms with wavelet transform, we develop the attentive wavelet transform to capture more important wavelet features, which enables the desired content to be enhanced and the redundant noise to be suppressed. To improve the image enhancement performance under extremely low-light environment, we extract illumination information from the input images and exploit it as the guidance for image enhancement through the frequency feature transform (FFT) layer. The proposed FFT layer generates frequency-aware affine transformation from the estimated illumination information, which can adaptively modulate the image features of different frequencies. Extensive experiments on synthetic and real-world datasets demonstrate that our IGAWN performs favorably against state-of-the-art LLIE methods.
Jingzhao Xu, Mengke Yuan, Dong-Ming Yan 0001, Tieru Wu
IEEE Trans. Multim.2
2022 GeoROS: Georeferenced Real-time Orthophoto Stitching with Unmanned Aerial Vehicle
abstract
Simultaneous orthophoto stitching during the flight of Unmanned Aerial Vehicles (UAV) can greatly promote the practicability and instantaneity of diverse applications such as emergency disaster rescue, digital agriculture, and cadastral survey, which is of remarkable interest in aerial photogrammetry. However, the inaccurately estimated camera poses and the intuitive fusion strategy of existing methods lead to misalignment and distortion artifacts in orthophoto mosaics. To address these issues, we propose a Georeferenced Real-time Orthophoto Stitching method (GeoROS), which can achieve efficient and accurate camera pose estimation through exploiting geolocation information in monocular visual simultaneous localization and mapping (SLAM) and fuse transformed images via orthogonality-preserving criterion. Specifically, in the SLAM process, georeferenced tracking is employed to acquire high-quality initial camera poses with a geolocation based motion model and facilitate non-linear pose optimization. Meanwhile, we design a georeferenced mapping scheme by introducing robust geolocation constraints in joint optimization of camera poses and the position of landmarks. Finally, aerial images warped with localized cameras are fused by considering both the orthogonality of camera orientation relative to the ground plane and the pixel centrality to fulfill global orthorectification. Besides, we construct two datasets with global navigation satellite system (GNSS) information of different scenarios and validate the superiority of our GeoROS method compared with state-of-the-art methods in accuracy and efficiency.
Guangze Gao, Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
IROS2
2022 Progressive polarization based reflection removal via realistic training data generation
Youxin Pang, Mengke Yuan, Qiang Fu 0002, Peiran Ren, Dong-Ming Yan 0001
Pattern Recognit.2
2022 LIST: low illumination scene text detector with automatic feature enhancement
Mengke Yuan, Tong Wang 0013, Peiran Ren, Dong-Ming Yan 0001
Vis. Comput.2
2022 Triple-strip attention mechanism-based natural disaster images classification and segmentation
Mengke Yuan, Jiaming Gu, Weiliang Meng, Shibiao Xu, Xiaopeng Zhang 0001
Vis. Comput.2
2021 Customized Summarizations of Visual Data Collections
abstract
Abstract We propose a framework to generate customized summarizations of visual data collections, such as collections of images, materials, 3D shapes, and 3D scenes. We assume that the elements in the visual data collections can be mapped to a set of vectors in a feature space, in which a fitness score for each element can be defined, and we pose the problem of customized summarizations as selecting a subset of these elements. We first describe the design choices a user should be able to specify for modeling customized summarizations and propose a corresponding user interface. We then formulate the problem as a constrained optimization problem with binary variables and propose a practical and fast algorithm based on the alternating direction method of multipliers (ADMM). Our results show that our problem formulation enables a wide variety of customized summarizations, and that our solver is both significantly faster than state‐of‐the‐art commercial integer programming solvers and produces better solutions than fast relaxation‐based solvers.
Mengke Yuan, Bernard Ghanem, Dong-Ming Yan 0001, Baoyuan Wu, Xiaopeng Zhang 0001, Peter Wonka
Comput. Graph. Forum1
2021 Tensor-Based Reliable Multiview Similarity Learning for Robust Spectral Clustering on Uncertain Data
abstract
Similarity graph learning is the most key technique for multiview spectral clustering. However, existing methods fail when applied to uncertain data contaminated with various types of noise in an open environment. Due to the damaged structure by noise, unreliable similar relationships are learned, which extends similarity inconsistency among views. Moreover, the high-order correlation hidden in graphs are ignored generally. To address these problems, we propose a reliable similarity learning scheme for multiview clustering on uncertain data. This method can significantly improve spectral clustering performance in a noisy environment, and the contributions of our scheme include the following three aspects: 1) Uncertain data subspace reconstruction and adaptive graph learning are combined to construct a view-specific graph from high-quality recovered data, thus improving robustness. 2) A low-rank tensor constraint is utilized to facilitate multiview fusion, where the latent high-order correlation among view graphs will be fully explored when learning the consensus graph structure. 3) Data recovery, view-specific graphs, and latent consensus tensor structure are assembled into a unified framework, to be optimized jointly for mutual benefit. Our study also develops an efficient algorithm for obtaining overall solutions. The experimental results on several datasets demonstrate that our proposed approach shows significant improvements in robustness and evaluation metrics over the comparison methods.
Ao Li 0002, Jiajia Chen 0004, Mengke Yuan, Shibiao Xu, Guanglu Sun
IEEE Trans. Reliab.5
2019 Fast and Error-Bounded Space-Variant Bilateral Filtering
Mengke Yuan, Longquan Dai, Dong-Ming Yan 0001, Liqiang Zhang 0001, Jun Xiao 0005, Xiaopeng Zhang 0001
J. Comput. Sci. Technol.1
2019 Interpreting and Extending the Guided Filter via Cyclic Coordinate Descent
abstract
The guided filter (GF) is a widely used smoothing tool in computer vision and image processing. However, to the best of our knowledge, few papers investigate the mathematical connection between this filter and the least-squares optimization. In this paper, we first interpret the guided filter as the cyclic coordinate descent (CCD) solver of a least-squares objective function. This discovery implies an extension approach to generalize the guided filter since we can change the least-squares objective function and define new filters as the first pass iteration of the CCD solver of modified objective functions. In addition, referring to the iterative minimizing procedure of the CCD, we can derive new rolling filtering schemes. So, we are reasonable to say that our discovery not only reveals an approach to design new GF-like filters adapting to specific requirements of applications but also offers thorough explanations for two rolling filtering schemes of the guided filter as well as the method to extend them. Experiments prove our new proposed filters and rolling filtering schemes could produce state-of-the-art results.
Longquan Dai, Mengke Yuan, Yuan Xie 0006, Xiaopeng Zhang 0001, Jinhui Tang 0001
IEEE Trans. Image Process.2
2017 Hardware-Efficient Guided Image Filtering for Multi-label Problem
abstract
The Guided Filter (GF) is well-known for its linear complexity. However, when filtering an image with an n-channel guidance, GF needs to invert an n × n matrix for each pixel. To the best of our knowledge existing matrix inverse algorithms are inefficient on current hardwares. This shortcoming limits applications of multichannel guidance in computation intensive system such as multi-label system. We need a new GF-like filter that can perform fast multichannel image guided filtering. Since the optimal linear complexity of GF cannot be minimized further, the only way thus is to bring all potentialities of current parallel computing hardwares into full play. In this paper we propose a hardware-efficient Guided Filter (HGF), which solves the efficiency problem of multichannel guided image filtering and yields competent results when applying it to multi-label problems with synthesized polynomial multichannel guidance. Specifically, in order to boost the filtering performance, HGF takes a new matrix inverse algorithm which only involves two hardware-efficient operations: element-wise arithmetic calculations and box filtering. In order to break the linear model restriction, HGF synthesizes a polynomial multichannel guidance to introduce nonlinearity. Benefiting from our polynomial guidance and hardware-efficient matrix inverse algorithm, HGF not only is more sensitive to the underlying structure of guidance but also achieves the fastest computing speed. Due to these merits, HGF obtains state-of-the-art results in terms of accuracy and efficiency in the computation intensive multi-label systems.
Longquan Dai, Mengke Yuan, Zechao Li, Xiaopeng Zhang 0001, Jinhui Tang 0001
CVPR2
2016 Speeding Up the Bilateral Filter: A Joint Acceleration Way
abstract
Computational complexity of the brute-force implementation of the bilateral filter (BF) depends on its filter kernel size. To achieve the constant-time BF whose complexity is irrelevant to the kernel size, many techniques have been proposed, such as 2D box filtering, dimension promotion, and shiftability property. Although each of the above techniques suffers from accuracy and efficiency problems, previous algorithm designers were used to take only one of them to assemble fast implementations due to the hardness of combining them together. Hence, no joint exploitation of these techniques has been proposed to construct a new cutting edge implementation that solves these problems. Jointly employing five techniques: kernel truncation, best N-term approximation as well as previous 2D box filtering, dimension promotion, and shiftability property, we propose a unified framework to transform BF with arbitrary spatial and range kernels into a set of 3D box filters that can be computed in linear time. To the best of our knowledge, our algorithm is the first method that can integrate all these acceleration techniques and, therefore, can draw upon one another's strong point to overcome deficiencies. The strength of our method has been corroborated by several carefully designed experiments. In particular, the filtering accuracy is significantly improved without sacrificing the efficiency at running time.
Longquan Dai, Mengke Yuan, Xiaopeng Zhang 0001
IEEE Trans. Image Process.2
2015 Fully Connected Guided Image Filtering
abstract
This paper presents a linear time fully connected guided filter by introducing the minimum spanning tree (MST) to the guided filter (GF). Since the intensity based filtering kernel of GF is apt to overly smooth edges and the fixed-shape local box support region adopted by GF is not geometric-adaptive, our filter introduces an extra spatial term, the tree similarity, to the filtering kernel of GF and substitutes the box window with the implicit support region by establishing all-pairs-connections among pixels in the image and assigning the spatial-intensity-aware similarity to these connections. The adaptive implicit support region composed by the pixels with large kernel weights in the entire image domain has a big advantage over the predefined local box window in presenting the structure of an image for the reason that: 1, MST can efficiently present the structure of an image, 2, the kernel weight of our filter considers the tree distance defined on the MST. Due to these reasons, our filter achieves better edge-preserving results. We demonstrate the strength of the proposed filter in several applications. Experimental results show that our method produces better results than state-of-the-art methods.
Longquan Dai, Mengke Yuan, Feihu Zhang, Xiaopeng Zhang 0001
ICCV2