VLDB 2026 Research / reviewers in the wild / expert
Menglong Yang
dblp:18/9181 · also Meng-Long Yang
· DBLP profile ↗
28ranked-venue papers
13as first author
11since 2021 · last 2025
0000-0003-0948-6847ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 9 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ISPDiffuser: Learning RAW-to-sRGB Mappings with Texture-Aware Diffusion Models and Histogram-Guided Color ConsistencyabstractRAW-to-sRGB mapping, or the simulation of the traditional camera image signal processor (ISP), aims to generate DSLR-quality sRGB images from raw data captured by smartphone sensors. Despite achieving comparable results to sophisticated handcrafted camera ISP solutions, existing learning-based methods still struggle with detail disparity and color distortion. In this paper, we present ISPDiffuser, a diffusion-based decoupled framework that separates the RAW-to-sRGB mapping into detail reconstruction in grayscale space and color consistency mapping from grayscale to sRGB. Specifically, we propose a texture-aware diffusion model that leverages the generative ability of diffusion models to focus on local detail recovery, in which a texture enrichment loss is further proposed to prompt the diffusion model to generate more intricate texture details. Subsequently, we introduce a histogram-guided color consistency module that utilizes color histogram as guidance to learn precise color information for grayscale to sRGB color consistency mapping, with a color consistency loss designed to constrain the learned color information. Extensive experimental results show that the proposed ISPDiffuser outperforms state-of-the-art competitors both quantitatively and visually. Yang Ren 0001, Hai Jiang 0006, Menglong Yang, Wei Li 0075, Shuaicheng Liu |
AAAI | 3 |
| 2025 | Learning Arbitrary-Scale RAW Image Downscaling with Wavelet-based Recurrent ReconstructionabstractImage downscaling is critical for efficient storage and transmission of high-resolution (HR) images. Existing learning-based methods focus on performing downscaling within the sRGB domain, which typically suffers from blurred details and unexpected artifacts. RAW images, with their unprocessed photonic information, offer greater flexibility but lack specialized downscaling frameworks. In this paper, we propose a wavelet-based recurrent reconstruction framework that leverages the information lossless attribute of wavelet transformation to fulfill the arbitrary-scale RAW image downscaling in a coarse-to-fine manner, in which the Low-Frequency Arbitrary-Scale Downscaling Module (LASDM) and the High-Frequency Prediction Module (HFPM) are proposed to preserve structural and textural integrity of the reconstructed low-resolution (LR) RAW images, alongside an energy-maximization loss to align high-frequency energy between HR and LR domain. Furthermore, we introduce the Realistic Non-Integer RAW Downscaling (Real-NIRD) dataset, featuring a non-integer downscaling factor of 1.3×, and incorporate it with publicly available datasets with integer factors (2×, 3×, 4×) for comprehensive benchmarking arbitrary-scale image downscaling purposes. Extensive experiments demonstrate that our method outperforms existing state-of-the-art competitors both quantitatively and visually. The code and dataset will be released at https://github.com/RenYangSCU/ASRD. Yang Ren 0001, Hai Jiang 0006, Wei Li 0075, Menglong Yang, Heng Zhang 0042, Zehua Sheng, Qingsheng Ye, Shuaicheng Liu |
ACM Multimedia | 4 |
| 2025 | An end-to-end robust feature learning method for face recognition
Menglong Yang, Hanyong Wang, Fangrui Wu, Xuebin Lv |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Multi-Modal Disordered Representation Learning Network for Description-Based Person SearchabstractDescription-based person search aims to retrieve images of the target identity via textual descriptions. One of the challenges for this task is to extract discriminative representation from images and descriptions. Most existing methods apply the part-based split method or external models to explore the fine-grained details of local features, which ignore the global relationship between partial information and cause network instability. To overcome these issues, we propose a Multi-modal Disordered Representation Learning Network (MDRL) for description-based person search to fully extract the visual and textual representations. Specifically, we design a Cross-modality Global Feature Learning Architecture to learn the global features from the two modalities and meet the demand of the task. Based on our global network, we introduce a Disorder Local Learning Module to explore local features by a disordered reorganization strategy from both visual and textual aspects and enhance the robustness of the whole network. Besides, we introduce a Cross-modality Interaction Module to guide the two streams to extract visual or textual representations considering the correlation between modalities. Extensive experiments are conducted on two public datasets, and the results show that our method outperforms the state-of-the-art methods on CUHK-PEDES and ICFG-PEDES datasets and achieves superior performance. Fan Yang 0104, Wei Li 0075, Menglong Yang, Binbin Liang |
AAAI | 3 |
| 2024 | Deformable registration framework for glioma images with absent correspondence based on auxiliary-image-aided intensity-consistency constraintabstractConsidering the tumor aggressive nature and the significant changes in anatomical structure, aligning the preoperative and follow up scans of glioma patients remains a challenge due to the presence of regions with absent correspondence. To address this challenge, this work proposed a novel bidirectional unsupervised deformable image registration framework for image pairs with missing correspondence based on an auxiliary-image-aided intensity-consistency constraint (ICC) strategy. Specifically, for any fixed and moving image pairs, we introduced an auxiliary image and warped it directly to fixed/moving image or warped it twice through a transition of moving/fixed image. By comparing the difference between these warped images, the weighting maps to identify and exclude regions with absent correspondence between fixed and moving image pairs can be generated. To verify the effectiveness of the proposed framework, we combined it with several deep learning-based registration models and tested it on BraTS-Reg challenge dataset, the results demonstrated that the proposed ICC strategy can improve the registration performance for all the models, with the improvement of average target registration error (TRE) and success rate (SR) being up to 44.9% and 66.7%, respectively. Comparing against the best existing forward-backward consistency strategy for dealing with missing correspondence registration, our auxiliary-image-aided ICC strategy can also decrease average TRE by 2.9%, demonstrating the superiority of the proposed framework. The present work is not limited to the glioma images, it can be used to address the registration problems for any image pairs with absent correspondence or inconsistent intensity. Lihui Wang 0002, Menglong Yang, Yue Min Zhu, Hongjiang Wei |
BIBM | 3 |
| 2024 | Space Networking Kit: A Novel Simulation Platform for Emerging LEO Mega-constellationsabstractFuturistic Space Networks (SN) present unprece-dented prospects for ubiquitous, low-latency Internet services. Yet, these networks also encounter unique challenges arising from the dynamic nature of satellites on a global scale. To comprehensively address emerging issues in SNs, researchers require the capability to conduct a diverse array of experiments. However, existing experimental approaches either implement visualization functionality but lack network functionality (e.g., the space simulator), or implement network functionality but lack visualization (e.g., the networking simulator). In this paper, we present SNK, a novel simulation platform with visualization and networking capabilities for evaluating the space network performance of global Internet services. SNK offers real-time communication visualization and supports the simulation of routing between edge node of network. The platform enables the evaluation of routing and network performance metrics such as latency, stretch, network capacity, and throughput under different network structures and density. The effectiveness of SNK is demonstrated through various simulation cases, including the routing between fixed edge stations or mobile edge stations and analysis of snace network structures. Xiangtong Wang, Xiaodong Han, Menglong Yang, Songchen Han, Wei Li 0075 |
ICC | 3 |
| 2024 | A Self-Attention Network for Stereo MatchingabstractLearning contextual information has been proved to be conducive to reducing mismatches in ill-posed regions, and many data-driven stereo matching algorithms achieve state-of-the-art performances. Global contextual information is difficult to learn by a shallow convolutional neural network, while a large network brings huge computational cost. This paper proposes a self-attention framework of stereo matching to learn both global and local contextual information. From a rough estimation, the disparity can be extremely refined by applying the proposed attention module. We present two self-attention modes to better learn global and local context synchronously. Experimental results show the proposed algorithm can well predict thin structures, large occluded or textureless regions, and achieves the comparable performance to state-of-the-art methods on the public stereo benchmark with a real-time speed. Menglong Yang, Hanyong Wang, Yang Ren 0001 |
ICME | 1 |
| 2024 | Feature Dimensionality Reduction With L2,p-Norm-Based Robust Embedding Regression for Classification of Hyperspectral ImagesabstractThe curse of dimensionality and noise corruption are two tough problems that need to be solved in hyperspectral image (HSI) classification. However, the current feature dimensionality reduction methods, including both feature extraction and feature selection ones, cannot simultaneously solve the above two problems well. To address this issue, this paper proposes a novel method calledL2,p-norm-based robust embedding regression (L2,p-RER) for robust feature dimensionality reduction of HSI, which can effectively suppress the impact of noises and reduce the feature dimensions. Specifically,L2,p-RER first integrates projection learning with robust principle component analysis (RPCA) to remove noise in a low-dimensional space. Secondly, an embedding regression regularization is proposed to improve the discriminability of the extracted low-dimensional features. Thirdly, aL2,1-norm constraint is imposed to improve the interpretability of the learned projection matrix, which can jointly extract the key features from all bands with their physical meanings certainly preserved. Last but most important, theL2,p-norm that can adaptively balance the sparsity and the convexity is employed to model the noise and regression residual in the embedded low-dimensional space, which can further enhance the robustness and generalization of the proposed method. In addition, extensive experiments conducted on three benchmark HSI datasets validated the effectiveness of the proposed method. Yangjun Deng, Menglong Yang, Heng-Chao Li 0001, Chen-Feng Long, Kui Fang, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Enabling High-Connectivity LEO Satellite Networks Via Encountering Inter-Satellite LinksabstractThe use of a mega-constellation comprised of thousands of Low Earth Orbit (LEO) satellites for global internet service has garnered significant attention. In the network layer of this system, geographical routing has been found to outperform centralized strategies due to the lower complexity and overhead. However, geographical routing can still result in “dead-ends” due to network gaps. To address this issue, we propose the use of encountering inter-satellite links (eISLs) to improve network connectivity and routing reachability. We further present a system model and analysis of eISLs, as well as our Dynamic eISLs Configuration (DeC) algorithm for establishing eISLs between encountering satellites. Our experimental results demonstrate that our proposed DeC under eISLs enabling in satellite networks can significantly reduce propagation latency by 22% and path stretch by 15% in centralized routing algorithms. Moreover, in geographical routing, DeC can effectively improve the reachable ratio from 55% to 100% while maintaining a 28% increase in throughput, outperforming schemes without eISLs. Our proposed eISL-enabled satellite network architecture shows promising results in improving routing efficiency and connectivity in LEO satellite systems. Xiangtong Wang, Wei Li 0075, Songchen Han, Menglong Yang, Zhiyun Jiang |
GLOBECOM | 4 |
| 2021 | Efficient Example Mining for Anchor-Free Face Detection
Siyi Hou, Ying Cai 0002, Menglong Yang |
ICIG (2) | 5 |
| 2021 | RLStereo: Real-Time Stereo Matching Based on Reinforcement LearningabstractMany state-of-the-art stereo matching algorithms based on deep learning have been proposed in recent years, which usually construct a cost volume and adopt cost filtering by a series of 3D convolutions. In essence, the possibility of all the disparities is exhaustively represented in the cost volume, and the estimated disparity holds the maximal possibility. The cost filtering could learn contextual information and reduce mismatches in ill-posed regions. However, this kind of methods has two main disadvantages: 1) cost filtering is very time-consuming, and it is thus difficult to simultaneously satisfy the requirements for both speed and accuracy; 2) thickness of the cost volume determines the disparity range which can be estimated, and the pre-defined disparity range may not meet the demand of practical application. This paper proposes a novel real-time stereo matching method called RLStereo, which is based on reinforcement learning and abandons the cost volume or the routine of exhaustive search. The trained RLStereo makes only a few actions iteratively to search the value of the disparity for each pair of stereo images. Experimental results show the effectiveness of the proposed method, which achieves comparable performances to state-of-the-art algorithms with real-time speed on the public large-scale testset, i.e., Scene Flow. Menglong Yang, Fangrui Wu, Wei Li 0075 |
IEEE Trans. Image Process. | 1 |
| 2020 | WaveletStereo: Learning Wavelet Coefficients of Disparity Map in Stereo MatchingabstractSome stereo matching algorithms based on deep learning have been proposed and achieved state-of-the-art performances since some public large-scale datasets were put online. However, the disparity in smooth regions and detailed regions is still difficult to accurately estimate simultaneously. This paper proposes a novel stereo matching method called WaveletStereo, which learns the wavelet coefficients of the disparity rather than the disparity itself. The WaveletStereo consists of several sub-modules, where the low-frequency sub-module generates the low-frequency wavelet coefficients, which aims at learning global context information and well handling the low-frequency regions such as textureless surfaces, and the others focus on the details. In addition, a densely connected atrous spatial pyramid block is introduced for better learning the multi-scale image features. Experimental results show the effectiveness of the proposed method, which achieves state-of-the-art performance on the large-scale test dataset Scene Flow. Menglong Yang, Fangrui Wu, Wei Li 0075 |
CVPR | 1 |
| 2019 | A fast and robust 3D face recognition approach based on deeply learned face representation
Ying Cai 0002, Yinjie Lei, Menglong Yang, Zhisheng You, Shiguang Shan |
Neurocomputing | 3 |
| 2019 | A feature learning approach for face recognition with robustness to noisy label based on top-N prediction
Menglong Yang, Feihu Huang 0002, Xuebin Lv |
Neurocomputing | 1 |
| 2018 | Learning both matching cost and smoothness constraint for stereo matching
Menglong Yang, Xuebin Lv |
Neurocomputing | 1 |
| 2018 | Multiscale overlapping blocks binarized statistical image features descriptor with flip-free distance for face verification in the wild
Tianyu Geng, Menglong Yang, Zhisheng You, Ying Cai 0002, Feihu Huang 0002 |
Neural Comput. Appl. | 2 |
| 2017 | The Euclidean embedding learning based on convolutional neural network for stereo matching
Menglong Yang, Yiguang Liu, Zhisheng You |
Neurocomputing | 1 |
| 2016 | Stereo matching based on classification of materials
Menglong Yang, Yiguang Liu, Ying Cai 0002, Zhisheng You |
Neurocomputing | 1 |
| 2016 | L1-Norm Low-Rank Matrix Decomposition by Neural Networks and MollifiersabstractThe L1-norm cost function of the low-rank approximation of the matrix with missing entries is not smooth, and also cannot be transformed into a standard linear or quadratic programming problem, and thus, the optimization of this cost function is still not well solved. To tackle this problem, first, a mollifier is used to smooth the cost function. High closeness of the smoothed function to the original one can be obtained by tuning the parameters contained in the mollifier. Next, a recurrent neural network is proposed to optimize the mollified function, which will converge to a local minimum. In addition, to boost the speed of the system, the mollifying process is implemented by a filtering procedure. The influence of two mollifier parameters is theoretically analyzed and experimentally confirmed, showing that one of the parameters is critical to computational efficiency and accuracy, while the other not. A large number of experiments on synthetic data show that the proposed method is competitive to the state-of-the-art methods. In particular, the experiments on large matrices and a real application in the structure from motion indicate that the memory requirement of the proposed algorithm is mild, making it suitable for real applications that often involve large-scale matrix decomposition. Yiguang Liu, Songfan Yang, Pengfei Wu 0002, Chunguang Li 0001, Menglong Yang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2015 | Multiclass classification based on a deep convolutional network for head pose estimationabstractHead pose estimation has been considered an important and challenging task in computer vision. In this paper we propose a novel method to estimate head pose based on a deep convolutional neural network (DCNN) for 2D face images. We design an effective and simple method to roughly crop the face from the input image, maintaining the individual-relative facial features ratio. The method can be used in various poses. Then two convolutional neural networks are set up to train the head pose classifier and then compared with each other. The simpler one has six layers. It performs well on seven yaw poses but is somewhat unsatisfactory when mixed in two pitch poses. The other has eight layers and more pixels in input layers. It has better performance on more poses and more training samples. Before training the network, two reasonable strategies including shift and zoom are executed to prepare training samples. Finally, feature extraction filters are optimized together with the weight of the classification component through training, to minimize the classification error. Our method has been evaluated on the CAS-PEAL-R1, CMU PIE, and CUBIC FacePix databases. It has better performance than state-of-the-art methods for head pose estimation. Ying Cai 0002, Menglong Yang |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2014 | A Probabilistic Framework for Multitarget Tracking with Mutual OcclusionsabstractMutual occlusions among targets can cause track loss or target position deviation, because the observation likelihood of an occluded target may vanish even when we have the estimated location of the target. This paper presents a novel probability framework for multitarget tracking with mutual occlusions. The primary contribution of this work is the introduction of a vectorial occlusion variable as part of the solution. The occlusion variable describes occlusion states of the targets. This forms the basis of the proposed probability framework, with the following further contributions: 1) Likelihood: A new observation likelihood model is presented, in which the likelihood of an occluded target is computed by referring to both of the occluded and oc-cluding targets. 2) Priori: Markov random field (MRF) is used to model the occlusion priori such that less likely "circular" or "cascading" types of occlusions have lower priori probabilities. Both the occlusion priori and the motion priori take into consideration the state of occlusion. 3) Optimization: A realtime RJMCMC-based algorithm with a newmove type called "occlusion state update" is presented. Experimental results show that the proposed framework can handle occlusions well, even including long-duration full occlusions, which may cause tracking failures in the traditional methods. Menglong Yang, Yiguang Liu, Longyin Wen, Zhisheng You, Stan Z. Li |
CVPR | 1 |
| 2014 | A homography transform based higher-order MRF model for stereo matching
Menglong Yang, Yiguang Liu, Zhisheng You, Yi Zhang 0018 |
Pattern Recognit. Lett. | 1 |
| 2013 | Classification by nearness in complementary subspaces
Menglong Yang, Yiguang Liu, Baojiang Zhong |
Pattern Anal. Appl. | 1 |
| 2013 | A robust face and ear based multimodal biometric system using sparse representation
Zengxi Huang, Yiguang Liu, Chunguang Li 0001, Menglong Yang |
Pattern Recognit. | 4 |
| 2012 | Online Multiple Instance Joint Model for Visual TrackingabstractAlthough numerous online learning strategies have been proposed to handle the appearance variation in visual tracking, the existing methods just perform well in certain cases since they lack effective appearance learning mechanism. In this paper, a joint model tracker (JMT) is presented, which consists of a generative model based on Multiple Subspaces and a discriminative model based on improved Multiple Instance Boosting (MIBoosting). The generative model utilizes a series of local constructed subspaces to update the Multiple Subspaces model and considers the energy dissipation of dimension reduction in updating step. The discriminative model adopts the Gaussian Mixture Model (GMM) to estimate the posterior probability of the likelihood function. These two parts supervise each other to update in multiple instance way which helps our tracker recover from drift. Extensive experiments on various databases validate the effectiveness of our proposed method over other state-of-the-art trackers. Longyin Wen, Zhaowei Cai, Menglong Yang, Zhen Lei 0001, Dong Yi, Stan Z. Li |
AVSS | 3 |
| 2012 | Video synchronization based on events alignment
Yiguang Liu, Menglong Yang, Zhisheng You |
Pattern Recognit. Lett. | 2 |
| 2011 | Estimating the fundamental matrix based on least absolute deviation
Menglong Yang, Yiguang Liu, Zhisheng You |
Neurocomputing | 1 |
| 2010 | The Reliability of Travel Time ForecastingabstractTravel time is a fundamental measure in transportation, and accurate travel time forecasting is crucial in intelligent transportation systems (ITSs). Currently, many techniques have been applied to travel time forecasting; however, the reliability of the prediction has not been studied in these approaches. In this paper, we propose an approach using the generalized autoregressive conditional heteroscedasticity (GARCH) model to study the volatility of travel time and supply the information about reliability for travel time forecasting. Three examples on real urban vehicular traffic data show the whole modeling process. In the experiments, we utilize the conditional predicted standard deviation (PSD) to express the reliability of travel time forecasting and screen out the sample points that are thought to be reliable forecasting. The results show that the root-mean-square error (RMSE), mean absolute error (MAE), and mean absolute percent error (MAPE) are all decreasing with an increase in the demand of the reliability. It proves that the model well depicts the reliability of travel time forecasting and that the proposed approach is feasible. Menglong Yang, Yiguang Liu, Zhisheng You |
IEEE Trans. Intell. Transp. Syst. | 1 |