VLDB 2026 Research / reviewers in the wild / expert
Miaogen Ling
dblp:195/8138
· DBLP profile ↗
13ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0002-6638-0152ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Corrigendum to "Motional Foreground Attention-based Video Crowd Counting" [Pattern Recognition 144 (2023) 109891]
Miaogen Ling, Tianhang Pan, Ke Wang 0047, Xin Geng 0001 |
Pattern Recognit. | 1 |
| 2026 | Cascaded Cross-Domain Interaction Network for Video Crowd CountingabstractMost existing video crowd counting methods focus on spatiotemporal correlations modeling to obtain more accurate and robust crowd number estimations. However, relying solely on the spatiotemporal domain limits the feature representation capability in complex scenes. In comparison to spatiotemporal domain features, frequency-domain features exhibit greater robustness to noise and illumination variations, while reducing feature redundancy. Therefore, in this paper, we propose a cascaded cross-domain feature interaction network that combines spatial and temporal-related frequency-domain features for video crowd counting. First, a unified representation of the frequency-domain feature for each frame is obtained by integrating the high- and low-frequency signals from Haar wavelet transform. Then, a bidirectional channel-wise cross attention is proposed to construct a temporal-related frequency representation of each frame. To obtain more robust spatial and frequency-domain features, a lightweight frequency-space mutual modulation method is proposed to enhance the representation capability of each domain feature. On the Venice dataset, CCINet achieves 5.4%/10.4% improvements in MAE/RMSE over the best existing method. The code is available at: https://github.com/sunjunyu777/CCINet. Miaogen Ling, Junyu Sun, Wei Fang 0007, Xin Geng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Dual-Branch Spatiotemporal Interaction Network for Video Crowd CountingabstractBy utilizing the spatiotemporal correlations of the consecutive frames, video crowd counting methods are usually more accurate and robust than image-based methods. However, they still suffer from two major problems. First, the crowd features are mostly extracted in a single and fixed temporal scale. Second, they often learn the spatial and spatiotemporal features in a serial approach and the interactions between them are mostly neglected. To address the above problems, we present a Dual-branch SpatioTemporal Interaction network (DSTI) for video crowd counting. Specifically, a purely video-based backbone called Meanformer is designed elaborately to establish a long-term, multi-scale and global temporal representation by combining the strengths of 3D convolution and transformer. Considering the lack of pre-trained weights, a spatiotemporal full-connected composition operation is proposed to boost the training of Meanformer with multi-scale features from an image-based backbone. Finally, a channel cross attention module is utilized to further improve the performance of DSTI by achieving cross-modal interaction. Experimental results show that the proposed method achieves advanced counting performance on six public datasets. Compared with the second-best method, the average reduction of MAE and RMSE reaches 5.4% and 5.5%, respectively. Miaogen Ling, Yongwen Liu, Jian Su 0001, Tianhang Pan, Xin Geng 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Decoupled Imbalanced Label Distribution LearningabstractLabel Distribution Learning (LDL) has been successfully implemented in numerous practical applications. However, the imbalance in label distributions presents a significant challenge due to the substantial variation in annotation information. To tackle this issue, we introduce Decoupled Imbalance Label Distribution Learning (DILDL), which decomposes the imbalanced label distribution into a dominant label distribution and a non-dominant label distribution. Our empirical findings reveal that an excessively high description degree of dominant labels can result in substantial gradient information attenuation for non-dominant labels during the learning process. Therefore, we employ the decoupling approach to balance the description degrees of both dominant and non-dominant labels independently. Furthermore, we align the feature representations with the representations of dominant and non-dominant labels separately, aiming to effectively mitigate the distribution shift problem. Experimental results demonstrate that our proposed DILDL outperforms other state-of-the-art methods for imbalance label distribution learning. Yongbiao Gao, Xiangcheng Sun, Miaogen Ling, Yi Zhai 0003, Guohua Lv |
IJCAI | 3 |
| 2025 | Dual-branch adjacent connection and channel mixing network for video crowd counting
Miaogen Ling, Jixuan Chen, Yongwen Liu, Wei Fang 0007, Xin Geng 0001 |
Pattern Recognit. | 1 |
| 2025 | Long and Recent Preference Learning With Recent-K Items Distribution for Recommender SystemabstractReinforcement learning (RL) aims to formulate the recommendation task as a Markov decision process (MDP) and trains an agent to automatically learn the optimal recommendation policy from interaction trajectories through trial-and-error and reward mechanisms. However, most existing RL-based approaches overlook the correlation between items and the dynamics of user interests implied in temporally close interactions. Therefore, in this paper, we propose a reinforcement learning method that incorporates a “recent-k items” distribution to capture users' local preferences. Specifically, we model the output layer as two distinct branches. The “recent-k items” branch, formulated with a Kullback-Leibler divergence loss, learns the recent interests of users, whereas the other branch utilizes a one-step temporal difference error to capture long-term preferences. The proposed structure is integrated into deep Q-learning and actor-critics, resulting in two enhanced methods named R$k$Q and R$k$AC, respectively. Furthermore, a novel soft inter-reward is carefully designed to enhance the proposed method, and we theoretically prove the convergence of the proposed algorithm. We perform extensive experiments on two large real-world datasets and conduct further analysis of the influences of different action sequences, time intervals, and enhancement capabilities for state-of-the-art models. The experimental results demonstrate the efficacy of our proposed methods. Yongbiao Gao, Sijie Niu, Guohua Lv, Miaogen Ling, Xin Geng 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Motional foreground attention-based video crowd counting
Miaogen Ling, Tianhang Pan, Ke Wang 0047, Xin Geng 0001 |
Pattern Recognit. | 1 |
| 2023 | MAMIQA: No-Reference Image Quality Assessment Based on Multiscale Attention Mechanism With Natural Scene StatisticsabstractNo-Reference Image Quality Assessment aims to evaluate the perceptual quality of an image, according to human perception. Many recent studies use Transformers to assign different self-attention mechanisms to distinguish regions of an image, simulating the perception of the human visual system (HVS). However, the quadratic computational complexity caused by the self-attention mechanism is time-consuming and expensive. Meanwhile, the image resizing in the feature extraction stage loses the full-size image quality. To address these issues, we propose a lightweight attention mechanism using decomposed large-kernel convolutions to extract multiscale features, and a novel feature enhancement module to simulate HVS. We also propose to compensate the information loss caused by image resizing, with supplementary features from natural scene statistics. Experimental results on five standard datasets show that the proposed method surpasses the SOTA, while significantly reducing the computational costs. Li Yu 0004, Farhad Pakdaman, Miaogen Ling, Moncef Gabbouj |
IEEE Signal Process. Lett. | 4 |
| 2023 | Fast Label Enhancement for Label Distribution LearningabstractLabel Distribution Learning (LDL) has attracted increasing research attentions due to its potential to address the label ambiguity problem in machine learning and success in many real-world applications. In LDL, it is usually expensive to obtain the ground-truth label distributions of data, but it is relatively easy to obtain the logical labels of data. How to use training instances only with logical labels to learn an effective LDL model is a challenging problem. In this paper, we propose a two-step framework to address this problem. Specifically, we firstly design an efficient recovery model to recover the latent label distributions of training instances, named Fast Label Enhancement (FLE). Our idea is to use non-negative matrix factorization (NMF) to mine the label distribution information from the feature space. Moreover, we take the instance-class similarities into consideration to discover the importance of each label to training instances, which is useful for learning precise label distributions. Then, we train a predictive model for testing instances based on generated label distributions of training instances and an existing LDL method (e.g., SA-BFGS). Experimental results on fifteen benchmark datasets show the effectiveness of the proposed two-step framework and verify the superiority of FLE over several state-of-the-art approaches. Ke Wang 0047, Ning Xu 0009, Miaogen Ling, Xin Geng 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Label distribution for multimodal machine learning
Ning Xu 0009, Miaogen Ling, Xin Geng 0001 |
Frontiers Comput. Sci. | 3 |
| 2019 | Soft video parsing by label distribution learning
Miaogen Ling, Xin Geng 0001 |
Frontiers Comput. Sci. | 1 |
| 2019 | Indoor Crowd Counting by Mixture of Gaussians Label Distribution LearningabstractIn this paper, we tackle the problem of crowd counting in indoor videos, where people often stay almost static for a long time. The label distribution, which covers a certain number of crowd counting labels, representing the degree to which each label describes the video frame, is previously adopted to model the label ambiguity of the crowd number. However, since the label ambiguity is significantly affected by the crowd number of the scene, we initialize the label distribution of each frame by the discretized Gaussian distribution with adaptive variance instead of the original single static Gaussian distribution. Moreover, considering the gradual change of crowd numbers in the adjacent frames, a mixture of Gaussian models is proposed to generate the final label distribution representation for each frame. The weights of the Gaussian models rely on the frame and feature distances between the current frame and the adjacent frames. The mixed l2,1-norm is adopted to restrict the weights of predicting the adjacent crowd numbers to be locally correlated. We collect three new indoor video datasets with frame number annotation for further research. The proposed approach achieves state-of-the-art performance on seven challenging indoor videos and cross-scene experiments. Miaogen Ling, Xin Geng 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Soft Video Parsing by Label Distribution LearningabstractIn this paper, we tackle the problem of segmenting out a sequence of actions from videos. The videos contain background and actions which are usually composed of ordered sub-actions. We refer the sub-actions and the background as semantic units. Considering the possible overlap between two adjacent semantic units, we utilize label distributions to annotate the various segments in the video. The label distribution covers a certain number of semantic unit labels, representing the degree to which each label describes the video segment. The mapping from a video segment to its label distribution is then learned by a Label Distribution Learning (LDL) algorithm. Based on the LDL model, a soft video parsing method with segmental regular grammars is proposed to construct a tree structure for the video. Each leaf of the tree stands for a video clip of background or sub-action. The proposed method shows promising results on the THUMOS'14 and MSR-II datasets and its computational complexity is much less than the state-of-the-art method. Xin Geng 0001, Miaogen Ling |
AAAI | 2 |