VLDB 2026 Research / reviewers in the wild / expert
Meng Xing
dblp:226/6915
· DBLP profile ↗
35ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Computer networks · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Collaborative Graph Agents for LLM-Based Graph Reasoning
Jindi Wang, Meng Xing, Qinhu Zhang, Dijun Gong |
ICIC (14) | 4 |
| 2026 | Multi-Image compressive encryption via a dynamic delayed feedback chaotic system and synchronized fractal diffusion
Zhiliang Zhu 0001, Wei Zhang 0150, Meng Xing, Hai Yu 0001 |
Expert Syst. Appl. | 4 |
| 2025 | From Gaze to Masks: Gaze-Based Weakly-Supervised Medical Image SegmentationabstractWeakly supervised medical image segmentation seeks to reduce reliance on exhaustive pixel-wise annotations by leveraging coarse or indirect supervision, thereby lowering annotation costs while maintaining segmentation quality. Existing methods rely on coarse heuristics, lacking spatial fidelity and semantic localization in complex medical imagery. To address this limitation, we propose a Gaze-based Weakly Supervised Medical Image Segmentation (GWMIS) framework that exploits gaze annotations as rich, low-cost semantic priors. GWMIS consists of two streams: a pseudo-label generation stream and a cross-modal fusion segmentation stream. The former fuses gaze-derived Gaussian priors with Segment Anything Model (SAM) for structure-consistent pseudo masks; the latter aligns multi-scale visual and gaze features via dual-encoder crossattention for anatomically faithful segmentation. Evaluations on colonic polyp (3,784 images, five public datasets) and brain tumor segmentation (3,929 MRI slices, LGG dataset) demonstrate superior performance over state-of-the-art weakly supervised approaches. Meng Xing, Wanlong Zhang, Yong Su 0003, Qifa Peng, Shuang Zhu |
BIBM | 1 |
| 2025 | Improving Adversarial Transferability through Channel-wise Scaling and Frequency-random DroppingabstractFor black-box attacks, most existing attack methods exhibit weak transferability due to the significant discrepancy between substitute model and victim model. We argue that the model-specific discriminative regions are a key factor causing overfitting to the source model. However, existing model augmentation methods focus on augmentations within a single domain, thereby restricting the diversity of the simulated models. In this paper, we present a novel model augmentation method named CSFD, which combines our proposed channel-wise scaling(CS) and frequency-random dropping(FD) to enhance the diversity of simulated model. Specifically, we first scale the image by channel with CS, which augments them in spatial domain. Then we randomly remove specific frequency patterns of the image with our FD, further augmenting the image in frequency domain by introducing a loss-preserving transformation. This enables us to fully exploit the properties of image in different domains and largely increases the diversity of simulated models. Additionally, we inject random noise perturbations into the sample, effectively exploring the decision boundary within an extended data distribution space. Extensive experiments on ImageNet dataset show that the proposed method has better performance than the existing methods. Zhiyong Feng 0002, Meng Xing, Jinqing Zheng |
ICASSP | 3 |
| 2025 | EfficientNet-BSFT-S: Dynamic Multi-Scale Modules for Robust Image Classification
Chengjie Guo, Minghong Dong, Xuewei Liu, Meng Xing, Yao Zhang 0019, Yude Bai |
ICIC (19) | 4 |
| 2025 | Gaze-and-Machine Dual-Driven Attention Fusion Network for Medical Image Classification
Qifa Peng, Shuang Zhu, Yong Su 0003, Meng Xing |
ICIC (28) | 4 |
| 2025 | PneumoNeXt: A Multi-Scale Attention and Contrastive Learning Approach for Pneumonia Diagnosis
Lirong Zhang, Meng Xing, Yao Zhang 0019, Yude Bai |
ICIC (19) | 2 |
| 2025 | Coming Out of the Dark: Human Pose Estimation in Low-light ConditionsabstractHuman pose estimation in low-light conditions is vital for applications such as surveillance and autonomous systems, yet the severe visual distortions hinder both manual annotation and estimation precision. Existing approaches typically rely on additional reference information to mitigate these issues, however, customized data collection equipment poses limitations on their scalability. To alleviate the issue, we construct a Low-Light Images and Poses (LLIP) dataset, which includes only paired low-light images and pose annotations obtained using off-the-shelf motion capture devices. Furthermore, we propose a Multi-grained High-frequency Feature Consistency Learning framework (MHFCL), which does not rely on additional reference information. MHFCL employs a Retinex-inspired restoration stream to recover high-frequency details and integrates them into pose estimation using a multi-grained consistency mechanism. Experiments demonstrate that our approach achieves a new benchmark in low-light pose estimation, while maintaining competitive performance in well-lit conditions. Yong Su 0003, Meng Xing, Changjae Oh, Xuewei Liu, Jieyang Li |
IJCAI | 3 |
| 2025 | Spatio-temporal graph-based self-labeling for video anomaly detection
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Changjae Oh, Valeriya V. Gribova, Vladimir Fedorovich Filaretoy, De-Shuang Huang |
Neurocomputing | 1 |
| 2025 | Semantic-driven dual consistency learning for weakly supervised video anomaly detection
Yong Su 0003, Yuyu Tan, Simin An, Meng Xing, Zhiyong Feng 0002 |
Pattern Recognit. | 4 |
| 2024 | Learning by Erasing: Conditional Entropy Based Transferable Out-of-Distribution DetectionabstractDetecting OOD inputs is crucial to deploy machine learning models to the real world safely. However, existing OOD detection methods require an in-distribution (ID) dataset to retrain the models. In this paper, we propose a Deep Generative Models (DGMs) based transferable OOD detection that does not require retraining on the new ID dataset. We first establish and substantiate two hypotheses on DGMs: DGMs exhibit a predisposition towards acquiring low-level features, in preference to semantic information; the lower bound of DGM's log-likelihoods is tied to the conditional entropy between the model input and target output. Drawing on the aforementioned hypotheses, we present an innovative image-erasing strategy, which is designed to create distinct conditional entropy distributions for each individual ID dataset. By training a DGM on a complex dataset with the proposed image-erasing strategy, the DGM could capture the discrepancy of conditional entropy distribution for varying ID datasets, without re-training. We validate the proposed method on the five datasets and show that, without retraining, our method achieves comparable performance to the state-of-the-art group-based OOD detection methods. The project codes will be open-sourced on our project website. Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Changjae Oh |
AAAI | 1 |
| 2024 | Exploring Imperceptible Adversarial Examples in YCbCr Color Space
Zhiyong Feng 0002, Meng Xing, Jinqing Zheng |
MMM (3) | 3 |
| 2024 | MIT: Multi-cue Injected Transformer for Two-Stage HOI Detection
Weilong Peng, Qingfeng Chen, Keke Tang, Meng Xing, Meie Fang |
PRCV (7) | 5 |
| 2024 | Anomalies cannot materialize or vanish out of thin air: A hierarchical multiple instance learning with position-scale awareness for video anomaly detection
Yong Su 0003, Yuyu Tan, Simin An, Meng Xing |
Expert Syst. Appl. | 4 |
| 2024 | VPE-WSVAD: Visual prompt exemplars for weakly-supervised video anomaly detection
Yong Su 0003, Yuyu Tan, Meng Xing, Simin An |
Knowl. Based Syst. | 3 |
| 2024 | Blockchain-based service recommendation and trust enhancement model
Chao Wang 0107, Shizhan Chen, Meng Xing, Hongyue Wu, Zhiyong Feng 0002 |
Knowl. Based Syst. | 3 |
| 2023 | Attention-based neural networks for trust evaluation in online social networks
Yanwei Xu 0003, Zhiyong Feng 0002, Meng Xing, Hongyue Wu, Xiao Xue 0001, Shizhan Chen, Chao Wang 0107, Lianyong Qi |
Inf. Sci. | 4 |
| 2023 | Prime: Privacy-preserving video anomaly detection via Motion Exemplar guidance
Yong Su 0003, Haohao Zhu, Yuyu Tan, Simin An, Meng Xing |
Knowl. Based Syst. | 5 |
| 2023 | Metapath-guided multi-headed attention networks for trust prediction in heterogeneous social networks
Yanwei Xu 0003, Zhiyong Feng 0002, Meng Xing, Hongyue Wu, Shizhan Chen, Xiao Xue 0001, Schahram Dustdar |
Knowl. Based Syst. | 3 |
| 2023 | Energy-Based Temporal Summarized Attentive Network for Zero-Shot Action RecognitionabstractRecently, Action Recognition (AR) is facing the scalability problem, since collecting and annotating data for the ever-growing action categories is exhausting and inappropriate. As an alternative to AR, Zero-Shot Action Recognition (ZSAR) is getting more and more attention in the community, as they could utilize a shared semantic/attribute space to recognize novel categories without annotated data. Different from the AR focuses on learning the correlation between actions, ZSAR needs to consider the correlation of action-action, label-label and action-label at the same time. However, as far as we know, there is no work to provide structural guidance for the framework design of ZSAR according to its task characteristics. In this paper, we demonstrate the rationality of using the Energy-Based Model (EBM) to guide the framework design of ZSAR based on their inference mechanism. Furthermore, under the guidance of EBM, we propose an Energy-based Temporal Summarized Attentive Network (ETSAN) to achieve ZSAR. Specifically, to ensure the effectiveness of cross-modal matching, EBM needs to capture the correlations of input-input, output-output and input-output, based on discriminative and focused input and output space. To this end, we first design the Temporal Summarized Attentive Mechanism (TSAM) to capture the correlation of action-action by constructing discriminative and focused input space. Then, a Label Semantic Adaptive Mechanism (LSAM) is proposed to learn the correlation of label-label by adjusting the semantic structure according to the target task. Finally, we devise an Energy Score Estimation Mechanism (ESEM) to measure the compatibility (i.e. energy score) between video representation and label semantic embedding. With end-to-end training, our framework can capture all three of the correlations mentioned above simultaneously by minimizing the energy score of the correct action-label pair. Experiments on the HMDB51 and UCF101 datasets show that the proposed architecture achieves comparable results among methods based on the spatial-temporal visual feature of sequence-level, which demonstrates the efficiency of the EBM in guiding the framework design of ZSAR. In addition, our code is available athttps://github.com/oOHCIOo/ETSAN. Cheng Qi, Zhiyong Feng 0002, Meng Xing, Yong Su 0003, Jinqing Zheng |
IEEE Trans. Multim. | 3 |
| 2021 | Marketing Logistics Cost Optimization and Application Research in Smart LivingabstractThis paper analyzes the cost control and optimization of marketing logistics by using intelligent communication technology, and reports the current situation in smart living. Taking the beer company affiliated to XX Brewery Group for example, factor analysis method and artificial intelligence network theory are adopted to optimize the marketing logistics cost in smart living, suggestions for optimizing the marketing logistics cost of the company under the background of smart living are given as follows: Using intelligent communication technology to unify logistics cost management and financial accounting methods, innovating logistics cost management mechanism through big data technology, comprehensive data mining and intelligent communication network and other technologies to reduce logistics cost. Xianhong Xu, Yuqing Zheng, Shanshan Tian, Meng Xing |
ICIS | 4 |
| 2021 | DVAMN: Dual Visual Attention Matching Network for Zero-Shot Action Recognition
Cheng Qi, Zhiyong Feng 0002, Meng Xing, Yong Su 0003 |
ICANN (5) | 3 |
| 2021 | MemTrust: Find Deep Trust in Your MindabstractTrust prediction is gaining significant interest since it could reduce the burden of user decision-makings effectively in various social activities. Existing works on trust prediction mainly based on trust networks, however, usually give little consideration to data sparsity and temporal continuity of user behavior. In order to solve these problems, we propose a comprehensive deep MemTrust model for trust prediction. With this model, we introduce a embedding layer to extend the feature space and alleviate the distinctive information oblivion caused by data sparsity. In addition, Long Short-Term Memory(LSTM) network is utilized to extract overall time series features through the multiple time slices of user features. Finally, the trust is estimated by pairwise time series features of users. Extensive experiments are validated on two real datasets, which demonstrate that the proposed model has superior performance compared with representative baseline approaches. Yanwei Xu 0003, Zhiyong Feng 0002, Xiao Xue 0001, Shizhan Chen, Hongyue Wu, Meng Xing, Hongqi Chen |
ICWS | 7 |
| 2021 | Spatio-temporal multi-factor model for individual identification from biological motion
Yong Su 0003, Weilong Peng, Meng Xing, Zhiyong Feng 0002 |
Ad Hoc Networks | 3 |
| 2021 | VDARN: Video Disentangling Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition
Yong Su 0003, Meng Xing, Simin An, Weilong Peng, Zhiyong Feng 0002 |
Ad Hoc Networks | 2 |
| 2021 | Disentangling style on dynamic aligned poses for individual identification
Yong Su 0003, Meng Xing, Weilong Peng, Zhiyong Feng 0002 |
Ad Hoc Networks | 3 |
| 2021 | Ventral & Dorsal Stream Theory based Zero-Shot Action Recognition
Meng Xing, Zhiyong Feng 0002, Yong Su 0003, Weilong Peng |
Pattern Recognit. | 1 |
| 2021 | Bayesian Covariance Representation with Global Informative Prior for 3D Action Recognition
Zhiyong Feng 0002, Yong Su 0003, Meng Xing |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Spatio-temporal metric learning for individual recognition from locomotion
Yong Su 0003, Simin An, Zhiyong Feng 0002, Meng Xing |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Cross-covariance matrix: Time-shifted correlations for 3D action recognition
Zhiyong Feng 0002, Yong Su 0003, Meng Xing |
Signal Process. | 4 |
| 2020 | An Image Cues Coding Approach for 3D Human Pose EstimationabstractAlthough Deep Convolutional Neural Networks (DCNNs) facilitate the evolution of 3D human pose estimation, ambiguity remains the most challenging problem in such tasks. Inspired by the Human Perception Mechanism (HPM), we propose an image-to-pose coding method to fill the gap between image cues and 3D poses, thereby alleviating the ambiguity of 3D human pose estimation. First, in 3D pose space, we divide the whole 3D pose space into multiple subregions named pose codes , turning a disambiguation problem into a classification problem. The proposed coding mechanism covers multiple camera views and provides a complete description for 3D pose space. Second, it is noteworthy that the articulated structure of the human body lies on a sophisticated product manifold and the error accumulation in the chain structure will undoubtedly affect the coding performance. Therefore, in image space, we extract the image cues from independent local image patches rather than the whole image. The mapping relationship between image cues and 3D pose codes is established by a set of DCNNs. The image-to-pose coding method transforms the implicit image cues into explicit constraints. Finally, the image-to-pose coding method is integrated into a linear matching mechanism to construct a 3D pose estimation method that effectively alleviates the ambiguity. We conduct extensive experiments on widely used public benchmarks. The experimental results show that our method effectively alleviates the ambiguity in 3D pose recovery and is robust to the variations of view. Meng Xing, Zhiyong Feng 0002, Yong Su 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2019 | Discriminative Saliency-pose-attention Covariance for Action RecognitionabstractMost covariance-based representations of actions are focused on the statistical features of poses by empirical averaging weighting. Note that these poses have a variety of saliency levels for different actions. Neglecting pose saliency could degrade the discriminative power of the covariance features, and further reduce the performance of action recognition. In this paper, we propose a novel saliency weighting covariance feature representation, Saliency-Pose-Attention Covariance(SPA-Cov), which reduces the negative effects from the ambiguous pose samples. Specifically, we utilize a discriminative approach to derive probability distribution of action categories for each pose, which is modeled by the uncertainty of information entropy to obtain the salient weighting. Experimental results show that our proposed method efficiently improves the discriminative power of the generated covariance. In some databases, the proposed SPA-Cov outperforms the state-of-the-art variant methods which are based on kernel matrix, Bayesian posterior features, temporal hierarchical features, etc. Zhiyong Feng 0002, Yong Su 0003, Meng Xing |
ICASSP | 4 |
| 2019 | Dynamic hand gesture recognition using motion pattern and shape descriptors
Meng Xing, Jing Hu 0007, Zhiyong Feng 0002, Yong Su 0003, Weilong Peng, Jinqing Zheng |
Multim. Tools Appl. | 1 |
| 2018 | Spatio-Temporal Large Margin Nearest Neighbor (ST-LMNN) Based on Riemannian Features for Individual IdentificationabstractThe major challenge for individual identification from periodical locomotion is to explore the unique spatio-temporal motion characteristic of each individual. In this paper, we present a novel spatio-temporal metric learning approach to activate the discriminating components of Riemannian features. Specifically, to deal with articulated motion embedded in a high dimensional space, we extend the Active Appearance Model to Riemannian manifold to measure the intrinsic variations of poses. Then we extract high order geometric features of each joint, which naturally suggests the biometric signatures of enrolled individuals. To model the most discriminant features in a linear subspace, we propose a Spatio-Temporal Large Margin Nearest Neighbor (ST-LMNN) algorithm to learn the low-dimensional linear embedding in the spatial and temporal domain, respectively. According to the experimental results from two public databases and a new built database, the proposed approach can achieve more accurate identification results in walking and running. Yong Su 0003, Zhiyong Feng 0002, Meng Xing |
ICME | 3 |
| 2018 | Sequential Articulated Motion Reconstruction from a Monocular Image SequenceabstractIn this article, we present a sequential approach for articulated motion estimation from a 2D skeleton sequence. This is a challenging task due to the complexity of human movements and the inherent depth ambiguities. The proposed approach models the human movement on a kinematic manifold with the tangent bundle, which is a natural geometrical representation of articulated motion. Combined with a second-order stochastic dynamic model based on the Markov hypothesis, we generalize the Extended Rauch Tung Striebel smoother to a Riemannian manifold to simulate the process of human movement. The human motor system might violate the Markov hypothesis when the human body is subject to external forces, and therefore a refinement stage is introduced to correct the estimation error. Specifically, the current estimation is refined in a feasible solution region consisting of a set of local estimations. This region is called a simplex, in which each element can be represented by a convex hull of all ingredients. We have proved that the refinement problem can be converted into a convex optimization problem with the simplicial constraint. Since the proposed formulation conforms to the principles of kinematic and spatio-temporal continuity of articulated motion, the reconstruction ambiguity can be alleviated essentially. The performance of the proposed algorithm is conducted on multiple synthetic sequences from the CMU and the HDM05 MoCap databases. The results show that, without requiring any training data, the proposed approach achieves greater accuracy over state-of-the-art baselines. Furthermore, the proposed approach outperforms two baselines on real sequences from the Human3.6m MoCap database. Yong Su 0003, Zhiyong Feng 0002, Weilong Peng, Meng Xing |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |