VLDB 2026 Research / reviewers in the wild / expert
Meng Jian
dblp:133/7412
· DBLP profile ↗
64ranked-venue papers
24as first author
37since 2021 · last 2026
0000-0001-5659-5128ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 26 · 11 first-author · 13 since 2021Databases, data management, data science and information retrieval · 9 · 6 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MSR-Rec: Multi-Step Reasoning-Enhanced LLM for Sequential RecommendationabstractSequential recommendation has become indispensable in modern digital services. Prevalent recommendation techniques formulate the recommendation task with a language instruction fed into large language models (LLMs) to generate recommendations. However, the implicit interaction scenario of recommendation task cannot provide explicit reasoning supervision to activate LLM's multi-step reasoning capability. Besides, the manner of reasoning for enhancing recommendation is still underexplored. Therefore, we investigate activating multi-step reasoning with users' interactions and propose a multi-step reasoning-enhanced LLM (MSR-Rec), which tightly integrates reasoning with recommendation from designing reasoning chain to reasoning-based recommendation. A task-decomposed reasoning chain is elaborately designed to imitate users' thinking process, seamlessly involving reasoning into recommendation. Following the reasoning chain, MSR-Rec synthesizes reasoning supervision and fine-tunes LLM to adapt for task-specific reasoning. In inference, bidirectional reasoning is implemented from user and item sides, performing a closed-loop reasoning for recommendation. Comprehensive experiments demonstrate that MSR-Rec achieves the state-of-the-art performance in both recommendation quality and reasoning interpretability, advancing the integration of reasoning and recommendation in LLM-based systems. Tuo Wang 0001, Meng Jian, Ge Shi 0002, Lifang Wu, Yashen Wang |
AAAI | 2 |
| 2026 | Enhancing Recommendations With Knowledge-Guided Interest ContrastabstractIn the digital age, the overwhelming amount of information necessitates advanced recommendation systems to deliver personalized content. However, these systems face significant challenges, such as sparse user-item interactions and long-tail bias. Recent studies construct structural learning or self-supervised learning on the interaction graph achieving a positive impact on alleviating the problems, but the interaction data itself may be far too little to solve the problems. While knowledge graphs (KGs) offer a promising solution by providing semantic depth to recommendations, their integration often introduces noise from redundant knowledge. Addressing these critical gaps, this study proposes a knowledge-guided interest contrast (KGIC) to enhance recommendations, which innovatively harmonizes collaborative filtering with semantic insights from KG. The KGIC model introduces three key innovations: (1) a knowledge filtering mechanism that selectively leverages interest-relevant signals from the knowledge graph to encode interest and avoid redundant knowledge interference; (2) an adaptive graph augmentation strategy that enhances the interaction graph based on semantic-aware interest propagation and interaction intensity estimation; and (3) a self-supervised contrastive learning task that mitigates long-tail bias and sparsity issues by homogenizing the embedding distribution between augmented views. The extensive evaluation reveals the superiority of KGIC with knowledge filtering and graph augmentation for recommendation. Meng Jian, Ruoxi Li, Yulong Bai 0002, Ge Shi 0002 |
IEEE Trans. Big Data | 1 |
| 2026 | Hierarchy-Aware Multimodal Distillation for RecommendationabstractBeyond behavioral interaction records, multimedia recommendation scenarios possess abundant semantic signals, which provide excellent data support for user interest mining. Recently, the multimodal enhanced interaction graph has been actively explored and has achieved great progress. However, these methods overlook the capability disparity of various modalities in learning users' interests and lack the ability to explore the hierarchical relationships of interests in modality, resulting in suboptimal recommendation performance. Therefore, this work investigates intra-modality hierarchical learning and inter-modality guidance, proposing a hyperbolic self-distillation (HSD) model for multimedia recommendation. In each modality space, HSD introduces a hyperbolic propagation to filter users' hierarchical interests from the interaction graph effectively. Inter-modality interests are aligned further by a two-level self-distillation strategy to designate multimodal interactions to teach single-modal learning, aiming at teaching and learning to promote each other. Extensive experiments on four public datasets demonstrate that the proposed HSD outperforms leading baselines for multimedia recommendation, verifying the effectiveness of hierarchical propagation and two-level self-distillation in mining users' hierarchical interests. Meng Jian, Tuo Wang 0001, Meijuan Yang, Lifang Wu |
IEEE Trans. Multim. | 1 |
| 2025 | Knowledge-Aware Intent Subgraph Learning for Recommendation
Langchen Lang, Meng Jian |
ICIG (1) | 2 |
| 2025 | Intent-Augmented Multimodal Graph Embedding for Multimedia RecommendationabstractTo address the interaction sparsity, recent recommendation techniques introduce multimodal semantics to enrich collaborative signals for modeling users' interests. However, these models neglect to depict interaction behaviors and omit to involve diverse behavioral patterns, which are the core characters of the recommendation scenario. Therefore, we propose an intent-augmented multimodal graph embedding (IMGE) model for multimodal recommendation, which constructs an interaction graph with multimodal interactions and augments the graph with behavioral interactions to promote recommendation. Unlike the conventional interaction graph, edges are encoded with multimodal semantics, and intent auxiliary nodes are injected into the neighbors of user/item nodes. IMGE innovatively modifies graph convolution by explicitly involving edge signals, which aggregates collaborative signals from the intent-augmented multimodal graph to embed both semantic and behavioral collaborative signals. Experiments on three public datasets demonstrate the superiority of the proposed IMGE, especially on the Baby dataset achieving 40.78% improvement by NDCG, verifying the effectiveness of encoding interactions with semantic and behavioral signals for recommendation. Ruoxi Li, Meng Jian, Lifang Wu |
ICMR | 2 |
| 2025 | Mitigating long-tail bias in recommendations via graph diffusion
Zhuoyang Xia, Meng Jian, Yulong Bai 0002, Lifang Wu, Shaona Wang |
Multim. Syst. | 2 |
| 2025 | Interest-Disentangled Contrastive Sample Generation for RecommendationabstractIn the domain of recommendations, previous works often retrieve items through sampling strategies from the database to gather negative signals for exploring implicit feedback. However, because of extremely sparse records, the existing items used as negative samples may not sufficiently support the interacted items in depicting the diverse interests of users. Consequently, the generation of negative samples needs to be explored in recommendation systems. In this study, we propose an interest-disentangled contrastive sample generation (IDCG) model to enhance interest modeling by contrasting interacted items with the generated samples for recommendation. Specifically, we decouple the interacted items of users into positively relevant and irrelevant factors of interest, providing a valuable clue to learn negatively relevant factors in personalized interests. Then, negative samples are generated by merging the learned negatively relevant factors and irrelevant factors. At this point, a two-level contrast is constructed between positive and negative samples and between the relevant factors of positives and negatives, providing auxiliary collaborative signals to debias and alleviate the interaction sparsity issue. Extensive experiments on three real datasets demonstrate the effectiveness of IDCG in generating targeted and meaningful negative samples from the perspective of disentangling relevant factors to promote interest modeling for recommendation. Meng Jian, Ruoxi Li, Meishan Liu, Meijuan Yang, Shaona Wang, Lifang Wu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2025 | Hierarchical Intent-Based Interest Disentanglement for Personalized RecommendationabstractTo address the data sparsity issue, conventional graph-based models leverage structural signals from the interaction graph to embed users' interests. However, these models learn a uniform representation for interest modeling, which blends users' diverse intents and inevitably biases interest learning, hindering recommendations. Although the fine-grained paradigm can learn the intents of interactions separately to alleviate learning bias, the relationships among intents and the disentangled manner require elaborate design. Existing fine-grained models emphasize intent diversity and employ additional data splitting for disentanglement, which ignores the hierarchical relationship, exacerbates data sparsity, and increases the computational burden. To address these issues, we explore hierarchical intents and adaptive intent learning, proposing a hierarchical intent-based interest disentanglement (HIID) model for personalized recommendation. HIID introduces learnable intent queries to guide interest disentanglement from global interactions in a split-free manner. It raises a hierarchical intent hypothesis to involve hierarchical CF signals for interest modeling, where intents within the same level appear relatively diverse, and the in-depth intents are abstracted from the superficial ones. Both adaptive intent learning and hierarchical hypothesis help extract significant CF signals to promote personalized recommendation. Extensive experiments on public datasets show that the proposed HIID outperforms the state-of-the-art CF models for recommendation. Furthermore, HIID implements adaptive interest disentanglement in a split-free manner, improving the training efficiency of the recommender model compared to the existing fine-grained interest models. Tuo Wang 0001, Meng Jian, Lifang Wu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Geometric-Augmented Self-Distillation for Graph-Based RecommendationabstractThe prevalent recommendation techniques explore the graph structure of interactions to alleviate the interaction sparsity issue for inferring users’ interests. These graph models focus on extracting local structural signals to model users’ interests, introducing grid-like distortion and ignoring the hierarchical tree-like structure when learning from the interaction graph. The learned interests lack significant hierarchical signals, resulting in suboptimal recommendation performance. In this article, we investigate geometric-augmented graph learning with hyperbolic and Euclidean geometries to delve into local structural and hierarchical knowledge from the interaction graph. A self-teaching network called geometric-augmented self-distillation (GASD) is proposed to transfer hierarchical knowledge from hyperbolic to Euclidean space. The transfer learning enables shrinking of the network into a primary student to implement effective and efficient inference in Euclidean space, preventing computational burden in hyperbolic space. Experiments on publicly available datasets demonstrate that the proposed GASD outperforms the state-of-the-art models, verifying the effectiveness and efficiency of knowledge transfer by self-distillation to aggregate knowledge adaptively for personalized recommendation. Meng Jian, Tuo Wang 0001, Zhuoyang Xia, Ge Shi 0002, Richang Hong, Lifang Wu |
ACM Trans. Inf. Syst. | 1 |
| 2025 | Dual Interest Learning with Context-Aware Adaptive Interaction for Social RecommendationabstractSocial recommendation utilizes social relations to extract auxiliary collaborative signals, effectively mitigating data sparsity issues. However, existing approaches predominantly focus on static influence from social friends while neglecting two critical aspects: the dynamic contextual patterns in user behaviors and the potential of collaborative users. To address these limitations and further alleviate data sparsity, we propose a context-aware dual graph attention network (CDGA) that simultaneously captures users’ static and dynamic interests through social relations and interaction records. The proposed CDGA model introduces a dynamic activation mechanism to simulate contextual influences, generating dynamic embeddings for users and items. Furthermore, we develop an adaptive fusion mechanism that integrates interaction channels across static and dynamic embeddings for interaction prediction. Extensive experiments on three benchmark datasets demonstrate that CDGA consistently outperforms state-of-the-art social recommendation methods, confirming its effectiveness. Meng Jian, Ruoxi Li, Xiaoyan Gao 0001, Liqiang Wei, Lifang Wu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Learning to compose diversified prompts for image emotion classificationabstractImage emotion classification (IEC) aims to extract the abstract emotions evoked in images. Recently, language-supervised methods such as contrastive language-image pretraining (CLIP) have demonstrated superior performance in image understanding. However, the underexplored task of IEC presents three major challenges: a tremendous training objective gap between pretraining and IEC, shared suboptimal prompts, and invariant prompts for all instances. In this study, we propose a general framework that effectively exploits the language-supervised CLIP method for the IEC task. First, a prompt-tuning method that mimics the pretraining objective of CLIP is introduced, to exploit the rich image and text semantics associated with CLIP. Subsequently, instance-specific prompts are automatically composed, conditioning them on the categories and image content of instances, diversifying the prompts, and thus avoiding suboptimal problems. Evaluations on six widely used affective datasets show that the proposed method significantly outperforms state-of-the-art methods (up to 9.29% accuracy gain on the EmotionROI dataset) on IEC tasks with only a few trained parameters. The code is publicly available at https://github.com/dsn0w/PT-DPC/for research purposes . Sinuo Deng, Lifang Wu, Ge Shi 0002, Lehao Xing, Meng Jian, Ye Xiang, Ruihai Dong |
Comput. Vis. Media | 5 |
| 2024 | Dynamic interest modeling via dual learning for recommendation
Meng Jian, Xinling Wang, Lifang Wu |
Multim. Tools Appl. | 1 |
| 2024 | Light dual hypergraph convolution for collaborative filtering
Meng Jian, Langchen Lang, Zun Li 0001, Tuo Wang 0001, Lifang Wu |
Pattern Recognit. | 1 |
| 2024 | Graph Contrastive Learning With Negative Propagation for RecommendationabstractPrevious recommendation models build interest embeddings heavily relying on the observed interactions and optimize the embeddings with a contrast between the interactions and randomly sampled negative instances. To our knowledge, the negative interest signals remain unexplored in interest encoding, which merely serves losses for backpropagation. Besides, the sparse undifferentiated interactions inherently bring implicit bias in revealing users’ interests, leading to suboptimal interest prediction. The negative interest signals would be a piece of promising evidence to support detailed interest modeling. In this work, we propose a perturbed graph contrastive learning with negative propagation (PCNP) for recommendation, which introduces negative interest to assist interest modeling in a contrastive learning (CL) architecture. An auxiliary channel of negative interest learning generates a contrastive graph by negative sampling and propagates complementary embeddings of users and items to encode negative signals. The proposed PCNP contrasts positive and negative embeddings to promote interest modeling for recommendation. Extensive experiments demonstrate the capability of PCNP using two-level CL to alleviate interaction sparsity and bias issues for recommendation. Meishan Liu, Meng Jian, Yulong Bai 0002, Jiancan Wu, Lifang Wu |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Counterfactual Graph Convolutional Learning for Personalized RecommendationabstractRecently, recommender systems have witnessed the fast evolution of Internet services. However, it suffers hugely from inherent bias and sparsity issues in interactions. The conventional uniform embedding learning policies fail to utilize the imbalanced interaction clue and produce suboptimal representations to users and items for recommendation. Towards the issue, this work is dedicated to bias-aware embedding learning in a decomposed manner and proposes a counterfactual graph convolutional learning (CGCL) model for personalized recommendation. Instead of debiasing with uniform interaction sampling, we follow the natural interaction bias to model users’ interests with a counterfactual hypothesis. CGCL introduces bias-aware counterfactual masking on interactions to distinguish the effects between majority and minority causes on the counterfactual gap. It forms multiple counterfactual worlds to extract users’ interests in minority causes compared to the factual world. Concretely, users and items are represented with a causal decomposed embedding of majority and minority interests for recommendation. Experiments show that the proposed CGCL is superior to the state-of-the-art baselines. The performance illustrates the rationality of the counterfactual hypothesis in bias-aware embedding learning for personalized recommendation. Meng Jian, Yulong Bai 0002, Xusong Fu, Ge Shi 0002, Lifang Wu |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2024 | Swarm Self-supervised Hypergraph Embedding for RecommendationabstractThe information era brings both opportunities and challenges to information services. Confronting information overload, recommendation technology is dedicated to filtering personalized content to meet users’ requirements. The extremely sparse interaction records and their imbalanced distribution become a big obstacle to building a high-quality recommendation model. In this article, we propose a swarm self-supervised hypergraph embedding (SHE) model to predict users’ interests by hypergraph convolution and self-supervised discrimination. SHE builds a hypergraph with multiple interest clues to alleviate the interaction sparsity issue and performs interest propagation to embed CF signals in hybrid learning on the hypergraph. It follows an auxiliary local view by similar hypergraph construction and interest propagation to restrain unnecessary propagation between user swarms. Besides, interest contrast further inserts self-discrimination to deal with long-tail bias issue and enhance interest modeling, which aid recommendation by a multi-task learning optimization. Experiments on public datasets show that the proposed SHE outperforms the state-of-the-art models demonstrating the effectiveness of hypergraph-based interest propagation and swarm-aware interest contrast to enhance embedding for recommendation. Meng Jian, Yulong Bai 0002, Lifang Wu |
ACM Trans. Knowl. Discov. Data | 1 |
| 2023 | Graph Contrastive Learning on Complementary Embedding for RecommendationabstractPrevious works build interest learning via mining deeply on interactions. However, the interactions come incomplete and insufficient to support interest modeling, even bringing severe bias into recommendations. To address the interaction sparsity and the consequent bias challenges, we propose a graph contrastive learning on complementary embedding (GCCE), which introduces negative interests to assist positive interests of interactions for interest modeling. To embed interest, we design a perturbed graph convolution by preventing embedding distribution from bias. Since negative samples are not available in the general scenario of implicit feedback, we elaborate a complementary embedding generation to depict users’ negative interests. Finally, we develop a new contrastive task to contrastively learn from the positive and negative interests to promote recommendation. We validate the effectiveness of GCCE on two real datasets, where it outperforms the state-of-the-art models for recommendation. Meishan Liu, Meng Jian, Ge Shi 0002, Ye Xiang, Lifang Wu |
ICMR | 2 |
| 2023 | A Novel Semantic-Enhanced Time-Aware Model for Temporal Knowledge Graph Completion
Yashen Wang, Meng Jian, Xiaoye Ouyang |
NLPCC (2) | 3 |
| 2023 | Preference Contrastive Learning for Personalized Recommendation
Yulong Bai 0002, Meng Jian, Shuyi Li 0003, Lifang Wu |
PRCV (9) | 2 |
| 2023 | Multimodal collaborative graph for image recommendation
Meng Jian, Ge Shi 0002, Lifang Wu, Zhangquan Wang |
Appl. Intell. | 1 |
| 2023 | Compatible intent-based interest modeling for personalized recommendation
Meng Jian, Tuo Wang 0001, Shenghua Zhou, Langchen Lang, Lifang Wu |
Appl. Intell. | 1 |
| 2023 | Non-pairwise Collaborative Filtering
Meng Jian, Chenlin Zhang, Tuo Wang 0001, Lifang Wu |
Neural Process. Lett. | 1 |
| 2022 | Boundary-Guided Probability HashingabstractDeep supervised hashing for Hamming space retrieval has recently attracted increasing attention because it enables large-scale image retrieval with constant-time cost. However, the existing Hamming space retrieval methods cannot effectively focus on different pairs simultaneously inside and outside the Hamming ball, making it difficult to push dissimilar pairs outside or pull similar pairs inside the Hamming ball. We propose a novel Boundary-Guided Probability Hashing (BGPH) method that introduces a boundary to guide probability distribution. It makes the probability of similar pairs within the Hamming ball greater than dissimilar pairs and vice versa, which fits the purpose of Hamming space retrieval well. Moreover, we propose a threshold weighting method to indicate when optimization should be stopped to avoid the problem that dissimilar data are pulled into the ball caused by over-optimization in multi-label retrieval scenarios. Comprehensive experiments on three benchmark datasets demonstrate that BGPH yields state-of-the-art retrieval performance. Wenjin Hu 0002, Lifang Wu, Ge Shi 0002, Meng Jian, Sinuo Deng |
ICME | 5 |
| 2022 | Multi-intent Compatible Transformer Network for Recommendation
Tuo Wang 0001, Meng Jian, Ge Shi 0002, Lifang Wu |
PRCV (1) | 2 |
| 2022 | Siamese Graph-Based Dynamic Matching for Collaborative Filtering
Meng Jian, Chenlin Zhang, Meishan Liu, Ge Shi 0002, Lifang Wu |
Inf. Sci. | 1 |
| 2022 | Key frame extraction based on global motion statistics for team-sport videos
Yuan Yuan 0021, Zhe Lu, Meng Jian, Lifang Wu, Xu Liu 0008 |
Multim. Syst. | 4 |
| 2022 | DPFL-Nets: Deep Pyramid Feature Learning Networks for Multiscale Change DetectionabstractDue to the complementary properties of different types of sensors, change detection between heterogeneous images receives increasing attention from researchers. However, change detection cannot be handled by directly comparing two heterogeneous images since they demonstrate different image appearances and statistics. In this article, we propose a deep pyramid feature learning network (DPFL-Net) for change detection, especially between heterogeneous images. DPFL-Net can learn a series of hierarchical features in an unsupervised fashion, containing both spatial details and multiscale contextual information. The learned pyramid features from two input images make unchanged pixels matched exactly and changed ones dissimilar and after transformed into the same space for each scale successively. We further propose fusion blocks to aggregate multiscale difference images (DIs), generating an enhanced DI with strong separability. Based on the enhanced DI, unchanged areas are predicted and used to train DPFL-Net in the next iteration. In this article, pyramid features and unchanged areas are updated alternately, leading to an unsupervised change detection method. In the feature transformation process, local consistency is introduced to constrain the learned pyramid features, modeling the correlations between the neighboring pixels and reducing the false alarms. Experimental results demonstrate that the proposed approach achieves superior or at least comparable results to the existing state-of-the-art change detection methods in both homogeneous and heterogeneous cases. Meijuan Yang, Licheng Jiao, Fang Liu 0001, Biao Hou, Shuyuan Yang 0001, Meng Jian |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2021 | Attribute-Level Interest Matching Network for Personalized Recommendation
Meng Jian, Ge Shi 0002, Lifang Wu, Ye Xiang |
PRCV (2) | 2 |
| 2021 | Latent label mining for group activity recognition in basketball videosabstractAbstract Motion information has been widely exploited for group activity recognition in sports video. However, in order to model and extract the various motion information between the adjacent frames, existing algorithms only use the coarse video‐level labels as supervision cues. This may lead to the ambiguity of extracted features and the omission of changing rules of motion patterns that are also important sports video recognition. In this paper, a latent label mining strategy for group activity recognition in basketball videos is proposed. The authors' novel strategy allows them to obtain the latent labels set for marking different frames in an unsupervised way, and build the frame‐level and video‐level representations with two separate levels of supervision signal. Firstly, the latent labels of motion patterns are digged using the unsupervised hierarchical clustering technique. The generated latent labels are then taken as the frame‐level supervision signal to train a deep CNN for the frame‐level features extraction. Lastly, the frame‐level features are fed into an LSTM network to build the spatio‐temporal representation for group activity recognition. Experimental results on the public NCAA dataset demonstrate that the proposed algorithm achieves state‐of‐the‐art performance. Lifang Wu, Ye Xiang, Meng Jian, Jialie Shen 0001 |
IET Image Process. | 4 |
| 2021 | Cosine metric supervised deep hashing with balanced similarity
Wenjin Hu 0002, Lifang Wu, Meng Jian, Hui Yu 0001 |
Neurocomputing | 3 |
| 2021 | Identity-constrained noise modeling with metric learning for face anti-spoofing
Yaowen Xu, Lifang Wu, Meng Jian, Wei-Shi Zheng 0001, Zhuming Wang |
Neurocomputing | 3 |
| 2021 | OS-LFFD: a light and fast face detector with Ommateum structure
Dezhong Xu, Lifang Wu, Yonghao He, Meng Jian, Junchi Yan |
Multim. Tools Appl. | 5 |
| 2021 | A siamese pedestrian alignment network for person re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Ying Chen 0005 |
Multim. Tools Appl. | 4 |
| 2021 | Semantic manifold modularization-based ranking for image recommendation
Meng Jian, Chenlin Zhang, Ting Jia, Lifang Wu, Xun Yang 0001, Lina Huo |
Pattern Recognit. | 1 |
| 2021 | Semi-supervised kernel matrix learning using adaptive constraint-based seed propagation
Meng Jian, Cheolkon Jung |
Pattern Recognit. | 1 |
| 2021 | Global motion estimation with iterative optimization-based independent univariate model for action recognition
Lifang Wu, Meng Jian, Jialie Shen 0001, Xianglong Lang |
Pattern Recognit. | 3 |
| 2021 | Weakly-supervised video object localization with attentive spatio-temporal correlation
Mingui Wang, Lifang Wu, Meng Jian, Xu Liu 0008 |
Pattern Recognit. Lett. | 4 |
| 2020 | Weakly-Supervised Video Object Grounding by Exploring Spatio-Temporal ContextsabstractGrounding objects in visual context from natural language queries is a crucial yet challenging vision-and-language task, which has gained increasing attention in recent years. Existing work has primarily investigated this task in the context of still images. Despite their effectiveness, these methods cannot be directly migrated into the video context, mainly due to 1) the complex spatio-temporal structure of videos and 2) the scarcity of fine-grained annotations of videos. To effectively ground objects in videos is profoundly more challenging and less explored. Xun Yang 0001, Xueliang Liu, Meng Jian, Xinjian Gao, Meng Wang 0001 |
ACM Multimedia | 3 |
| 2020 | Person image synthesis through siamese generative adversarial network
Ying Chen 0005, Shixiong Xia, Jiaqi Zhao 0001, Meng Jian, Yong Zhou 0003, Qiang Niu, Rui Yao 0006, Dongjun Zhu |
Neurocomputing | 4 |
| 2020 | Fusing motion patterns and key visual information for semantic event recognition in basketball videos
Lifang Wu, Qi Wang 0076, Meng Jian, Boxuan Zhao, Junchi Yan, Chang Wen Chen |
Neurocomputing | 4 |
| 2020 | Diverse sample generation with multi-branch conditional generative adversarial network for remote sensing objects detection
Dongjun Zhu, Shixiong Xia, Jiaqi Zhao 0001, Yong Zhou 0003, Meng Jian, Qiang Niu, Rui Yao 0006, Ying Chen 0005 |
Neurocomputing | 5 |
| 2020 | Content-Based Bipartite User-Image Correlation for Image Recommendation
Meng Jian, Ting Jia, Lifang Wu |
Neural Process. Lett. | 1 |
| 2020 | Visual Sentiment Analysis by Combining Global and Local Information
Lifang Wu, Mingchao Qi, Meng Jian |
Neural Process. Lett. | 3 |
| 2020 | Ontology-Based Global and Collective Motion Patterns for Event Classification in Basketball VideosabstractIn multi-person videos, especially team sport videos, a semantic event is usually represented as a confrontation between two teams of players, which can be represented as collective motion. In broadcast basketball videos, specific camera motions are used to present specific events. Therefore, a semantic event in broadcast basketball videos is closely related to both the global motion (camera motion) and the collective motion. A semantic event in basketball videos can be generally divided into three stages: pre-event, event occurrence (event-occ), and post-event. By analyzing the influence of different stages of video segments to semantic events discrimination, it is observed that the pre-event and event-occ segments are effective for classification, while the post-events are effective for event success/failure classification. In this paper, we propose an ontology-based global and collective motion pattern (On_GCMP) algorithm for the basketball event classification. First, a two-stage GCMP-based event classification scheme is proposed. The GCMP is extracted using the optical flow. The two-stage scheme progressively combines a five-class event classification algorithm on event-occs and a two-class event classification algorithm on pre-events. Both algorithms utilize the sequential convolutional neural networks (CNNs) and the long short-term memory (LSTM) networks to extract the spatial and temporal features of GCMP for event classification. Second, we utilize the post-event segments to predict success/failure using deep features of images in the video frames (RGB_DF_VF)-based algorithms. Finally, the event classification results and success/failure classification results are integrated to obtain the final results. To evaluate the proposed scheme, we collected a new dataset called NCAA+, which is automatically obtained from the NCAA dataset by extending the fixed length of video clips forward and backward of the corresponding semantic events. The experimental results demonstrate that the proposed scheme achieves the mean average precision of 58.10% on NCAA+. It is higher by 6.50% than the state of the art on NCAA. Lifang Wu, Jiaoyu He, Meng Jian, Yaowen Xu, Dezhong Xu, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | MF-SORT: Simple Online and Realtime Tracking with Motion Features
Heng Fu, Lifang Wu, Meng Jian |
ICIG (1) | 3 |
| 2019 | Prosodic Characteristics of Mandarin Declarative and Interrogative Utterances in Parkinson's Disease
Meng Jian, Wentao Gu |
INTERSPEECH | 2 |
| 2019 | Cross-modal Collaborative Manifold Propagation for Image RecommendationabstractWith the rapid evolution of social networks, the increasing user intention gap and visual semantic gap both bring great challenge for users to access satisfied contents. It becomes promising to investigate users' customized multimedia recommendation. In this paper, we propose cross-modal collaborative manifold propagation (CMP) for image recommendation. CMP leverages users' interest distribution to propagate images' user records, which lets users know the trend from others and produces interest-aware image candidates upon users' interests. Visual distribution is investigated simultaneously to propagate users' visual records along dense semantic visual manifold. Visual manifold propagation helps to estimate semantic accurate user-image correlations for the candidate images in recommendation ranking. Experimental performance demonstrate the collaborative user-image inferring ability of CMP with effective user interest manifold propagation and semantic visual manifold propagation in personalized image recommendation. Meng Jian, Ting Jia, Xun Yang 0001, Lifang Wu, Lina Huo |
ICMR | 1 |
| 2019 | A Siamese Pedestrian Alignment Network for Person Re-identification
Yong Zhou 0003, Jiaqi Zhao 0001, Meng Jian, Rui Yao 0006, Bing Liu 0016, Xuning Liu |
PRCV (1) | 4 |
| 2019 | Deep key frame extraction for sport training
Meng Jian, Lifang Wu, Yonghao He |
Neurocomputing | 1 |
| 2018 | Visual Tracking by Combining the Structure-Aware Network and Spatial-Temporal RegressionabstractIn this paper, we propose a novel visual tracking algorithm by combining the structure-aware network (SA-Net) and spatial-temporal regression model. We first use SA-Net to obtain the initial location proposal, and the deep features are extracted using a fine-tuned convolutional neural network model. Finally, both the location proposal and deep features, including historical information, are input into the long short-term memory (LSTM) for end-to-end spatial temporal regression to adjust the initial location proposal from SA-Net. The experimental results on the challenging OTB dataset demonstrate that the proposed scheme is robust to missing tracking caused by occlusion or object deformation. Additionally, the compared experiments show that the proposed scheme is more competitive than state-of-the-art algorithms. Dezhong Xu, Lifang Wu, Meng Jian, Qi Wang 0076 |
ICPR | 3 |
| 2018 | Establishing a Large Scale Dataset for Image Emotion Analysis Using Chinese Emotion Ontology
Lifang Wu, Mingchao Qi, Meng Jian |
PRCV (4) | 4 |
| 2018 | Multimodal Joint Representation for User Interest Analysis on Content Curation Social Networks
Lifang Wu, Meng Jian |
PRCV (3) | 3 |
| 2018 | Shot Boundary Detection with Spatial-Temporal Convolutional Neural Networks
Lifang Wu, Meng Jian, Zhijia Zhao 0004 |
PRCV (2) | 3 |
| 2018 | Visual saliency estimation using constraints
Meng Jian, Lifang Wu, Cheolkon Jung, Qingtao Fu, Ting Jia |
Neurocomputing | 1 |
| 2018 | A fast hybrid retargeting scheme with seam context and content aware strip partition
Lifang Wu, Chuncan Yan, Meng Jian, Weiming Dong, Chang Wen Chen |
Neurocomputing | 3 |
| 2018 | Multi-perspective User2Vec: Exploiting re-pin activity for user representation learning in content curation social network
Lifang Wu, Meng Jian, Xiuzhen Zhang 0001 |
Signal Process. | 4 |
| 2017 | A Quality Evaluation Scheme to 3D Printing Objects Using Stereovision Measurement
Lifang Wu, Xiao-hua Guo, Lidong Zhao, Meng Jian |
ICIG (3) | 4 |
| 2017 | Reducing noisy labels in weakly labeled data for visual sentiment analysisabstractDeep learning-based visual sentiment analysis requires a large dataset for training. Dataset from social networks is popular but noisy because some images collected in this manner are mislabeled. Therefore, it is necessary to refine the dataset. Based on observations to such datasets, we propose a refinement algorithm based on the sentiments of adjective-noun pairs (ANPs) and tags. We first determine the unreliably labeled images through the sentiment contradiction between the ANPs and tags. These images are removed if the numbers of tags with positive and negative sentiments are equal. The remaining images are labeled again based on the majority vote of the tags' sentiments. Furthermore, we improve the traditional deep learning model by combining the softmax and Euclidean loss functions. Additionally, the improved model is trained using the refined dataset. Experiments demonstrate that both the dataset refinement algorithm and improved deep learning model are beneficial. The proposed algorithms outperform the benchmark results. Lifang Wu, Meng Jian, Jiebo Luo 0001, Xiuzhen Zhang 0001, Mingchao Qi |
ICIP | 3 |
| 2016 | Interactive Image Segmentation Using Adaptive Constraint PropagationabstractIn this paper, we propose interactive image segmentation using adaptive constraint propagation (ACP), called ACP Cut. In interactive image segmentation, the interactive inputs provided by users play an important role in guiding image segmentation. However, these simple inputs often cause bias that leads to failure in preserving object boundaries. To effectively use this limited interactive information, we employ ACP for semisupervised kernel matrix learning which adaptively propagates the interactive information into the whole image, while successfully keeping the original data coherence. Moreover, ACP Cut adopts seed propagation to achieve discriminative structure learning and reduce the computational complexity. Experimental results demonstrate that the ACP Cut extracts foreground objects successfully from the background and outperforms the state-of-the-art methods for interactive image segmentation in terms of both effectiveness and efficiency. Meng Jian, Cheolkon Jung |
IEEE Trans. Image Process. | 1 |
| 2016 | Semi-Supervised Bi-Dictionary Learning for Image Classification With Smooth Representation-Based Label PropagationabstractIn this paper, we propose semi-supervised bi-dictionary learning for image classification with smooth representation-based label propagation (SRLP). Natural images contain complex contents of multiple objects with complicated background, clutter, and occlusions, which prevents image features from belonging to a specific category. Therefore, we employ reconstruction-based classification to implement discriminative dictionary learning in a probabilistic manner. We jointly learn a discriminative dictionary called anchor in the feature space and its corresponding soft label called anchor label in the label space, where the combination of anchor and anchor label is referred to as bi-dictionary. The learnt bi-dictionary is utilized to bridge the semantic gap in image classification. First, SRLP constructs smoothed reconstruction problems for bi-dictionary learning. Then, SRLP produces the reconstruction coefficients in the feature space over the anchor to infer soft labels of samples in the label space. Experimental results demonstrate that the proposed method is capable of learning a pair of discriminative dictionaries for image classification in the feature and label spaces and outperforms the-state-of-the-art reconstruction-based classification ones. Meng Jian, Cheolkon Jung |
IEEE Trans. Multim. | 1 |
| 2015 | Interactive image retrieval using constraints
Meng Jian, Cheolkon Jung, Yanbo Shen |
Neurocomputing | 1 |
| 2015 | Adaptive Constraint Propagation for Semi-Supervised Kernel Matrix Learning
Meng Jian, Cheolkon Jung, Yanbo Shen, Licheng Jiao |
Neural Process. Lett. | 1 |
| 2014 | Interactive image segmentation via kernel propagation
Cheolkon Jung, Meng Jian, Licheng Jiao, Yanbo Shen |
Pattern Recognit. | 2 |
| 2014 | Discriminative Structure Learning for Semantic Concept Detection With Graph EmbeddingabstractSemantic concept detection is a very promising way to manage huge amounts of personal contents. In this paper, we propose discriminative structure learning for semantic concept detection with graph embedding. We focus on the task of whole-image categorization and employ graphical model inference based semi-supervised learning (SSL) to detect the semantic category of an image. To effectively extract global features from images, we utilize the spatial pyramid image representation. Then, we perform data warping over the histogram intersection kernel-based graph to learn discriminative features and make image distributions more discriminative for both labeled and unlabeled images. By data warping, each cluster of images is mapped into a relatively compact cluster as well as clusters become well-separated. Moreover, we adopt low-rank representation (LRR) in the embedded space to capture the global discriminative structure from the learned features for label propagation due to its good ability of capturing the global structure of data distributions and robustness against noise and outliers. Finally, we design a smooth nonlinear detector on the captured global discriminative structure to effectively propagate the concepts of labeled images to unlabeled images. Extensive experiments are conducted on four publicly available databases to verify the superiority of the proposed method compared to the state-of-the-art methods. Meng Jian, Cheolkon Jung, Yaoguo Zheng |
IEEE Trans. Multim. | 1 |