VLDB 2026 Research / reviewers in the wild / expert
Junyang Chen 0001
dblp:196/7893
· DBLP profile ↗
86ranked-venue papers
21as first author
76since 2021 · last 2026
0000-0002-1139-8654ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 10 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 5 first-author · 24 since 2021Databases, data management, data science and information retrieval · 17 · 6 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 4 first-author · 14 since 2021Computer networks · 4 · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Multimodal Fake News Detection by Multi-perspective Rationale Generation and VerificationabstractThe rapid proliferation of social media platforms has led to a surge in multimodal fake news, where deceptive content often combines text and images to mislead audiences. Traditional unimodal detection methods struggle to address the complexity of such content, necessitating holistic multimodal approaches. While the latest advancements in Multimodal Large Language Models (MLLMs) offer new opportunities for enhancing detection performance by analyzing multi-dimensional features, including source credibility, cross-modal contradictions, emotional bias, and manipulative writing patterns, these methods suffer from a key flaw: a susceptibility to hallucinations or erroneous reasoning, which can lead to flawed conclusions and ultimately biased detection results. We propose the Multimodal Fake News Detection via Multi-perspective Rationale Generation and Verification (MMRGV) model to mitigate this challenge. Our method employs a cross-verification mechanism to screen and reconcile contradictions among different rationales, thereby preserving the LLM's analytical advantages while mitigating the impact of erroneous reasoning or hallucinations on the final detection. Subsequently, these optimized rationales are fused via an adaptive weighting strategy to output a robust final prediction. Extensive experiments on three benchmark datasets (Twitter, Weibo, and GossipCop) demonstrate the superiority of our method, achieving state-of-the-art accuracy of 0.9972, 0.9663, and 0.8772, respectively, and significantly outperforming existing baselines. These results validate the effectiveness of multi-perspective rationale generation and cross-verification in enhancing multimodal fake news detection, offering a resilient solution to combat misinformation in the era of generative AI. Junyang Chen 0001, Yueqian Li, Ka Chung Ng, Huan Wang 0005, Liang-Jie Zhang |
AAAI | 1 |
| 2026 | Zero-shot Recommendation: Towards Class Semantic Relation Learning for Inferring Labels of Unseen Micro-videosabstractMicro-video label prediction plays a pivotal role on contemporary video-sharing platforms, such as Kwai and Tiktok. The emergence of video content lacking labels presents a formidable challenge for conventional user interest prediction methods. This paper addresses the challenge of micro-video label prediction, particularly for unseen videos, by proposing a zero-shot method called Class Semantic Relation Learning (CSRL). Unlike traditional user interest prediction models, CSRL leverages the pre-trained Large Language Model (LLM) to enhance prediction accuracy for unlabeled videos. The novelty of CSRL lies in its integration of three key components: a raw feature autoencoder, LLM-enhanced features, and a decomposed graph network. The decomposed graph network is specifically designed to disentangle the relationships between labeled and unlabeled videos, offering a significant improvement over previous methods. By fusing hidden topics with LLM-enhanced text, CSRL effectively handles sparse video features. Experiments on large-scale datasets from the Kwai platform show that CSRL achieves state-of-the-art results, with up to 44.64% improvement in Hit Ratio (HR), highlighting its superiority over existing zero-shot recommendation models in predicting user interests within the user-video network. Junyang Chen 0001, Huan Wang 0005, Yirui Wu, Qiuzhen Lin, Yunfeng Diao, Junkai Ji |
AAAI | 1 |
| 2026 | Multi-Source Unsupervised Graph Domain Adaptation via Concise Propagation-Transformation PipelineabstractUnsupervised graph domain adaptation (UGDA) aims to transfer knowledge from a labeled source graph to an unlabeled target graph, addressing the performance degradation caused by distributional shifts in node attributes and graph structures across domains. Despite recent progress, existing UGDA approaches still face two key challenges: (C1) Data-level: Most methods rely on a single source domain, overlooking the complementary knowledge that could be leveraged from multiple sources. (C2) Model-level: Many UGDA models emphasize complex, handcrafted Graph neural network (GNN) architectures, while simpler yet effective designs with propagation (P) & transformation (T) pipeline remain underexplored. To address these challenges, in this paper, we propose a novel approach, which leverages Concise Propagation–Transformation pipeline for multi-source unsupervised Graph Domain Adaptation, dubbed as CPT-GDA, to better capture complementary knowledge from multiple sources in an efficient manner. Specifically, the proposed CPT-GDA adopts a dual-branch GNN architecture with different depths of propagation but the same P-T patterns, which enables the model to efficiently learn node representations to mitigate domain discrepancy. Meanwhile, to facilitate effective knowledge transfer across graphs, we derive three optimization objectives: (1) the classifier loss to learn discriminative representations; (2) the alignment loss weighted by the graph Wasserstein distance to align the structure and feature distribution; and (3) the pseudo-label loss to refine target node representations. Extensive experiments on real-world datasets confirm that the proposed method outperforms recent state-of-the-art baselines, demonstrating its effectiveness. Yi Li 0018, Xin Zheng 0008, Junyang Chen 0001, Yanqing Guo, Alan Wee-Chung Liew, Shirui Pan |
WWW | 4 |
| 2026 | Enhancing news classification: domain-specific guided pretraining based on adaptive selective masking
Qiao Ding, Heng Ding, Jian Wang 0078, Yantuan Xian, Nanyu Li, Junyang Chen 0001 |
Knowl. Based Syst. | 8 |
| 2026 | Alignment-aware fine-tuning of vision-language models for out-of-distribution generalization
Yirui Wu, Mohammed A.-M. Salem, Lixin Yuan, Junyang Chen 0001, Huan Wang 0005, Shaohua Wan 0001 |
Multim. Syst. | 5 |
| 2026 | PRISM: Link Prediction in Attributed Networks With Uncertain ModalitiesabstractLink prediction for attributed graphs has garnered significant attention due to its ability to enhance predictive performance by leveraging multi-modal node attributes. However, real-world challenges such as privacy concerns, content restrictions, and attribute constraints often result in nodes facing varying degrees of missing modalities in their attributes, significantly limiting the effectiveness of existing approaches. Building on this fact, we propose a model for linkPRediction in attrIbuted networkSwith uncertainModalities (PRISM), which learns the shared representations across various scenarios of missing modalities through dual-level adversarial training.PRISMcomprises four modules,i.e.,a GCN extractor, an adversarial extractor, an attentive fusion, and an adaptive aggregator. The GCN extractor leverages graph convolutional networks (GCN) to extract fundamental representations from the network topology. The adversarial extractor employs dual-level adversarial training to acquire the shared representations across various multi-modal scenarios at the node-level and link-level, respectively. The attentive fusion applies the multi-head attention mechanism to integrate the shared representations and the fundamental representations. The adaptive aggregator comprehensively considers both node-level and link-level representations to predict the existence of links. Experimental evaluation using real-world datasets demonstrates thatPRISMsignificantly outperforms existing state-of-the-art link prediction methods for multi-modal attributed graphs under missing modalities by improving the Recall@50 metric (R@50) by up to 38.79%. Muhammad Asif Ali, Huan Wang 0005, Zhongfei Zhang, Junyang Chen 0001, Di Wang 0015 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | Semi-Supervised Breast Lesion Segmentation Using Confidence-Ranked Features and Bi-Level PrototypesabstractAutomated lesion segmentation through breast ultrasound (BUS) images is an essential prerequisite in computer-aided diagnosis. However, the task of breast segmentation remains challenging, due to the time-consuming and labor-intensive process of acquiring precise labeled data, as well as severely ambiguous lesion boundaries and low contrast in BUS images. In this article, we propose a novel semi-supervised breast segmentation framework based on confidence-ranked features and bi-level prototypes (CoBiNet) to alleviate these issues. Our outputs are derived from two branches: classifier and projector. In the projector branch, we first rank the features by multilevel sampling to obtain multiple feature sets with different confidence levels. Then, these sets are progressed in two directions. One is to acquire local prototypes at each level by local sampling and perform trans-confidence level (TCL) contrastive learning. This encourages the low-confidence features to converge to the high-confidence features, which enhances the model's ability to recognize ambiguous regions. The other process is to generate more representative global prototypes by global sampling, followed by generating more reliable predictions and performing cross-guidance (CG) consistency learning with the classifier output predictions, facilitating knowledge transfer between the structure-aware projector and the category-discriminative classifier branches. Extensive experiments on two well-known public datasets, BUSI and UDIAT, demonstrate the superiority of our method over state-of-the-art approaches. Codes will be released upon publication. Siyao Jiang, Huisi Wu, Yu Zhou 0027, Junyang Chen 0001, Harry Qin |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2026 | Generative Regularities in Multi-Layer Networks: A Shared-Latent Space Representation ApproachabstractUnderstanding structural regularities across layers in multi-layer networks is essential for uncovering their underlying generative mechanisms. While link prediction has been widely explored in multi-layer networks, it is typically treated as an isolated technical problem, often missing its broader implications for network structure and the mechanisms driving edge formation. In this article, we investigate the extent to which network layers exhibit shared generative regularities. By examining the alignment of latent representations across layers, we assess the similarity of their underlying mechanisms and leverage this alignment to improve predictive performance. To facilitate this, we introduce a new metric, C ross- L ayer G enerative C onsistency ( CLGC ), which quantitatively captures the degree of structural and generative alignment between network layers. CLGC is grounded in the shared-latent space framework, positing that layers generated by similar mechanisms will produce compatible latent representations. To realize this approach, we present SupportNet – Support prediction and consistency analysis in multi-layer Net works–a GCN-based model augmented with adversarial training to effectively learn robust shared-latent space representations. These representations support both accurate link prediction and interpretable evaluation of cross-layer generative consistency. Experiments on real-world multi-layer networks demonstrate that SupportNet delivers strong link prediction results improving AUC by 17.47%, AP by 40.41% and AUPR by 39.59% on the Kapferer dataset, while CLGC reveals significant patterns of structural and generative alignment among layers. Muhammad Asif Ali, Anyu Xue, Huan Wang 0005, Junyang Chen 0001, Di Wang 0015 |
ACM Trans. Web | 5 |
| 2025 | Deconfound Semantic Shift and Incompleteness in Incremental Few-shot Semantic SegmentationabstractIncremental few-shot semantic segmentation (IFSS) expands segmentation capacity of the trained model to segment new-class images with few samples. However, semantic meanings may shift from background to object class or vice versa during incremental learning. Moreover, new-class samples often lack representative attribute features when the new class greatly differs from the pre-learned old class. In this paper, we propose a causal framework to discuss the cause of semantic shift and incompleteness in IFSS, and we deconfound the revealed causal effects from two aspects. First, we propose a Causal Intervention Module (CIM) to resist semantic shift. CIM progressively and adaptively updates prototypes of old class, and removes the confounder in an intervention manner. Second, a Prototype Refinement Module (PRM) is proposed to complete the missing semantics. In PRM, knowledge gained from the episode learning scheme assists in fusing features of new-class and old-class prototypes. Experiments on both PASCAL-VOC 2012 and ADE20k benchmarks demonstrate the outstanding performance of our method. Yirui Wu, Yuhang Xia, Lixin Yuan, Junyang Chen 0001, Jun Liu 0036, Shaohua Wan 0001 |
AAAI | 5 |
| 2025 | HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video UnderstandingabstractMultimodal large language models have become a popular topic in deep visual understanding due to many promising real-world applications. However, hour-long video understanding, spanning over one hour and containing tens of thousands of visual frames, remains under-explored because of 1) challenging long-term video analyses, 2) inefficient large-model approaches, and 3) lack of large-scale benchmark datasets. Among them, in this paper, we focus on building a large-scale hour-long long video benchmark, HLV-1K1, designed to evaluate long video understanding models. HLV-1K comprises 1009 hour-long videos with 14,847 high-quality question answering (QA) and multi-choice question asnwering (MCQA) pairs with time-aware query and diverse annotations, covering frame-level, within-event-level, cross-event-level, and long-term reasoning tasks. We evaluate our benchmark using existing state-of-the-art methods and demonstrate its value for testing deep long video understanding capabilities at different levels and for various tasks. This includes promoting future long video understanding tasks at a granular level, such as deep understanding of long live videos, meeting recordings, and movies. Heqing Zou, Tianze Luo, Guiyang Xie, Victor Xiao Jie Zhang, Fengmao Lv, Guangcong Wang, Junyang Chen 0001, Zhuochen Wang, Hansheng Zhang, Huaijian Zhang |
ICME | 7 |
| 2025 | ABNet: Mitigating Sample Imbalance in Anomaly Detection Within Dynamic GraphsabstractIn dynamic graphs, detecting anomalous nodes faces challenges due to sample imbalance, stemming from the scarcity of anomalous samples and feature representation bias. Existing methods often use unsupervised or semi-supervised learning to extract anomalous samples from unlabeled data, but struggle to obtain enough anomalous instances due to their low occurrence. Moreover, GNN-based approaches often prioritize normal samples, neglecting rare anomalies. To address these issues, we propose the Anomaly Balance Network (ABNet), designed to alleviate sample imbalance and enhance anomaly detection. ABNet includes three key components: a feature extractor that compares node features across time points to avoid bias, an anomaly augmenter that amplifies anomaly details and generates diverse anomalous samples, and an anomaly detector using meta-learning to adapt to graph evolution. Experimental results show that ABNet outperforms existing methods on three real-world datasets, effectively addressing sample imbalance. Yifan Hong 0001, Muhammad Asif Ali, Huan Wang 0005, Junyang Chen 0001, Di Wang 0015 |
IJCAI | 4 |
| 2025 | Diffuse&Refine: Intrinsic Knowledge Generation and Aggregation for Incremental Object DetectionabstractIncremental Object Detection(IOD) targets at progressively extending capability of object detectors to recognize new classes. However, representation confusion between old and new classes leads to catastrophic forgetting. To alleviate this problem, we propose DiffKA, with intrinsic knowledge generated and aggregated by forward and backward diffusion, gradually establishing rigid class boundary. With incremental streaming data, forward diffusion spreads information to generate potential inter-class associations among new- and old-class prototypes within a hierarchical tree, named as Intrinsic Correlation Tree(ICTree), to store intrinsic knowledge. Afterwards, backward diffusion refines and aggregates the generated knowledge in ICTree, explicitly establishing rigid class boundary to mitigate representation confusion. To keep semantic consistency with extreme IOD settings, we reorganize semantic relevance of old- and new-class prototypes in paradigms to adaptively and effectively update DiffKA. Experiments on MS COCO dataset show DiffKA achieves state-of-the-art performance on IOD tasks with significant advantages. Yirui Wu, Lixin Yuan, Jun Liu 0036, Junyang Chen 0001, Huan Wang 0005, Wenhai Wang |
IJCAI | 6 |
| 2025 | Deduction with Induction: Combining Knowledge Discovery and Reasoning for Interpretable Deep Reinforcement LearningabstractDeep reinforcement learning (DRL) has achieved remarkable success in dynamic decision-making tasks. However, its inherent opacity and cold start problem hinder transparency and training efficiency. To address these challenges, we propose HRL-ID, a neural-symbolic framework that combines automated rule discovery with logical reasoning within a hierarchical DRL structure. HRL-ID dynamically extracts first-order logic rules from environmental interactions, iteratively refines them through success-based updates, and leverages these rules to guide action execution during training. Extensive experiments on Atari benchmarks demonstrate that HRL-ID outperforms state-of-the-art methods in training efficiency and interpretability, achieving higher reward rates and successful knowledge transfer between domains. Junyang Chen 0001, Yuanfeng Song, Rui Mao 0001, Fangzhen Lin |
IJCAI | 3 |
| 2025 | LUSTER: Link Prediction Utilizing Shared-Latent Space Representation in Multi-Layer NetworksabstractLink prediction in multi-layer networks is a longstanding issue that predicts missing links based on the observed structures across all layers. Existing link prediction methods in multi-layer network typically merge the multi-layer network into a single-layer network and/or perform explicit calculations using intra-layer and inter-layer similarity metrics. However, these approaches often overlook the role of coupling in multi-layer networks, specifically the shared information and latent relationships between layers, which in turn limits prediction performance. This calls the need for methods that can extract representations in a shared-latent space to enhance inter-layer information sharing and prediction performance. In this paper, we propose a novel end-to-end framework namely: Link prediction Utilizing Shared-laTent spacE Representation (LUSTER) in multi-layer networks. LUSTER consists of four key modules: the representation extractor, the latent space learner, the complementary enhancer, and the link predictor. The representation extractor focuses on learning the intra-layer representations of each layer, capturing the data characteristics within the layer. The latent space learner extracts representations from the shared-latent space across different network layers through adversarial training. The complementary enhancer combines the intra-layer representations and the shared-latent space representations through orthogonal fusion, providing comprehensive information. Finally, the link predictor uses the enhanced representations to predict missing links. Extensive experimental analyses demonstrate that LUSTER outperforms state-of-the-art methods for link prediction in multi-layer networks, improving the AUC metric by up to 15.87%. Muhammad Asif Ali, Huan Wang 0005, Junyang Chen 0001, Di Wang 0015 |
WWW | 4 |
| 2025 | Traffic prediction and load balancing routing algorithm based on deep Q-network for SD-IoT
Qiao Ding, Nanyu Li, Heng Ding, Jian Wang 0078, Yongqing Chen, Yantuan Xian, Junyang Chen 0001 |
Adv. Eng. Informatics | 8 |
| 2025 | Unveiling user interests: A deep user interest exploration network for sequential location recommendation
Junyang Chen 0001, Jingcai Guo, Qin Zhang 0011, Kaishun Wu, Liangjie Zhang, Victor C. M. Leung, Huan Wang 0005, Zhiguo Gong |
Inf. Sci. | 1 |
| 2025 | CGraphNet: Contrastive Graph Context Prediction for Sparse Unlabeled Short Text Representation Learning on Social MediaabstractUnlabeled text representation learning (UTRL), encompassing static word embeddings such as Word2Vec and contextualized word embeddings such as bidirectional encoder representations from transformer (BERT), aims to capture semantic word relationships in a low-dimensional space without the need for manual labeling. These word embeddings are invaluable for downstream tasks such as document classification and clustering. However, the surge of short texts generated daily on social media platforms results in sparse word cooccurrences, compromising UTRL outcomes. Contextualized models such as recurrent neural network (RNN) and BERT, while impressive, often struggle with predicting the next word due to sparse word sequences in short texts. To address this, we introduce CGraphNet, a contrastive graph context prediction model designed for UTRL. This approach converts short texts into graphs, establishing links between sequentially occurring words. Information from the next word and its neighbors informs the target prediction, a process referred to as graph context prediction, mitigating sparse word cooccurrence issues in brief sentences. To minimize noise, an attention mechanism assigns importance to neighbors, while a contrastive objective encourages more distinctive representations by comparing the target word with its neighbors. Our experiments demonstrate CGraphNet's superior performance over other baselines, particularly in classification and clustering tasks on real-world datasets. Junyang Chen 0001, Jingcai Guo, Xueliang Li 0002, Huan Wang 0005, Zhenghua Xu 0001, Zhiguo Gong, Liang-Jie Zhang, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | A Review of Few-Shot and Zero-Shot Learning for Node Classification in Social NetworksabstractNode classification tasks aim to assign labels or categories to entire graphs based on their structural properties or node attributes. It can be adopted for various types of graph systems, including but not limited to network traffic, biological networks, knowledge graphs, etc., especially to social networks. This problem is well-studied, and solutions have demonstrated significant success in numerous real-world applications. However, in the situation where emerging categories are scarce or even have no labeled data, classical methods perform poorly on the whole, which has attracted growing attention. Based on this, in this article, we divide researches for node classification in social networks into two broad categories: traditional methods and novel strategies (few-shot/zero-shot learning). In traditional node classification methods, we summarize some classical methods for both homogeneous and heterogeneous networks, which includes unsupervised classifier, matrix factorization techniques, supervised methods, random-walk, and meta-path. Meanwhile, we introduce novel methods in few-shot or zero-shot learning. The article outlines the technical principles of various methods and analyzes their performance across different classes. It further summarizes the benchmark datasets used for evaluating node classification tasks. Finally, the major opportunities, challenges, and future research directions in few-shot and zero-shot learning for node classification in graph scenarios are discussed. Junyang Chen 0001, Rui Mi, Huan Wang 0005, Huisi Wu, Jiqian Mo, Jingcai Guo, Zhihui Lai 0001, Liang-Jie Zhang, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2025 | Multitask Asynchronous Metalearning for Few-Shot Anomalous Node Detection in Dynamic NetworksabstractFew-shot anomalous node detection in dynamic networks has been extensively investigated in the field of research. In this few-shot scenario, the detection of these anomalous nodes is particularly challenging due to the continuously evolving network topology and data distribution over time, which is known as concept drift. Concept drift refers to the phenomenon where the underlying concepts or patterns in the data generation process change over time, leading to varying data distributions across different periods. Due to these changes in data distribution, the patterns learned during training may become invalid under the new data distribution. Existing models primarily aim to enhance the representation of evolving node attributes and relationships to mitigate the impact of concept drift in few-shot scenarios. However, the scarcity of anomalous samples further limits the model's ability to learn new patterns, thereby reducing its effectiveness in addressing concept drift in few-shot scenarios. To address this challenge, we propose the multitask asynchronous metalearning framework (MAMF), which aims to mitigate bias induced by concept drift in few-shot anomalous node detection. Our framework consists of four main components: a feature extractor, an anomaly simulator, an asynchronous learner, and a type detector. The feature extractor captures the relative variations of each node in an evolving graph stream. The anomaly simulator uses generative adversarial models to learn anomaly distributions and generate samples at different time intervals. The asynchronous learner samples from various time distributions to create metatasks for anomalous node detection, allowing it to adapt to changes between these distributions. To aid in few-shot anomalous node detection, the type detector is used for anomaly type recognition. Our framework achieves AUC improvements of 5.12%, 6.87%, and 1.91% over the best existing methods on Wikipedia, Reddit, and Mooc datasets, respectively, demonstrating its effectiveness and robustness in adapting to concept drift and detecting anomalous nodes. Yifan Hong 0001, Chuanqi Shi, Junyang Chen 0001, Huan Wang 0005, Di Wang 0015 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2025 | CIG2S: A Cross-View Image Geo-Localization Model Based on G2S Transform Suitable for Center-Misaligned ScenariosabstractIn multimedia social networks, the user's geo-location can be inferred by matching his shared images with the referenced satellite images, viz. cross-view image geo-localization. Although the existing most cross-view image geo-localization methods perform well in the center-misaligned scenario, in practical application, the shooting location of the query ground image is most likely not aligned with the center point of satellite images. Then, their geo-localization accuracy would drastically decrease. Therefore, we propose a novel cross-view image geo-localization model based on ground-to-satellite (G2S) transform, named CIG2S. First, the queried ground image is transformed into the aerial-view by spherical transform, generating G2S images, which could improve the similarity between ground and satellite images. Second, multiscale features are extracted from the original ground image, G2S images, and satellite images by twins-PCPVT. Furthermore, a dynamic similarity weighted loss function is designed to measure the distance between the query ground image and the referenced satellite image. Experimental results on three center-misaligned datasets, including VIGOR and the center-misaligned versions of CVUSA and CVACT, demonstrate that the proposed CIG2S model can significantly improve the geo-localization accuracy. For example, when compared with another vision-transformer-based model L2LTR-polar, CIG2S can outperform about 6.6% and 15.8% in the center-misaligned datasets CVUSA_CM and CVACT_CM. Jiangshan Li, Chunfang Yang, Baojun Qi, Ma Zhu, Junyang Chen 0001, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | Adaptive Density Estimation for Personalized Recommendations Across Varied User Activity LevelsabstractTop-N recommendation systems are recognized as highly effective for delivering personalized services that cater to the varied interests of users. Nonetheless, current state-of-the-art (SOTA) analyses reveal a marked variability in their performance across users with differing levels of activity, which substantially undermines the quality of personalized recommendation services. Prevailing research tends to overlook this discrepancy, often presuming a uniform probability distribution in user preferences and employing a static model (such as a single latent vector) for user representation. This oversimplification impedes the adaptability of existing models to accommodate the spectrum of user activity levels. In our research, we introduce the variational kernel density estimation (VKDE) approach, an innovative nonparametric method designed to accurately capture the unique preference distributions of individual users. The VKDE framework integrates multiple local distributions to construct a comprehensive global preference profile for each user. We have developed a novel variational kernel function that delineates user-specific interests and constructs each local distribution accordingly. Additionally, we present a tailored sampling strategy that simplifies the complexity of the training process while preserving the efficacy of the recommendations. Empirical evaluations conducted on four widely recognized public datasets demonstrate that our VKDE model achieves superior performance over the SOTA alternatives, significantly enhancing accuracy for users with a broad range of activity levels. Wei Liu 0061, Huaijie Zhu, Jianxing Yu, Libin Zheng 0001, Jian Yin 0001, Ruishi Liang, Xin Liu 0101, Junyang Chen 0001, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2025 | EPM: Evolutionary Perception Method for Anomaly Detection in Noisy Dynamic GraphsabstractWith the rapid expansion of interactions across various domains such as knowledge graphs and social networks, anomaly detection in dynamic graphs has become increasingly critical for mitigating potential risks. However, existing anomaly detection methods often assume noise-free dynamic graphs, overlooking the prevalence of noisy dynamic graphs in real-world applications. Specifically, noisy dynamic graphs affected by structural noises-such as spurious and missing nodes and edges-struggle to consistently provide reliable structural evidence for anomaly detection. To tackle this challenge, we propose an Evolutionary Perception Method (EPM) for identifying anomalous nodes in noisy dynamic graphs by resisting the interference of structural noises. EPM primarily consists of two components: a dynamic fitter and a filtering reviser. The dynamic fitter characterizes the interaction dynamics of nodes that removes and generates links at each period as a multiple superposition state, utilizing various link prediction algorithms to fit evolutionary mechanisms. Additionally, the filtering reviser designs evolutional entropies to quantify the evolutional uncertainty in multiple superposition states, further designing the Kalman filter to optimize these entropies. Extensive experiments show that the proposed EPM method surpasses state-of-the-art approaches in detecting anomalous nodes in noisy dynamic graphs. Huan Wang 0005, Junyang Chen 0001, Yirui Wu, Victor C. M. Leung, Di Wang 0015 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Hybrid Reinforced Medical Report Generation With M-Linear Attention and Repetition PenaltyabstractTo reduce doctors' workload, deep-learning-based automatic medical report generation has recently attracted more and more research efforts, where deep convolutional neural networks (CNNs) are employed to encode the input images, and recurrent neural networks (RNNs) are used to decode the visual features into medical reports automatically. However, these state-of-the-art methods mainly suffer from three shortcomings: 1) incomprehensive optimization; 2) low-order and unidimensional attention; and 3) repeated generation. In this article, we propose a hybrid reinforced medical report generation method with m-linear attention and repetition penalty mechanism (HReMRG-MR) to overcome these problems. Specifically, a hybrid reward with different weights is employed to remedy the limitations of single-metric-based rewards, and a local optimal weight search algorithm is proposed to significantly reduce the complexity of searching the weights of the rewards from exponential to linear. Furthermore, we use m-linear attention modules to learn multidimensional high-order feature interactions and to achieve multimodal reasoning, while a new repetition penalty is proposed to apply penalties to repeated terms adaptively during the model's training process. Extensive experimental studies on two public benchmark datasets show that HReMRG-MR greatly outperforms the state-of-the-art baselines in terms of all metrics. The effectiveness and necessity of all components in HReMRG-MR are also proved by ablation studies. Additional experiments are further conducted and the results demonstrate that our proposed local optimal weight search algorithm can significantly reduce the search time while maintaining superior medical report generation performances. Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Chang Qi, Thomas Lukasiewicz |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Reliable Service Recommendation: A Multi-Modal Adversarial Method for Personalized Recommendation Under Uncertain Missing ModalitiesabstractPersonalized recommendation is of paramount importance in online content platforms like Kuai and Tencent. To ensure accurate recommendations, it is crucial to consider multi-modal information in both items and user-user/item interactions. While existing works on multimedia recommendation have made strides in leveraging multi-modal contents to enrich item representations, many of them overlook the practical scenario of multiple modality missing. As a result, the performance of recommendation systems can be significantly compromised in such cases. In this paper, we introduce a novel multi-modal adversarial method called$MMAM$, which aims to provide reliable personalized recommendation services even in the presence of uncertain missing modalities. The core idea behind$MMAM$is to design a generator that can effectively encode both user-user/item interactions and multi-modal contents, taking into account various missing cases. The generator is trained to learn transferable features from different combinations of missing modalities in order to deceive a discriminative classifier. Additionally, we propose a modal discriminator that can classify the missing cases of multi-modalities, further enhancing the capability of the model. Moreover, a well-equipped predictor utilizes the transferable features to predict potential user interests. To improve the prediction accuracy, we design a type discriminator that enhances the classification of link types. By employing a mini-max game between the generator and the discriminators,$MMAM$successfully obtains transferable features that encompass multi-modal contents, even when facing uncertain missing modalities. We conduct extensive experiments on industrial datasets, including Kuai and Tencent. Comparing with state-of-the-art approaches, MMAM achieves improvements in personalized recommendation tasks under uncertain missing modalities. MMAM holds promise for enhancing multi-modal personalized recommendations in real-world applications. Junyang Chen 0001, Jingcai Guo, Huan Wang 0005, Kaishun Wu, Liang-Jie Zhang |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | Knowledge Transfer Service for Multi-Granular Traffic Prediction in Smart CitiesabstractTraffic prediction is essential for the efficient resource allocation and management of smart cities. However, many cities face challenges in accessing sufficient traffic data due to high collection costs and privacy concerns, which impacts the accuracy of predictions. To address the issue of data sparsity, transfer learning methods have been explored with promising outcomes. Nonetheless, key challenges remain in determining which knowledge to transfer and how to transfer it effectively. In this paper, we propose a Knowledge Transfer Service for multi-granular traffic prediction in smart cities, leveraging Multi-Granular Spatiotemporal Pattern Transfer Learning (MGSTPAT). MGSTPAT utilizes meta-learning to acquire rich, multi-granularity spatiotemporal knowledge from data-abundant source cities and initializes robust models. Additionally, our Embedded Adaptive Clustering module generates task-specific multi-granularity patterns, while the Granularity-Aware Pattern Attention mechanism enables target cities with limited data to select the most relevant patterns, enhancing prediction accuracy. Extensive experiments on real-world datasets demonstrate the effectiveness of our approach in addressing traffic prediction challenges across cities. Jiqian Mo, Junyang Chen 0001, Zhiguo Gong |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Sparse Enhanced Network: An Adversarial Generation Method for Robust Augmentation in Sequential RecommendationabstractSequential Recommendation plays a significant role in daily recommendation systems, such as e-commerce platforms like Amazon and Taobao. However, even with the advent of large models, these platforms often face sparse issues in the historical browsing records of individual users due to new users joining or the introduction of new products. As a result, existing sequence recommendation algorithms may not perform well. To address this, sequence-based data augmentation methods have garnered attention. Existing sequence enhancement methods typically rely on augmenting existing data, employing techniques like cropping, masking prediction, random reordering, and random replacement of the original sequence. While these methods have shown improvements, they often overlook the exploration of the deep embedding space of the sequence. To tackle these challenges, we propose a Sparse Enhanced Network (SparseEnNet), which is a robust adversarial generation method. SparseEnNet aims to fully explore the hidden space in sequence recommendation, generating more robust enhanced items. Additionally, we adopt an adversarial generation method, allowing the model to differentiate between data augmentation categories and achieve better prediction performance for the next item in the sequence. Experiments have demonstrated that our method achieves a remarkable 4-14% improvement over existing methods when evaluated on the real-world datasets. (https://github.com/junyachen/SparseEnNet) Junyang Chen 0001, Guoxuan Zou, Pan Zhou 0001, Yirui Wu, Zhenghan Chen, Houcheng Su, Huan Wang 0005, Zhiguo Gong |
AAAI | 1 |
| 2024 | Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageabstractGiven an untrimmed video and a sentence query, video moment retrieval using language (VMR) aims to locate a target query-relevant moment. Since the untrimmed video is overlong, almost all existing VMR methods first sparsely down-sample each untrimmed video into multiple fixed-length video clips and then conduct multi-modal interactions with the query feature and expensive clip features for reasoning, which is infeasible for long real-world videos that span hours. Since the video is downsampled into fixed-length clips, some query-related frames may be filtered out, which will blur the specific boundary of the target moment, take the adjacent irrelevant frames as new boundaries, easily leading to cross-modal misalignment and introducing both boundary-bias and reasoning-bias. To this end, in this paper, we propose an efficient approach, SpotVMR, to trim the query-relevant clip. Besides, our proposed SpotVMR can serve as plug-and-play module, which achieves efficiency for state-of-the-art VMR methods while maintaining good retrieval performance. Especially, we first design a novel clip search model that learns to identify promising video regions to search conditioned on the language query. Then, we introduce a set of low-cost semantic indexing features to capture the context of objects and interactions that suggest where to search the query-relevant moment. Also, the distillation loss is utilized to address the optimization issues arising from end-to-end joint training of the clip selector and VMR model. Extensive experiments on three challenging datasets demonstrate its effectiveness. Daizong Liu, Wanlong Fang, Pan Zhou 0001, Zichuan Xu, Wenzheng Xu, Junyang Chen 0001, Renfu Li |
AAAI | 7 |
| 2024 | Sharpness-Aware Model-Agnostic Long-Tailed Domain GeneralizationabstractDomain Generalization (DG) aims to improve the generalization ability of models trained on a specific group of source domains, enabling them to perform well on new, unseen target domains. Recent studies have shown that methods that converge to smooth optima can enhance the generalization performance of supervised learning tasks such as classification. In this study, we examine the impact of smoothness-enhancing formulations on domain adversarial training, which combines task loss and adversarial loss objectives. Our approach leverages the fact that converging to a smooth minimum with respect to task loss can stabilize the task loss and lead to better performance on unseen domains. Furthermore, we recognize that the distribution of objects in the real world often follows a long-tailed class distribution, resulting in a mismatch between machine learning models and our expectations of their performance on all classes of datasets with long-tailed class distributions. To address this issue, we consider the domain generalization problem from the perspective of the long-tail distribution and propose using the maximum square loss to balance different classes which can improve model generalizability. Our method's effectiveness is demonstrated through comparisons with state-of-the-art methods on various domain generalization datasets. Code: https://github.com/bamboosir920/SAMALTDG. Houcheng Su, Weihao Luo, Daixian Liu, Mengzhu Wang, Junyang Chen 0001, Cong Wang 0018, Zhenghan Chen |
AAAI | 6 |
| 2024 | ROG_PL: Robust Open-Set Graph Learning via Region-Based Prototype LearningabstractOpen-set graph learning is a practical task that aims to classify the known class nodes and to identify unknown class samples as unknowns. Conventional node classification methods usually perform unsatisfactorily in open-set scenarios due to the complex data they encounter, such as out-of-distribution (OOD) data and in-distribution (IND) noise. OOD data are samples that do not belong to any known classes. They are outliers if they occur in training (OOD noise), and open-set samples if they occur in testing. IND noise are training samples which are assigned incorrect labels. The existence of IND noise and OOD noise is prevalent, which usually cause the ambiguity problem, including the intra-class variety problem and the inter-class confusion problem. Thus, to explore robust open-set learning methods is necessary and difficult, and it becomes even more difficult for non-IID graph data. To this end, we propose a unified framework named ROG_PL to achieve robust open-set learning on complex noisy graph data, by introducing prototype learning. In specific, ROG_PL consists of two modules, i.e., denoising via label propagation and open-set prototype learning via regions. The first module corrects noisy labels through similarity-based label propagation and removes low-confidence samples, to solve the intra-class variety problem caused by noise. The second module learns open-set prototypes for each known class via non-overlapped regions and remains both interior and border prototypes to remedy the inter-class confusion problem. The two modules are iteratively updated under the constraints of classification loss and prototype diversity loss. To the best of our knowledge, the proposed ROG_PL is the first robust open-set node classification method for graph data with complex noise. Experimental evaluations of ROG_PL on several benchmark graph datasets demonstrate that it has good performance. Qin Zhang 0011, Jiexin Lu, Liping Qiu, Shirui Pan, Xiaojun Chen 0006, Junyang Chen 0001 |
AAAI | 7 |
| 2024 | PH-Net: Semi-Supervised Breast Lesion Segmentation via Patch-Wise HardnessabstractWe present a novel semi-supervised framework for breast ultrasound (BUS) image segmentation, which is a very challenging task owing to (1) large scale and shape variations of breast lesions and (2) extremely ambiguous boundaries caused by massive speckle noise and artifacts in BUS images. While existing models achieved certain progress in this task, we believe the main bottleneck nowadays for further improvement is that we still cannot deal with hard cases well. Our framework aims to break through this bottleneck, which includes two innovative components: an adaptive patch augmentation scheme and a hard-patch contrastive learning module. We first identify hard patches by computing the average entropy of each patch and then shield hard patches to prevent them from being cropped out while performing random patch cutmix. Such a scheme is able to prevent hard regions from being inadequately trained under strong augmentation. We further develop a new hard-patch contrastive learning algorithm to direct model attention to hard regions by applying extra contrast to pixels in hard patches, further improving segmentation performance on hard cases. We demonstrate the superior-ity of our framework to state-of-the-art approaches on two famous BUS datasets, achieving better performance under different labeling conditions. The code is available at https://github.com/jjjsyyy/PH-Net. Siyao Jiang, Huisi Wu, Junyang Chen 0001, Qin Zhang 0011, Harry Qin |
CVPR | 3 |
| 2024 | Hiding Imperceptible Noise in Curvature-Aware Patches for 3D Point Cloud Attack
Daizong Liu, Keke Tang, Pan Zhou 0001, Lixing Chen, Junyang Chen 0001 |
ECCV (30) | 6 |
| 2024 | CONC: Complex-noise-resistant Open-set Node Classification with Adaptive Noise Detection
Qin Zhang 0011, Jiexin Lu, Huisi Wu, Shirui Pan, Junyang Chen 0001 |
IJCAI | 6 |
| 2024 | ParsNets: A Parsimonious Composition of Orthogonal and Low-Rank Linear Networks for Zero-Shot Learning
Jingcai Guo, Qihua Zhou, Xiaocheng Lu, Ruibin Li, Jie Zhang 0076, Junyang Chen 0001, Xin Xie 0001, Song Guo 0001 |
IJCAI | 8 |
| 2024 | EGonc : Energy-based Open-Set Node Classification with substitute UnknownsabstractOpen-set Classification (OSC) is a critical requirement for safely deploying machine learning models in the open world, which aims to classify samples from known classes and reject samples from out-of-distribution (OOD).
Existing methods exploit the feature space of trained network and attempt at estimating the uncertainty in the predictions.
However, softmax-based neural networks are found to be overly confident in their predictions even on data they have never seen before and
the immense diversity of the OOD examples also makes such methods fragile.
To this end, we follow the idea of estimating the underlying density of the training data to decide whether a given input is close to the in-distribution (IND) data and adopt Energy-based models (EBMs) as density estimators.
A novel energy-based generative open-set node classification method, \textit{EGonc}, is proposed to achieve open-set graph learning.
Specifically, we generate substitute unknowns to mimic the distribution of real open-set samples firstly, based on the information of graph structures.
Then, an additional energy logit representing the virtual OOD class is learned from the residual of the feature against the principal space, and matched with the original logits by a constant scaling. This virtual logit serves as the indicator of OOD-ness.
EGonc has nice theoretical properties that guarantee an overall distinguishable margin between the detection scores for IND and OOD samples.
Comprehensive experimental evaluations of EGonc also demonstrate its superiority. Qin Zhang 0011, Zelin Shi, Shirui Pan, Junyang Chen 0001, Huisi Wu, Xiaojun Chen 0006 |
NeurIPS | 4 |
| 2024 | Open-world structured sequence learning via dense target encoding
Qin Zhang 0011, Qincai Li, Haolong Xiang, Zhizhi Yu, Junyang Chen 0001, Peng Zhang 0001, Xiaojun Chen 0006 |
Inf. Sci. | 6 |
| 2024 | CUPID: Improving Battle Fairness and Position Satisfaction in Online MOBA Games with a Re-matchmaking SystemabstractThe multiplayer online battle arena (MOBA) genre has gained significant popularity and economic success, attracting considerable research interest within the Human-Computer Interaction community. Enhancing the gaming experience requires a deep understanding of player behavior, and a crucial aspect of MOBA games is matchmaking, which aims to assemble teams of comparable skill levels. However, existing matchmaking systems often neglect important factors such as players' position preferences and team assignment, resulting in imbalanced matches and reduced player satisfaction. To address these limitations, this paper proposes a novel framework called CUPID, which introduces a novel process called ''re-matchmaking'' to optimize team and position assignments to improve both fairness and player satisfaction. CUPID incorporates a pre-filtering step to ensure a minimum level of matchmaking quality, followed by a pre-match win-rate prediction model that evaluates the fairness of potential assignments. By simultaneously considering players' position satisfaction and game fairness, CUPID aims to provide an enhanced matchmaking experience. Extensive experiments were conducted on two large-scale, real-world MOBA datasets to validate the effectiveness of CUPID. The results surpass all existing state-of-the-art baselines, with an average relative improvement of 7.18% in terms of win prediction accuracy. Furthermore, CUPID has been successfully deployed in a popular online MOBA game. The deployment resulted in significant improvements in match fairness and player satisfaction, as evidenced by critical Human-Computer Interaction (HCI) metrics covering usability, accessibility, and engagement, observed through A/B testing. To the best of our knowledge, Cupid is the first re-matchmaking system designed specifically for large-scale MOBA games. Ge Fan, Chaoyun Zhang, Junyang Chen 0001, Zenglin Xu |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2024 | SiamATTRPN: Enhance Visual Tracking With Channel and Spatial AttentionabstractVisual tracking is an important research topic in the field of computer vision. The current Siamese tracker based on the region proposal network (SiamRPN) has achieved promising tracking results in terms of efficiency and performance. However, through our empirical study, we have observed that deep features learned by SiamRPN are of substandard quality, as the salient regions within the deep features fail to correspond accurately with meaningful objects. To address this limitation, we propose an approach to enhance the quality of the learned deep features through the incorporation of an attention mechanism. Attention mechanisms have been shown to be effective in distinguishing similar objects, as they suppress background objects while highlighting target information that is most relevant. As a result, a new tracking method with channel and spatial attention termed SiamATTRPN is explored. To verify the effectiveness of SiamATTRPN, experiments on benchmark datasets demonstrate that our proposed tracker outperforms the baseline tracker significantly. Huayue Cai, Xiang Zhang 0008, Long Lan, Wenxin Shen, Junyang Chen 0001, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Compressing the Multiobject Tracking Model via Knowledge DistillationabstractRecent multiobject tracking (MOT) methods usually use very deep neural networks to achieve competitive accuracy, which inevitably results in degraded inference speed. To strike a better balance between tracking accuracy and speed, in this work, we propose to compress the MOT model via knowledge distillation (KD), enabling the more lightweight student model to obtain similar performance as the teacher model. Nonetheless, despite KD has been well studied for simpler tasks such as image classification, the complexity of MOT poses new challenges because the MOT model is more sensitive to foreground information than the classification model. To deal with that, we first propose attention-guided feature distillation, which focuses the student model on the crucial region (foreground and the region with strong discrepancy against itself) of the teacher’s feature map. Moreover, we propose foreground mask, which leverages the knowledge from the teacher model to filter out the low-quality soft labels from the background, thereby reducing their negative effects for distillation. Evaluations on several benchmarks demonstrate that the proposed KD method can make the student network achieve leading performance, meanwhile running faster than the teacher network 20.0%–27.4% and reducing the parameters 28.5%–87.1%. To the best of our knowledge, this is the first work to compress the MOT model via KD. Tianyi Liang 0001, Mengzhu Wang, Junyang Chen 0001, Dingyao Chen, Zhigang Luo, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | TFC: Transformer Fused Convolution for Adversarial Domain AdaptationabstractIn unsupervised domain adaptation (UDA), a classifier is applied to the target domain without or with limited labels, when the target domain has no or few labels. Recently, inspired by their capabilities of long-distance feature dependencies, vision transformer (ViT)-based methods have been used in UDA, however, they ignore the fact that ViT lacks strength in extracting local feature details. To handle the above problems, the purpose of this article is to demonstrate how to take advantage of both convolutional operations and transformer mechanisms for adversarial UDA by using a hybrid network structure called transformer fused convolution (TFC). TFC integrates local features with global features to boost the representation capacity for UDA which can enhance the discrimination between foreground and background. Moreover, to improve the robustness of the TFC, we leverage an uncertainty penalty loss to make incorrect classes have consistently lower scores. Extensive experiments validate the significant performance gains compared to conditional adversarial domain adaptation (CDAN) on all five datasets including DomainNet ($\uparrow ~8.5$%), VisDA-2017 ($\uparrow ~14.9$%), Office-Home ($\uparrow ~18.9$%), Office-31 ($\uparrow ~11.5$%), and ImageCLEF-DA ($\uparrow ~5.5$%). Mengzhu Wang, Junyang Chen 0001, Ye Wang 0023, Zhiguo Gong, Kaishun Wu, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Multidocument Aspect Classification for Aspect-Based Abstractive SummarizationabstractMultidocument aspect-based summarization (AspSumm) aims to generate focused summaries based on the target aspects from a cluster of relevant documents. Generating such summaries can better satisfy readers’ specific points of interest, as readers may have different concerns about the same articles. However, previous methods usually generate aspect-based summaries based on the given aspects without using the relationship among aspects to assist in the summarization. In this work, we propose a two-stage general framework for multidocument AspSumm. The model first discovers the latent relationship among aspects and then uses relevant sentences selected by aspect discovery to generate abstractive summaries. We exploit latent dependencies among aspects using a tag mask training (TMT) strategy, which increases the interpretability of the model. In addition to improvements in summarization over aspect-based strong baselines, experimental results show that our proposed model can accurately discover multidomain aspects on the WikiAsp dataset. Ye Wang 0023, Mengzhu Wang, Zhenghan Chen, Zhiping Cai, Junyang Chen 0001, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2024 | Quaternion Cross-Modality Spatial Learning for Multi-Modal Medical Image SegmentationabstractRecently, the Deep Neural Networks (DNNs) have had a large impact on imaging process including medical image segmentation, and the real-valued convolution of DNN has been extensively utilized in multi-modal medical image segmentation to accurately segment lesions via learning data information. However, the weighted summation operation in such convolution limits the ability to maintain spatial dependence that is crucial for identifying different lesion distributions. In this paper, we propose a novel Quaternion Cross-modality Spatial Learning (Q-CSL) which explores the spatial information while considering the linkage between multi-modal images. Specifically, we introduce to quaternion to represent data and coordinates that contain spatial information. Additionally, we propose Quaternion Spatial-association Convolution to learn the spatial information. Subsequently, the proposed De-level Quaternion Cross-modality Fusion (De-QCF) module excavates inner space features and fuses cross-modality spatial dependency. Our experimental results demonstrate that our approach compared to the competitive methods perform well with only 0.01061 M parameters and 9.95G FLOPs. Junyang Chen 0001, Guoheng Huang, Xiaochen Yuan, Guo Zhong, Zewen Zheng, Chi-Man Pun, Jian Zhu 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Resisting the Edge-Type Disturbance for Link Prediction in Heterogeneous NetworksabstractThe rapid development of heterogeneous networks has proposed new challenges to the long-standing link prediction problem. Existing models trained on the verified edge samples from different types usually learn type-specific knowledge, and their type-specific predictions may be contradictory for unverified edge samples with uncertain types. This challenge is termed edge-type disturbance in link prediction in heterogeneous networks. To address this challenge, we develop a disturbance-resilient prediction method ( DRPM ) comprising a structural characterizer, a type differentiator, and a resilient predictor. The structural characterizer is responsible for learning edge representations for link prediction. Concurrently, the type differentiator distinguishes type-specific edge representations to generate diverse type experts while maximizing their link prediction performances on specific types. Furthermore, the resilient predictor evaluates the reliability weights of different type experts to develop a resilient prediction mechanism to aggregate discriminable predictions. Extensive experiments conducted on various real-world datasets demonstrate the importance of the explainable introduction of the edge-type disturbance and the superiority of DRPM over state-of-the-art methods. Huan Wang 0005, Ruigang Liu, Chuanqi Shi, Junyang Chen 0001, Lei Fang 0001, Zhiguo Gong |
ACM Trans. Knowl. Discov. Data | 4 |
| 2024 | Satisfying Energy-Efficiency Constraints for Mobile SystemsabstractEnergy-efficiency is one of the most important design criteria for mobile systems, such as smartphones and tablets. But current mobile systems always over-provision resources to satisfy users. The root cause is that, we have no knowledge on how much of system performance/energy will exactly satisfy users. Psychophysics defines the quantified link between physical stimuli and human-perceived stimuli. So, we will leverage psychophysics to study the quantified correlation between computer architecture resources (i.e., physical stimuli) and user satisfaction (i.e., human-perceived stimuli). We then exploit such correlation to precisely apportion resources to operate tasks and accurately satisfy users. Benefiting from our precisely-defined user satisfaction criteria and well-designed algorithms, we can reduce energy consumption of computer architectures by up to 42.9% without harming user experience. To the best of our knowledge, we for the first time theoretically and accurately model such substantial correlation. Our work opens a new research domain for fundamentally improving mobiles’ energy-efficiency. Xueliang Li 0002, Shicong Hong, Junyang Chen 0001, Junkai Ji, Chengwen Luo 0001, Guihai Yan, Zhibin Yu 0001, Jianqiang Li 0001 |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Data-Driven Pick-Up Location Recommendation for Ride-Hailing ServicesabstractRide-hailing service (RHS) has become an important transportation mode in our daily life. Although many works have been proposed to improve RHS from different aspects, only few works focus on the selections of pick-up locations, where rider and driver meet and start a trip. In this paper, we presentMPLRec, a data-driven pick-up location recommendation system that exploits riders' specific mobility demands,e.g.destination, and historical experiences to meet riders' travel requirements.MPLRecgenerates potential pick-up locations over the road network and characterizes them with rich features that describe a location from the riders' perspective. We also build spatio-temporal indexes to organize potential pick-up locations and historical data for facilitating online recommending. When processing an online recommendation request,MPLRecderives candidate pick-up locations and investigates them with materialized features, which are computed from historical order and trajectory data while considering rider's mobility demands. Based on these features, a novel scoring function is used to derive the best pick-up location for each request. Moreover, we implement an RHS simulator to evaluateMPLRecusing large-scale practical ride-hailing datasets. Extensive experiments and simulations demonstrate the effectiveness and efficiency ofMPLRec, which can complete each request within 0.5 s and largely reduce the ride-hailing costs when compared to baseline methods. Zhidan Liu 0001, Hongquan Zhang, Guofeng Ouyang, Junyang Chen 0001, Kaishun Wu |
IEEE Trans. Mob. Comput. | 4 |
| 2023 | A Topic-Aware Graph-Based Neural Network for User Interest Summarization and Item Recommendation in Social Media
Junyang Chen 0001, Ge Fan, Zhiguo Gong, Xueliang Li 0002, Victor C. M. Leung, Mengzhu Wang |
DASFAA (2) | 1 |
| 2023 | MPS-AMS: Masked Patches Selection and Adaptive Masking Strategy Based Self-Supervised Medical Image SegmentationabstractExisting self-supervised learning methods based on contrastive learning and masked image modeling have demonstrated impressive performances. However, current masked image modeling methods are mainly utilized in natural images, and their applications in medical images are relatively lacking. Besides, their fixed high masking strategy limits the upper bound of conditional mutual information, and the gradient noise is considerable, making less the learned representation information. Motivated by these limitations, in this paper, we propose masked patches selection and adaptive masking strategy based self-supervised medical image segmentation method, named MPS-AMS. We leverage the masked patches selection strategy to choose masked patches with lesions to obtain more lesion representation information, and the adaptive masking strategy is utilized to help learn more mutual information and improve performance further. Extensive experiments on three public medical image segmentation datasets (BUSI, Hecktor, and Brats2018) show that our proposed method greatly outperforms the state-of-the-art self-supervised baselines. Xiangtao Wang, Shuo Zhang 0017, Junyang Chen 0001, Thomas Lukasiewicz, Zhenghua Xu 0001 |
ICASSP | 6 |
| 2023 | MvCo-DoT: Multi-View Contrastive Domain Transfer Network for Medical Report GenerationabstractIn clinical scenarios, multiple medical images with different views are usually generated at the same time, and they have high semantic consistency. However, the existing medical report generation methods cannot exploit the rich multi-view mutual information of medical images. Therefore, in this work, we propose the first multi-view medical report generation model, called MvCo-DoT. Specifically, MvCo-DoT first propose a multi-view contrastive learning (MvCo) strategy to help the deep reinforcement learning based model utilize the consistency of multi-view inputs for better model learning. Then, to close the performance gaps of using multi-view and single-view inputs, a domain transfer network is further proposed to ensure MvCo-DoT achieve almost the same performance as multi-view inputs using only single-view inputs. Extensive experiments on the IU X-Ray public dataset show that MvCo-DoT outperforms the SOTA medical report generation baselines in all metrics. Xiangtao Wang, Zhenghua Xu 0001, Wenting Xu, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 5 |
| 2023 | Multi-Head Feature Pyramid Networks for Breast Mass DetectionabstractAnalysis of X-ray images is one of the main tools to diagnose breast cancer. The ability to quickly and accurately detect the location of masses from the huge amount of image data is the key to reducing the morbidity and mortality of breast cancer. Currently, the main factor limiting the accuracy of breast mass detection is the unequal focus on the mass boxes, leading the network to focus too much on larger masses at the expense of smaller ones. In the paper, we propose the multi-head feature pyramid module (MHFPN) to solve the problem of unbalanced focus of target boxes during feature map fusion and design a multi-head breast mass detection network (MBMDnet). Experimental studies show that, comparing to the SOTA detection baselines, our method improves by 6.58% (in AP@50) and 5.4% (in TPR@50) on the commonly used IN-breast dataset, while about 6-8% improvements (in AP@20) are also observed on the public MIAS and BCS-DBT datasets. Hexiang Zhang, Zhenghua Xu 0001, Shuo Zhang 0017, Junyang Chen 0001, Thomas Lukasiewicz |
ICASSP | 5 |
| 2023 | Hierarchical Crowdsourcing for Data Labeling with Heterogeneous CrowdabstractWith the rapid and continuous development of data-driven technologies such as supervised learning, high-quality labeled data sets are commonly required by many applications. Due to the easiness of crowdsourcing small tasks with low cost, a straightforward solution for label quality improvement is to collect multiple labels from a crowd, and then aggregate the answers. The aggregation strategies include majority voting and its many variants, EM-based approaches, Graph Neural Nets and so on. However, due to the uncertainty information loss and commonly existing task correlations, the aggregated labels usually contain errors and may damnify the downstream model training.To address the above problem, we propose a hierarchical crowdsourcing framework1for data labeling with noisy answers about correlated data. We make use of the heterogeneity of the labeling crowd and form an initialization-checking-update loop to improve the quality of labeled data. We formalize and successfully solve the core optimization problem, namely, selecting a proper set of checking tasks for each round. We prove that maximizing the expected quality improvement is equivalent to minimizing the conditional entropy of the observations given the crowdsourced answer families for the selected task set, which is NP-hard to solve. Therefore, we design an efficient approximation algorithm and conduct a series of experiments on real data. The experimental results show that the proposed method effectively improves the quality of the labeled data sets as well as the SOTA performance, yet without extra human labor costs. Wenxi Huang, Zhenhan Su, Junyang Chen 0001, Di Jiang 0004, Lixin Fan, Chen Zhang 0013, Defu Lian, Kaishun Wu |
ICDE | 4 |
| 2023 | Zero-shot Micro-video Classification with Neural Variational Inference in Graph Prototype NetworkabstractMicro-video classification plays a central role in online content recommendation platforms, such as Kwai and Tik-Tok. Existing works on video classification largely exploit the interactions between users and items as well as the item labels to provide quality recommendation services. However, scarce or even no labeled data of emerging videos is a great challenge for existing classification methods. In this paper, we propose a zero-shot micro-video classification model (NVIGPN) by exploiting the hidden topics behind items to guide the representation learning in user-item interactions. Specifically, we study this zero-shot classification in two stages: (1) exploiting a generalized semantic hidden topic descriptions for transferable knowledge learning, and (2) designing a graph-based learning model for guiding the minor seen class information to the unseen ones. Through mining the transferable knowledge between the hidden topics and the small number of the seen classes, NVIGPN can achieves state-of-the-art performances in predicting the unseen classes of micro-videos. We conduct extensive experiments to demonstrate the effectiveness of our method. Junyang Chen 0001, Zhijiang Dai, Huisi Wu, Mengzhu Wang, Qin Zhang 0011, Huan Wang 0005 |
ACM Multimedia | 1 |
| 2023 | A Closer Look at Classifier in Adversarial Domain GeneralizationabstractThe task of domain generalization is to learn a classification model from multiple source domains and generalize it to unknown target domains. The key to domain generalization is learning discriminative domain-invariant features. Invariant representations are achieved using adversarial domain generalization as one of the primary techniques. For example, generative adversarial networks have been widely used, but suffer from the problem of low intra-class diversity, which can lead to poor generalization ability. To address this issue, we propose a new method called auxiliary classifier in adversarial domain generalization (CloCls). CloCls improve the diversity of the source domain by introducing auxiliary classifier. Combining typical task-related losses, e.g., cross-entropy loss for classification and adversarial loss for domain discrimination, our overall goal is to guarantee the learning of condition-invariant features for all source domains while increasing the diversity of source domains. Further, inspired by smoothing optima have improved generalization for supervised learning tasks like classification. We leverage that converging to a smooth minima with respect task loss stabilizes the adversarial training leading to better performance on unseen target domain which can effectively enhances the performance of domain adversarial methods. We have conducted extensive image classification experiments on benchmark datasets in domain generalization, and our model exhibits sufficient generalization ability and outperforms state-of-the-art DG methods. Ye Wang 0023, Junyang Chen 0001, Mengzhu Wang, Hao Li 0058, Wei Wang 0335, Houcheng Su, Zhihui Lai 0001, Wei Wang 0077, Zhenghan Chen |
ACM Multimedia | 2 |
| 2023 | Interpolation Normalization for Contrast Domain GeneralizationabstractDomain generalization refers to the challenge of training a model from various source domains that can generalize well to unseen target domains. Contrastive learning is a promising solution that aims to learn domain-invariant representations by utilizing rich semantic relations among sample pairs from different domains. One simple approach is to bring positive sample pairs from different domains closer, while pushing negative pairs further apart. However, in this paper, we find that directly applying contrastive-based methods is not effective in domain generalization. To overcome this limitation, we propose to leverage a novel contrastive learning approach that promotes class-discriminative and class-balanced features from source domains. Essentially, clusters of sample representations from the same category are encouraged to cluster, while those from different categories are spread out, thus enhancing the model's generalization capability. Furthermore, most existing contrastive learning methods use batch normalization, which may prevent the model from learning domain-invariant features. Inspired by recent research on universal representations for neural networks, we propose a simple emulation of this mechanism by utilizing batch normalization layers to distinguish visual classes and formulating a way to combine them for domain generalization tasks. Our experiments demonstrate a significant improvement in classification accuracy over state-of-the-art techniques on popular domain generalization benchmarks, including Digits-DG, PACS, Office-Home and DomainNet. Mengzhu Wang, Junyang Chen 0001, Huan Wang 0005, Huisi Wu, Zhidan Liu 0001, Qin Zhang 0011 |
ACM Multimedia | 2 |
| 2023 | PromptRestorer: A Prompting Image Restoration Method with Degradation PerceptionabstractWe show that raw degradation features can effectively guide deep restoration models, providing accurate degradation priors to facilitate better restoration. While networks that do not consider them for restoration forget gradually degradation during the learning process, model capacity is severely hindered. To address this, we propose a Prompting image Restorer, termed as PromptRestorer. Specifically, PromptRestorer contains two branches: a restoration branch and a prompting branch. The former is used to restore images, while the latter perceives degradation priors to prompt the restoration branch with reliable perceived content to guide the restoration process for better recovery. To better perceive the degradation which is extracted by a pre-trained model from given degradation observations, we propose a prompting degradation perception modulator, which adequately considers the characters of the self-attention mechanism and pixel-wise modulation, to better perceive the degradation priors from global and local perspectives. To control the propagation of the perceived content for the restoration branch, we propose gated degradation perception propagation, enabling the restoration branch to adaptively learn more useful features for better recovery. Extensive experimental results show that our PromptRestorer achieves state-of-the-art results on 4 image restoration tasks, including image deraining, deblurring, dehazing, and desnowing. Cong Wang 0018, Jinshan Pan, Wei Wang 0335, Jiangxin Dong, Mengzhu Wang, Yakun Ju, Junyang Chen 0001 |
NeurIPS | 7 |
| 2023 | A Neural Inference of User Social Interest for Item RecommendationabstractAbstract User-generated content is daily produced in social media, as such user interest summarization is critical to distill salient information from massive information for recommendation tasks. While the interested messages (e.g., tags or posts) from a single user are usually sparse becoming a bottleneck for existing methods, we propose a neural inference method (NIGraphNet) by mining user social interest for item recommendation. It can unearth user latent topics combined with user relation learning. Specifically, we exploit a neural variational inference approach to learn the distributions between user interests and hidden topics. (We denote it as interest-topic distributions in the following.) Then, we adopt a unified graph-based training loss that jointly learns the hidden topics and user relations for item recommendation. Experiments on two datasets collected from well-known social media platforms demonstrate the superior performance of our model in the tasks of user interest summarization and item recommendation. Further discussions also show that exploiting the latent topic representations and user relations is conducive to the user’s automatic language understanding. Junyang Chen 0001, Mengzhu Wang, Ge Fan, Guo Zhong, Ou Liu, Wenfeng Du, Zhenghua Xu 0001, Zhiguo Gong |
Data Sci. Eng. | 1 |
| 2023 | Boosting unsupervised domain adaptation: A Fourier approach
Mengzhu Wang, Shanshan Wang 0008, Ye Wang 0023, Wei Wang 0335, Tianyi Liang 0001, Junyang Chen 0001, Zhigang Luo |
Knowl. Based Syst. | 6 |
| 2023 | IRLM: Inductive Representation Learning Model for Personalized POI RecommendationabstractWith the rapid development of the Internet of Things technology, the concept of smart cities that aims to help residents improve their quality of life has raised much attention in several application areas. In the context of smart cities, the provision of point of interest (POI) recommendations become an important requirement because a wide range of POIs are available for urban dwellers. Location-based social networks (LBSNs) such as Foursquare and Gowalla provide a massive volume of user check-in records that can assist users in choosing new POIs. However, user trajectories are mostly sparse in the real world. For example, users only check in a few POIs, and this makes it difficult to provide recommendations based on limited history trajectories. Though some attempts have adopted auxiliary geographical information to enhance POI recommendation, they still encounter the following problems: 1) the geographical trajectories of users are usually sparse in real-world datasets; 2) users may be more interested in the remote POIs; and 3) the previous models inherently perform transductive learning that cannot handle well the recommendation of unseen users and POIs. To address these problems, we propose an inductive representation learning model (IRLM) for location recommendation. IRLM contains two parts, namely geographic feature extraction and inductive representation learning. IRLM first captures global geographical influences among POIs through a standard Gaussian mixture model (GMM). Then IRLM adopts an attention neural network for the recommendation. Experimental results indicate that our proposed model can achieve superior performance over state-of-the-art models. Junyang Chen 0001, Mengzhu Wang, Zhenghua Xu 0001, Xueliang Li 0002, Zhiguo Gong, Kaishun Wu, Victor C. M. Leung |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Rethinking Maximum Mean Discrepancy for Visual Domain AdaptationabstractExisting domain adaptation approaches often try to reduce distribution difference between source and target domains and respect domain-specific discriminative structures by some distribution [e.g., maximum mean discrepancy (MMD)] and discriminative distances (e.g., intra-class and inter-class distances). However, they usually consider these losses together and trade off their relative importance by estimating parameters empirically. It is still under insufficient exploration so far to deeply study their relationships to each other so that we cannot manipulate them correctly and the model's performance degrades. To this end, this article theoretically proves two essential facts: 1) minimizing MMD equals to jointly minimizing their data variance with some implicit weights but, respectively, maximizing the source and target intra-class distances so that feature discriminability degrades and 2) the relationship between intra-class and inter-class distances is as one falls and another rises. Based on this, we propose a novel discriminative MMD with two parallel strategies to correctly restrain the degradation of feature discriminability or the expansion of intra-class distance; specifically: 1) we directly impose a tradeoff parameter on the intra-class distance that is implicit in the MMD according to 1) and 2) we reformulate the inter-class distance with special weights that are analogical to those implicit ones in the MMD and maximizing it can also lead to the intra-class distance falling according to 2). Notably, we do not consider the two strategies in one model due to 2). The experiments on several benchmark datasets not only prove the validity of our revealed theoretical results but also demonstrate that the proposed approach could perform better than some compared state-of-art methods substantially. Our preliminary MATLAB code will be available at https://github.com/WWLoveTransfer/. Wei Wang 0335, Zhengming Ding, Feiping Nie 0001, Junyang Chen 0001, Zhihui Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Combatting Energy Issues for Mobile ApplicationsabstractEnergy efficiency is an important criterion to judge the quality of mobile apps, but one third of our arbitrarily sampled apps suffer from energy issues that can quickly drain battery power. To understand these issues, we conduct an empirical study on 36 well-maintained apps such as Chrome and Firefox, whose issue tracking systems are publicly accessible. Our study involves issue causes, manifestation, fixing efforts, detection techniques, reasons of no-fixes, and debugging techniques. Inspired by the empirical study, we propose a novel testing framework for detecting energy issues in real-world mobile apps. Our framework examines apps with well-designed input sequences and runtime context. We develop leading edge technologies, e.g., pre-designing input sequences with potential energy overuse and tuning tests on-the-fly, to achieve high efficacy in detecting energy issues. A large-scale evaluation shows that 90.4% of the detected issues in our experiments were previously unknown to developers. On average, these issues can double the energy consumption of the test cases where the issues were detected. And our test achieves a low number of false positives. Finally, we show how our test reports can help developers fix the issues. Xueliang Li 0002, Junyang Chen 0001, Yepang Liu 0001, Kaishun Wu, John P. Gallagher |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2022 | Using Psychophysics to Guide Power Adaptation for Input Methods on Mobile ArchitecturesabstractThe predominant user activities on mobile architectures (e.g., smartphones) involve entering text in instant messaging apps, short message services, and social networking services. Recent research reveals that the normal use of input methods drains approximately half of the battery capacity due to their above-average power requirements and frequent use. In this paper, we first study the power characteristics of mobile input methods and find that they consistently over-provision resources to satisfy users. For example, the psychophysical evidence available indicates the response of spell-checking features within a time threshold makes users feel they have instant feedback. However, current systems perform it very quickly, which is imperceptible to users and costly in terms of energy use. Given this over- provisioning, the system can be slowed down to save energy while retaining the feeling of instant response. Inspired by this observation, we also exploit several other psychophysical facts to identify the exact criteria to satisfy users. As a result, we present a user experience-oriented technology, utexia, to optimize the energy use of mobile input methods. The evaluation shows that utexia conserves up to 42.9% in energy use while strictly ensuring a good user experience. Xueliang Li 0002, Shicong Hong, Junyang Chen 0001, Guihai Yan, Kaishun Wu |
HPCA | 3 |
| 2022 | Attention-based Adversarial Partial Domain AdaptationabstractWith the rapid development of vision-based deep learning (DL), it is an effective method to generate large-scale synthetic data to supplement real data to train the DL models for domain adaptation. However, previous vanilla domain adaptation methods generally assume the same label space, and such an assumption is no longer valid for a more realistic scenario where it requires adaptation from a larger and more diverse source domain to a smaller target domain with less number of classes. To handle this problem, we propose an attention-based adversarial partial domain adaptation (AADA). Specifically, we leverage adversarial domain adaptation to augment the target domain by using source domain, then we can readily turn this task into a vanilla domain adaptation. Meanwhile, to accurately focus on the transferable features, we apply attention-based method to train the adversarial networks to obtain better transferable semantic features. Experiments on four benchmarks demonstrate that the proposed method outperforms existing methods by a large margin, especially on the tough domain adaptation tasks, e.g. VisDA-2017. Mengzhu Wang, Shan An, Xiao Luo 0001, Xiong Peng, Wei Yu 0029, Junyang Chen 0001, Zhigang Luo |
ICASSP | 6 |
| 2022 | Field-aware Variational Autoencoders for Billion-scale User Representation LearningabstractUser representation learning plays an essential role in Internet applications, such as recommender systems. Though developing a universal embedding for users is demanding, only few previous works are conducted in an unsupervised learning manner. The unsupervised method is however important as most of the user data is collected without specific labels. In this paper, we harness the unsupervised advantages of Variational Autoencoders (VAEs), to learn user representation from large-scale, high-dimensional, and multi-field data. We extend the traditional VAE by developing Field-aware VAE (FVAE) to model each feature field with an independent multinomial distribution. To reduce the complexity in training, we employ dynamic hash tables, a batched softmax function, and a feature sampling strategy to improve the efficiency of our method. We conduct experiments on multiple datasets, showing that the proposed FVAE significantly outperforms baselines on several tasks of data reconstruction and tag prediction. Moreover, we deploy the proposed method in real-world applications and conduct online A/B tests in a look-alike system. Results demonstrate that our method can effectively improve the quality of recommendation. To the best of our knowledge, it is the first time that the VAE-based user representation learning model is applied to real-world recommender systems. Ge Fan, Chaoyun Zhang, Junyang Chen 0001, Baopu Li, Zenglin Xu, Luyu Peng, Zhiguo Gong |
ICDE | 3 |
| 2022 | MV-HAN: A Hybrid Attentive Networks based Multi-View Learning Model for Large-scale Contents RecommendationabstractIndustrial recommender systems usually employ multi-source data to improve the recommendation quality, while effectively sharing information between different data sources remain a challenge. In this paper, we introduce a novel Multi-View Approach with Hybrid Attentive Networks (MV-HAN) for contents retrieval at the matching stage of recommender systems. The proposed model enables high-order feature interaction from various input features while effectively transferring knowledge between different types. By employing a well-placed parameters sharing strategy, the MV-HAN substantially improves the retrieval performance in sparse types. The designed MV-HAN inherits the efficiency advantages in the online service from the two-tower model, by mapping users and contents of different types into the same features space. This enables fast retrieval of similar contents with an approximate nearest neighbor algorithm. We conduct offline experiments on several industrial datasets, demonstrating that the proposed MV-HAN significantly outperforms baselines on the content retrieval tasks. Importantly, the MV-HAN is deployed in a real-world matching system. Online A/B test results show that the proposed method can significantly improve the quality of recommendations. Ge Fan, Chaoyun Zhang, Junyang Chen 0001 |
ASE | 4 |
| 2022 | Realizing Emotional Interactions to Learn User Experience and Guide Energy Optimization for Mobile ArchitecturesabstractIn the age of AI, mobile architectures such as smartphones are still “cold machines”; machines do not feel. If the architecture is able to feel users’ feelings and runtime user experience (UX), it will accordingly adapt performance/energy to find the optimal system-operating state that consumes the least energy to satisfy users. In this paper, we will utilize users’ facial expressions (FEs) to learn their runtime UX. We know that FEs are the natural and direct way for humans to convey their emotions and feelings. Our study reveals that FEs also reflect UX. Our research for the first time quantifies the link between FEs and UX. Leveraging this link, the architecture will be able to use the front camera to see FEs and feel users’ UX. Based on UX, the architecture can appropriately provision computing resources. We propose Vi-energy system to realize the above idea. Our evaluation shows that Vi-energy reduces energy consumption by 52.9% at maximum and secures UX. Xueliang Li 0002, Zhuobin Shi, Junyang Chen 0001, Yepang Liu 0001 |
MICRO | 3 |
| 2022 | Single image rain removal using recurrent scale-guide networks
Cong Wang 0018, Honghe Zhu, Wanshu Fan, Xiao-Ming Wu 0003, Junyang Chen 0001 |
Neurocomputing | 5 |
| 2022 | ω-net: Dual supervised medical image segmentation with multi-dimensional self-attention and diversely-connected multi-scale convolution
Zhenghua Xu 0001, Junyang Chen 0001, Thomas Lukasiewicz, Zhigang Fu |
Neurocomputing | 5 |
| 2022 | DnSwin: Toward real-world denoising via a continuous Wavelet Sliding Transformer
Hao Li 0058, Zhijing Yang, Ziying Zhao, Junyang Chen 0001, Yukai Shi, Jinshan Pan |
Knowl. Based Syst. | 5 |
| 2022 | Attentive differential convolutional neural networks for crowd flow prediction
Jiqian Mo, Zhiguo Gong, Junyang Chen 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Informative pairs mining based adaptive metric learning for adversarial domain adaptation
Mengzhu Wang, Paul Li, Li Shen 0008, Ye Wang 0023, Shanshan Wang 0008, Wei Wang 0335, Xiang Zhang 0008, Junyang Chen 0001, Zhigang Luo |
Neural Networks | 8 |
| 2022 | Robust Matrix Factorization via Minimum Weighted Error Entropy CriterionabstractLearning the intrinsic low-dimensional subspace from high-dimensional data is a key step for many social systems of artificial intelligence. In practical scenarios, the observed data are usually corrupted by many types of noise, which brings a great challenge for social systems to analyze data. As a commonly utilized subspace learning technique, robust low-rank matrix factorization (LRMF) focuses on recovering the underlying subspaces in a noisy environment. However, most of the existing approaches simply assume that the noise contaminating the data is independent identically distributed (i.i.d.), such as Gaussian and Laplacian noises. This assumption, though greatly simplifies the underlying learning problem, may not hold for more complex non-i.i.d. noise widely existed in social systems. In this work, we suggest a robust LRMF approach to deal with various types of noise in a unified manner. Different from traditional algorithms, noise in our framework is modeled using an independent and piecewise identically distributed (i.p.i.d.) source, which employs a collection of distributions, instead of a single one to characterize the statistical behavior of the underlying noise. Assisted by the generic noise model, we then design a robust LRMF algorithm under the information-theoretic learning (ITL) framework through a new minimization criterion. By adopting the half-quadratic optimization paradigm, we further deliver an optimization strategy for our proposed method. Experimental results on both synthetic and real data are provided to demonstrate the superiority of our proposed scheme. Yuanman Li, Jiantao Zhou 0001, Junyang Chen 0001, Jinyu Tian 0001, Li Dong 0006, Xia Li 0006 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Self-Training Enhanced: Network Embedding and Overlapping Community Detection With Adversarial LearningabstractNetwork embedding (NE) aims to encode the relations of vertices into a low-dimensional space. After NE, we can obtain the learned vectors of vertices that preserve the proximity of network structures for subsequent applications, e.g., vertex classification and link prediction. In existing NE models, they usually exploit the skip-gram with a negative sampling method to optimize their objective functions. Generally, this method learns the vertex representation only from the local connectivity of vertices (i.e., neighbors). However, there is a larger scope of vertex connectivity in real-world scenarios: a vertex may have multifaceted aspects and should belong to overlapping communities. Taking a social network as the overlapping example, a user may subscribe to the channels of politics, economy, and sports simultaneously, but the politics share more common attributes with the economy and less with the sports. In this article, we propose an adversarial learning approach (ACNE) for modeling overlapping communities of vertices. Specifically, we map the association between communities and vertices into an embedding space. Moreover, we take further research on enhancing our ACNE with the following two operations. First, in the initialization stage, we adopt a walking strategy with perception to obtain paths containing more possible boundary vertices to improve overlapping community detection. Then, after representation learning with ACNE, we use soft community assignments from a simple classifier as supervision to update the weights of ACNE. This self-training mechanism referred to as ACNE-ST can help ACNE to achieve better performance. Experimental results demonstrate that the proposed methods, including ACNE and ACNE-ST, can outperform the state-of-the-art models on the subsequent tasks of vertex classification and overlapping community detection. Junyang Chen 0001, Zhiguo Gong, Jiqian Mo, Wei Wang 0077, Wei Wang 0335, Cong Wang 0018, Weiwen Liu, Kaishun Wu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | CRL: Collaborative Representation Learning by Coordinating Topic Modeling and Network EmbeddingsabstractNetwork representation learning (NRL) has shown its effectiveness in many tasks, such as vertex classification, link prediction, and community detection. In many applications, vertices of social networks contain textual information, e.g., citation networks, which form a text corpus and can be applied to the typical representation learning methods. The global context in the text corpus can be utilized by topic models to discover the topic structures of vertices. Nevertheless, most existing NRL approaches focus on learning representations from the local neighbors of vertices and ignore the global structure of the associated textual information in networks. In this article, we propose a unified model based on matrix factorization (MF), named collaborative representation learning (CRL), which: 1) considers complementary global and local information simultaneously and 2) models topics and learns network embeddings collaboratively. Moreover, we incorporate the Fletcher-Reeves (FR) MF, a conjugate gradient method, to optimize the embedding matrices in an alternative mode. We call this parameter learning method as AFR in our work that can achieve convergence after a few numbers of iterations. Also, by evaluating CRL on topic coherence and vertex classification using several real-world data sets, our experimental study shows that this collaborative model not only can improve the performance of topic discovery over the baseline topic models but also can learn better network representations than the state-of-the-art context-aware NRL models. Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Weiwen Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Adversarial Caching Training: Unsupervised Inductive Network Representation Learning on Large-Scale GraphsabstractNetwork representation learning (NRL) has far-reaching effects on data mining research, showing its importance in many real-world applications. NRL, also known as network embedding, aims at preserving graph structures in a low-dimensional space. These learned representations can be used for subsequent machine learning tasks, such as vertex classification, link prediction, and data visualization. Recently, graph convolutional network (GCN)-based models, e.g., GraphSAGE, have drawn a lot of attention for their success in inductive NRL. When conducting unsupervised learning on large-scale graphs, some of these models employ negative sampling (NS) for optimization, which encourages a target vertex to be close to its neighbors while being far from its negative samples. However, NS draws negative vertices through a random pattern or based on the degrees of vertices. Thus, the generated samples could be either highly relevant or completely unrelated to the target vertex. Moreover, as the training goes, the gradient of NS objective calculated with the inner product of the unrelated negative samples and the target vertex may become zero, which will lead to learning inferior representations. To address these problems, we propose an adversarial training method tailored for unsupervised inductive NRL on large networks. For efficiently keeping track of high-quality negative samples, we design a caching scheme with sampling and updating strategies that has a wide exploration of vertex proximity while considering training costs. Besides, the proposed method is adaptive to various existing GCN-based models without significantly complicating their optimization process. Extensive experiments show that our proposed method can achieve better performance compared with the state-of-the-art models. Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Cong Wang 0018, Zhenghua Xu 0001, Jianming Lv, Xueliang Li 0002, Kaishun Wu, Weiwen Liu |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Dense Feature Pyramid Grids Network for Single Image DerainingabstractRainy images degrade the visional performance that may bring down the accuracy of various applications. In this paper, we propose a novel densely connected network with Dense Feature Pyramid Grids Modules, called DFPGN, to solve the rain removal task. Specifically, in the proposed DFPG, there are five operations from different layers with various pathways and scales as the input of the current layer so that each layer can fuse various features from shallower and deeper ones to improve the deraining ability of the network. Extensive experiments on real and synthetic rainy images are conducted to demonstrate the proposed method achieves superior rain removal performance over state-of-the-art approaches. Cong Wang 0018, Zhixun Su, Junyang Chen 0001 |
ICASSP | 4 |
| 2021 | HNS: Hierarchical negative sampling for network representation learning
Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Weiwen Liu |
Inf. Sci. | 1 |
| 2021 | TAM: Targeted Analysis Model With Reinforcement Learning on Short TextsabstractMining topics on social media (e.g., Twitter and Facebook) is an important task for various applications, such as hot topic discovery, advertising, and promotion activities. Topic modeling techniques are helpful to find out topics that people are talking about. However, current full-analysis models cannot perform well on a focused analysis task-find out all topics related to one particular area in short documents. One reason is that the targeted topic is usually sparse in the corpus of short texts. Another one is, during clustering, even minor errors may compound and render the model useless. This article studies these problems and proposes a targeted analysis model (TAM) with reinforcement learning (RL) to extract any specific topic in a given corpus and perform fine-grained topic generation. In this work, we design a reward function of RL to prevent the false propagation problem induced by Gibbs sampling during the clustering. We amend the targeted topic modeling techniques to the case of RL and use policy search combined with the Gibbs EM algorithm for parameter estimation. Metrics of F1 score and the proposed normalized mutual information-F1 are exploited for the evaluation of clustering and topic generation, respectively. Our experiments have demonstrated that TAM can outperform state-of-the-art models-specifically achieving 25.7% improvement on the F1 score for binary clustering on average. Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Weiwen Liu, Cong Wang 0018 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | A Multi-graph Convolutional Network Framework for Tourist Flow PredictionabstractWith the advancement of Cyber Physic Systems and Social Internet of Things, the tourism industry is facing challenges and opportunities. We can now able to collect, store, and analyze large amounts of travel data. With the help of data science and artificial intelligence, smart tourism enables tourists with great autonomy and convenience for an intelligent trip. It is of great significance to make full use of these massive data to provide better services for smart tourism. However, due to the skewed and imbalanced visiting for point of interest located at different places, it is of great significance to predict the tourist flow of each place, which can help the service providers for designing a better schedule visiting strategy in advance. Against this background, this article proposes a multi-graph convolutional network framework, named AMOUNT, for tourist flow prediction. To capture the diverse relationships among POIs, AMOUNT first constructs three subgraphs, including the geographical graph, interaction graph, and the co-relation graph. Then, a multi-graph convolution network is utilized to predict the future tourist flow. Experimental results on two real-world datasets indicate that the proposed AMOUNT model outperforms all other baseline tourist flow prediction approaches. Wei Wang 0077, Junyang Chen 0001, Yushu Zhang 0001, Zhiguo Gong, Neeraj Kumar 0001, Wei Wei 0006 |
ACM Trans. Internet Techn. | 2 |
| 2020 | Personalized Re-ranking with Item Relationships for E-commerceabstractRe-ranking is a critical task for large-scale commercial recommender systems. Given the initial ranked lists, top candidates are re-ranked to improve the accuracy of the ranking results. However, existing re-ranking strategies are sub-optimal due to (i) most prior works do not consider explicit item relationships, like being substitutable or complementary, which may mutually influence the user satisfaction on other items in the lists, and (ii) they usually apply an identical re-ranking strategy for all users, with personalized user preferences and intents ignored. To resolve the problem, we construct a heterogeneous graph to fuse the initial scoring information and item relationships information. We develop a graph neural network based framework, IRGPR, to explicitly model transitive item relationships by recursively aggregating relational information from multi-hop neighborhoods. We also incorporate a novel intent embedding network to embed personalized user intents into the propagation. We conduct extensive experiments on real-world datasets, demonstrating the effectiveness of IRGPR in re-ranking. Further analysis reveals that modeling the item relationships and personalized intents are particularly useful for improving the performance of re-ranking. Weiwen Liu, Qing Liu 0020, Ruiming Tang, Junyang Chen 0001, Xiuqiang He 0001, Pheng-Ann Heng |
CIKM | 4 |
| 2020 | Adversarial Learning for Overlapping Community Detection and Network EmbeddingabstractNetwork Embedding (NE) aims at modeling network graph by encoding vertices and edges into a low-dimensional space. These learned vectors which preserve proximities can be used for subsequent applications, such as vertex classification and link prediction. Skip-gram with negative sampling is the most widely used method for existing NE models to approximate their objective functions. However, this method only focuses on learning representation from the local connectivity of vertices (i.e., neighbors). In real-world scenarios, a vertex may have multifaceted aspects and should belong to overlapping communities. For example, in a social network, a user may subscribe to political, economic and sports channels simultaneously, but the politics share more common attributes with the economy and less with the sports. In this paper, we propose an adversarial learning approach for modeling overlapping communities of vertices. Each community and vertex are mapped into an embedding space, while we also learn the association between each pair of community and vertex. The experimental results show that our proposed model not only can outperform the state-of-the-art (including GANs-based) models on vertex classification tasks but also can achieve superior performances on overlapping community detection. Junyang Chen 0001, Zhiguo Gong, Quanyu Dai, Chunyuan Yuan, Weiwen Liu |
ECAI | 1 |
| 2020 | Joint Self-Attention and Scale-Aggregation for Self-Calibrated Deraining NetworkabstractIn the field of multimedia, single image deraining is a basic pre-processing work, which can greatly improve the visual effect of subsequent high-level tasks in rainy conditions. In this paper, we propose an effective algorithm, called JDNet, to solve the single image deraining problem and conduct the segmentation and detection task for applications. Specifically, considering the important information on multi-scale features, we propose a Scale-Aggregation module to learn the features with different scales. Simultaneously, Self-Attention module is introduced to match or outperform their convolutional counterparts, which allows the feature aggregation to adapt to each channel. Furthermore, to improve the basic convolutional feature transformation process of Convolutional Neural Networks (CNNs), Self-Calibrated convolution is applied to build long-range spatial and inter-channel dependencies around each spatial location that explicitly expand fields-of-view of each convolutional layer through internal communications and hence enriches the output features. By designing the Scale-Aggregation and Self-Attention modules with Self-Calibrated convolution skillfully, the proposed model has better deraining results both on real-world and synthetic datasets. Extensive experiments are conducted to demonstrate the superiority of our method compared with state-of-the-art methods. The source code will be available at https://supercong94.wixsite.com/supercong94. Cong Wang 0018, Yutong Wu 0002, Zhixun Su, Junyang Chen 0001 |
ACM Multimedia | 4 |
| 2020 | DCSFN: Deep Cross-scale Fusion Network for Single Image Rain RemovalabstractRain removal is an important but challenging computer vision task as rain streaks can severely degrade the visibility of images that may make other visions or multimedia tasks fail to work. Previous works mainly focused on feature extraction and processing or neural network structure, while the current rain removal methods can already achieve remarkable results, training based on single network structure without considering the cross-scale relationship may cause information drop-out. In this paper, we explore the cross-scale manner between networks and inner-scale fusion operation to solve the image rain removal task. Specifically, to learn features with different scales, we propose a multi-sub-networks structure, where these sub-networks are fused via a cross-scale manner by Gate Recurrent Unit to inner-learn and make full use of information at different scales in these sub-networks. Further, we design an inner-scale connection block to utilize the multi-scale information and features fusion way between different scales to improve rain representation ability and we introduce the dense block with skip connection to inner-connect these blocks. Experimental results on both synthetic and real-world datasets have demonstrated the superiority of our proposed method, which outperforms over the state-of-the-art methods. The source code will be available at https://supercong94.wixsite.com/supercong94. Cong Wang 0018, Xiaoying Xing, Yutong Wu 0002, Zhixun Su, Junyang Chen 0001 |
ACM Multimedia | 5 |
| 2020 | Exploiting Heterogeneous Artist and Listener Preference Graph for Music Genre ClassificationabstractMusic genres are useful for indexing, organizing, searching, and recommending songs and albums. Therefore, the automatic classification of music genres is an essential part of almost all kinds of music applications. Recent works focus on exploiting text, audio, or multi-modal information for genre classification, without considering the influence of the artists' and listeners' preference. However, intuitively, artists have their composing preferences, and listeners also have their music tastes. Both of them provide helpful hints to the music genre from different views, which are crucial to improve classification performance. Chunyuan Yuan, Qianwen Ma, Junyang Chen 0001, Wei Zhou 0019, Xiaodan Zhang 0004, Xuehai Tang, Jizhong Han, Songlin Hu 0001 |
ACM Multimedia | 3 |
| 2020 | Inductive Document Representation Learning for Short Text Clustering
Junyang Chen 0001, Zhiguo Gong, Wei Wang 0077, Wei Wang 0335, Weiwen Liu, Cong Wang 0018 |
ECML/PKDD (3) | 1 |
| 2020 | A Dirichlet process biterm-based mixture model for short text stream clustering
Junyang Chen 0001, Zhiguo Gong, Weiwen Liu |
Appl. Intell. | 1 |
| 2020 | Geography-Aware Inductive Matrix Completion for Personalized Point-of-Interest Recommendation in Smart CitiesabstractWith the development of Internet of Things technology, the will to make cities smarter is growing. In the context of smart cities, where people are surrounded by a tremendous number of points of interests (POIs), POI recommendation is of great significance. The massive amount of user check-in data collected by location-based social networks (LBSNs) can help users explore new places via POI recommendations. With rich auxiliary data becomeing available in LBSNs, purely exploiting users’ check-in information for POI recommendation is not sufficient. Although several attempts have been done to employ geographical influence to enhance POI recommendation, they simply use a certain distribution function to measure the geographical influence between POIs, which may lead to biased results. To this end, this article proposes a geography-aware inductive matrix completion (GAIMC) approach for personalized POI recommendation. Specifically, the GAIMC model consists of two parts, including geographic feature extraction via a Gaussian mixture model (GMM) and inductive matrix completion for recommendation. The GAIMC model first captures the geographical influence among users and POIs by the GMM, which can mine clustering information with hierarchical structures based on the users’ check-in data and POIs’ location information. Then, a matrix-completion method called inductive matrix completion, which can incorporate geographical features with the user-POI association metric, is utilized to recommend POIs. The experimental results on two real-world LBSN data sets demonstrate that our proposed model can achieve the best recommendation performance in comparison with the state-of-the-art counterparts. Wei Wang 0077, Junyang Chen 0001, Jinzhong Wang, Junxin Chen 0001, Zhiguo Gong |
IEEE Internet Things J. | 2 |
| 2020 | Trust-Enhanced Collaborative Filtering for Personalized Point of Interests RecommendationabstractPredicting the user's trajectory behavior sequence based on point of interests (POIs) recommendation is of great significance in the realization of the smart city with the emerging of social Internet of Things technology. One of the widely adopted frameworks is the user-based collaborative filtering, where the explicit POI rating is calculated based on similar users’ preference. However, the trust between users is seldom considered. We believe that if two users show similar preferences or personality traits, the trust level between them should be high. To this end, we propose to calculate the trust-enhanced user similarity in user-based collaborative filtering based on network representation learning. Meanwhile, due to the significance of geographic influence and temporal influence, we integrate these two factors into POI recommendation by a fusion model. Therefore, our proposed POI recommendation system is unified collaborative recommendation framework, which fuses trust-enhanced users’ preferences to potential POIs with geographic influences and temporal influence for POI recommendation. Finally, we conduct extensive experiments on two real-world datasets by comparing with several state-of-the-art methods in terms of precision@k and recall@k. Experimental results indicate that our proposed trust-enhanced collaborative filtering method outperforms other recommendation approaches. Wei Wang 0077, Junyang Chen 0001, Jinzhong Wang, Junxin Chen 0001, Jinquan Liu, Zhiguo Gong |
IEEE Trans. Ind. Informatics | 2 |
| 2019 | A nonparametric model for online topic discovery with word embeddings
Junyang Chen 0001, Zhiguo Gong, Weiwen Liu |
Inf. Sci. | 1 |