EDBT 2026 Demo / reviewers in the wild / expert
Qi Zhang 0020
dblp:52/323-20
· DBLP profile ↗
36ranked-venue papers in the field
1as first author
34since 2021 · last 2026
0000-0002-1037-1361ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 18Data Mining & Knowledge Discovery · 9Database Systems & Data Management · 3 (1 first)Big Data, Cloud & Distributed Data Systems · 3Knowledge Engineering, Semantic Web & Information Systems · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Cognitive Neuroscience and Graph Intelligence: Hippocampus-Inspired Multi-View Hypergraph Learning for Web Finance FraudabstractOnline financial services constitute an essential component of contemporary web ecosystems, yet their openness introduces substantial exposure to fraud that harms vulnerable users and weakens trust in digital finance. Such threats have become a significant web harm that erodes societal fairness and affects the well-being of online communities. However, existing detection methods based on graph neural networks (GNNs) struggle with two persistent challenges: (1) long-tailed data distributions, which obscure rare but critical fraudulent cases, and (2) fraud camouflage, where malicious transactions mimic benign behaviors to evade detection. To fill these gaps, we propose HIMVH, a Hippocampus-Inspired Multi-View Hypergraph learning model for web finance fraud detection. Specifically, drawing inspiration from the scene conflict monitoring role of the hippocampus, we design a cross-view inconsistency perception module that captures subtle discrepancies and behavioral heterogeneity across multiple transaction views. This module enables the model to identify subtle cross-view conflicts for detecting online camouflaged fraudulent behaviors. Furthermore, inspired by the match-mismatch novelty detection mechanism of the CA1 region, we introduce a novelty-aware hypergraph learning module that measures feature deviations from neighborhood expectations and adaptively reweights messages, thereby enhancing sensitivity to online rare fraud patterns in the long-tailed settings. Extensive experiments on six web-based financial fraud datasets demonstrate that HIMVH achieves 6.42% improvement in AUC, 9.74% in F1 and 39.14% in AP on average over 15 SOTA models. Rongkun Cui, Kun Zhu 0024, Qi Zhang 0020 |
WWW | 4 |
| 2026 | Towards Fair Large Language Model-based Recommender Systems without Costly Retraining
Jin Li 0028, Huilin Gu, Shoujin Wang, Qi Zhang 0020, Shui Yu 0001, Chen Wang 0008, Xiwei Xu 0001, Fang Chen 0001 |
WWW | 4 |
| 2026 | Counterfactual Augmented Causal Reasoning for Aspect-Based Sentiment Analysis
Liang Hu 0004, Mingzhu Zhou, Tangwei Ye, Xuejie Yang, Zhongyuan Lai, Qi Zhang 0020, Usman Naseem |
WWW | 8 |
| 2026 | AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information MinimizationabstractThe escalating energy consumption and carbon footprint of training large-scale Web AI models pose urgent challenges for sustainable development. Dataset Distillation (DD) offers a promising avenue for green AI by compressing large datasets into small synthetic ones for efficient training. However, most existing DD methods overfit to the inductive biases of specific source architectures (e.g., CNNs or ViTs), resulting in poor cross-model generalization. This limitation necessitates redundant re-distillation processes for different architectures, severely undermining the energy-saving potential of DD. To address this, we introduce Adversarial Mutual Information Distillation (AMID), a rigorous framework designed to create highly reusable and robust synthetic datasets. From an information-theoretic perspective, we cast model-agnosticism as minimizing the mutual information (MI) between the synthetic data and the specific identity of the distillation model. We convert this intractable objective into a tractable two-player adversarial game, which unifies knowledge preservation with adversarial unlearning of architectural bias. Extensive experiments on CIFAR-10 and Tiny ImageNet demonstrate that AMID achieves state-of-the-art cross-architecture generalization across diverse CNNs and ViTs. Crucially, our analysis confirms that AMID significantly reduces the computational overhead and CO2 emissions of downstream training while maintaining robust performance, paving the way for energy-efficient, transferable, and sustainable Web AI ecosystems. Aoqi Wu, Weiquan Huang, Liang Hu 0004, Yifan Yang 0004, Qi Zhang 0020, Jiaxing Miao, Yuhan Tang, Zhongyuan Lai |
WWW | 7 |
| 2026 | Accurate Trajectory Recovery in Underserved Areas via Location Inference from Web Crowdsourced Data
Tangwei Ye, Liang Hu 0004, Zhongyuan Lai, Qi Zhang 0020, Jiaxing Miao, Kun Yi 0001 |
WWW | 4 |
| 2026 | QDDR: Quality-Driven Intent Disentanglement with Dual-Path Modeling for Recommendation
Nengjun Zhu, Yixun Lu, Qi Zhang 0020 |
WWW | 3 |
| 2026 | Dynamic Min-Max Multi-Dimensional Reinforcement Backdoor Attacks and Orchestrated Closed-Loop Defense in Fairness-Aware Web Federated FinanceabstractIn the rapidly evolving web-based financial ecosystem where digital banking services become critical infrastructure for underserved communities, credit card fraud disproportionately affects vulnerable populations relying on financial platforms. However, previous studies overlook extreme data scarcity conditions, particularly at small-to-medium web banks that serve as crucial gateways for vulnerable communities. This paper addresses the fundamental challenge of building inclusive and secure financial systems operable at true web scale. To overcome this deficiency, we propose a novel web-based fairness-aware federated fraud detection model, CLARF, which utilizes the designed privacy-enhanced representation fusion and fraud-aware contrastive learning modules to enhance detection performance under conditions of data scarcity and label imbalance. Furthermore, current federated fraud detection systems critically neglect vulnerability to backdoor attacks, where malicious actors can implant hidden triggers during model aggregation, compromising system integrity. We propose a novel dynamic web Min-Max adversarial game framework where attackers employ hybrid multi-stage reinforcement learning with multi-dimensional reward mechanisms to dynamically evolve triggers that achieve excellent tradeoff between stealthiness and effectiveness. Defender adapts a closed-loop Selection-Evaluation-Suppression framework where high-reliability clients are selected via Fisher information to carry out reverse trigger engineering. Then clients' confidence scores are calculated as weights to minimize Attack Success Rate (ASR) during aggregation. Extensive experiments on six financial fraud datasets demonstrate the superiority of CLARF model and Min-Max adversarial game paradigm compared with multiple SOTA models. Ruixiao Zhu, Kun Zhu 0024, Qi Zhang 0020, Changjun Jiang 0002 |
WWW | 4 |
| 2026 | Medication mapping and diagnosis enhancement for fine-grained medication recommendation
Yishuo Li, Qi Zhang 0020, Shoujin Wang, Weiyu Zhang 0001, Jiasheng Si, Wenpeng Lu |
Inf. Sci. | 2 |
| 2026 | Knowledge-Enhanced Multi-Level Session Graph Model for Interactive Recommendation through Deep Reinforcement LearningabstractDeep reinforcement learning (DRL), has shown promise in solving intractable challenges in interactive recommendation systems (IRS). In DRL-based interactive recommendation, state modeling is vital for well-capturing users’ continuous interaction behaviors with recommendation systems. To effectively capture the behavior of users, existing works for state modeling have evolved from sequential-based modeling to session-based modeling. However, existing session-based state modeling works in IRS are still not fully explored with premature session models and insufficient fusion for different session features. As a result, they cannot capture complicated session patterns during interaction, leading to significant information loss. In this article, we propose a Knowledge-enhanced Multi-Level Session Graph (KMSG) model for interactive recommendation to address the above challenge. KMSG models the user’s interactive data into multi-level session graphs and effectively encodes the states via graph neural networks. Specifically, a novel 3-level item transition graph is designed to capture the common session patterns and intra-session item transitions. We further utilize the information from the knowledge graph to enhance the item relations in KMSG. We then design an attention-based graph neural network to propagate the information in KMSG. Extensive experiments on four real-world benchmark datasets demonstrate the superiority of KMSG over state-of-the-art baselines and the rationality of our design in KMSG. Longxiang Shi, Shoujin Wang, Qi Zhang 0020, Kui Su, Shijian Li |
ACM Trans. Knowl. Discov. Data | 5 |
| 2025 | A Multimodal Prompt-based Framework for Analyzing Code-Mixed and Low-Resource MemesabstractThe emergence of social media has led memes to become a powerful mode of communication, blending text, images, and emojis. However, this surge in meme usage has also seen a rise in offensive material. With manual content moderation proving impractical due to the sheer volume of data, there's a pressing need for automated methods to identify harmful memes. Yet, existing research predominantly targets high-resource languages such as English, neglecting low-resource ones like Nepali. To bridge this gap, we introduce the first Nepali meme dataset annotated for hate speech and sentiment. Our contributions are threefold: (1) We create and release NeMeme, a unique dataset featuring Nepali and code-mixed Nepali memes (combining Nepali and English). (2) We evaluate NeMeme using cutting-edge unimodal and multimodal models to establish initial performance benchmarks. (3) We introduce MemeNePAL, a novel multimodal framework employing prompt-assisted learning to effectively categorize Nepali memes. MemeNePAL overcomes the shortcomings of prior state-of-the-art (SOTA) techniques, which were designed for high-resource languages and struggle with Nepali's linguistic differences and cultural subtleties. This work not only promotes inclusivity in content moderation research but also aligns with UN Sustainable Development Goals such as promoting well-being, reducing inequalities, and fostering peace. We adhere to FAIR principles by making the dataset publicly available. Surendrabikram Thapa, Hariram Veeramani, Liang Hu 0004, Qi Zhang 0020, Wei Wang 0077, Usman Naseem |
ICWSM | 4 |
| 2025 | A Survey on Deep Learning based Time Series Analysis with Frequency TransformationabstractRecently, frequency transformation (FT) has been increasingly incorporated into deep learning models to significantly enhance state-of-the-art accuracy and efficiency in time series analysis. The advantages of FT, such as high efficiency and a global view, have been rapidly explored and exploited in various time series tasks and applications, demonstrating the promising potential of FT as a new deep learning paradigm for time series analysis. Despite the growing attention and the proliferation of research in this emerging field, there is currently a lack of a systematic review and in-depth analysis of deep learning-based time series models with FT. It is also unclear why FT can enhance time series analysis and what its limitations are in the field. To address these gaps, we present a comprehensive review that systematically investigates and summarizes the recent research advancements in deep learning-based time series analysis with FT. Specifically, we explore the primary approaches used in current models that incorporate FT, the types of neural networks that leverage FT, and the representative FT-equipped models in deep time series analysis. We propose a novel taxonomy to categorize the existing methods in this field, providing a structured overview of the diverse approaches employed in incorporating FT into deep learning models for time series analysis. Finally, we highlight the advantages and limitations of FT for time series modeling and identify potential future research directions that can further contribute to the community of time series analysis. Kun Yi 0001, Qi Zhang 0020, Wei Fan 0010, Longbing Cao, Shoujin Wang, Guodong Long, Liang Hu 0004, Qingsong Wen, Hui Xiong 0001 |
KDD (2) | 2 |
| 2025 | IN-Flow: Instance Normalization Flow for Non-stationary Time Series ForecastingabstractDue to the non-stationarity of time series, the distribution shift problem largely hinders the performance of time series forecasting. Existing solutions either rely on using certain statistics to specify the shift, or developing specific mechanisms for certain network architectures. However, the former would fail for the unknown shift beyond simple statistics, while the latter has limited compatibility on different forecasting models. To overcome these problems, we first propose a decoupled formulation for time series forecasting, with no reliance on fixed statistics and no restriction on forecasting architectures. This formulation regards the removing-shift procedure as a special transformation between a raw distribution and a desired target distribution and separates it from the forecasting. Such a formulation is further formalized into a bi-level optimization problem, to enable the joint learning of the transformation (outer loop) and forecasting (inner loop). Moreover, the special requirements of expressiveness and bi-direction for the transformation motivate us to propose instance normalization flow (IN-Flow), a novel invertible network for time series transformation. Different from the classic ''normalizing flow'' models, IN-Flow does not aim for normalizing input to the prior distribution (e.g., Gaussian distribution) for generation, but creatively transforms time series distribution by stacking normalization layers and flow-based invertible networks, which is thus named ''normalization'' flow. Finally, we have conducted extensive experiments on both synthetic data and real-world data, which demonstrate the superiority of our method. Wei Fan 0010, Shun Zheng 0001, Pengyang Wang, Rui Xie 0002, Kun Yi 0001, Qi Zhang 0020, Jiang Bian 0002, Yanjie Fu |
KDD (1) | 6 |
| 2025 | ALSA: Context-Sensitive Prompt Privacy Preservation in Large Language ModelsabstractThe remarkable prompting capability of large language models (LLMs) offers substantial convenience to users across diverse backgrounds. Nevertheless, as the sensitive information within prompts is inevitably exposed to LLMs, caution must be exercised to preserve privacy. Among various studies, text anonymization is considered an effective approach to preventing privacy leakage in prompts through text substitution. However, existing works overemphasize privacy while overlooks preserving contextual integrity, degrading semantic consistency. To address these concerns, this paper introduces a context-sensitive prompt privacy-preserving framework, namely Adaptive Linguistic Sanitization and Anonymization (ALSA). In specific, ALSA incorporates a three-dimensional scoring mechanism to dynamically quantify the substitutability of each word within a prompt by integrating the Privacy Leakage Risk Score (PLRS), the Contextual Information Importance Score (CIIS), and the Task Relevance Score (TRS). Subsequently, a clustering technique is adopted to dynamically determine the threshold for assigning an anonymization action (i.e., Retain, Replace, Encrypt, or Delete) by balancing privacy, semantics, and task relevance. Extensive experiments on five benchmark datasets validate the superiority of ALSA over state-of-the-art baselines in terms of accuracy, privacy preservation, and semantic integrity. Hongru Ma, Wenpeng Lu, Tianyi Wang 0006, Qi Zhang 0020, Yingjie Zhu, Jiasheng Si |
KDD (2) | 5 |
| 2025 | Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal RecommendationsabstractIncomplete scenario is a prevalent, practical, yet challenging setting in Multimodal Recommendations (MMRec), where some item modalities are missing due to various factors. Recently, a few efforts have sought to improve the recommendation accuracy by exploring generic structures from incomplete data. However, two significant gaps persist: 1) the difficulty in accurately generating missing data due to the limited ability to capture modality distributions; and 2) the critical but overlooked visibility bias, where items with missing modalities are more likely to be disregarded due to the prioritization of items' multimodal data over user preference alignment. This bias raises serious concerns about the fair treatment of items. To bridge these two gaps, we propose a novel Modality-Diffused Counterfactual (MoDiCF) framework for incomplete multimodal recommendations. MoDiCF features two key modules: a novel modality-diffused data completion module and a new counterfactual multimodal recommendation module. The former, equipped with a particularly designed multimodal generative framework, accurately generates and iteratively refines missing data from learned modality-specific distribution spaces. The latter, grounded in the causal perspective, effectively mitigates the negative causal effects of visibility bias and thus assures fairness in recommendations. Both modules work collaboratively to address the two aforementioned significant gaps for generating more accurate and fair results. Extensive experiments on three real-world datasets demonstrate the superior performance of MoDiCF in terms of both recommendation accuracy and fairness. The code and processed datasets are released at https://github.com/JinLi-i/MoDiCF. Jin Li 0002, Shoujin Wang, Qi Zhang 0020, Shui Yu 0001, Fang Chen 0001 |
WWW | 3 |
| 2025 | Time-aware Medication Recommendation via Intervention of Dynamic Treatment RegimesabstractMedication recommendation aims to suggest personalized drug combinations to patients based on their longitudinal medical histories stored in electronic health record (EHR) datasets. Patients' Dynamic Treatment Regimes (DTRs) determine how patients' drug combinations change along with the evolution of disease treatment. DTRs are effective for comprehending disease-treatment dynamics and for recommending a timely and personalized combination of medications for patients. However, existing medication recommender systems (MRSs) overlook the multiple treatment pathways generated by the intervention of DTRs and can only recommend a single treatment paradigm, ignoring the fact that patients may be at different treatment stages and thus require different treatment regime. Such disregard leads to a significant limitation in recommending personalized medication combinations tailored to different treatment stages, yielding greatly compromised accuracy and applicability of MRSs. Moreover, existing methods often overlook the time interval information over patients' successive visits, which is critical to indicate patients' treatment evolution. To address these significant gaps, we propose a Time-aware Medication Recommendation Framework via Intervention of Dynamic Treatment Regimes, called MR-DTR. To explicitly illustrate the intervention processes of DTRs on similar patients, we employ a co-guided graph to connect various patient sequences. In addition, to fully utilize the time interval information, we design a time-aware guidance mechanism dedicated to the co-guided graph to efficiently learn medication representation using the patient's guidance information. We also introduce relative time intervals in the encoder to act as positional information. Extensive experiments on two real-world datasets demonstrate that MR-DTR surpasses state-of-the-art models in terms of recommendation performance. Our code is available at: https://github.com/liyifo/MR-DTR. Yishuo Li, Qi Zhang 0020, Wenpeng Lu, Xueping Peng, Weiyu Zhang 0001, Jiasheng Si, Yongshun Gong, Liang Hu 0004 |
WWW | 2 |
| 2025 | DiGrI: Distorted Greedy Approach for Human-Assisted Online Suicide Ideation DetectionabstractUser-generated content on social media platforms provides a valuable resource for developing automated computational methods to detect mental health issues online leading to suicidal thoughts automatically. Although current fully automated methods show promise, they may produce uncertain predictions, leading to flawed conclusions. To address this, we propose a novel model called DiGrI, or Distorted Greedy Approach for Human-Assisted Online Suicide Ideation Detection, which reformulates suicide ideation assessment as a selective, prioritized prediction problem. The model incorporates a novel multi-classifier distorted greedy model that is optimized to operate under various levels of automation and abstains from making uncertain predictions with theoretical guarantees. Our results show that DiGrI outperforms strong comparative models including large language models in detecting mental health issues on a publicly available Reddit dataset. We discuss the empirical and practical implications, including the ethical considerations of using DiGrI for online automatic suicide ideation detection involving humans, if it were to be translated for use in clinical and public health practice. Usman Naseem, Liang Hu 0008, Qi Zhang 0020, Shoujin Wang, Shoaib Jameel |
WWW | 3 |
| 2025 | RevGNN: Negative Sampling Enhanced Contrastive Graph Learning for Academic Reviewer RecommendationabstractAcquiring reviewers for academic submissions is a challenging recommendation scenario. Recent graph learning-driven models have made remarkable progress in the field of recommendation, but their performance in the academic reviewer recommendation task may suffer from a significant false negative issue. This arises from the assumption that unobserved edges represent negative samples. In fact, the mechanism of anonymous review results in inadequate exposure of interactions between reviewers and submissions, leading to a higher number of unobserved interactions compared to those caused by reviewers declining to participate. Therefore, investigating how to better comprehend the negative labeling of unobserved interactions in academic reviewer recommendations is a significant challenge. This study aims to tackle the ambiguous nature of unobserved interactions in academic reviewer recommendations. Specifically, we propose an unsupervised Pseudo Neg-Label strategy to enhance graph contrastive learning (GCL) for recommending reviewers for academic submissions, which we call RevGNN. RevGNN utilizes a two-stage encoder structure that encodes both scientific knowledge and behavior using Pseudo Neg-Label to approximate review preference. Extensive experiments on three real-world datasets demonstrate that RevGNN outperforms all baselines across four metrics. Additionally, detailed further analyses confirm the effectiveness of each component in RevGNN. Weibin Liao, Yifan Zhu 0001, Qi Zhang 0020, Zhonghong Ou, Xuesong Li 0003 |
ACM Trans. Inf. Syst. | 4 |
| 2024 | Shrink: Data Compression by Semantic Extraction and Residuals EncodingabstractThe distributed data infrastructure in Internet of Things (IoT) ecosystems requires efficient data-series compression methods, as well as the capability to meet different accuracy demands. However, the compression performance of existing compression methods degrades sharply when calling for ultra-accurate data recovery. In this paper, we introduce Shrink, a novel highly accurate data compression method that offers a higher compression ratio and lower runtime than prior compressors. Shrink extracts data semantics in the form of linear segments to construct a compact knowledge base, using a dynamic error threshold which can adapt to data characteristics. Then, it captures the remaining data details as residuals to support lossy compression at diverse resolutions as well as lossless compression. As Shrink effectively identifies repeated semantics, its compression ratio increases with data size. Our experimental evaluation demonstrates that Shrink outperforms state-of-art methods, achieving a twofold to fivefold improvement in compression ratio depending on the dataset. Guoyou Sun, Panagiotis Karras, Qi Zhang 0020 |
IEEE Big Data | 3 |
| 2024 | THYMES: A Framework for Detecting Suicidal Ideation from Social Media Posts Using Hyperbolic LearningabstractMental health concerns are a critical issue in today’s digital age, posing a threat to both individual and societal well-being and making the identification of at-risk individuals crucial. Analyzing an individual’s social media post history can offer insights into their mental health state and help identify the presence of suicidal ideation. However, the complexity of linguistic and temporal data, along with sparsity and time irregularities, poses a formidable challenge in machine learning. Previous methods in this domain either rely on Euclidean space for processing which does not adequately model the power-law properties of social media posts, or lose information due to the discretization of the time axis. To address these challenges, we propose a novel framework, THYMES, which leverages pre-trained encoders and a rich representation learning paradigm with hyperbolic learning to model power-law features for enhanced sequence modeling. We perform experiments on two datasets and demonstrate that THYMES outperforms previously proposed methods while maintaining classification fairness under heavy data imbalances. Additionally, we qualitatively analyze commonly misclassified samples to reveal the shortcomings of models in this domain. Surendrabikram Thapa, Mohammad Salman, Siddhant Bikram Shah, Shuvam Shiwakoti, Qi Zhang 0020, Liang Hu 0004, Muhammad Imran Razzak, Usman Naseem |
IEEE Big Data | 5 |
| 2024 | SAFENet: Towards a Robust Suicide Assessment in Social Media Using Selective Prediction FrameworkabstractThe rising rate of mental health issues in the digital age underscores the critical need for proactive interventions to assess an individual’s well-being. This problem is further exacerbated by the social stigma surrounding the subject, which suppresses the willingness of victims to seek help. Social media can serve as an outlet for such individuals to express their negative emotions or thoughts of self-harm. The social media account of an individual can offer a plethora of valuable information that can be used to predict their mental health. By unifying principles of robust classifier training and selective classification, we propose a novel framework, SAFENet, to predict the suicide risk of users by using their historical social media posts. When the confidence of prediction is low or the individual is classified as a high-risk user, SAFENet delegates the analysis of the posts to a human evaluator for further intervention. Our experiments show that SAFENet outperforms existing state-of-the-art frameworks. We further qualitatively analyze predictions from SAFENet and demonstrate that it performs robustly on difficult samples that may cause contemporary methods to make errors. Our system addresses the urgent need for efficient and effective mental health intervention in the digital era. Surendrabikram Thapa, Mohammad Salman, Siddhant Bikram Shah, Qi Zhang 0020, Junaid Rashid, Liang Hu 0004, Muhammad Imran Razzak, Usman Naseem |
IEEE Big Data | 4 |
| 2024 | Structural Representation Learning and Disentanglement for Evidential Chinese Patent Approval PredictionabstractAutomatic Chinese patent approval prediction is an emerging and valuable task in patent analysis. However, it involves a rigorous and transparent decision-making process that includes patent comparison and examination to assess its innovation and correctness. This resultant necessity of decision evidentiality, coupled with intricate patent comprehension presents significant challenges and obstacles for the patent analysis community. Consequently, few existing studies are addressing this task. This paper presents the pioneering effort on this task using a retrieval-based classification approach. We propose a novel framework called DiSPat, which focuses on structural representation learning and disentanglement to predict the approval of Chinese patents and offer decision-making evidence. DiSPat comprises three main components: base reference retrieval to retrieve the Top-k most similar patents as a reference base; structural patent representation to exploit the inherent claim hierarchy in patents for learning a structural patent representation; disentangled representation learning to learn disentangled patent representations that enable the establishment of an evidential decision-making process. To ensure a thorough evaluation, we have meticulously constructed three datasets of Chinese patents. Extensive experiments on these datasets unequivocally demonstrate our DiSPat surpasses state-of-the-art baselines on patent approval prediction, while also exhibiting enhanced evidentiality. Jinzhi Shan, Qi Zhang 0020, Chongyang Shi 0001, Mengting Gui, Shoujin Wang, Usman Naseem |
CIKM | 2 |
| 2024 | Exploitation or Exploration Next? User Behavior Decoupling and Emerging Intent Modeling for Next-Item RecommendationabstractRecent trends in next-item recommendation systems have focused on modeling user intents. Traditional methods often extract users' inherent intents from the most representative items in a session, overlooking “unexpected items” that deviate from the majority in various contextual aspects. These unexpected items, frequently present, can be crucial indicators of a user's inclination towards exploring new options, signaling emerging intents that warrant significant attention. In response, we introduce DbMei, a novel approach that decouples user behaviors and emphasizes the modeling of emerging intents. DbMei distinguishes between two user behavior types: “focused shopping”, which aligns with users' inherent intents, and”wandering shopping”, which aligns with emerging intents. Focused shopping is analyzed using topic modeling and hypergraph learning while wandering shopping is explored through session neighbor retrieval. An exploitation-exploration mechanism is employed to determine the behavioral probability distribution for upcoming items. This integrated modeling of focused and wandering shopping behaviors drives our recommendation process. Extensive empirical studies on two real-world datasets, Amazon-KDD and Beauty, showcase DbMei's superiority over leading methods regarding Recall and MRR metrics. Our code is publicly available at https://github.com/sunlingdan-123/DbMei. Nengjun Zhu, Lingdan Sun, Xiangfeng Luo, Jian Cao 0001, Qi Zhang 0020, Xinjiang Lu |
ICDM | 5 |
| 2024 | Reinforced Subject-Aware Graph Neural Network for Related Work Generation
Luyao Yu, Qi Zhang 0020, Chongyang Shi 0001, An Lao, Liang Xiao 0010 |
KSEM (1) | 2 |
| 2024 | MSynFD: Multi-hop Syntax Aware Fake News DetectionabstractThe proliferation of social media platforms has fueled the rapid dissemination of fake news, posing threats to our real-life society. Existing methods use multimodal data or contextual information to enhance the detection of fake news by analyzing news content and/or its social context. However, these methods often overlook essential textual news content (articles) and heavily rely on sequential modeling and global attention to extract semantic information. These existing methods fail to handle the complex, subtle twists1 in news articles, such as syntax-semantics mismatches and prior biases, leading to lower performance and potential failure when modalities or social context are missing. To bridge these significant gaps, we propose a novel multi-hop syntax aware fake news detection (MSynFD) method, which incorporates complementary syntax information to deal with subtle twists in fake news. Specifically, we introduce a syntactical dependency graph and design a multi-hop subgraph aggregation mechanism to capture multi-hop syntax. It extends the effect of word perception, leading to effective noise filtering and adjacent relation enhancement. Subsequently, a sequential relative position-aware Transformer is designed to capture the sequential information, together with an elaborate keyword debiasing module to mitigate the prior bias. Extensive experimental results on two public benchmark datasets verify the effectiveness and superior performance of our proposed MSynFD over state-of-the-art detection models. Liang Xiao 0010, Qi Zhang 0020, Chongyang Shi 0001, Shoujin Wang, Usman Naseem, Liang Hu 0004 |
WWW | 2 |
| 2024 | Attentive multi-granularity perception network for person search
Qixian Zhang, Jun Wu 0006, Duoqian Miao 0001, Cairong Zhao, Qi Zhang 0020 |
Inf. Sci. | 5 |
| 2024 | Learning Informative Representation for Fairness-Aware Multivariate Time-Series Forecasting: A Group-Based PerspectiveabstractMultivariate time series (MTS) forecasting penetrates various aspects of our economy and society, whose roles become increasingly recognized. However, often MTS forecasting is unfair, not only degrading their practical benefits but even incurring potential risk. Unfair MTS forecasting may be attributed to disparities relating to advantaged and disadvantaged variables, which has rarely been studied in the MTS forecasting. In this work, we formulate the MTS fairness modeling problem as learning informative representations attending to both advantaged and disadvantaged variables. Accordingly, we propose a novel framework, namedFairFor, for fairness-aware MTS forecasting, i.e.,fair MTS forecasting.FairForuses adversarial learning to generate both group-irrelevant and -relevant representations for downstream forecasting.FairForfirst adopts recurrent graph convolution to capture spatio-temporal variable correlations and to group variables by leveraging a spectral relaxation of the K-means objective. Then, it utilizes a novel filtering$\&$fusion module to filter group-relevant information and generate group-irrelevant representations by orthogonality regularization. The group-irrelevant and -relevant representations form highly informative representations, facilitating to share the knowledge from advantaged variables to disadvantaged variables and guarantee the fairness of forecasting. Extensive experiments on four public datasets demonstrate theFairForeffectiveness for fair forecasting and significant performance improvement. Qi Zhang 0020, Shoujin Wang, Kun Yi 0001, Zhendong Niu, Longbing Cao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Deep Coupling Network for Multivariate Time Series ForecastingabstractMultivariate time series (MTS) forecasting is crucial in many real-world applications. To achieve accurate MTS forecasting, it is essential to simultaneously consider both intra- and inter-series relationships among time series data. However, previous work has typically modeled intra- and inter-series relationships separately and has disregarded multi-order interactions present within and between time series data, which can seriously degrade forecasting accuracy. In this article, we reexamine intra- and inter-series relationships from the perspective of mutual information and accordingly construct a comprehensive relationship learning mechanism tailored to simultaneously capture the intricate multi-order intra- and inter-series couplings. Based on the mechanism, we propose a novel deep coupling network for MTS forecasting, named DeepCN, which consists of a coupling mechanism dedicated to explicitly exploring the multi-order intra- and inter-series relationships among time series data concurrently, a coupled variable representation module aimed at encoding diverse variable patterns, and an inference module facilitating predictions through one forward step. Extensive experiments conducted on seven real-world datasets demonstrate that our proposed DeepCN achieves superior performance compared with the state-of-the-art baselines. Kun Yi 0001, Qi Zhang 0020, Kaize Shi, Liang Hu 0004, Ning An 0001, Zhendong Niu |
ACM Trans. Inf. Syst. | 2 |
| 2023 | PatSTEG: Modeling Formation Dynamics of Patent Citation Networks via The Semantic-Topological Evolutionary GraphabstractPatent documents in the patent database (PatDB) are crucial for research, development, and innovation as they contain valuable technical information. However, PatDB presents a multifaceted challenge in comparison to publicly available preprocessed databases due to the intricate nature of patent text and the inherent sparsity within the patent citation network. Although patent text analysis and citation analysis bring new opportunities to explore patent data mining, no existing work exploits the complementation of them. To this end, we propose a joint semantic-topological evolutionary graph learning approach (PatSTEG) to model the formation dynamics of patent citation networks. More specifically, we first create a real-world dataset of Chinese patents named CNPat, and leveraging its patent texts and citations to construct a patent citation network. Then, PatSTEG is modeled to study the evolutionary dynamics of patent citation formation by jointly considering the semantic and topological information. Extensive experiments are conducted on both CNPat and public datasets to prove the superiority of PatSTEG over other state-of-the-art methods. All the results provide valuable references for patent literature research and technical exploration. Ran Miao, Xueyu Chen, Liang Hu 0004, Minghua Wan, Qi Zhang 0020, Cairong Zhao |
ICDM | 6 |
| 2023 | Session-based Interactive Recommendation via Deep Reinforcement LearningabstractDeep reinforcement learning (DRL), has shown promise in solving intractable challenges in interactive recommendation systems. In DRL-based interactive recommendation, state modeling is crucial for well-capturing users’ continuous interaction behaviors with shopping systems. A user’s multiple continuous interactions in a given time period (e.g., the time from login to log out) naturally constitute a session. However, existing studies often overlook such valuable session structure and characteristics and instead simply treat them as sequences. As a result, they are not able to capture the complex transitions over users’ interactions within or between sessions, leading to significant information loss. To bridge this significant gap, in this paper, we propose Session-based Interactive Recommendation with Graph Neural Networks (SIR-GNN). SIR-GNN models interaction data as sessions and employs novel graph neural networks to capture rich transition patterns among interactions. Specifically, a novel 3-level transition module is well designed to effectively capture common patterns from all sessions, intra-session transitions, and adjacent-item transitions respectively, followed by an attention-based gated graph neural network to model the state representation for SIR well. Extensive experiments on 3 real-world benchmark datasets demonstrate the superiority of SIR-GNN over state-of-the-art baselines and the rationality of our design in SIR-GNN. Longxiang Shi, Shoujin Wang, Qi Zhang 0020, Shijian Li |
ICDM | 4 |
| 2023 | MDKG: Graph-Based Medical Knowledge-Guided Dialogue Generation
Usman Naseem, Surendrabikram Thapa, Qi Zhang 0020, Liang Hu 0004, Mehwish Nasim |
SIGIR | 3 |
| 2023 | Show Me The Best Outfit for A Certain Scene: A Scene-aware Fashion Recommender SystemabstractFashion recommendation (FR) has received increasing attention in the research of new types of recommender systems. Existing fashion recommender systems (FRSs) typically focus on clothing item suggestions for users in three scenarios: 1) how to best recommend fashion items preferred by users; 2) how to best compose a complete outfit, and 3) how to best complete a clothing ensemble. However, current FRSs often overlook an important aspect when making FR, that is, the compatibility of the clothing item or outfit recommendations is highly dependent on the scene context. To this end, we propose the scene-aware fashion recommender system (SAFRS), which uncovers a hitherto unexplored avenue where scene information is taken into account when constructing the FR model. More specifically, our SAFRS addresses this problem by encoding scene and outfit information in separation attention encoders and then fusing the resulting feature embeddings via a novel scene-aware compatibility score function. Extensive qualitative and quantitative experiments are conducted to show that our SAFRS model outperforms all baselines for every evaluated metric. Tangwei Ye, Liang Hu 0004, Qi Zhang 0020, Zhongyuan Lai, Usman Naseem, Dora D. Liu |
WWW | 3 |
| 2023 | Modeling User Demand Evolution for Next-Basket PredictionabstractUsers’ purchase behaviors are complex and dynamic, which are usually driven by various personal demands evolving with time. According to psychology and economic theories, user demands can be satisfied with a sequence of purchase behaviors, resulting in a basket of items. However, most of the existing works simply predict the next basket from a shallow perspective of (purchase) sequence data modeling without deep insight into the underlying factors which drive user purchase behaviors. In fact, filling a basket with multiple items is a process to incrementally satisfy a user's demand. Therefore, the key challenges to predict a user's next basket lie in (1) how to track the changes of the user's demand, and (2) how to satisfy her demand at a given moment. To this end, we propose an Evolving DEmand SAtisfaction (EvoDESA) model to model a user's demand evolution for next-basket prediction. In EvoDESA, a demand evolution module learns the dynamics of user demand over a sequence of basket-purchase behaviors. Then, a next-basket planning module effectively packs an optimal combination of items to best satisfy the user's current demand. Extensive experiments on three real-world transaction datasets demonstrate the considerable superiority of EvoDESA over the state-of-the-art approaches. Shoujin Wang, Yan Wang 0002, Liang Hu 0004, Xiuzhen Zhang 0001, Qi Zhang 0020, Quan Z. Sheng, Mehmet A. Orgun, Longbing Cao, Defu Lian |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2022 | An Ion Exchange Mechanism Inspired Story Ending Generator for Different Characters
Qi Zhang 0020, Chongyang Shi 0001, Kaiying Jiang, Liang Hu 0004, Shoujin Wang |
ECML/PKDD (2) | 2 |
| 2022 | Sequential/Session-based Recommendations: Challenges, Approaches, Applications and OpportunitiesabstractIn recent years, sequential recommender systems (SRSs) and session-based recommender systems (SBRSs) have emerged as a new paradigm of RSs to capture users' short-term but dynamic preferences for enabling more timely and accurate recommendations. Although SRSs and SBRSs have been extensively studied, there are many inconsistencies in this area caused by the diverse descriptions, settings, assumptions and application domains. There is no work to provide a unified framework and problem statement to remove the commonly existing and various inconsistencies in the area of SR/SBR. There is a lack of work to provide a comprehensive and systematic demonstration of the data characteristics, key challenges, most representative and state-of-the-art approaches, typical real- world applications and important future research directions in the area. This work aims to fill in these gaps so as to facilitate further research in this exciting and vibrant area. Shoujin Wang, Qi Zhang 0020, Liang Hu 0004, Xiuzhen Zhang 0001, Yan Wang 0002, Charu C. Aggarwal |
SIGIR | 2 |
| 2020 | Mix2Vec: Unsupervised Mixed Data RepresentationabstractUnsupervised representation learning on mixed data is highly challenging but rarely explored. It has to tackle significant challenges related to common issues in real-life mixed data, including sparsity, dynamics and heterogeneity of attributes and values. This work introduces an effective and efficient unsupervised deep representer called Mix2Vec to automatically learn a universal representation of dynamic mixed data with the above complex characteristics. Mix2Vec is empowered with three effective mechanisms: random shuffling prediction, prior distribution matching, and structural informativeness maximization, to tackle the aforementioned challenges. These mechanisms are implemented as an unsupervised deep neural representer Mix2Vec. Mix2Vec converts complex mixed data into vector space-based representations that are universal and comparable to all data objects and transparent and reusable for both unsupervised and supervised learning tasks. Extensive experiments on four large mixed datasets demonstrate that Mix2Vec performs significantly better than state-of-the-art deep representation methods. We also empirically verify the designed mechanisms in terms of representation quality, visualization and capability of enabling better performance of downstream tasks. Chengzhang Zhu, Qi Zhang 0020, Longbing Cao, Arman Abrahamyan |
DSAA | 2 |
| 2019 | HCBC: A Hierarchical Case-Based Classifier Integrated with Conceptual ClusteringabstractThe structured case representation improves case-based reasoning (CBR) by exploring structures in the case base and the relevance of case structures. Recent CBR classifiers have mostly been built upon the attribute-value case representation rather than structured case representation, in which the structural relations embodied in their representation structure are accordingly overlooked in improving the similarity measure. This results in retrieval inefficiency and limitations on the performance of CBR classifiers. This paper proposes a hierarchical case-based classifier, HCBC, which introduces a concept lattice to hierarchically organize cases. By exploiting structural case relations in the concept lattice, a novel dynamic weighting model is proposed to enhance the concept similarity measure. Based on this similarity measure, HCBC retrieves the top-K concepts that are most similar to a new case by using a bottom-up pruning-based recursive retrieval (PRR) algorithm. The concepts extracted in this way are applied to suggest a class label for the case by a weighted majority voting. Experimental results show that HCBC outperforms other classifiers in terms of classification performance and robustness on categorical data, and also works confidently well on numeric datasets. In addition, PRR effectively reduces the search space and greatly improves the retrieval efficiency of HCBC. Qi Zhang 0020, Chongyang Shi 0001, Zhendong Niu, Longbing Cao |
IEEE Trans. Knowl. Data Eng. | 1 |