Liang Hu 0004

dblp:48/5388-4 · DBLP profile ↗
← Back
24ranked-venue papers in the field
4as first author
16since 2021 · last 2026
0000-0001-8588-2177ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 14 (3 first)Data Mining & Knowledge Discovery · 6 (1 first)Big Data, Cloud & Distributed Data Systems · 2Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Counterfactual Augmented Causal Reasoning for Aspect-Based Sentiment Analysis
Liang Hu 0004, Mingzhu Zhou, Tangwei Ye, Xuejie Yang, Zhongyuan Lai, Qi Zhang 0020, Usman Naseem
WWW2
2026 AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information Minimization
abstract
The escalating energy consumption and carbon footprint of training large-scale Web AI models pose urgent challenges for sustainable development. Dataset Distillation (DD) offers a promising avenue for green AI by compressing large datasets into small synthetic ones for efficient training. However, most existing DD methods overfit to the inductive biases of specific source architectures (e.g., CNNs or ViTs), resulting in poor cross-model generalization. This limitation necessitates redundant re-distillation processes for different architectures, severely undermining the energy-saving potential of DD. To address this, we introduce Adversarial Mutual Information Distillation (AMID), a rigorous framework designed to create highly reusable and robust synthetic datasets. From an information-theoretic perspective, we cast model-agnosticism as minimizing the mutual information (MI) between the synthetic data and the specific identity of the distillation model. We convert this intractable objective into a tractable two-player adversarial game, which unifies knowledge preservation with adversarial unlearning of architectural bias. Extensive experiments on CIFAR-10 and Tiny ImageNet demonstrate that AMID achieves state-of-the-art cross-architecture generalization across diverse CNNs and ViTs. Crucially, our analysis confirms that AMID significantly reduces the computational overhead and CO2 emissions of downstream training while maintaining robust performance, paving the way for energy-efficient, transferable, and sustainable Web AI ecosystems.
Aoqi Wu, Weiquan Huang, Liang Hu 0004, Yifan Yang 0004, Qi Zhang 0020, Jiaxing Miao, Yuhan Tang, Zhongyuan Lai
WWW5
2026 Accurate Trajectory Recovery in Underserved Areas via Location Inference from Web Crowdsourced Data
Tangwei Ye, Liang Hu 0004, Zhongyuan Lai, Qi Zhang 0020, Jiaxing Miao, Kun Yi 0001
WWW2
2025 A Multimodal Prompt-based Framework for Analyzing Code-Mixed and Low-Resource Memes
abstract
The emergence of social media has led memes to become a powerful mode of communication, blending text, images, and emojis. However, this surge in meme usage has also seen a rise in offensive material. With manual content moderation proving impractical due to the sheer volume of data, there's a pressing need for automated methods to identify harmful memes. Yet, existing research predominantly targets high-resource languages such as English, neglecting low-resource ones like Nepali. To bridge this gap, we introduce the first Nepali meme dataset annotated for hate speech and sentiment. Our contributions are threefold: (1) We create and release NeMeme, a unique dataset featuring Nepali and code-mixed Nepali memes (combining Nepali and English). (2) We evaluate NeMeme using cutting-edge unimodal and multimodal models to establish initial performance benchmarks. (3) We introduce MemeNePAL, a novel multimodal framework employing prompt-assisted learning to effectively categorize Nepali memes. MemeNePAL overcomes the shortcomings of prior state-of-the-art (SOTA) techniques, which were designed for high-resource languages and struggle with Nepali's linguistic differences and cultural subtleties. This work not only promotes inclusivity in content moderation research but also aligns with UN Sustainable Development Goals such as promoting well-being, reducing inequalities, and fostering peace. We adhere to FAIR principles by making the dataset publicly available.
Surendrabikram Thapa, Hariram Veeramani, Liang Hu 0004, Qi Zhang 0020, Wei Wang 0077, Usman Naseem
ICWSM3
2025 A Survey on Deep Learning based Time Series Analysis with Frequency Transformation
abstract
Recently, frequency transformation (FT) has been increasingly incorporated into deep learning models to significantly enhance state-of-the-art accuracy and efficiency in time series analysis. The advantages of FT, such as high efficiency and a global view, have been rapidly explored and exploited in various time series tasks and applications, demonstrating the promising potential of FT as a new deep learning paradigm for time series analysis. Despite the growing attention and the proliferation of research in this emerging field, there is currently a lack of a systematic review and in-depth analysis of deep learning-based time series models with FT. It is also unclear why FT can enhance time series analysis and what its limitations are in the field. To address these gaps, we present a comprehensive review that systematically investigates and summarizes the recent research advancements in deep learning-based time series analysis with FT. Specifically, we explore the primary approaches used in current models that incorporate FT, the types of neural networks that leverage FT, and the representative FT-equipped models in deep time series analysis. We propose a novel taxonomy to categorize the existing methods in this field, providing a structured overview of the diverse approaches employed in incorporating FT into deep learning models for time series analysis. Finally, we highlight the advantages and limitations of FT for time series modeling and identify potential future research directions that can further contribute to the community of time series analysis.
Kun Yi 0001, Qi Zhang 0020, Wei Fan 0010, Longbing Cao, Shoujin Wang, Guodong Long, Liang Hu 0004, Qingsong Wen, Hui Xiong 0001
KDD (2)8
2025 Time-aware Medication Recommendation via Intervention of Dynamic Treatment Regimes
abstract
Medication recommendation aims to suggest personalized drug combinations to patients based on their longitudinal medical histories stored in electronic health record (EHR) datasets. Patients' Dynamic Treatment Regimes (DTRs) determine how patients' drug combinations change along with the evolution of disease treatment. DTRs are effective for comprehending disease-treatment dynamics and for recommending a timely and personalized combination of medications for patients. However, existing medication recommender systems (MRSs) overlook the multiple treatment pathways generated by the intervention of DTRs and can only recommend a single treatment paradigm, ignoring the fact that patients may be at different treatment stages and thus require different treatment regime. Such disregard leads to a significant limitation in recommending personalized medication combinations tailored to different treatment stages, yielding greatly compromised accuracy and applicability of MRSs. Moreover, existing methods often overlook the time interval information over patients' successive visits, which is critical to indicate patients' treatment evolution. To address these significant gaps, we propose a Time-aware Medication Recommendation Framework via Intervention of Dynamic Treatment Regimes, called MR-DTR. To explicitly illustrate the intervention processes of DTRs on similar patients, we employ a co-guided graph to connect various patient sequences. In addition, to fully utilize the time interval information, we design a time-aware guidance mechanism dedicated to the co-guided graph to efficiently learn medication representation using the patient's guidance information. We also introduce relative time intervals in the encoder to act as positional information. Extensive experiments on two real-world datasets demonstrate that MR-DTR surpasses state-of-the-art models in terms of recommendation performance. Our code is available at: https://github.com/liyifo/MR-DTR.
Yishuo Li, Qi Zhang 0020, Wenpeng Lu, Xueping Peng, Weiyu Zhang 0001, Jiasheng Si, Yongshun Gong, Liang Hu 0004
WWW8
2024 THYMES: A Framework for Detecting Suicidal Ideation from Social Media Posts Using Hyperbolic Learning
abstract
Mental health concerns are a critical issue in today’s digital age, posing a threat to both individual and societal well-being and making the identification of at-risk individuals crucial. Analyzing an individual’s social media post history can offer insights into their mental health state and help identify the presence of suicidal ideation. However, the complexity of linguistic and temporal data, along with sparsity and time irregularities, poses a formidable challenge in machine learning. Previous methods in this domain either rely on Euclidean space for processing which does not adequately model the power-law properties of social media posts, or lose information due to the discretization of the time axis. To address these challenges, we propose a novel framework, THYMES, which leverages pre-trained encoders and a rich representation learning paradigm with hyperbolic learning to model power-law features for enhanced sequence modeling. We perform experiments on two datasets and demonstrate that THYMES outperforms previously proposed methods while maintaining classification fairness under heavy data imbalances. Additionally, we qualitatively analyze commonly misclassified samples to reveal the shortcomings of models in this domain.
Surendrabikram Thapa, Mohammad Salman, Siddhant Bikram Shah, Shuvam Shiwakoti, Qi Zhang 0020, Liang Hu 0004, Muhammad Imran Razzak, Usman Naseem
IEEE Big Data6
2024 SAFENet: Towards a Robust Suicide Assessment in Social Media Using Selective Prediction Framework
abstract
The rising rate of mental health issues in the digital age underscores the critical need for proactive interventions to assess an individual’s well-being. This problem is further exacerbated by the social stigma surrounding the subject, which suppresses the willingness of victims to seek help. Social media can serve as an outlet for such individuals to express their negative emotions or thoughts of self-harm. The social media account of an individual can offer a plethora of valuable information that can be used to predict their mental health. By unifying principles of robust classifier training and selective classification, we propose a novel framework, SAFENet, to predict the suicide risk of users by using their historical social media posts. When the confidence of prediction is low or the individual is classified as a high-risk user, SAFENet delegates the analysis of the posts to a human evaluator for further intervention. Our experiments show that SAFENet outperforms existing state-of-the-art frameworks. We further qualitatively analyze predictions from SAFENet and demonstrate that it performs robustly on difficult samples that may cause contemporary methods to make errors. Our system addresses the urgent need for efficient and effective mental health intervention in the digital era.
Surendrabikram Thapa, Mohammad Salman, Siddhant Bikram Shah, Qi Zhang 0020, Junaid Rashid, Liang Hu 0004, Muhammad Imran Razzak, Usman Naseem
IEEE Big Data6
2024 MSynFD: Multi-hop Syntax Aware Fake News Detection
abstract
The proliferation of social media platforms has fueled the rapid dissemination of fake news, posing threats to our real-life society. Existing methods use multimodal data or contextual information to enhance the detection of fake news by analyzing news content and/or its social context. However, these methods often overlook essential textual news content (articles) and heavily rely on sequential modeling and global attention to extract semantic information. These existing methods fail to handle the complex, subtle twists1 in news articles, such as syntax-semantics mismatches and prior biases, leading to lower performance and potential failure when modalities or social context are missing. To bridge these significant gaps, we propose a novel multi-hop syntax aware fake news detection (MSynFD) method, which incorporates complementary syntax information to deal with subtle twists in fake news. Specifically, we introduce a syntactical dependency graph and design a multi-hop subgraph aggregation mechanism to capture multi-hop syntax. It extends the effect of word perception, leading to effective noise filtering and adjacent relation enhancement. Subsequently, a sequential relative position-aware Transformer is designed to capture the sequential information, together with an elaborate keyword debiasing module to mitigate the prior bias. Extensive experimental results on two public benchmark datasets verify the effectiveness and superior performance of our proposed MSynFD over state-of-the-art detection models.
Liang Xiao 0010, Qi Zhang 0020, Chongyang Shi 0001, Shoujin Wang, Usman Naseem, Liang Hu 0004
WWW6
2024 Deep Coupling Network for Multivariate Time Series Forecasting
abstract
Multivariate time series (MTS) forecasting is crucial in many real-world applications. To achieve accurate MTS forecasting, it is essential to simultaneously consider both intra- and inter-series relationships among time series data. However, previous work has typically modeled intra- and inter-series relationships separately and has disregarded multi-order interactions present within and between time series data, which can seriously degrade forecasting accuracy. In this article, we reexamine intra- and inter-series relationships from the perspective of mutual information and accordingly construct a comprehensive relationship learning mechanism tailored to simultaneously capture the intricate multi-order intra- and inter-series couplings. Based on the mechanism, we propose a novel deep coupling network for MTS forecasting, named DeepCN, which consists of a coupling mechanism dedicated to explicitly exploring the multi-order intra- and inter-series relationships among time series data concurrently, a coupled variable representation module aimed at encoding diverse variable patterns, and an inference module facilitating predictions through one forward step. Extensive experiments conducted on seven real-world datasets demonstrate that our proposed DeepCN achieves superior performance compared with the state-of-the-art baselines.
Kun Yi 0001, Qi Zhang 0020, Kaize Shi, Liang Hu 0004, Ning An 0001, Zhendong Niu
ACM Trans. Inf. Syst.5
2023 PatSTEG: Modeling Formation Dynamics of Patent Citation Networks via The Semantic-Topological Evolutionary Graph
abstract
Patent documents in the patent database (PatDB) are crucial for research, development, and innovation as they contain valuable technical information. However, PatDB presents a multifaceted challenge in comparison to publicly available preprocessed databases due to the intricate nature of patent text and the inherent sparsity within the patent citation network. Although patent text analysis and citation analysis bring new opportunities to explore patent data mining, no existing work exploits the complementation of them. To this end, we propose a joint semantic-topological evolutionary graph learning approach (PatSTEG) to model the formation dynamics of patent citation networks. More specifically, we first create a real-world dataset of Chinese patents named CNPat, and leveraging its patent texts and citations to construct a patent citation network. Then, PatSTEG is modeled to study the evolutionary dynamics of patent citation formation by jointly considering the semantic and topological information. Extensive experiments are conducted on both CNPat and public datasets to prove the superiority of PatSTEG over other state-of-the-art methods. All the results provide valuable references for patent literature research and technical exploration.
Ran Miao, Xueyu Chen, Liang Hu 0004, Minghua Wan, Qi Zhang 0020, Cairong Zhao
ICDM3
2023 MDKG: Graph-Based Medical Knowledge-Guided Dialogue Generation
Usman Naseem, Surendrabikram Thapa, Qi Zhang 0020, Liang Hu 0004, Mehwish Nasim
SIGIR4
2023 Show Me The Best Outfit for A Certain Scene: A Scene-aware Fashion Recommender System
abstract
Fashion recommendation (FR) has received increasing attention in the research of new types of recommender systems. Existing fashion recommender systems (FRSs) typically focus on clothing item suggestions for users in three scenarios: 1) how to best recommend fashion items preferred by users; 2) how to best compose a complete outfit, and 3) how to best complete a clothing ensemble. However, current FRSs often overlook an important aspect when making FR, that is, the compatibility of the clothing item or outfit recommendations is highly dependent on the scene context. To this end, we propose the scene-aware fashion recommender system (SAFRS), which uncovers a hitherto unexplored avenue where scene information is taken into account when constructing the FR model. More specifically, our SAFRS addresses this problem by encoding scene and outfit information in separation attention encoders and then fusing the resulting feature embeddings via a novel scene-aware compatibility score function. Extensive qualitative and quantitative experiments are conducted to show that our SAFRS model outperforms all baselines for every evaluated metric.
Tangwei Ye, Liang Hu 0004, Qi Zhang 0020, Zhongyuan Lai, Usman Naseem, Dora D. Liu
WWW2
2023 Modeling User Demand Evolution for Next-Basket Prediction
abstract
Users’ purchase behaviors are complex and dynamic, which are usually driven by various personal demands evolving with time. According to psychology and economic theories, user demands can be satisfied with a sequence of purchase behaviors, resulting in a basket of items. However, most of the existing works simply predict the next basket from a shallow perspective of (purchase) sequence data modeling without deep insight into the underlying factors which drive user purchase behaviors. In fact, filling a basket with multiple items is a process to incrementally satisfy a user's demand. Therefore, the key challenges to predict a user's next basket lie in (1) how to track the changes of the user's demand, and (2) how to satisfy her demand at a given moment. To this end, we propose an Evolving DEmand SAtisfaction (EvoDESA) model to model a user's demand evolution for next-basket prediction. In EvoDESA, a demand evolution module learns the dynamics of user demand over a sequence of basket-purchase behaviors. Then, a next-basket planning module effectively packs an optimal combination of items to best satisfy the user's current demand. Extensive experiments on three real-world transaction datasets demonstrate the considerable superiority of EvoDESA over the state-of-the-art approaches.
Shoujin Wang, Yan Wang 0002, Liang Hu 0004, Xiuzhen Zhang 0001, Qi Zhang 0020, Quan Z. Sheng, Mehmet A. Orgun, Longbing Cao, Defu Lian
IEEE Trans. Knowl. Data Eng.3
2022 An Ion Exchange Mechanism Inspired Story Ending Generator for Different Characters
Qi Zhang 0020, Chongyang Shi 0001, Kaiying Jiang, Liang Hu 0004, Shoujin Wang
ECML/PKDD (2)5
2022 Sequential/Session-based Recommendations: Challenges, Approaches, Applications and Opportunities
abstract
In recent years, sequential recommender systems (SRSs) and session-based recommender systems (SBRSs) have emerged as a new paradigm of RSs to capture users' short-term but dynamic preferences for enabling more timely and accurate recommendations. Although SRSs and SBRSs have been extensively studied, there are many inconsistencies in this area caused by the diverse descriptions, settings, assumptions and application domains. There is no work to provide a unified framework and problem statement to remove the commonly existing and various inconsistencies in the area of SR/SBR. There is a lack of work to provide a comprehensive and systematic demonstration of the data characteristics, key challenges, most representative and state-of-the-art approaches, typical real- world applications and important future research directions in the area. This work aims to fill in these gaps so as to facilitate further research in this exciting and vibrant area.
Shoujin Wang, Qi Zhang 0020, Liang Hu 0004, Xiuzhen Zhang 0001, Yan Wang 0002, Charu C. Aggarwal
SIGIR3
2018 FastPM: An approach to pattern matching via distributed stream processing
Dingyu Yang, Jianmei Guo, Zhi-Jie Wang 0009, Yuan Wang 0003, Jingsong Zhang, Liang Hu 0004, Jian Yin 0001, Jian Cao 0001
Inf. Sci.6
2017 Perceiving the Next Choice with Comprehensive Transaction Embeddings for Online Recommendation
Shoujin Wang, Liang Hu 0004, Longbing Cao
ECML/PKDD (2)2
2017 Improving the Quality of Recommendations for Users and Items in the Tail of Distribution
abstract
Short-head and long-tail distributed data are widely observed in the real world. The same is true of recommender systems (RSs), where a small number of popular items dominate the choices and feedback data while the rest only account for a small amount of feedback. As a result, most RS methods tend to learn user preferences from popular items since they account for most data. However, recent research in e-commerce and marketing has shown that future businesses will obtain greater profit from long-tail selling. Yet, although the number of long-tail items and users is much larger than that of short-head items and users, in reality, the amount of data associated with long-tail items and users is much less. As a result, user preferences tend to be popularity-biased. Furthermore, insufficient data makes long-tail items and users more vulnerable to shilling attack. To improve the quality of recommendations for items and users in the tail of distribution, we propose a coupled regularization approach that consists of two latent factor models: C-HMF, for enhancing credibility, and S-HMF, for emphasizing specialty on user choices. Specifically, the estimates learned from C-HMF and S-HMF recurrently serve as the empirical priors to regularize one another. Such coupled regularization leads to the comprehensive effects of final estimates, which produce more qualitative predictions for both tail users and tail items. To assess the effectiveness of our model, we conduct empirical evaluations on large real-world datasets with various metrics. The results prove that our approach significantly outperforms the compared methods.
Liang Hu 0004, Longbing Cao, Jian Cao 0001, Zhiping Gu, Guandong Xu, Jie Wang 0006
ACM Trans. Inf. Syst.1
2016 Learning Informative Priors from Heterogeneous Domains to Improve Recommendation in Cold-Start User Domains
abstract
In the real-world environment, users have sufficient experience in their focused domains but lack experience in other domains. Recommender systems are very helpful for recommending potentially desirable items to users in unfamiliar domains, and cross-domain collaborative filtering is therefore an important emerging research topic. However, it is inevitable that the cold-start issue will be encountered in unfamiliar domains due to the lack of feedback data. The Bayesian approach shows that priors play an important role when there are insufficient data, which implies that recommendation performance can be significantly improved in cold-start domains if informative priors can be provided. Based on this idea, we propose a Weighted Irregular Tensor Factorization (WITF) model to leverage multi-domain feedback data across all users to learn the cross-domain priors w.r.t. both users and items. The features learned from WITF serve as the informative priors on the latent factors of users and items in terms of weighted matrix factorization models. Moreover, WITF is a unified framework for dealing with both explicit feedback and implicit feedback. To prove the effectiveness of our approach, we studied three typical real-world cases in which a collection of empirical evaluations were conducted on real-world datasets to compare the performance of our model and other state-of-the-art approaches. The results show the superiority of our model over comparison models.
Liang Hu 0004, Longbing Cao, Jian Cao 0001, Zhiping Gu, Guandong Xu, Dingyu Yang
ACM Trans. Inf. Syst.1
2015 Document similarity analysis via involving both explicit and implicit semantic couplings
abstract
Document similarity analysis is increasingly critical since roughly 80% of big data is unstructured. Accordingly, semantic couplings (relatedness) have been recognized valuable for capturing the relationships between terms (words or phrases). Existing work focuses more on explicit relatedness, with respective models built. In this paper, we propose a comprehensive semantic similarity measure: Semantic Coupling Similarity (SCS), which (1) captures intra-term pair couplings within term pairs represented by patterns of explicit term co-occurrences in a document set, (2) extracts inter-term pair couplings between term pairs indicated by implicit couplings between term pairs through indirectly linked terms and paths between terms after term connections are converted to a graph presentation; and (3) semantic coupling similarity, integrating intra- and inter-term pair couplings towards a comprehensive capturing of explicit and implicit couplings between terms across documents. SCS caters for both synonymy and polysemy, and outperforms baseline methods consistently on all real data sets.
Liang Hu 0004, Wei Liu 0007, Longbing Cao
DSAA2
2014 Bayesian Heteroskedastic Choice Modeling on Non-identically Distributed Linkages
abstract
Choice modeling (CM) aims to describe and predict choices according to attributes of subjects and options. If we presume each choice making as the formation of link between subjects and options, immediately CM can be bridged to link analysis and prediction (LAP) problem. However, such a mapping is often not trivial and straightforward. In LAP problems, the only available observations are links among objects but their attributes are often inaccessible. Therefore, we extend CM into a latent feature space to avoid the need of explicit attributes. Moreover, LAP is usually based on binary linkage assumption that models observed links as positive instances and unobserved links as negative instances. Instead, we use a weaker assumption that treats unobserved links as pseudo negative instances. Furthermore, most subjects or options may be quite heterogeneous due to the long-tail distribution, which is failed to capture by conventional LAP approaches. To address above challenges, we propose a Bayesian heteroskedastic choice model to represent the non-identically distributed linkages in the LAP problems. Finally, the empirical evaluation on real-world datasets proves the superiority of our approach.
Liang Hu 0004, Wei Cao 0012, Jian Cao 0001, Guandong Xu, Longbing Cao, Zhiping Gu
ICDM1
2013 GEAM: A General and Event-Related Aspects Model for Twitter Event Detection
Yue You, Guangyan Huang, Jian Cao 0001, Enhong Chen, Jing He 0004, Yanchun Zhang, Liang Hu 0004
WISE (2)7
2013 Personalized recommendation via cross-domain triadic factorization
abstract
Collaborative filtering (CF) is a major technique in recommender systems to help users find their potentially desired items. Since the data sparsity problem is quite commonly encountered in real-world scenarios, Cross-Domain Collaborative Filtering (CDCF) hence is becoming an emerging research topic in recent years. However, due to the lack of sufficient dense explicit feedbacks and even no feedback available in users' uninvolved domains, current CDCF approaches may not perform satisfactorily in user preference prediction. In this paper, we propose a generalized Cross Domain Triadic Factorization (CDTF) model over the triadic relation user-item-domain, which can better capture the interactions between domain-specific user factors and item factors. In particular, we devise two CDTF algorithms to leverage user explicit and implicit feedbacks respectively, along with a genetic algorithm based weight parameters tuning algorithm to trade off influence among domains optimally. Finally, we conduct experiments to evaluate our models and compare with other state-of-the-art models by using two real world datasets. The results show the superiority of our models against other comparative models.
Liang Hu 0004, Jian Cao 0001, Guandong Xu, Longbing Cao, Zhiping Gu, Can Zhu
WWW1