Songlin Hu 0001

dblp:67/4108-1 · DBLP profile ↗
← Back
27ranked-venue papers in the field
1as first author
13since 2021 · last 2026
0000-0002-7170-3809ORCID · conflict

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 10 (1 first)Data Mining & Knowledge Discovery · 8Information Retrieval & Web Search · 8Knowledge Engineering, Semantic Web & Information Systems · 1
YearPublicationVenuePosition
2026 Mitigating Adversarial Attacks by Transferring LLM-generated Narrative Reasoning for Robust Fake News Detection
abstract
Propagation-based fake news detectors primarily extract structural patterns from news propagation trees via graph neural networks (GNNs), which are crucial for trustworthy information access on social platforms. However, these systems remain vulnerable to adversarial message injection, increasingly enabled by large language models (LLMs). Such attacks pollute both semantic and structural signals, causing GNN-based aggregators to fuse logically conflicting content and yield unreliable representations. To address this, we propose LLM-TKT, a novel framework that distills LLM-based narrative reasoning into lightweight GNNs for robust fake news detection. The framework operates in two stages. First, we construct an offline LLM-driven narrative hub to synthesize global propagation narratives and diagnose local node-level coherence. Second, we design a dual-level narrative alignment to learn the semantic invariance of propagation with the guidance of propagation narratives. It filters unreliable neighbor nodes via local consistency and optimizes graph representations via global anchoring. Experiments on three real-world datasets demonstrate that LLM-TKT significantly outperforms existing methods, particularly in defending against sophisticated LLM-driven injection attacks without incurring runtime LLM inference costs.
Mengyang Chen, Lingwei Wei, Wei Zhou 0019, Songlin Hu 0001
SIGIR4
2024 Leveraging Evolution Patterns to Enhance Script Event Prediction by Large Language Models
Shuchong Wei, Liangjun Zang, Songlin Hu 0001
DASFAA (5)5
2024 Pyramidal Cross-Modal Transformer with Sustained Visual Guidance for Multi-Label Image Classification
abstract
Multi-label image classification poses a formidable challenge due to the presence of multiple objects in each image, rendering it notably complex to decipher the visual content comprehensively. Discriminating between multiple objects necessitates the establishment of robust visual label dependencies. Previous methods attempt to formulate cross-modal interaction or one-shot co-occurrence relationship guidance. However, it not only exhibits limitations when handling occluded or blurry objects but also fails to fully leverage the diverse hierarchical properties for sustainably guiding the learning process of label dependencies. To sustainably establish hierarchical visual label dependencies, this paper introduces a Pyramidal Cross-modal Transformer framework for MLIC tasks. Specifically, the pyramidal visual guidance layer parses the visual features into a multi-resolution pyramid structure, allowing the updated visual-related information to provide sustained guidance for label semantics. This surpasses the conventional pre-processing of co-occurrence relationships. Besides, the hybrid modal interaction layer is proposed to effectively mitigate the semantic disparities between visual and label information with modal-blended indiscriminate attention, replacing vanilla self-attention. Several combination blocks consisting of these two layers are integrated and embedded within the encoder-decoder structure to facilitate the exploration of meticulous visual label dependencies. Extensive experiments on two widely-used benchmarks, including MS-COCO and PASCAL VOC 2007, consistently demonstrate that PCMT could provide state-of-the-art results.
Ruyun Wang, Fuqing Zhu, Jizhong Han, Songlin Hu 0001
ICMR5
2024 Propagation Structure-Semantic Transfer Learning for Robust Fake News Detection
Mengyang Chen, Lingwei Wei, Wei Zhou 0019, Zhou Yan, Songlin Hu 0001
ECML/PKDD (7)6
2024 Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval
abstract
Masked auto-encoder pre-training has emerged as a prevalent technique for initializing and enhancing dense retrieval systems. It generally utilizes additional Transformer decoder blocks to provide sustainable supervision signals and compress contextual information into dense representations. However, the underlying reasons for the effectiveness of such a pre-training technique remain unclear. The usage of additional Transformer-based decoders also incurs significant computational costs. In this study, we aim to shed light on this issue by revealing that masked auto-encoder (MAE) pre-training with enhanced decoding significantly improves the term coverage of input tokens in dense representations, compared to vanilla BERT checkpoints. Building upon this observation, we propose a modification to the traditional MAE by replacing the decoder of a masked auto-encoder with a completely simplified Bag-of-Word prediction task. This modification enables the efficient compression of lexical signals into dense representations through unsupervised pre-training. Remarkably, our proposed method achieves state-of-the-art retrieval performance on several large-scale retrieval benchmarks without requiring any additional parameters, which provides a 67% training speed-up compared to standard masked auto-encoder pre-training with enhanced decoding.
Guangyuan Ma, Xing Wu 0002, Zijia Lin, Songlin Hu 0001
SIGIR4
2023 Improving Event Representation with Supervision from Available Semantic Resources
Shuchong Wei, Liangjun Zang, Songlin Hu 0001
DASFAA (3)5
2023 Orthrus: A Dual-Branch Model for Time Series Forecasting with Multiple Exogenous Series
Ziang Yang, Biyu Zhou, Xuehai Tang, Ruixuan Li 0001, Songlin Hu 0001
DASFAA (1)5
2023 ContE: contextualized knowledge graph embedding for circular relations
Shangwen Lv, Fuqing Zhu, Longtao Huang, Songlin Hu 0001
Data Min. Knowl. Discov.6
2023 KR-GCN: Knowledge-Aware Reasoning with Graph Convolution Network for Explainable Recommendation
abstract
Incorporating knowledge graphs (KGs) into recommender systems to provide explainable recommendation has attracted much attention recently. The multi-hop paths in KGs can provide auxiliary facts for improving recommendation performance as well as explainability. However, existing studies may suffer from two major challenges: error propagation and weak explainability. Considering all paths between every user-item pair might involve irrelevant ones, which leads to error propagation of user preferences. Defining meta-paths might alleviate the error propagation, but the recommendation performance would heavily depend on the pre-defined meta-paths. Some recent methods based on graph convolution network (GCN) achieve better recommendation performance, but fail to provide explainability. To tackle the above problems, we propose a novel method named K nowledge-aware R easoning with G raph C onvolution N etwork (KR-GCN). Specifically, to alleviate the effect of error propagation, we design a transition-based method to determine the triple-level scores and utilize nucleus sampling to select triples within the paths between every user-item pair adaptively. To improve the recommendation performance and guarantee the diversity of explanations, user-item interactions and knowledge graphs are integrated into a heterogeneous graph, which is performed with the graph convolution network. A path-level self-attention mechanism is adopted to discriminate the contributions of different selected paths and predict the interaction probability, which improves the relevance of the final explanation. Extensive experiments conducted on three real-world datasets show that KR-GCN consistently outperforms several state-of-the-art baselines. And human evaluation proves the superiority of KR-GCN on explainability.
Longtao Huang, Qianqian Lu, Songlin Hu 0001
ACM Trans. Inf. Syst.4
2022 CLZT: A Contrastive Learning Based Framework for Zero-Shot Text Classification
Songlin Hu 0001, Ruixuan Li 0001
DASFAA (2)3
2022 Cascade-Enhanced Graph Convolutional Network for Information Diffusion Prediction
Lingwei Wei, Chunyuan Yuan, Yinan Bao, Wei Zhou 0019, Xian Zhu, Songlin Hu 0001
DASFAA (1)7
2021 Entity and Relation Matching Consensus for Entity Alignment
abstract
Entity alignment aims to match synonymous entities across different knowledge graphs, which is a fundamental task for knowledge integration. Recently, researchers have devoted to leveraging rich information within relations to enhance entity alignment. They explicitly incorporate relations in entity representation and alignment, demonstrating remarkable results. However, affected by the semantic assumptions from early works, these works represent a relation by combining all the entities it connects, ignoring the semantic independence between entity and relation. Moreover, since these works perform alignment by comparing embedding similarity, they fail to consider a graph level alignment and tend to find local false correspondences.
Jinzhu Yang, Wei Zhou 0019, Wanhui Qian, Xin Wang 0086, Jizhong Han, Songlin Hu 0001
CIKM7
2021 PEN4Rec: Preference Evolution Networks for Session-Based Recommendation
Dou Hu 0001, Lingwei Wei, Wei Zhou 0019, Xiaoyong Huai, Zhiqi Fang, Songlin Hu 0001
KSEM6
2020 An Event-Oriented Neural Ranking Model for News Retrieval
abstract
Event-oriented news retrieval (ENR) is the task of retrieving news articles related to the specific event in response to the event-oriented query. Previous approaches usually focus on optimizing traditional retrieval models through hand-crafted features from the perspective of new articles. However, these approaches often fail to work well in reality, as they do not consider the essential natures of the event, i.e., dynamics, coupling. In this paper, we propose a novel and effective event-oriented neural ranking model for news retrieval (ENRMNR). Our model exploits a deep attention mechanism to tackle the dynamics and coupling derived from event evolution. Specifically, the word-level bidirectional attention allows the model to identify which query words about the subevent are related to the news article words, and vice-versa, in order to tackle the dynamics. Moreover, the hierarchical attention at passage-level and document-level allows it to capture fine-grained event representations for the coupling between different events within a news article. Experimental results on real-world datasets demonstrate that ENRMNR model significantly outperforms competitive models.
Wanhui Qian, Liangjun Zang, Fuqing Zhu, Ruixuan Li 0001, Jizhong Han, Songlin Hu 0001
CIKM8
2020 RE-GCN: Relation Enhanced Graph Convolutional Network for Entity Alignment in Heterogeneous Knowledge Graphs
Jinzhu Yang, Wei Zhou 0019, Lingwei Wei, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001
DASFAA (2)6
2020 AutoSUM: Automating Feature Extraction and Multi-user Preference Simulation for Entity Summarization
Dongjun Wei, Fuqing Zhu, Liangjun Zang, Wei Zhou 0019, Songlin Hu 0001
PAKDD (2)7
2020 Hierarchical Interaction Networks with Rethinking Mechanism for Document-Level Sentiment Analysis
Lingwei Wei, Dou Hu 0001, Wei Zhou 0019, Xuehai Tang, Xiaodan Zhang 0004, Xin Wang 0086, Jizhong Han, Songlin Hu 0001
ECML/PKDD (3)8
2020 DyHGCN: A Dynamic Heterogeneous Graph Convolutional Network to Learn Users' Dynamic Preferences for Information Diffusion Prediction
Chunyuan Yuan, Wei Zhou 0019, Xiaodan Zhang 0004, Songlin Hu 0001
ECML/PKDD (3)6
2020 Beyond Statistical Relations: Integrating Knowledge Relations into Style Correlations for Multi-Label Music Style Classification
abstract
Automatically labeling multiple styles for every song is a comprehensive application in all kinds of music websites. Recently, some researches explore review-driven multi-label music style classification and exploit style correlations for this task. However, their methods focus on mining the statistical relations between different music styles and only consider shallow style relations. Moreover, these statistical relations suffer from the underfitting problem because some music styles have little training data. To tackle these problems, we propose a novel knowledge relations integrated framework (KRF) to capture the complete style correlations, which jointly exploits the inherent relations between music styles according to external knowledge and their statistical relations. Based on the two types of relations, we use graph convolutional network to learn the deep correlations between styles automatically. Experimental results show that our framework significantly outperforms the state-of-the-art methods. Further studies demonstrate that our framework can effectively alleviate the underfitting problem and learn meaningful style correlations.
Qianwen Ma, Chunyuan Yuan, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001
WSDM5
2019 A Time-Series Sockpuppet Detection Method for Dynamic Social Relationships
Wei Zhou 0019, Jingli Wang, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001
DASFAA (1)6
2019 Jointly Embedding the Local and Global Relations of Heterogeneous Graph for Rumor Detection
abstract
The development of social media has revolutionized the way people communicate, share information and make decisions, but it also provides an ideal platform for publishing and spreading rumors. Existing rumor detection methods focus on finding clues from text content, user profiles, and propagation patterns. However, the local semantic relation and global structural information in the message propagation graph have not been well utilized by previous works. In this paper, we present a novel global-local attention network (GLAN) for rumor detection, which jointly encodes the local semantic and global structural information. We first generate a better integrated representation for each source tweet by fusing the semantic information of related retweets with the attention mechanism. Then, we model the global relationships among all source tweets, retweets, and users as a heterogeneous graph to capture the rich structural information for rumor detection. We conduct experiments on three real-world datasets, and the results demonstrate that GLAN significantly outperforms the state-of-the-art models in both rumor detection and early detection scenarios.
Chunyuan Yuan, Qianwen Ma, Wei Zhou 0019, Jizhong Han, Songlin Hu 0001
ICDM5
2019 Learning Review Representations from user and Product Level Information for Spam Detection
abstract
Opinion spam has become a widespread problem in social media, where hired spammers write deceptive reviews to promote or demote products to mislead the consumers for profit or fame. Existing works mainly focus on manually designing discrete textual or behavior features, which cannot capture complex global semantics of reviews. Although recent works apply deep learning methods to learn review-level semantic features, their models ignore the impact of the user-level and product-level information on learning review semantics and the inherent user-review-product relationship information. In this paper, we propose a Hierarchical Fusion Attention Network (HFAN) to automatically learn the semantics of reviews from user and product level. Specifically, we design a multiattention unit to extract user(product)-related review information. Then, we use orthogonal decomposition and fusion attention to learn a user, review, and product representation from the review information. Finally, we take the review as a relation between user and product entity and apply TransH to jointly encode this relationship into review representation. Experimental results obtained more than 10% absolute precision improvement over the state-of-the-art performances on four real-world datasets, which show the effectiveness and versatility of the model.
Chunyuan Yuan, Wei Zhou 0019, Qianwen Ma, Shangwen Lv, Jizhong Han, Songlin Hu 0001
ICDM6
2019 A Multimodal Text Matching Model for Obfuscated Language Identification in Adversarial Communication?
abstract
Obfuscated language is created to avoid censorship in adversarial communication such as sensitive information conveying, strong sentiment expression, secret actions plan, and illegal trading. The obfuscated sentences are usually generated by replacing one word with another to conceal the textual content. Intelligence and security agencies identify such adversarial messages by scanning with a watch-list of red-flagged terms. Though semantic expansion techniques are adopted, the precision and recall of the identification is limited due to the ambiguity and the unbounded creation way. To this end, this paper frames the obfuscated language identification problem as a text matching task, where each message is checked whether matches a red-flagged term. We propose a multimodal text matching model which combining textual and visual features. The proposed model extends a Bi-directional Long Short Term Memory network with a visual-level representation component to achieve the given task. Comparative experiments on real-world dataset demonstrate that the proposed method could achieve a better performance than the previous methods.
Longtao Huang, Junyu Lin 0002, Jizhong Han, Songlin Hu 0001
WWW5
2017 KIEM: A Knowledge Graph based Method to Identify Entity Morphs
abstract
An entity on the web can be referred by numerous morphs that are always ambiguous, implicit and informal, which makes it challenging to accurately identify all the morphs corresponding to a specific entity. In this paper, we introduce a novel method based on knowledge graph, which takes advantage of both knowledge reasoning and statistic learning. First, we present a model to build a knowledge graph for the given entity. The knowledge graph integrates the fragmented knowledge on how humans create morphs. Then, the candidate morphs are generated based on the rules summarized from the knowledge graph. At last, we use a classification method to filter the useless candidates and identify the target morphs. The experiments conducted on real world dataset demonstrate efficiency of our proposed method in terms of precision and recall.
Longtao Huang, Shangwen Lv, Fangzhou Lu, Yue Zhai, Songlin Hu 0001
CIKM6
2015 DualTable: A hybrid storage model for update optimization in Hive
abstract
Hive is the most mature and prevalent data warehouse tool providing SQL-like interface in the Hadoop ecosystem. It is successfully used in many Internet companies and shows its value for big data processing in traditional industries. However, enterprise big data processing systems as in Smart Grid applications usually require complicated business logics and involve many data manipulation operations like updates and deletes. Hive cannot offer sufficient support for these while preserving high query performance. Hive using the Hadoop Distributed File System (HDFS) for storage cannot implement data manipulation efficiently and Hive on HBase suffers from poor query performance even though it can support faster data manipulation. There is a project based on Hive issue Hive-5317 to support update operations, but it has not been finished in Hive's latest version. Since this ACID compliant extension adopts same data storage format on HDFS, the update performance problem is not solved. In this paper, we propose a hybrid storage model called DualTable, which combines the efficient streaming reads of HDFS and the random write capability of HBase. Hive on DualTable provides better data manipulation support and preserves query performance at the same time. Experiments on a TPC-H data set and on a real smart grid data set show that Hive on DualTable is up to 10 times faster than Hive when executing update and delete operations.
Songlin Hu 0001, Wantao Liu, Tilmann Rabl, Hans-Arno Jacobsen, Xubin Pei, Jiye Wang
ICDE1
2015 QMapper for Smart Grid: Migrating SQL-based Application to Hive
abstract
Apache Hive has been widely used by Internet companies for big data analytics applications. It can provide the capability of compiling high-level languages into efficient MapReduce workflows, which frees users from complicated and time consuming programming. The popularity of Hive and its HiveQL-compatible systems like Impala and Shark attracts attentions from traditional enterprises as well. However, enterprise big data processing systems such as Smart Grid applications often have to migrate their RDBMS-based legacy applications to Hive rather than directly writing new logic in HiveQL. Considering their differences in syntax and cost model, manual translation from SQL in RDBMS to HiveQL is very difficult, error-prone, and often leads to poor performance.
Yingzhong Xu, Yue Liu 0006, Songlin Hu 0001
SIGMOD Conference5
2014 DGFIndex for Smart Grid: Enhancing Hive with a Cost-Effective Multidimensional Range Index
abstract
In Smart Grid applications, as the number of deployed electric smart meters increases, massive amounts of valuable meter data is generated and collected every day. To enable reliable data collection and make business decisions fast, high throughput storage and high-performance analysis of massive meter data become crucial for grid companies. Considering the advantage of high efficiency, fault tolerance, and price-performance of Hadoop and Hive systems, they are frequently deployed as underlying platform for big data processing. However, in real business use cases, these data analysis applications typically involve multidimensional range queries (MDRQ) as well as batch reading and statistics on the meter data. While Hive is high-performance at complex data batch reading and analysis, it lacks efficient indexing techniques for MDRQ. In this paper, we propose DGFIndex, an index structure for Hive that efficiently supports MDRQ for massive meter data. DGFIndex divides the data space into cubes using the grid file technique. Unlike the existing indexes in Hive, which stores all combinations of multiple dimensions, DGFIndex only stores the information of cubes. This leads to smaller index size and faster query processing. Furthermore, with pre-computing user-defined aggregations of each cube, DGFIndex only needs to access the boundary region for aggregation query. Our comprehensive experiments show that DGFIndex can save significant disk space in comparison with the existing indexes in Hive and the query performance with DGFIndex is 2-50 times faster than existing indexes in Hive and HadoopDB for aggregation query, 2-5 times faster than both for non-aggregation query, 2-75 times faster than scanning the whole table in different query selectivity.
Yue Liu 0006, Songlin Hu 0001, Tilmann Rabl, Wantao Liu, Hans-Arno Jacobsen, Kaifeng Wu, Jintao Li 0001
Proc. VLDB Endow.2