EDBT 2026 Demo / reviewers in the wild / expert
Huailiang Peng
dblp:157/8982
· DBLP profile ↗
23ranked-venue papers
4as first author
16since 2021 · last 2026
0009-0000-4943-8853ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graphs for Logic and Texts for Context: A Multi-Agent Orchestrated Hybrid RAG with Stepwise Question DecompositionabstractMulti-hop question answering is a fundamental challenge for RAG, as it requires stepwise reasoning and progressive integration of evidence from multiple sources. Most existing approaches decompose complex questions into multiple sub-questions at once and perform retrieval in parallel, which neglects dependencies among sub-questions and frequently results in broken reasoning chains. Furthermore, the use of a single retrieval source—either unstructured text or structured knowledge graphs—cannot adequately support the heterogeneous information requirements of different reasoning steps. In this work, we propose a multi-agent iterative RAG framework that enables dependency-aware reasoning through evidence-driven sub-question generation. A task-planning agent incrementally generates single-hop sub-questions by identifying evidence gaps from an evolving evidence set, ensuring that each sub-question is well-posed and contextually grounded. Meanwhile, the framework dynamically selects between Graph-RAG Agent and Text-RAG Agent at the sub-question level, exploiting their complementary strengths in relational reasoning and factual coverage. Extensive experiments on HotpotQA, 2WikiMultihopQA, and MuSiQue show that our approach consistently outperforms strong baselines, achieving state-of-the-art accuracy while maintaining greater robustness and stability as reasoning depth increases. Huailiang Peng, Qiong Dai |
ICMR | 2 |
| 2025 | RelationalCoder: Rethinking Complex Tables via Programmatic Relational TransformationabstractSemi-structured tables, with their varied layouts and formatting artifacts, remain a major obstacle for automated data processing and analytics.To address these challenges, we propose RELATIONALCODER, which uniformly converts semi-structured tables into relational data, enabling smooth integration with the rich ecosystem of data processing and analytics tools.By leveraging SQL code, RELATIONAL-CODER prevents schema errors and markedly improves normalization quality across multiple relational tables. Haoyu Dong 0001, Yue Hu 0002, Huailiang Peng, Yanan Cao 0001 |
ACL (1) | 3 |
| 2025 | Adaptive Social Bot Detection through Bridging the Feature Bias Between Source and Target Users
Huailiang Peng, Yanan Cao 0001, Qiong Dai |
ICMR | 2 |
| 2025 | LayerNavigator: Finding Promising Intervention Layers for Efficient Activation Steering in Large Language ModelsabstractActivation steering is an efficient technique for aligning the behavior of large language models (LLMs) by injecting steering vectors directly into a model’s residual stream during inference.
A pivotal challenge in this approach lies in choosing the right layers to intervene, as inappropriate selection can undermine behavioral alignment and even impair the model’s language fluency and other core capabilities.
While single-layer steering allows straightforward evaluation on held-out data to identify the "best" layer, it offers only limited alignment improvements.
Multi-layer steering promises stronger control but faces a combinatorial explosion of possible layer subsets, making exhaustive search impractical.
To address these challenges, we propose LayerNavigator, which provides a principled and promising layer selection strategy.
The core innovation of LayerNavigator lies in its novel, quantifiable criterion that evaluates each layer's steerability by jointly considering two key aspects: discriminability and consistency.
By reusing the activations computed during steering vector generation, LayerNavigator requires no extra data and adds negligible overhead.
Comprehensive experiments show that LayerNavigator achieves not only superior alignment but also greater scalability and interpretability compared to existing strategies.
Our code is available at https://github.com/Bryson-Arrot/LayerNavigator Huailiang Peng, Qiong Dai, Yanan Cao 0001 |
NeurIPS | 2 |
| 2025 | Reinforced GNNs for Multiple Instance LearningabstractMultiple instance learning (MIL) trains models from bags of instances, where each bag contains multiple instances, and only bag-level labels are available for supervision. The application of graph neural networks (GNNs) in capturing intrabag topology effectively improves MIL. Existing GNNs usually require filtering low-confidence edges among instances and adapting graph neural architectures to new bag structures. However, such asynchronous adjustments to structure and architecture are tedious and ignore their correlations. To tackle these issues, we propose a reinforced GNN framework for MIL (RGMIL), pioneering the exploitation of multiagent deep reinforcement learning (MADRL) in MIL tasks. MADRL enables the flexible definition or extension of factors that influence bag graphs or GNNs and provides synchronous control over them. Moreover, MADRL explores structure-to-architecture correlations while automating adjustments. Experimental results on multiple MIL datasets demonstrate that RGMIL achieves the best performance with excellent explainability. The code and data are available at https://github.com/RingBDStack/RGMIL. Xusheng Zhao, Qiong Dai, Jia Wu 0001, Hao Peng 0001, Huailiang Peng, Zhengtao Yu 0001, Philip S. Yu |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Coarse-to-fine label propagation with hybrid representation for deep semi-supervised bot detection
Huailiang Peng, Yujun Zhang 0001, Qiong Dai |
Wirel. Networks | 1 |
| 2024 | Interest-Aware Social Bot Detection with Contrastive Hard Sample MiningabstractSocial bots frequently engage in malicious activities like spreading misinformation and phishing on major social media platforms, significantly impacting the fairness and security of these platforms. Therefore, detecting social bots has become a very critical task. However, we observe two challenges for bot detection methods: neglected discrepancies under various interests (e.g., politics, entertainment) and challenging cases in the real world (e.g., carefully camouflaged bots, individualized genuine users). To tackle these issues, we propose BotCHMIA, a novel interest-aware social bot detection method enhanced with challenging cases. Specifically, to enhance feature representations by various user interests, we propose an interest-aware feature collaboration that utilizes a series of expert networks and an interest adapter to acquire user interest-specific information and fuse it with task-specific feature representations extracted by a bot detection projection. Additionally, we estimate sample hardness during the training process based on the model’s classification confidence and improve existing supervised contrastive loss with randomly selected challenging cases, namely hard samples, to enhance the discriminability of user feature representations. We conduct extensive experiments on two real social bot datasets, and the results demonstrate the practical benefits gained from our proposed detection method. Huailiang Peng, Yujun Zhang 0001, Qiong Dai |
HPCC | 1 |
| 2024 | RDGCN: Reinforced Dependency Graph Convolutional Network for Aspect-based Sentiment AnalysisabstractAspect-based sentiment analysis (ABSA) is dedicated to forecasting the sentiment polarity of aspect terms within sentences. Employing graph neural networks to capture structural patterns from syntactic dependency parsing has been confirmed as an effective approach for boosting ABSA. In most works, the topology of dependency trees or dependency-based attention coefficients is often loosely regarded as edges between aspects and opinions, which can result in insufficient and ambiguous syntactic utilization. To address these problems, we propose a new reinforced dependency graph convolutional network (RDGCN) that improves the importance calculation of dependencies in both distance and type views. Initially, we propose an importance calculation criterion for the minimum distances over dependency trees. Under the criterion, we design a distance-importance function that leverages reinforcement learning for weight distribution search and dissimilarity control. Since dependency types often do not have explicit syntax like tree distances, we use global attention and mask mechanisms to design type-importance functions. Finally, we merge these weights and implement feature aggregation and classification. Comprehensive experiments show the effectiveness of the criterion and importance functions. RDGCN yields excellent analysis results. Xusheng Zhao, Hao Peng 0001, Qiong Dai, Huailiang Peng, Yanbing Liu 0007, Qinglang Guo, Philip S. Yu |
WSDM | 5 |
| 2023 | Multi-omics Sampling-based Graph Transformer for Synthetic Lethality PredictionabstractSynthetic lethality (SL) prediction is used to identify if the co-mutation of two genes results in cell death. The prevalent strategy is to abstract SL prediction as an edge classification task on gene nodes within SL data and achieve it through graph neural networks (GNNs). However, GNNs suffer from limitations in their message passing mechanisms, including over-smoothing and over-squashing issues. Moreover, harnessing the information of non-SL gene relationships within large-scale multi-omics data to facilitate SL prediction poses a non-trivial challenge. To tackle these issues, we propose a new multi-omics sampling-based graph transformer for SL prediction (MSGT-SL). Concretely, we introduce a shallow multi-view GNN to acquire local structural patterns from both SL and multi-omics data. Further, we input gene features that encode multi-view information into the standard self-attention to capture long-range dependencies. Notably, starting with batch genes from SL data, we adopt parallel random walk sampling across multiple omics gene graphs encompassing them. Such sampling effectively and modestly incorporates genes from omics in a structure-aware manner before using self-attention. We showcase the effectiveness of MSGT-SL on real-world SL tasks, demonstrating the empirical benefits gained from the graph transformer and multi-omics data. Xusheng Zhao, Hao Liu 0007, Qiong Dai, Hao Peng 0001, Huailiang Peng |
BIBM | 6 |
| 2023 | Learning Discriminative Text Representation for Streaming Social Event DetectionabstractEvent detection on social platforms can help people perceive essential events and make actionable decisions. Existing document-pivot streaming social event detection methods generally embed documents and perform text clustering. They face the challenges of constantly changing context and unknown event categories and struggle by designing compound text representation methods and various similarity measures. However, phased, well-designed methods are excessively fragile and unable to utilize the potential of text representations fully. Meanwhile, their complex threshold settings result in clustering-based event detection suffering the pain of ever-changing environments. We propose a text representation learning method namely Text Similarity Contrastive Learning Neural Network (Text-SimCLNN) to tackle these challenges. Text-SimCLNN uses smaller parts to learn the similarity probability of text pairs from semantic and structural perspectives, effectively bridging the gap between text representation learning and similarity measure in streaming event detection. Event discovery and merging in streams can be easily performed based on the learned representations, and we use various techniques to speed up such processes. Furthermore, we introduce an online update mechanism that uses heterogeneous graphs to generate high-quality samples to enable stable and reliable inductive learning. Extensive experiments on two real-world datasets demonstrate that our method far exceeds state-of-the-art (SOTA). Chaodong Tong, Huailiang Peng, Qiong Dai, Ruitong Zhang 0001, Hanjie Xu, Xian-Ming Gu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Domain-Aware Federated Social Bot Detection with Multi-Relational Graph Neural NetworksabstractSocial networks have been the widespread popular tools for communication and socialization, and it also been the ideal platform for bots to publish malicious information. Therefore, social bot detection is essential for the social network's security. Existing methods almost ignore the differences in bot behaviors in multiple domains. Thus, we first propose a DomainAware detection method with Multi-Relational Graph neural networks (DA-MRG) to improve detection performance. Specifically, DA-MRG constructs multi-relational graphs with users' features and relationships, obtains the user presentations with graph embedding and distinguishes bots from humans with domainaware classifiers. Meanwhile, considering the similarity between bot behaviors in different social networks, we believe that sharing data among them could boost detection performance. However, the data privacy of users needs to be strictly protected. To overcome the problem, we implement a study of federated learning framework for DA-MRG to achieve data sharing between different social networks and protect data privacy simultaneously. We conduct extensive experiments on TwiBot-20, and the results demonstrate that the proposed method can effectively achieve federated social bot detection. Huailiang Peng, Yujun Zhang 0001, Shuhai Wang |
IJCNN | 1 |
| 2022 | Aspect Is Not You Need: No-aspect Differential Sentiment Framework for Aspect-based Sentiment AnalysisabstractJiahao Cao, Rui Liu, Huailiang Peng, Lei Jiang, Xu Bai. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jiahao Cao 0002, Rui Liu 0032, Huailiang Peng, Lei Jiang 0003 |
NAACL-HLT | 3 |
| 2021 | Multi-order Proximity Graph Structure Embedding
Lei Jiang 0003, Huailiang Peng, Qiong Dai |
CollaborateCom (2) | 3 |
| 2021 | Cross-Network Community Sensing for Anchor Link PredictionabstractAnchor link prediction focuses on finding accounts related to the same natural person in different online platforms and can benefit several services like user modeling and cross-network recommendation systems. Recently, community structure information has attracted interest from researchers, while most of existing methods only treat community structure information as additional information and lack utilizing the interaction between communities. To address these limitations, we propose a cross-network community sensing model for anchor link prediction (CCALP). Our CCALP model regards communities as special nodes in each network and further models community level inter-network relationships by existing anchor links. We further utilize an attention mechanism to integrate cross-network community neighbors' vectors. This attention mechanism can help us balance the degree of utilization of community knowledge. Through extensive experiments on two real-world datasets, we demonstrate that CCALP outperforms the existing baseline approaches. Lintao Lan, Huailiang Peng, Chaodong Tong, Qiong Dai |
IJCNN | 2 |
| 2021 | Dual Adversarial Network Based on BERT for Cross-domain Sentiment Classification
Shaokang Zhang, Lei Jiang 0003, Huailiang Peng |
NLPCC (1) | 4 |
| 2021 | Discriminative Representation Learning for Cross-Domain Sentiment Classification
Shaokang Zhang, Lei Jiang 0003, Huailiang Peng, Qiong Dai, Jianlong Tan |
PAKDD (2) | 3 |
| 2020 | HEAM: Heterogeneous Network Embedding with Automatic Meta-path Construction
Ruicong Shi, Huailiang Peng, Lei Jiang 0003, Qiong Dai |
KSEM (1) | 3 |
| 2020 | Category-Level Adversarial Network for Cross-Domain Sentiment Classification
Shaokang Zhang, Huailiang Peng, Yanan Cao 0001, Lei Jiang 0003, Qiong Dai, Jianlong Tan |
KSEM (2) | 2 |
| 2020 | An Interactive Two-Pass Decoding Network for Joint Intent Detection and Slot Filling
Huailiang Peng, Mengjun Shen, Lei Jiang 0003, Qiong Dai, Jianlong Tan |
NLPCC (2) | 1 |
| 2019 | Chinese Social Media Entity Linking Based on Effective Context with Topic SemanticsabstractOn social media, entity linking is very important for natural language processing tasks, such as Sentiment Analysis, Question Answering (QA) and Machine Translation. Compared to English-oriented entity linking, Chinese entity linking has its special difficulties. Just like the entity linking for short text, Chinese microblogs have lots of noise and the mention lacks effective context information. In order to solve these problems, we present a new model for Chinese microblogs entity linking. Entity linking usually includes two steps: candidate entities generation and candidate entities ranking. First, based on the characteristics of Chinese, we put forward multi-method fusion strategies for candidate generation to improve the recall rate of candidate entities. Second, we propose a new neural network model called TAS (Topic attention Siamese) for candidate entities ranking. In TAS model, we add effective topic semantics on Siamese network to learn representations of context, mention and entity, and rank the mention-entity similarity. The representation of mention incorporates information from multiple sentences on the same topic, which can effectively solve the problem of the lack of contextual information. We also use Character-enhanced Word Embedding model (CWE) to pre-train both word embedding and characters embedding to work out noise and word segmentation impact. Experimental results demonstrate that our method significantly outperforms the state-of-the-art results for entity linking on Chinese social media. Chengfang Ma, Ying Sha, Jianlong Tan, Li Guo 0001, Huailiang Peng |
COMPSAC (1) | 5 |
| 2019 | Improving Natural Language Understanding by Reverse Mapping Bytepair EncodingabstractRecently, language models (LMs) or language representation models are widely used in natural language understanding (NLU) tasks.However, these LMs are usually trained on large unlabeled text corpora, while the finetuning process simply takes words or wordpieces as model input.Because of the differences between language model and NLU task objectives, the problem of lack of concern on some key words exists.Thus in this paper, we propose a method called reverse mapping bytepair encoding, which maps named-entity information and other word-level linguistic features back to subwords during the encoding procedure of bytepair encoding (BPE).We employ this method to the Generative Pre-trained Transformer (OpenAI GPT) (Radford et al., 2018) by adding a weighted linear layer after the embedding layer.We also propose a new model architecture named as the multi-channel separate transformer to evaluate the effectiveness of the newly introduced information by employing a training process without parameter-sharing.Experiments on Story Cloze, RTE, SciTail and SST-2 datasets demonstrate the effectiveness of our approach.Compared with the original results in GPT, our approach gains 1.58% absolute increase on Stories Cloze, 6.4% on RTE, 0.69% on SciTail and 0.8% on SST-2. Chaodong Tong, Huailiang Peng, Qiong Dai, Lei Jiang 0003, Jianghua Huang |
CoNLL | 2 |
| 2019 | Improving Transformer with Sequential Context Representations for Abstractive Text Summarization
Tian Cai, Mengjun Shen, Huailiang Peng, Lei Jiang 0003, Qiong Dai |
NLPCC (1) | 3 |
| 2018 | A High-Performance Round-Robin Regular Expression Matching Architecture Based on FPGAabstractState-of-the-art Network Intrusion Detection Systems (NIDSs) use regular expressions to detect attacks or vulnerabilities. In order to keep up with the ever-increasing speed, more and more NIDSs need to be implemented by dedicated hardware. A major bottleneck is that NIDSs scan incoming packets just byte by byte, which greatly limits their throughput. In this paper, we propose a novel architecture for regular expression (RE) matching that consumes multiple characters per time. This architecture contains all the advantages of three FPGA-based algorithms to improve RE matching speed: Simple State Merge Tree (SSMT), Distribute Data in Round-Robin (DDRR), and Multi-path Speculation. Our architecture was tested on several real-life RE rulesets. It could yield a performance of 140Gbps processing rates on a single FPGA chip, while maintaining memory efficiency. This makes it a very practical solution for NIDS in 100G Ethernet standard network, which is currently the fastest approved standard of Ethernet. The experimental results also show that the throughput is about 108 times better than that of the original DFA, while the memory consumption is only about110of the original DFA. Lei Jiang 0003, Huailiang Peng, Qiong Dai |
ISCC | 4 |