Menglong Lu

dblp:228/1517 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Tree-based Publish-Subscribe Model and Load Partition for Distributed Object System
abstract
Object-oriented systems are commonly built by composing local components, while the composition relation is independent of their physical deployment. Existing message-queue-based middleware usually adopts a flat publish-subscribe model, making objects potentially reachable from one another and causing redundant cross-node communication. This paper proposes a communication-aware partitioning method based on a hierarchical publish-subscribe architecture. The system is modeled as a hierarchical communication tree, in which communication load is quantified through message propagation paths. Deployment mappings, cut-edge variables, and capacity constraints are encoded as an optimization model based on satisfiability modulo theories (SMT). Because directly solving the complete SMT model incurs high overhead in large-scale scenarios, we further design a bottom-up subtree partitioning algorithm as a highly scalable approximate solution. Experimental results show that, compared with the baseline algorithms, the proposed method reduces the weighted cross-node communication cost and provides an efficient near-optimal approximate partitioning scheme for large-scale hierarchical workloads within an acceptable accuracy range. It is suitable for distributed systems with hierarchical characteristics, such as EDA simulation.
Lujia Yin, Zhongxiang Dai, Xiufen Fu, Menglong Lu, Chuan Ai
SIGCOMM5
2026 LLM-Driven Adversarial Example Synthesis for Emerging Topic Rumor Detection on Social Media
abstract
Rumor detection is essential for building a responsible web and internet ecosystem, which has attracted significant attention from the research community. However,emerging topic rumor detection, i.e., identify rumors at the early stages of a topic's emergence where only limited discussions can be observed, still remains a challenge. Technically, this scenario is accompanied by the issues ofdata scarcityon emerging topics and thedata distribution discrepancybetween old topics and emerging new topic. In this paper, we propose a new framework termedLLM-drivenADversarialExampleSynthesis (LADES) for emerging topic rumor detection. LADES utilizes Large Language Models (LLMs) for generating readable and contextually coherent adversarial examples. The generated adversarial examples not only expand the training set to tackle the data scarcity issue, but also act as a bridge to connect the data distribution of old and new topics. To overcome training instability in adversarial example generation, LADES introduces a gradient-free Markov Chain Monte Carlo (MCMC) sampling method. This method ensures adversarial examples are readable and contextually coherent by harnessing LLMs, while promoting effective attacks through entropy-based sampling that targets model uncertainty. To mitigate the impact of potential mislabeling in synthetic data, LADES implements a meta-mixed-learning mechanism. This mechanism dynamically adjusts the weights of synthetic adversarial examples, guided by limited labeled data from emerging topics, thereby alleviating the data noise.
Menglong Lu, Zejiang He, Yaohui Guo, Zhiliang Tian, Chengcheng Shao, Dongsheng Li 0001, Zhen Huang 0006
IEEE Trans. Knowl. Data Eng.1
2025 LLM-based Rumor Detection via Influence Guided Sample Selection and Game-based Perspective Analysis
abstract
Zhiliang Tian, Jingyuan Huang, Zejiang He, Zhen Huang, Menglong Lu, Linbo Qiao, Songzhu Mei, Yijie Wang, Dongsheng Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhiliang Tian, Zejiang He, Zhen Huang 0006, Menglong Lu, Linbo Qiao, Songzhu Mei, Yijie Wang 0001
ACL (1)5
2025 GCML: Gradient Coherence Guided Meta-Learning for Cross-Domain Emerging Topic Rumor Detection
abstract
With the emergence of new topics on social media as sources of rumor propagation, addressing the domain shift between the source and target domain and the target domain samples scarcity remains a crucial task in cross-domain rumor detection.Traditional deep learningbased methods and LLM-based methods are mostly focused on the in-domain condition, thus having poor performance in cross-domain setting.Existing domain adaptation rumor detection approaches ignore the data generalization differences and rely on a large amount of unlabeled target domain samples to achieve domain adaptation, resulting in less effective on emerging topic rumor detection.In this paper, we propose a Gradient Coherence guided Meta-Learning approach (GCML) for emerging topics rumor detection.Firstly, we calculate the task generalization score of each source task (sampled from source domain) from a gradient coherence perspective, and selectively learn more "generalizable" tasks that are more beneficial in adapting to the target domain.Secondly, we leverage meta-learning to alleviate the target domain samples scarcity, which utilizes task generalization scores to re-weight metatest gradients and adaptively updates learning rate.Extensive experimental results on realworld datasets show that our method substantially outperforms SOTA baselines.
Zejiang He, Menglong Lu, Zhen Huang 0006, Zhiliang Tian, Dongsheng Li 0001
EMNLP3
2025 A Mixture-of-Experts Framework with Fake Review Detection for Robust Recommendation Systems
Yaohui Guo, Menglong Lu, Zhilong Lv, Jinhui Zhao, Zhen Huang 0006, Dongsheng Li 0001
ICIC (7)2
2025 ST-QAT: Leveraging Self-Training to Enhance Quantization-Aware Training
abstract
In recent years, the scale and computational demand for deep neural network models have been continuously increasing, leading to a growing need for efficient model deployment methods. Model quantization technology is an efficient way for model compression. However, low-bit quantization of models may result in diminished model accuracy. Quantization-aware training (QAT) is a representative training method for alleviating the decline of model accuracy. Nevertheless, most QAT methods require a large amount of labeled datasets and lack generalizability for tasks with limited data. We propose a new approach by using self-training, a semi-supervised method, to meet the dataset requirements of QAT. Based on the analysis of the characteristics of self-training and quantization-aware training, we use a meta-learning module for pseudo-label dataset selection and employ quantization operations to introduce noise into the model to enhance training effectiveness. We conducted experimental evaluations of traditional quantization training methods and our method on four network models (ResNet-50, ResNet-101, MobileNet-v2, MobileNet-v3) with the Cifar100 dataset. Our method enhances the training performance of both full-precision and quantized models compared to traditional quantization training methods. Utilizing 30% of the labeled dataset, compared to QAT, our method demonstrates an average accuracy enhancement of 2.19% after model training, while the average accuracy decline during the transition from a full-precision model to an 8-bit model is diminished by 0.55%.
Menglong Lu, Zhilin Wang
IJCNN2
2025 STARTS: Simulation Traits Assisted Random Test Selection for Multiprocessor Verification
abstract
Test selection is vital for accelerating multiprocessor design verification. Current methods focus on the similarities among random test cases but overlook critical runtime characteristics that significantly impact outcomes. This limitation hinders their ability to capture patterns generated at runtime in multiprocessor systems. To address this, we propose Simulation Traits Assisted Random Test Selection (STARTS), which utilizes features obtained from fast model simulation of the target multiprocessor during verification. STARTS rapidly simulates test cases to predict hardware behavior and employs a GRU-based variational autoencoder with a self-attention mechanism to embed test cases, capturing both data and control dependencies between stimulus actions. The distance in latent space serves as a key selection criterion. By integrating hardware-related information with intrinsic multiprocessor characteristics, our method enhances test selection effectiveness. Experimental results show that STARTS reduces the time needed to achieve coverage goals compared to other unsupervised learning methods, demonstrating its superior performance in selecting optimal stimuli.
Li Zhou 0009, Menglong Lu, Junbo Tie
ITC2
2025 Hippocampal-like Sequential Editing for Continual Knowledge Updates in Large Language Models
abstract
Large language models (LLMs) are now pivotal in real-world applications. Model editing has emerged as a promising paradigm for efficiently modifying LLMs without full retraining. However, current editing approaches face significant limitations due to parameter drift, which stems from inconsistencies between newly edited knowledge and the model's existing knowledge. In sequential editing scenarios, cumulative drifts progressively lead to model collapse characterized by general capability degradation and balance between acquiring new knowledge and catastrophic forgetting of existing knowledge. Drawing inspiration from the hippocampal trisynaptic circuit for continual memorizing and forgetting, we propose a Hippocampal-like Sequential Editing (HSE) framework that designs the unlearning of obsolete knowledge, domain-specific knowledge update separation and replay for edited knowledge. Specifically, the HSE framework designs three core mechanisms: (1) Machine unlearning selectively erases outdated knowledge to facilitate integration of new information, (2) Fisher Information Matrix-guided parameter updates prevents cross-domain knowledge interference, and (3) Parameter replay consolidates long-term editing memory through lightweight and global replay of editing data in a parametric form. Theoretical analysis demonstrates that HSE achieves smaller generalization error bounds, more stable convergence and higher computational efficiency. Experimental results validate its effective balance between acquiring new knowledge and mitigating catastrophic forgetting, maintaining or even slightly enhancing general capabilities. In practical applications, experiments confirm its effectiveness in multi-domain hallucination mitigation, healthcare knowledge injecting, and societal bias reduction.
Quntian Fang, Zhen Huang 0006, Zhiliang Tian, Minghao Hu 0001, Dongsheng Li 0001, Yiping Yao, Xinyue Fang, Menglong Lu, Guotong Geng
NeurIPS8
2025 A cause fusion framework with information bottleneck for conversational causal emotion entailment
Xinxin Su, Zhen Huang 0006, Menglong Lu, Sisi Dai, Yong Dou
Neural Networks3
2024 Auto-Prox: Training-Free Vision Transformer Architecture Search via Automatic Proxy Discovery
abstract
The substantial success of Vision Transformer (ViT) in computer vision tasks is largely attributed to the architecture design. This underscores the necessity of efficient architecture search for designing better ViTs automatically. As training-based architecture search methods are computationally intensive, there’s a growing interest in training-free methods that use zero-cost proxies to score ViTs. However, existing training-free approaches require expert knowledge to manually design specific zero-cost proxies. Moreover, these zero-cost proxies exhibit limitations to generalize across diverse domains. In this paper, we introduce Auto-Prox, an automatic proxy discovery framework, to address the problem. First, we build the ViT-Bench-101, which involves different ViT candidates and their actual performance on multiple datasets. Utilizing ViT-Bench-101, we can evaluate zero-cost proxies based on their score-accuracy correlation. Then, we represent zero-cost proxies with computation graphs and organize the zero-cost proxy search space with ViT statistics and primitive operations. To discover generic zero-cost proxies, we propose a joint correlation metric to evolve and mutate different zero-cost proxy candidates. We introduce an elitism-preserve strategy for search efficiency to achieve a better trade-off between exploitation and exploration. Based on the discovered zero-cost proxy, we conduct a ViT architecture search in a training-free manner. Extensive experiments demonstrate that our method generalizes well to different datasets and achieves state-of-the-art results both in ranking correlation and final accuracy. Codes can be found at https://github.com/lilujunai/Auto-Prox-AAAI24.
Zimian Wei, Peijie Dong, Zheng Hui, Anggeng Li, Lujun Li 0001, Menglong Lu, Hengyue Pan, Dongsheng Li 0001
AAAI6
2024 ESC-CoT: Easy-to-Hard Self-Comparative Chain-of-Thought for News Discourse Profiling
abstract
News Discourse Profiling is a discourse task that aims to recognize the semantic role of each sentence in an article. Within this task, the model not only requires comprehending the news content but also needs to analyze the logical relation in a discourse, posing a great challenge to traditional deep learning techniques. Recently, large language model (LLM) technology has demonstrated significant potential for enhancing the task of news discourse profiling. These models, characterized by their vast number of parameters, excel in logical reasoning and contextual understanding, allowing them to grasp the intricate details and relationships within the text. When paired with the effective chain-of-thought (CoT) technique, LLMs have shown impressive performance in many complex reasoning tasks. However, the discourse structure significantly differentiates between news, thus it is difficult to analyze all news with a unified fixed CoT. In this paper, We propose an adaptive CoT technique for the News Discourse Profiling task, namely Easy-to-Hard Self-Comparative Chain-of- Thought (ESC-CoT). ESC-CoT formulates the task as a multiple-iteration process, i.e., ESC-CoT handles the easiest part at the current iteration and gradually transfers to more complex parts, which can reduce the risk of early error propagation. To alleviate error propagation along the reasoning chain, ESC-CoT compares the conflicting results from different CoT processes, analyzing the context and meaning of sentences more deeply, thereby improving LLM's reasoning ability and reducing the possibility of error. Experiment results demonstrate that ESC-CoT substantially surpasses the traditional news discourse profiling methods. Compared with state-of-the-art (SOTA) methods on two benchmark datasets, the ESC-CoT technique shows an improvement of 7.8 and 19.6 on F1 score, respectively.
Zejiang He, Menglong Lu, Zhen Huang 0006, Jinhui Zhao
ICTAI4
2024 MoveFormer: Spatial Graph Periodic Injection Network for Next POI Recommendation
Yongheng Li, Zhen Huang 0006, Tianfu He, Menglong Lu, Zeyun Zhao
KSEM (2)6
2024 Meta Learning Based Rumor Detection with Awareness of Social Bot
Zhilong Lv, Zhen Huang 0006, Menglong Lu, Zhiliang Tian, Xin Niu 0002, Dongsheng Li 0001
KSEM (3)3
2023 DaMSTF: Domain Adversarial Learning Enhanced Meta Self-Training for Domain Adaptation
abstract
Self-training emerges as an important research line on domain adaptation.By taking the model's prediction as the pseudo labels of the unlabeled data, self-training bootstraps the model with pseudo instances in the target domain.However, the prediction errors of pseudo labels (label noise) challenge the performance of self-training.To address this problem, previous approaches only use reliable pseudo instances, i.e., pseudo instances with high prediction confidence, to retrain the model.Although these strategies effectively reduce the label noise, they are prone to miss the hard examples.In this paper, we propose a new self-training framework for domain adaptation, namely Domain adversarial learning enhanced Self-Training Framework (DaMSTF).Firstly, DaMSTF involves meta-learning to estimate the importance of each pseudo instance, so as to simultaneously reduce the label noise and preserve hard examples.Secondly, we design a meta constructor for constructing the meta validation set, which guarantees the effectiveness of the meta-learning module by improving the quality of the meta validation set.Thirdly, we find that the meta-learning module suffers from the training guidance vanishment and tends to converge to an inferior optimal.To this end, we employ domain adversarial learning as a heuristic neural network initialization method, which can help the meta-learning module converge to a better optimal.Theoretically and experimentally, we demonstrate the effectiveness of the proposed DaMSTF.On the cross-domain sentiment classification task, DaMSTF improves the performance of BERT with an average of nearly 4%.
Menglong Lu, Zhen Huang 0006, Zhiliang Tian, Yang Liu 0259, Dongsheng Li 0001
ACL (1)1
2023 DMFormer: Closing the gap Between CNN and Vision Transformers
abstract
Vision transformers have shown excellent performance in computer vision tasks. As the computation cost of their self-attention mechanism is expensive, recent works tried to replace the self-attention mechanism in vision transformers with convolutional operations, which is more efficient with built-in inductive bias. However, these efforts either ignore multi-level features or lack dynamic prosperity, leading to sub-optimal performance. In this paper, we propose a Dynamic Multi-level Attention mechanism (DMA), which captures different patterns of input images by multiple kernel sizes and enables input-adaptive weights with a gating mechanism. Based on DMA, we present an efficient backbone network named DMFormer. DMFormer adopts the overall architecture of vision transformers, while replacing the self-attention mechanism with our proposed DMA. Extensive experimental results on ImageNet-1K and ADE20K datasets demonstrated that DMFormer achieves state-of-the-art performance, which outperforms similar-sized vision transformers(ViTs) and convolutional neural networks (CNNs).
Zimian Wei, Hengyue Pan, Lujun Li 0001, Menglong Lu, Xin Niu 0002, Peijie Dong, Dongsheng Li 0001
ICASSP4
2023 Meta-Tsallis-Entropy Minimization: A New Self-Training Approach for Domain Adaptation on Text Classification
abstract
Text classification is a fundamental task for natural language processing, and adapting text classification models across domains has broad applications. Self-training generates pseudo-examples from the model's predictions and iteratively trains on the pseudo-examples, i.e., minimizes the loss on the source domain and the Gibbs entropy on the target domain. However, Gibbs entropy is sensitive to prediction errors, and thus, self-training tends to fail when the domain shift is large. In this paper, we propose Meta-Tsallis Entropy minimization (MTEM). MTEM uses an instance adaptive Tsallis entropy to replace the Gibbs entropy and a meta-learning algorithm to optimize the instance adaptive Tsallis entropy on the target domain. To reduce the computation cost of MTEM, we propose an approximation technique to approximate the second-order derivation involved in the meta-learning. To efficiently generate pseudo labels, we propose an annealing sampling mechanism for exploring the model's prediction probability. Theoretically, we prove the convergence of the meta-learning algorithm in MTEM and analyze the effectiveness of MTEM in achieving domain adaptation. Experimentally, MTEM improves the adaptation performance of BERT with an average of 4 percent on the benchmark dataset.
Menglong Lu, Zhen Huang 0006, Zhiliang Tian, Xuanyu Fei, Dongsheng Li 0001
IJCAI1
2022 Social Bot-Aware Graph Neural Network for Early Rumor Detection
abstract
Early rumor detection is a key challenging task to prevent rumors from spreading widely. Sociological research shows that social bots’ behavior in the early stage has become the main reason for rumors’ wide spread. However, current models do not explicitly distinguish genuine users from social bots, and their failure in identifying rumors timely. Therefore, this paper aims at early rumor detection by accounting for social bots’ behavior, and presents a Social Bot-Aware Graph Neural Network, named SBAG. SBAG firstly pre-trains a multi-layer perception network to capture social bot features, and then constructs multiple graph neural networks by embedding the features to model the early propagation of posts, which is further used to detect rumors. Extensive experiments on three benchmark datasets show that SBAG achieves significant improvements against the baselines and also identifies rumors within 3 hours while maintaining more than 90% accuracy.
Zhen Huang 0006, Zhilong Lv, Xiaoyun Han, Binyang Li, Menglong Lu, Dongsheng Li 0001
COLING5
2022 SIFTER: A Framework for Robust Rumor Detection
abstract
With the development of online social media, the fabrication and dissemination of rumors are much easier than before. As a result, automatic rumor detection becomes more urgent for internet governors, and has received great interest from the AI community. Even though many initial successes have been achieved in terms of detection accuracy and timeliness, existing rumor detection methods are still not robust due to thefrequent domain shiftproblem and theinconsistent propagationproblem.Frequent domain shiftrefers that rumors’ topic changes frequently on online social media, so the trained model may expire quickly in real applications.Inconsistent propagationis another practical challenge, referring to the fact that the spread of rumors is often accompanied by misleading comments. Therefore, a vulnerable model may output inconsistent predictions for the same instance. In this paper, we propose a new framework to improve the robustness of existing rumor detection methods, named SIFTER (SubjectiveInFormaTionEnhancedReinforcement learning). To address thefrequent domain shift, SIFTER employs multi-task learning to introduce external knowledge that can explicitly describe rumors, and thus makes models more adaptive across domains. In particular, SIFTER involves a multi-task learning module to capture subjective information, which is found to be discriminative in rumors’ propagations. To address theinconsistent propagation, SIFTER employs reinforcement learning to realize a new training schema tailed for rumor detection, namely sequential training, which can reduce the attendance of noisy comments in the feature extraction process. Experimental results on two benchmark datasets, Twitter and Weibo, validate the effectiveness of SIFTER. After migrating existing methods to SIFTER, their accuracy score is improved up to 9.4% in cross-domain inference, and their reverse ratio decreases by nearly 4% in continuous prediction.
Menglong Lu, Zhen Huang 0006, Binyang Li, Zheng Qin 0002, Dongsheng Li 0001
IEEE ACM Trans. Audio Speech Lang. Process.1
2021 Rumor Verification on Social Media with Stance-Aware Recursive Tree
Xiaoyun Han, Zhen Huang 0006, Menglong Lu, Dongsheng Li 0001, Jinyan Qiu
KSEM3
2019 A Distributed Topic Model for Large-Scale Streaming Text
Yicong Li 0001, Menglong Lu, Dongsheng Li 0001
KSEM (2)3
2019 Correction to: A Distributed Topic Model for Large-Scale Streaming Text
Yicong Li 0001, Menglong Lu, Dongsheng Li 0001
KSEM (2)3
2019 Mini-batch cutting plane method for regularized risk minimization
abstract
Although concern has been recently expressed with regard to the solution to the non-convex problem, convex optimization is still important in machine learning, especially when the situation requires an interpretable model. Solution to the convex problem is a global minimum, and the final model can be explained mathematically. Typically, the convex problem is re-casted as a regularized risk minimization problem to prevent overfitting. The cutting plane method (CPM) is one of the best solvers for the convex problem, irrespective of whether the objective function is differentiable or not. However, CPM and its variants fail to adequately address large-scale dataintensive cases because these algorithms access the entire dataset in each iteration, which substantially increases the computational burden and memory cost. To alleviate this problem, we propose a novel algorithm named the mini-batch cutting plane method (MBCPM), which iterates with estimated cutting planes calculated on a small batch of sampled data and is capable of handling large-scale problems. Furthermore, the proposed MBCPM adopts a “sink” operation that detects and adjusts noisy estimations to guarantee convergence. Numerical experiments on extensive real-world datasets demonstrate the effectiveness of MBCPM, which is superior to the bundle methods for regularized risk minimization as well as popular stochastic gradient descent methods in terms of convergence speed.
Menglong Lu, Linbo Qiao, Dongsheng Li 0001, Xicheng Lu
Frontiers Inf. Technol. Electron. Eng.1
2018 Asynchronous Bundle Method for Large-Scale Regularized Risk Minimization
abstract
Bundle method for regularized risk minimization (BMRM) is a variant of Cutting Plane Method (CPM). It performs efficiently in solving a convex minimization problem, which is a core part in a plethora of machine learning applications. Nonetheless, while exposed to the challenge of large-scale learning, the synchronous parallel implementation of BMRM easily encounters the straggler problem due to the diversity among heterogeneous working nodes' capability and unevenness in the inherent data distribution. In this paper, we propose a novel asynchronous distributed BMRM implementation, which employs an asynchronous computing window to fully explore the fast nodes' computational capabilities while reserving the good convergence of the BMRM. Extensive experiments show that the asynchronous BMRM algorithm has significant improvement of performance over its synchronous counterpart, and owns the ability to solve large-scale problems efficiently.
Menglong Lu, Linbo Qiao, Dawen Ding, Dongsheng Li 0001
IJCNN1