VLDB 2026 Research / reviewers in the wild / expert
Minghua He
dblp:30/1609
· DBLP profile ↗
30ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0003-4439-9810ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 5 first-author · 4 since 2021Software engineering, systems software and programming languages · 8 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Computer networks · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language ModelsabstractLeyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Aiwei Liu, Lijie Wen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Leyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu, Liancheng Fang, Minghua He, Lingzhe Zhang, Zhaoyang Liu 0003, Bolin Ding, Aiwei Liu, Lijie Wen 0001 |
ACL (1) | 6 |
| 2026 | VarParser: Unleashing the Neglected Power of Variables for LLM-based Log ParsingabstractLogs serve as a primary source of information for engineers to diagnose failures in large-scale online service systems. Log parsing, which extracts structured events from massive unstructured log data, is a critical first step for downstream tasks like anomaly detection and failure diagnosis. With advances in large language models (LLMs), leveraging their strong text understanding capabilities has proven effective for accurate log parsing. However, existing LLM-based log parsers all focus on the constant part of logs, ignoring the potential contribution of the variable part to log parsing. This constant-centric strategy brings four key problems. First, inefficient log grouping and sampling with only constant information. Second, a relatively large number of LLM invocations due to constant-based cache, leading to low log parsing accuracy and efficiency. Third, a relatively large number of consumed constant tokens in prompts leads to high LLM invocation costs. At last, these methods only retain placeholders in the results, losing the system visibility brought by variable information in logs. Jinrui Sun, Minghua He, Ying Li 0012 |
WWW | 3 |
| 2025 | ExeCoder: Empowering Large Language Models with Executability Representation for Code TranslationabstractMinghua He, Yue Chen, Fangkai Yang, Pu Zhao, Wenjie Yin, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Minghua He, Yue Chen 0014, Fangkai Yang, Pu Zhao 0004, Yu Kang 0006, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001 |
EMNLP | 1 |
| 2025 | Weakly-Supervised Log-Based Anomaly Detection with Inexact Labels via Multi-Instance LearningabstractLog-based anomaly detection is essential for maintaining software availability. However, existing log-based anomaly detection approaches heavily rely on fine-grained exact labels of log entries which are very hard to obtain in real-world systems. This brings a key problem that anomaly detection models require supervision signals while labeled log entries are unavailable. Facing this problem, we propose a new labeling strategy called inexact labeling that instead of labeling an log entry, system experts can label a bag of log entries in a time span. Furthermore, we propose MIDLog, a weakly supervised log-based anomaly detection approach with inexact labels. We leverage the multiinstance learning paradigm to achieve explicit separation of anomalous log entries from the inexact labeled anomalous log set so as to deduce exact anomalous log labels from inexact labeled log sets. Extensive evaluation on three public datasets shows that our approach achieves an F1 score of over 85% with inexact labels. Minghua He, Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001 |
ICSE | 1 |
| 2025 | CSLParser: A Collaborative Framework Using Small and Large Language Models for Log ParsingabstractLog parsing is a prerequisite for log analysis. Recently, large language models (LLMs) have demonstrated high accuracy in log parsing. However, their frequent invocations incur substantial costs. To address this issue, some methods have turned to small language models (SLMs), which offer improved efficiency but suffer from reduced accuracy due to limited model capacity. To achieve both high accuracy and efficiency, we propose CSLParser, a collaborative log parsing framework using SLMs and LLMs. CSLParser delegates most log parsing tasks to SLMs and selectively invokes LLMs to correct parsing results generated by SLMs, thereby effectively reducing the invocation cost of LLMs while maintaining high accuracy. Specifically, to enhance the accuracy of SLMs, we propose a diversified sampling strategy to select diverse samples for training, enabling SLMs to effectively handle diverse log patterns. To efficiently invoke LLMs, we design a rule-based selection strategy to identify hard cases that are challenging for SLMs to correctly parse, which are subsequently corrected by LLMs. Additionally, we propose a dynamic template updating mechanism that merges similar templates based on structural and semantic information to further enhance parsing accuracy. Extensive experiments on public large-scale log datasets show that CSLParser outperforms state-of-the-art baselines in both accuracy and efficiency. Weijie Hong, Yifan Wu 0002, Lingzhe Zhang, Chiming Duan, Pei Xiao 0005, Minghua He, Xixuan Yang, Ying Li 0012 |
ISSRE | 6 |
| 2025 | ZeroLog: Zero-Label Generalizable Cross-System Log-based Anomaly DetectionabstractLog-based anomaly detection is an important task in ensuring the stability and reliability of software systems. One of the key problems in this task is the lack of labeled logs. Existing works usually leverage large-scale labeled logs from mature systems to train an anomaly detection model of a target system based on the idea of transfer learning. However, these works still require a certain number of labeled logs from the target system. In this paper, we take a step forward and study a valuable yet underexplored setting: zero-label cross-system log-based anomaly detection, that is, no labeled logs are available in the target system. Specifically, we propose ZeroLog, a system-agnostic representation meta-learning method that enables cross-system log-based anomaly detection under zero-label conditions. To achieve this, we leverage unsupervised domain adaptation to perform adversarial training between the source and target domains, aiming to learn system-agnostic general feature representations. By employing meta-learning, the learned representations are further generalized to the target system without any target labels. Experimental results on three public log datasets from different systems show that ZeroLog reaches over $\mathbf{8 0 \%}$ F1-score without labels, comparable to state-of-the-art cross-system methods trained with labeled logs, and outperforms existing methods under zero-label conditions. Xinlong Zhao, Minghua He, Ying Li 0012, Gang Huang 0001 |
ISSRE | 3 |
| 2025 | LogAction: Consistent Cross-system Anomaly Detection through Logs via Active Domain AdaptationabstractLog-based anomaly detection is a essential task for ensuring the reliability and performance of software systems. However, the performance of existing anomaly detection methods heavily relies on labeling, while labeling a large volume of logs is highly challenging. To address this issue, many approaches based on transfer learning and active learning have been proposed. Nevertheless, their effectiveness is hindered by issues such as the gap between source and target system data distributions and cold-start problems. In this paper, we propose LogAction, a novel log-based anomaly detection model based on active domain adaptation. LogAction integrates transfer learning and active learning techniques. On one hand, it uses labeled data from a mature system to train a base model, mitigating the cold-start issue in active learning. On the other hand, LogAction utilize free energy-based sampling and uncertainty-based sampling to select logs located at the distribution boundaries for manual labeling, thus addresses the data distribution gap in transfer learning with minimal human labeling efforts. Experimental results on six different combinations of datasets demonstrate that LogAction achieves an average 93.01% F1 score with only 2% of manual labels, outperforming some state-of-the-art methods by 26.28%. Website: https://logaction.github.io Chiming Duan, Minghua He, Pei Xiao 0005, Zhewei Zhong, Yan Niu, Lingzhe Zhang, Siyu Yu, Yifan Wu 0002, Weijie Hong, Ying Li 0012, Gang Huang 0001 |
ASE | 2 |
| 2025 | United We Stand: Towards End-to-End Log-based Fault Diagnosis via Interactive Multi-Task LearningabstractLog-based fault diagnosis is essential for maintaining software system availability. However, existing fault diagnosis methods are built using a task-independent manner, which fails to bridge the gap between anomaly detection and root cause localization in terms of data form and diagnostic objectives, resulting in three major issues: 1) Diagnostic bias accumulates in the system; 2) System deployment relies on expensive monitoring data; 3) The collaborative relationship between diagnostic tasks is overlooked. Facing this problems, we propose a novel end-to-end log-based fault diagnosis method, Chimera, whose key idea is to achieve end-to-end fault diagnosis through bidirectional interaction and knowledge transfer between anomaly detection and root cause localization. Chimera is based on interactive multitask learning, carefully designing interaction strategies between anomaly detection and root cause localization at the data, feature, and diagnostic result levels, thereby achieving both sub-tasks interactively within a unified end-to-end framework. Evaluation on two public datasets and one industrial dataset shows that Chimera outperforms existing methods in both anomaly detection and root cause localization, achieving improvements of over 2.92%~5.00% and 19.01% ~ 37.09%, respectively. It has been successfully deployed in production, serving an industrial cloud platform. Minghua He, Chiming Duan, Pei Xiao 0005, Siyu Yu, Lingzhe Zhang, Weijie Hong, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001 |
ASE | 1 |
| 2025 | Walk the Talk: Is Your Log-based Software Reliability Maintenance System Really Reliable?abstractLog-based software reliability maintenance systems are crucial for sustaining stable customer experience. However, existing deep learning-based methods represent a black box for service providers, making it impossible for providers to understand how these methods detect anomalies, thereby hindering trust and deployment in real production environments. To address this issue, this paper defines a trustworthiness metric—diagnostic faithfulness—for models to gain service providers’ trust, based on surveys of SREs at a major cloud provider. We design two evaluation tasks: attention-based root cause localization and event perturbation. Empirical studies demonstrate that existing methods perform poorly in diagnostic faithfulness. Consequently, we propose FaithLog, a faithful log-based anomaly detection system, which achieves faithfulness through a carefully designed causality-guided attention mechanism and adversarial consistency learning. Evaluation results on two public datasets and one industrial dataset demonstrate that the proposed method achieves state-of-the-art performance in diagnostic faithfulness. Minghua He, Chiming Duan, Pei Xiao 0005, Lingzhe Zhang, Kangjin Wang, Yifan Wu 0002, Ying Li 0012, Gang Huang 0001 |
ASE | 1 |
| 2025 | CoorLog: Efficient-Generalizable Log Anomaly Detection via Adaptive Coordinator in Software EvolutionabstractFrequent software updates lead to log evolution, posing generalization challenges for current log anomaly detection. Traditional log anomaly detection research focuses on using small deep learning models (SMs), but these models inherently lack generalization due to their closed-world assumption. Large language models (LLMs) exhibit strong semantic understanding and generalization capabilities, making them promising for log anomaly detection. However, they suffer from computational inefficiencies. To balance efficiency and generalization, we propose a collaborative log anomaly detection scheme (CoorLog) that uses an adaptive coordinator to integrate SM and LLM. The coordinator determines if incoming logs have evolved. Non-evolved logs are routed to the SM, while evolved logs are directed to the LLM for detailed inference using the constructed Evol-CoT. To gradually adapt to evolution, we introduce the adaptive evolution mechanism (AEM), which updates the coordinator to redirect evolved logs identified by the LLM to the SM. Simultaneously, the SM is fine-tuned to inherit the LLM’s judgment on these logs. Extensive experiments on real-world datasets demonstrate that CoorLog achieves superior F1-scores in both intra-version and inter-version anomaly detection. Additionally, CoorLog reduces processing time by 91.63% and token consumption by 85.59% compared to using an LLM alone. Pei Xiao 0005, Chiming Duan, Minghua He, Yifan Wu 0002, Gege Gao, Lingzhe Zhang, Weijie Hong, Ying Li 0012, Gang Huang 0001 |
ASE | 3 |
| 2024 | LLMeLog: An Approach for Anomaly Detection based on LLM-enriched Log EventsabstractLog-based anomaly detection is an essential task in maintaining software reliability. Existing log-based anomaly detection approaches often consist of three key phases: log parsing, event embedding, and model construction. Event embedding efficiently extracts semantic information from log events and produces vector representations of log events. However, existing event embedding methods suffer from two key problems. First, semantic noises are buried in log events leading to inevitable gaps between the obtained semantics from log events and their essential meanings. Second, there exists a gap between general semantic embedding and the specific embedding requirement of anomaly detection tasks. To mitigate these problems and improve the quality of representations of log events, we propose a novel anomaly detection approach named LLMeLog. It leverages the capabilities of large language models (LLMs) to enrich the contents of log events with in-context learning techniques. Then it utilizes the enriched log events to fine-tune a pre-trained BERT model. At last, it trains a transformer-based anomaly detection model with the event representations produced by the pre-trained BERT model. Evaluation results on three public log datasets show that LLMeLog achieves the best performance across all datasets, boasting F1-scores exceeding 99%. Besides, when using only 10% of labeled data as training data, our approach can still achieve over 90% F1-scores. Minghua He, Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001 |
ISSRE | 1 |
| 2023 | Mining and Injecting Legal Prior Knowledge to Improve the Generalization Ability of Neural Networks in Chinese Judgments
Yaying Chen, Nanfei Gu, Minghua He |
ICANN (5) | 5 |
| 2023 | FAIR: A Causal Framework for Accurately Inferring Judgments Reversals
Minghua He, Nanfei Gu, Yuntao Shi, Qionghui Zhang, Yaying Chen |
ICANN (1) | 1 |
| 2020 | A hybrid improved kernel LDA and PNN algorithm for efficient face recognition
Aijia Ouyang, Yanmin Liu, Shengyu Pei, Xuyu Peng, Minghua He |
Neurocomputing | 5 |
| 2014 | Clustering web documents using hierarchical representation with multi-granularity
Faliang Huang, Shichao Zhang 0001, Minghua He, Xindong Wu 0001 |
World Wide Web | 3 |
| 2013 | An Intelligent Broker Agent for Energy Trading: An MDP Approach
Rodrigue Talla Kuate, Minghua He, Maria Chli, Hai H. Wang |
IJCAI | 2 |
| 2013 | A Knowledge Based System of Principled Negotiation for Complex Business Contract
Xudong Luo 0001, Kwang Mong Sim 0001, Minghua He |
KSEM | 3 |
| 2013 | A Two-Stage Win-Win Multiattribute Negotiation Model: Optimization and then ConcessionabstractMany automated negotiation models have been developed to solve the conflict in many distributed computational systems. However, the problem of finding win–win outcome in multiattribute negotiation has not been tackled well. To address this issue, based on an evolutionary method of multiobjective optimization, this paper presents a negotiation model that can find win–win solutions of multiple attributes, but needs not to reveal negotiating agents’ private utility functions to their opponents or a third‐party mediator. Moreover, we also equip our agents with a general type of utility functions of interdependent multiattributes, which captures human intuitions well. In addition, we also develop a novel time‐dependent concession strategy model, which can help both sides find a final agreement among a set of win–win ones. Finally, lots of experiments confirm that our negotiation model outperforms the existing models developed recently. And the experiments also show our model is stable and efficient in finding fair win–win outcomes, which is seldom solved in the existing models. Li Pan 0001, Xudong Luo 0001, Xiangxu Meng, Chunyan Miao, Minghua He, Xingchen Guo |
Comput. Intell. | 5 |
| 2012 | Kemnad: a Knowledge Engineering Methodology for Negotiating Agent DevelopmentabstractAutomated negotiation is widely applied in various domains. However, the development of such systems is a complex knowledge and software engineering task. So, a methodology there will be helpful. Unfortunately, none of existing methodologies can offer sufficient, detailed support for such system development. To remove this limitation, this paper develops a new methodology made up of (1) a generic framework (architectural pattern) for the main task, and (2) a library of modular and reusable design pattern (templates) of subtasks. Thus, it is much easier to build a negotiating agent by assembling these standardized components rather than reinventing the wheel each time. Moreover, because these patterns are identified from a wide variety of existing negotiating agents (especially high impact ones), they can also improve the quality of the final systems developed. In addition, our methodology reveals what types of domain knowledge need to be input into the negotiating agents. This in turn provides a basis for developing techniques to acquire the domain knowledge from human users. This is important because negotiation agents act faithfully on the behalf of their human users and thus the relevant domain knowledge must be acquired from the human users. Finally, our methodology is validated with one high impact system. Xudong Luo 0001, Chunyan Miao, Nicholas R. Jennings, Minghua He, Zhiqi Shen 0001, Minjie Zhang 0001 |
Comput. Intell. | 4 |
| 2011 | AstonCAT-Plus: An Efficient Specialist for the TAC Market Design Tournament
Meng Chang, Minghua He, Xudong Luo 0001 |
IJCAI | 2 |
| 2010 | Designing a Successful Adaptive Agent for TAC Ad AuctionabstractThis paper describes the design and evaluation of Aston-TAC, the runner-up in the Ad Auction Game of 2009 International Trading Agent Competition. In particular, we focus on how Aston-TAC generates adaptive bid prices according to the Market-based Value Per Click and how it selects a set of keyword queries to bid on to maximise the expected profit under limited conversion capacity. Through evaluation experiments, we show that AstonTAC performs well and stably not only in the competition but also across a broad range of environments. Meng Chang, Minghua He |
ECAI | 2 |
| 2006 | A heuristic bidding strategy for buying multiple goods in multiple english auctionsabstractThis paper presents the design, implementation, and evaluation of a novel bidding algorithm that a software agent can use to obtain multiple goods from multiple overlapping English auctions. Specifically, an Earliest Closest First heuristic algorithm is proposed that uses neurofuzzy techniques to predict the expected closing prices of the auctions and to adapt the agent's bidding strategy to reflect the type of environment in which it is situated. This algorithm first identifies the set of auctions that are most likely to give the agent the best return and then, according to its attitude to risk, it bids in some other auctions that have approximately similar expected returns, but which finish earlier than those in the best return set. We show through empirical evaluation against a number of methods proposed in the multiple auction literature that our bidding strategy performs effectively and robustly in a wide range of scenarios. Minghua He, Nicholas R. Jennings, Adam Prügel-Bennett |
ACM Trans. Internet Techn. | 1 |
| 2004 | An adaptive bidding agent for multiple English auctions: a neuro-fuzzy approachabstractThis work presents the design, implementation and evaluation of a novel bidding strategy for obtaining goods in multiple overlapping English auctions. The strategy uses fuzzy sets to express trade-offs between multi-attribute goods and exploits neuro-fuzzy techniques to predict the expected closing prices of the auctions and to adapt the agent's bidding strategy to reflect the type of environment in which it is situated. We show, through empirical evaluation against a number of methods proposed in the multiple auction literature, that our strategy performs effectively and robustly in a wide range of scenarios. Minghua He, Nicholas R. Jennings, Adam Prügel-Bennett |
FUZZ-IEEE | 1 |
| 2004 | Designing a successful trading agent: A fuzzy set approachabstractSoftware agents are increasingly being used to represent humans in online auctions. Such agents have the advantages of being able to systematically monitor a wide variety of auctions and then make rapid decisions about what bids to place in what auctions. They can do this continuously and repetitively without losing concentration. To provide a means of evaluating and comparing (benchmarking) research methods in this area the trading agent competition (TAC) was established. This competition involves a number of agents bidding against one another in a number of related auctions (operating different protocols) to purchase travel packages for customers. Against this background, this paper describes the design, implementation and evaluation of SouthamptonTAC, one of the most successful participants in both the Second and the Third International Competitions. Our agent uses fuzzy techniques at the heart of its decision making: to make bidding decisions in the face of uncertainty, to make predictions about the likely outcomes of auctions, and to alter the agent's bidding strategy in response to the prevailing market conditions. Minghua He, Nicholas R. Jennings |
IEEE Trans. Fuzzy Syst. | 1 |
| 2003 | On Agent-Mediated Electronic CommerceabstractThis paper surveys and analyzes the state of the art of agent-mediated electronic commerce (e-commerce), concentrating particularly on the business-to-consumer (B2C) and business-to-business (B2B) aspects. From the consumer buying behavior perspective, agents are being used in the following activities: need identification, product brokering, buyer coalition formation, merchant brokering, and negotiation. The roles of agents in B2B e-commerce are discussed through the business-to-business transaction model that identifies agents as being employed in partnership formation, brokering, and negotiation. Having identified the roles for agents in B2C and B2B e-commerce, some of the key underpinning technologies of this vision are highlighted. Finally, we conclude by discussing the future directions and potential impediments to the wide-scale adoption of agent-mediated e-commerce. Minghua He, Nicholas R. Jennings, Ho-fung Leung |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | A Fuzzy-Logic Based Bidding Strategy for Autonomous Agents in Continuous Double AuctionsabstractIncreasingly, many systems are being conceptualized, designed, and implemented as marketplaces in which autonomous software entities (agents) trade services. These services can be commodities in e-commerce applications or data and knowledge services in information economies. In many of these cases, there are both multiple agents that are looking to procure services and multiple agents that are looking to sell services at any one time. Such marketplaces are termed continuous double auctions (CDAs). Against this background, this paper develops new algorithms that buyer and seller agents can use to participate in CDAs. These algorithms employ heuristic fuzzy rules and fuzzy reasoning mechanisms in order to determine the best bid to make given the state of the marketplace. Moreover, we show how an agent can dynamically adjust its bidding behavior to respond effectively to changes in the supply and demand in the marketplace. We then show, by empirical evaluations, how our agents outperform four of the most prominent algorithms previously developed for CDAs (several of which have been shown to outperform human bidders in experimental studies). Minghua He, Ho-fung Leung, Nicholas R. Jennings |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2003 | Southampton TAC: An adaptive autonomous trading agentabstractSoftware agents are increasingly being used to represent humans in on-line auctions. Such agents have the advantages of being able to systematically monitor a wide variety of auctions and then make rapid decisions about what bids to place in what auctions. They can do this continuously and repetitively without losing concentration. Moreover, in complex multiple auction settings, agents may need to modify their behavior in one auction depending on what is happening in another. To provide a means of evaluating and comparing (benchmarking) research methods in this area, the Trading Agent Competition (TAC) was established. This competition involves a number of agents bidding against one another in a number of related auctions (operating different protocols) to purchase travel packages for customers. Against this background, this artcle describes the design, implementation and evaluation of our adaptive autonomous trading agent, SouthamptonTAC, one of the most successful participants in TAC 2002. Minghua He, Nicholas R. Jennings |
ACM Trans. Internet Techn. | 1 |
| 2002 | SouthamptonTAC: Designing a Successful Trading Agent
Minghua He, Nicholas R. Jennings |
ECAI | 1 |
| 2002 | Agents in E-Commerce: State of the Art
Minghua He, Ho-fung Leung |
Knowl. Inf. Syst. | 1 |
| 2001 | An agent bidding strategy based on fuzzy logic in a continuous double auctionabstractThis paper presents the design, implementation and evaluation of a fuzzy logic based bidding strategy, the FL-strategy, for an agent in a Continuous Double Auction (CDA). According to the outstanding bid, the outstanding ask and the reference price, the FL-strategy employs heuristic fuzzy rules and a reasoning mechanism to find the best ask/bid for an agent. Minghua He, Ho-fung Leung |
SMC | 1 |