Huaqian Cai

dblp:163/1747 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Leveraging BERT and Large Language Models for Mapping Heterogeneous Scientific and Technological Resources to Their Identifiers
abstract
Currently, there are various scientific and technological resource retrieval databases in the world. The resources stored in these databases may be identified by different identification systems. How to determine whether scientific and technological resources identified by different identification systems are the same resource is an urgent problem to be solved. This paper proposes a software service that leverages BERT and large language models to perform semantic analysis and similarity matching of scientific and technological resource content, and then maps the resources to their respective identifiers. The service effectively solves the problem of how to quickly retrieve the same resource from a large number of scientific and technological resources with diverse identification types, and improves the efficiency and quality of the retrieval. Implemented as a Chrome plugin, the service facilitates seamless mapping heterogeneous scientific and technological resources to their identifiers. We conduct a series of experiments. Their results demonstrate the effectiveness, scalability and stability of the service. To the best of our knowledge, we are the first to propose the service integrating BERT with large language models to extract and identify the key content of scientific and technological resources from web pages.
Yanchun Sun, Xiaohan Zhao, Huizhen Jiang, Huaqian Cai, Changfa Lu, Gang Huang 0001
SSE5
2025 Weakly-Supervised Log-Based Anomaly Detection with Inexact Labels via Multi-Instance Learning
abstract
Log-based anomaly detection is essential for maintaining software availability. However, existing log-based anomaly detection approaches heavily rely on fine-grained exact labels of log entries which are very hard to obtain in real-world systems. This brings a key problem that anomaly detection models require supervision signals while labeled log entries are unavailable. Facing this problem, we propose a new labeling strategy called inexact labeling that instead of labeling an log entry, system experts can label a bag of log entries in a time span. Furthermore, we propose MIDLog, a weakly supervised log-based anomaly detection approach with inexact labels. We leverage the multiinstance learning paradigm to achieve explicit separation of anomalous log entries from the inexact labeled anomalous log set so as to deduce exact anomalous log labels from inexact labeled log sets. Extensive evaluation on three public datasets shows that our approach achieves an F1 score of over 85% with inexact labels.
Minghua He, Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001
ICSE4
2024 LLMeLog: An Approach for Anomaly Detection based on LLM-enriched Log Events
abstract
Log-based anomaly detection is an essential task in maintaining software reliability. Existing log-based anomaly detection approaches often consist of three key phases: log parsing, event embedding, and model construction. Event embedding efficiently extracts semantic information from log events and produces vector representations of log events. However, existing event embedding methods suffer from two key problems. First, semantic noises are buried in log events leading to inevitable gaps between the obtained semantics from log events and their essential meanings. Second, there exists a gap between general semantic embedding and the specific embedding requirement of anomaly detection tasks. To mitigate these problems and improve the quality of representations of log events, we propose a novel anomaly detection approach named LLMeLog. It leverages the capabilities of large language models (LLMs) to enrich the contents of log events with in-context learning techniques. Then it utilizes the enriched log events to fine-tune a pre-trained BERT model. At last, it trains a transformer-based anomaly detection model with the event representations produced by the pre-trained BERT model. Evaluation results on three public log datasets show that LLMeLog achieves the best performance across all datasets, boasting F1-scores exceeding 99%. Besides, when using only 10% of labeled data as training data, our approach can still achieve over 90% F1-scores.
Minghua He, Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001
ISSRE4
2024 LogCAE: An Approach for Log-based Anomaly Detection with Active Learning and Contrastive Learning
abstract
Log-based anomaly detection plays a crucial role in maintaining the reliability of software systems. Unsupervised models are more suitable for real-world usage because they do not rely on huge data labeling efforts. However, their effectiveness is limited because of the lack of supervision of data labels. To balance model effectiveness and labeling efforts, existing approaches enhance model capabilities by incorporating relatively few but key human labels as a golden signal, thereby improving the model ability with acceptable labeling efforts. However, these methods still face limitations of complex human labels and insufficient utilization of human knowledge. In this paper, we introduce LogCAE, a two-stage log anomaly detection approach based on active learning and contrastive learning. It utilizes an unsupervised model to learn from unlabeled log data without human labels and incorporates human knowledge through active learning during online optimization. We employ contrastive learning to optimize the representation of log samples in feature space for more efficient usage of human labels. We conducted experiments on three distinct public log datasets (Thunderbird, BGL, and Zookeeper). The results show that our method improves 12.93% F1-score on average with 6.06% labeled data samples. Besides, our approach is more effective in utilizing human labels than state-of-the-art approaches.
Pei Xiao 0005, Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001
ISSRE4
2023 AFALog: A General Augmentation Framework for Log-based Anomaly Detection with Active Learning
abstract
Log-based anomaly detection is becoming more and more important for maintaining the availability of modern microservice systems. Existing supervised/semi-supervised log anomaly detection models require a large amount of human-labeled logs for training which are hard to collect in real-world systems. Unsupervised models often perform poorly without explicit anomaly labels. To improve the performance of unsupervised models, in this paper, we first make an empirical study of existing unsupervised models to tackle the reason why they often produce unsatisfied results. We find that anomaly detection results produced by existing unsupervised models are significantly affected by two key problems including Not-Cover (NC) problem and Suspicious-Noise (SN) problem. To solve these problems, we propose a novel augmentation framework called AFALog. AFALog leverages the idea of active learning to incorporate human knowledge so as to augment data quality. It can support almost all existing unsupervised models and improve their performance. Our experiments on two open datasets and one dataset collected from a real-world microservice system demonstrate that DALog improves the F1-score by an average of 6.61%, with only 5.9% labeled training data.
Chiming Duan, Huaqian Cai, Ying Li 0012, Gang Huang 0001
ISSRE3
2021 BDLedger: A Scalable Distributed Ledger for Large-Scale Data Recording
Gang Huang 0001, Kaidong Wu, Chaoran Luo, Huaqian Cai, Xiang Jing, Yun Ma 0002
BlockSys5
2020 SmartPipe: Towards Interoperability of Industrial Applications via Computational Reflection
Huaqian Cai, Yun Ma 0003, Tian-Yue Fan, Ying Zhang 0012, Gang Huang 0001
J. Comput. Sci. Technol.2
2019 TransDroid: Automatic Client-based Service Evolving in Android Apps
Huaqian Cai, Ying Zhang 0012
Internetware2
2018 LogPruner: detect, analyze and prune logging calls in Android apps
Xin Zhou 0008, Kaidong Wu, Huaqian Cai, Shuai Lou, Ying Zhang 0012, Gang Huang 0001
Sci. China Inf. Sci.3
2017 CollaDroid: Automatic Augmentation of Android Application with Lightweight Interactive Collaboration
abstract
Collaborative work supported by mobile applications has become more and more popular. Mobile collaboration in some cases needs to be conducted in an interactive way to allow the sharing of the requester screen with the collaborator. Existing interactive screen sharing techniques, however, may cause heavy network traffic and high latency and lack fine-grained control of the scope of collaboration. In this paper, we propose CollaDroid, a lightweight and UI Description based technique for interactive collaboration of Android applications. CollaDroid can automatically transform an Android application to a collaboration augmented application with which a requester can interactively collaborate with a remote collaborator by synchronizing UI (User Interface) content and events. The results of our experimental study show that CollaDroid is applicable for a large part of applications in the Android Market and can provide an efficient collaboration mechanism with low network traffic and latency. And the results of our user study show that the collaboration mechanism implemented by CollaDroid is well accepted by users.
Jiahuan Zheng, Xin Peng 0001, Huaqian Cai, Gang Huang 0001, Ying Zhang 0012, Wenyun Zhao
CSCW4
2017 LogPruner: A Tool for Pruning Logging Call in Android Apps
abstract
The prevalence of mobile platforms, especially the large market share of Android, has promoted the popularity of mobile applications (a.k.a. apps). In developing the apps, logging acts as a crucial tool to help developers debug their app before publishing. In this paper, we present an empirical study on how logging is used in current popular Android apps and reveal the security risks of deactivating the log call instead of removing the call and its associated instructions. To this end, we propose a static analysis scheme to remove the logging call as well as those associated instructions that construct the parameters for the call. We then implement the scheme as a tool called LogPruner and evaluate it with a set of 10 top apps collected from Google Play and Wandoujia. The results show that LogPruner can outperform the naive logging removal approach by 11.8% to 512.5% on pruned instructions in the collected apps.
Huaqian Cai, Xin Zhou 0008, Shuai Lou, Ying Zhang 0012, Gang Huang 0001
Internetware1
2017 DelayDroid: an instrumented approach to reducing tail-time energy of Android apps
Gang Huang 0001, Huaqian Cai, Maciej Swiech, Ying Zhang 0012, Xuanzhe Liu, Peter A. Dinda
Sci. China Inf. Sci.2
2016 Prospects for Shaping User-Centric Mobile Application Workloads to Benefit the Cloud
abstract
Approaches to making cloud operation more efficient, for example through scheduling and power management, largely assume that the workload offered from mobile, user-facing applications is a given and that the cloud must simply adapt to it. We flip this assumption 180 degrees and ask to what extent can we instead shape the user-centric workload into a form that would benefit such approaches. Using a toolchain hat allows us to interpose on frontend/backend interactions in popular Android applications, we add the ability to introduce delays and collect information about user satisfaction. We conduct an "in the wild" user study using this capability, and report on its results. Delays of up to 750 ms can be introduced with little effect on most users, although this is very much user and application dependent. Finally, given our study results, we consider reshaping the application requests by selective delays to have exponential interarrival times (Poisson arrivals), and find that we are often able to do so without exceeding the user's delay tolerance.
Maciej Swiech, Huaqian Cai, Peter A. Dinda, Gang Huang 0001
MASCOTS2
2015 DelayDroid: Reducing Tail-Time Energy by Refactoring Android Apps
abstract
Mobile devices with 3G/4G networking often waste energy in the so-called "tail time" during which the radio is kept on even though no communication is occurring. Prior work has proposed policies to reduce this energy waste by batching network requests. However, this work is challenging to apply in practice due to a lack of mechanisms. In response, we have developed DelayDroid, a framework that allows a developer to add the needed policy to existing, unmodified Android applications (apps) with no human effort. This allows such prior work (as well as our own policies) to be readily deployed and evaluated. The DelayDroid compile-time uses static analysis and bytecode refactoring to identify method calls that send network requests and modify such calls to detour them to the DelayDroid run-time. The run-time then applies a policy to batch them, avoiding the tail time energy waste. DelayDroid also includes a cross-app communication mechanism that supports policies that optimize across multiple apps running together, and we propose a policy that does so. We evaluated the correctness and universality of the DelayDroid mechanisms on 14 popular Android apps chosen from the Google App Store. To evaluate our proposed policy, we studied three DelayDroid-enabled apps (weather forecasting, email client, and news client) running together, finding that the DelayDroid mechanisms combined with our policy can reduce 3G/4G tail time energy waste by 36%.
Huaqian Cai, Ying Zhang 0012, Zhi Jin 0001, Xuanzhe Liu, Gang Huang 0001
Internetware1