VLDB 2026 Research / reviewers in the wild / expert
Miles Q. Li
dblp:249/2321
· DBLP profile ↗
10ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0001-7091-3268ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Security and privacy · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Training Dynamics of a 1.7B LLaMa Model: A Data-Efficient ApproachabstractPretraining large language models is a complex endeavor influenced by multiple factors, including model architecture, data quality, training continuity, and hardware constraints. In this paper, we share insights gained from the experience of training DMaS-LLaMa-Lite, a fully open source, 1.7-billionparameter, LLaMa-based model, on approximately 20 billion tokens of carefully curated data. We chronicle the full training trajectory, documenting how evolving validation loss levels and downstream benchmarks reflect transitions from incoherent text to fluent, contextually grounded output. Beyond pretraining, we extend our analysis to include a post-training phase focused on instruction tuning, where the model was refined to produce more contextually appropriate, user-aligned responses. We highlight practical considerations such as the importance of restoring optimizer states when resuming from checkpoints, and the impact of hardware changes on training stability and throughput. While qualitative evaluation provides an intuitive understanding of model improvements, our analysis extends to various performance benchmarks, demonstrating how high-quality data and thoughtful scaling enable competitive results with significantly fewer training tokens. By detailing these experiences and offering training logs, checkpoints, and sample outputs, we aim to guide future researchers and practitioners in refining their pretraining strategies. Miles Q. Li, Benjamin C. M. Fung, Shih-Chia Huang |
IJCNN | 1 |
| 2025 | Security concerns for Large Language Models: A survey
Miles Q. Li, Benjamin C. M. Fung |
J. Inf. Secur. Appl. | 1 |
| 2024 | Survey on Explainable AI: Techniques, challenges and open issues
Adel Abusitta 0001, Miles Q. Li, Benjamin C. M. Fung |
Expert Syst. Appl. | 2 |
| 2022 | VDGraph2Vec: Vulnerability Detection in Assembly Code using Message Passing Neural NetworksabstractSoftware vulnerability detection is one of the most challenging tasks faced by reverse engineers. Recently, vulnerability detection has received a lot of attention due to a drastic increase in the volume and complexity of software. Reverse engineering is a time-consuming and labor-intensive process for detecting malware and software vulnerabilities. However, with the advent of deep learning and machine learning, it has become possible for researchers to automate the process of identifying potential security breaches in software by developing more intelligent technologies. In this research, we propose VDGraph2Vec, an automated deep learning method to generate representations of assembly code for the task of vulnerability detection. Previous approaches failed to attend to topological characteristics of assembly code while discovering the weakness in the software. VDGraph2Vec embeds the control flow and semantic information of assembly code effectively using the expressive capabilities of message passing neural networks and the RoBERTa model. Our model is able to learn the important features that help distinguish between vulnerable and non-vulnerable software. We carry out our experimental analysis for performance benchmark on three of the most common weaknesses and demonstrate that our model can identify vulnerabilities with high accuracy and outperforms the current state-of-the-art binary vulnerability detection models. Ashita Diwan, Miles Q. Li, Benjamin C. M. Fung |
ICMLA | 2 |
| 2022 | Interpretable Malware Classification based on Functional Analysis
Miles Q. Li, Benjamin C. M. Fung |
ICSOFT | 1 |
| 2022 | On the Effectiveness of Interpretable Feedforward Neural NetworkabstractDeep learning models have achieved state-of-the-art performance in many classification tasks. However, most of them cannot provide an explanation for their classification results. Machine learning models that are interpretable are usually linear or piecewise linear and yield inferior performance. Non-linear models achieve much better classification performance, but it is usually hard to explain their classification results. As a counter-example, an interpretable feedforward neural network (IFFNN) is proposed to achieve both high classification performance and interpretability for malware detection. If the IFFNN can perform well in a more flexible and general form for other classification tasks while providing meaningful explanations, it may be of great interest to the applied machine learning community. In this paper, we propose a way to generalize the interpretable feedforward neural network to multi-class classification scenarios and any type of feedforward neural networks, and evaluate its classification performance and interpretability on interpretable datasets. We conclude by finding that the generalized IFFNNs achieve comparable classification performance to their normal feedforward neural network counterparts and provide meaningful explanations. Thus, this kind of neural network architecture has great practical use. Miles Q. Li, Benjamin C. M. Fung, Adel Abusitta 0001 |
IJCNN | 1 |
| 2022 | DyAdvDefender: An instance-based online machine learning model for perturbation-trial-based black-box adversarial defense
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland |
Inf. Sci. | 1 |
| 2021 | A Novel and Dedicated Machine Learning Model for Malware Classification
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland, Steven H. H. Ding |
ICSOFT | 1 |
| 2021 | I-MAD: Interpretable malware detector using Galaxy Transformer
Miles Q. Li, Benjamin C. M. Fung, Philippe Charland, Steven H. H. Ding |
Comput. Secur. | 1 |
| 2021 | Malware classification and composition analysis: A survey of recent developments
Adel Abusitta 0001, Miles Q. Li, Benjamin C. M. Fung |
J. Inf. Secur. Appl. | 2 |