VLDB 2026 Research / reviewers in the wild / expert
Hao Zhang 0085
dblp:55/2270-85
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2022
0000-0003-4419-660XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | CASMS: Combining clustering with attention semantic model for identifying security bug reports
Jacky W. Keung, Zhen Yang 0022, Xiao Yu 0008, Yishu Li, Hao Zhang 0085 |
Inf. Softw. Technol. | 6 |
| 2021 | OPE: Transforming Programs with Clean and Precise Separation of Tested Intraprocedural Program Paths with Path ProfilingabstractExecuting program paths outside the ones tested means that the program is executing scenarios not tested before deployment. No existing technique can produce a program that precisely contains an arbitrary set of tested program paths in each procedure of a tested program. This paper presents the first work, a novel technique called OPE, to address this problem. OPE first builds a transformed procedure that contains the target set of tested paths for every procedure in a tested program. It extends the transformed procedure with additional branches and basic blocks of code to include all remaining paths of the given procedure. The resultant transformed program is functionally equivalent to the tested program. OPE achieves an inherent strict separation of the tested paths from the rest ready for deployment or follow-up program testing and analysis tasks. The experiment confirms that OPE generates programs with clean path separations and outperforms the previous state-of-the-art path encoding technique when applied to path profiling. Chunbai Yang, Imran Ashraf 0001, Hao Zhang 0085, Wing Kwong Chan |
QRS | 4 |
| 2021 | RegionTrack: A Trace-Based Sound and Complete Checker to Debug Transactional Atomicity Violations and Non-Serializable TracesabstractAtomicity is a correctness criterion to reason about isolated code regions in a multithreaded program when they are executed concurrently. However, dynamic instances of these code regions, called transactions , may fail to behave atomically, resulting in transactional atomicity violations. Existing dynamic online atomicity checkers incur either false positives or false negatives in detecting transactions experiencing transactional atomicity violations. This article proposes RegionTrack. RegionTrack tracks cross-thread dependences at the event, dynamic subregion, and transaction levels. It maintains both dynamic subregions within selected transactions and transactional happens-before relations through its novel timestamp propagation approach. We prove that RegionTrack is sound and complete in detecting both transactional atomicity violations and non-serializable traces. To the best of our knowledge, it is the first online technique that precisely captures the transitively closed set of happens-before relations over all conflicting events with respect to every running transaction for the above two kinds of issues. We have evaluated RegionTrack on 19 subjects of the DaCapo and the Java Grande Forum benchmarks. The empirical results confirm that RegionTrack precisely detected all those transactions which experienced transactional atomicity violations and identified all non-serializable traces. The overall results also show that RegionTrack incurred 1.10x and 1.08x lower memory and runtime overheads than Velodrome and 2.10x and 1.21x lower than Aerodrome, respectively. Moreover, it incurred 2.89x lower memory overhead than DoubleChecker. On average, Velodrome detected about 55% fewer violations than RegionTrack, which in turn reported about 3%–70% fewer violations than DoubleChecker. Shangru Wu, Ernest Bota Pobee, Xiupei Mei, Hao Zhang 0085, Bo Jiang 0001, Wing Kwong Chan |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2021 | DeepEutaxy: Diversity in Weight Search Direction for Fixing Deep Learning Model Training Through Batch PrioritizationabstractDeveloping a deep learning (DL) based software system is slow. One of the critical issues is to conduct many trials and errors in developing a DL model that usually serves as the major component of such a system. A major reason for this inefficiency is the progress of gradual reduction of the gap between the DL model under training and the ground truths. Prior techniques commonly focus on optimizing such errors after the errors have formed. They are insensitive to how a training dataset is provided to the DL model under training in batches, making their approaches nonproactive to deal with such errors. In this article, we propose DeepEutaxy, the first work to repair the model convergence problem from the batch prioritization perspective. Our key insight is that increasing the diversity (i.e., dissimilarity) of the corresponding weights of complex DL models before and after each training step can make the models learn faster and optimize the training errors quicker. DeepEutaxy first trains a DL model with several epochs for initialization. It then partitions and continually prioritizes the training batches for subsequent training epochs based on our novel notion of diversity between the pair of models before and after training on each batch, capturing the strength of the search direction to deal with training errors impacted by that batch. The experiment on six DL models over the MNIST and CIFAR-10 datasets shows that DeepEutaxy can accelerate the convergence of DL models on these two datasets with speedups of 1.75-8.45 and 2.67-15.15 times with respect to the training and test accuracies, respectively. DeepEutaxy can also be integrated into the existing techniques and compare favorably with the prior art in the experiment. Hao Zhang 0085, Wing Kwong Chan |
IEEE Trans. Reliab. | 1 |
| 2019 | Apricot: A Weight-Adaptation Approach to Fixing Deep Learning ModelsabstractA deep learning (DL) model is inherently imprecise. To address this problem, existing techniques retrain a DL model over a larger training dataset or with the help of fault injected models or using the insight of failing test cases in a DL model. In this paper, we present Apricot, a novel weight-adaptation approach to fixing DL models iteratively. Our key observation is that if the deep learning architecture of a DL model is trained over many different subsets of the original training dataset, the weights in the resultant reduced DL model (rDLM) can provide insights on the adjustment direction and magnitude of the weights in the original DL model to handle the test cases that the original DL model misclassifies. Apricot generates a set of such reduced DL models from the original DL model. In each iteration, for each failing test case experienced by the input DL model (iDLM), Apricot adjusts each weight of this iDLM toward the average weight of these rDLMs correctly classifying the test case and/or away from that of these rDLMs misclassifying the same test case, followed by training the weight-adjusted iDLM over the original training dataset to generate a new iDLM for the next iteration. The experiment using five state-of-the-art DL models shows that Apricot can increase the test accuracy of these models by 0.87%-1.55% with an average of 1.08%. The experiment also reveals the complementary nature of these rDLMs in Apricot. Hao Zhang 0085, Wing Kwong Chan |
ASE | 1 |