VLDB 2026 Research / reviewers in the wild / expert
Anthony Bellotti
dblp:180/9055 · also Anthony Graham Bellotti
· DBLP profile ↗
13ranked-venue papers
0as first author
11since 2021 · last 2027
0000-0001-6317-5877ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 5 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | High-fidelity text refinement for ControlNet-guided latent diffusion in document inpaintingabstractBlind document image inpainting (DII) aims to restore degraded document scans without prior knowledge of noise locations, yet existing methods either produce over-smoothed text or introduce pseudo-characters. We propose a novel high-fidelity, text-guided restoration framework based on a ControlNet-guided Latent Diffusion Model (CGLDM). First, we extract and refine noisy OCR outputs using a two-stage Vision Language Model (VLM) and Large Language Model (LLM) pipeline, leveraging both global document context and local text cues to deliver near ground-truth textual fidelity. Next, these refined text features, together with an initial visually restored image, condition a latent diffusion process that progressively denoises and reconstructs clean image latents. To suppress patch-wise background inconsistencies inherent in high-resolution processing, we introduce an explicit feature-alignment loss of diffusion model training that enforces agreement between the predicted features and the VAE-encoded features of the ground-truth image. Extensive experiments on the FUNSD-ZH dataset demonstrate that our approach outperforms state-of-the-art methods in OCR legibility and maintains competitive image-level fidelity. Qinglin Mao, Shengzhe Xu, Songliang Chen, Pushpendu Kar, Anthony Bellotti |
Expert Syst. Appl. | 6 |
| 2025 | ALDII: Adaptive Learning-based Document Image Inpainting to enhance the handwritten Chinese character legibility of human and machineabstractDocument Image Inpainting (DII) has been applied to degraded documents, including financial and historical documents, to enhance the legibility of images for: (1) human readers by providing high visual quality images; and (2) machine recognizers such as Optical Character Recognition (OCR), thereby reducing recognition errors. With the advent of Deep Learning (DL), DL-based DII methods have achieved remarkable enhancements in terms of either human or machine legibility. However, focusing on improving machine legibility causes visual image degradation, affecting human readability. To address this contradiction, we propose an adaptive learning-based DII method, namely ALDII, that applies domain adaptation strategy, our approach acts like a plug-in module that is capable of constraining a total feature space before optimizing legibility of human and machine, respectively. We evaluate our ALDII on a Chinese handwritten character dataset, which includes single-character and text-line images. Compared to other state-of-the-art approaches, experimental results demonstrated superior performance of our ALDII with metrics of both human and machine legibility. Qinglin Mao, Jingjin Li, Pushpendu Kar, Anthony Bellotti |
Neurocomputing | 5 |
| 2025 | Pruning convolutional neural networks for inductive conformal predictionabstractNeural network pruning is a popular approach to reducing model storage size and inference time by removing redundant parameters in the neural network. However, the uncertainty of predictions from pruned models is unexplored. In this paper we study neural network pruning in the context of conformal predictors (CP). The conformal prediction framework built on top of machine learning algorithms supplements their predictions with reliable uncertainty measure in the form of prediction sets, under the independent and identically distributed assumption on the data. Convolutional neural networks (CNNs) have complicated architectures and are widely used in various applications nowadays. Therefore, we focus on pruning CNNs and, in particular, filter-level pruning. We first propose a brute force method that estimates the contribution of a filter to the CP’s predictive efficiency and removes those with the least contribution. Given the computation inefficiency of the brute force method, we also propose the Taylor expansion to approximate the filter’s contribution. Furthermore, we improve the global pruning method by protecting the most important filters within each layer from being pruned. In addition, we explore the ConfTr loss function which is optimized to yield maximal CP efficiency in the context of neural network pruning. We have conducted extensive experimental studies and compared the results regarding the trade-offs between predictive efficiency, computational efficiency, and network sparsity. These results are instructive for deploying pruned neural networks with applications using conformal prediction where reliable predictions and reduced computational cost are relevant, such as in safety-critical applications. Xindi Zhao, Amin Farjudian, Anthony Bellotti |
Neurocomputing | 3 |
| 2025 | Metamorphic Testing and exploration for Machine Learning credit score modelsabstractContext: The rapid development of Machine Learning (ML) has led to the proposal of various ML models to improve credit score assessment, creating a need for effective validation methods to ensure their performance aligns with business expectations. Objective: This paper introduces a novel approach for validating credit scoring models by focusing on user-hypothesized business expectations, enabling testers to predict how input changes affect outputs and assess alignment with business intuition. Methods: The approach uses Metamorphic Testing (MT), applying Metamorphic Relations (MRs) to examine input–output relationships, and Metamorphic Exploration (ME), an advanced extension of MT that constructs MRs based on user expectations. A case study evaluates and contrasts three popular ML models, neural networks, random forests, and gradient boosting tree, using both traditional evaluation metrics in credit scoring and ME. The study investigates how models selected based on traditional metrics perform when evaluated against MRs. Results: Empirical findings reveal that all three models often violate MRs, with violations becoming more extensive as model complexity increases. Neural networks have low number of MR violations on average but tends to be less robust. Interestingly, random forests exhibit most MR violations relative to the other two models. Traditional metrics fail to capture these violations, highlighting their limitations in ensuring alignment with business expectations. Conclusions: ME is proposed as a complementary validation method for model selection and post-deployment monitoring, ensuring models adhere to business intuition. The study underscores the importance of combining traditional metrics with ME, particularly for complex models like neural networks, to improve reliability in real-world applications. Zhihao Ying, Anthony Bellotti, Joseph L. Breeden, Dave Towey |
Inf. Softw. Technol. | 2 |
| 2025 | MRGS-ART: Metamorphic Relation and Group Selection Based on Adaptive Random TestingabstractABSTRACT Metamorphic testing (MT) is effective in detecting software failures; it detects failures by examining the metamorphic relations (MRs) among source test cases (STCs), follow‐up test cases (FTCs) and their respective outputs. The STCs together with the corresponding FTCs, considered as a whole, are called metamorphic groups (MGs). MT performance relies heavily on the MRs and MGs. Previous studies have mainly focused on improving MT performance by identifying effective MRs, or through generation of MGs with high quality, but have somewhat neglected the selection of MRs and MGs from existing ones. In this paper, we address this issue by introducing a new metric for guiding the selection of effective MR‐MG pairs from a new perspective: The MR‐MG pair is chosen such that the MR makes the current MG as far away as possible from the executed MGs. We design an MR‐MG pair selection algorithm, named metamorphic relation and group selection based on adaptive random testing (MRGS‐ART), to implement our metric. The intuition behind MRGS‐ART is that we attempt to improve MT performance by achieving an even distribution of STCs and FTCs in their corresponding input domains for all the MRs used. Experimental results indicate that MRGS‐ART can enhance MT performance. We believe that this is the first comprehensive and systematic demonstration, from the perspective of both MRs and MGs, that making STCs and FTCs evenly distributed in their corresponding input domains can improve MT performance. Finally, by analysing the experimental results, we provide guidance on how to most effectively implement MRGS‐ART. Zhihao Ying, Dave Towey, Anthony Bellotti, Zhiquan Zhou 0001 |
Softw. Test. Verification Reliab. | 3 |
| 2025 | Identifying crowdfunding storytellers who deliver successful projects: a machine learning approachabstractAbstract Crowdfunding plays a key role in financial technology to provide individuals and enterprises with funding opportunities to establish start-ups and/or new business ventures. It is mainly used to link projects’ creators and backers, collect money and plan fundraising projects via social networks. This paper proposes a machine learning-enabled approach to analyse Kickstarter numerical and textual data and predict the successful funding and delivery of crowdfunding projects. It offers crowdfunding stakeholders benefits including creator credibility assessment, project risk reduction, and backer confidence enhancement. This research proposes a data preprocessing approach to prepare the dataset and extract the relevant features for the predictions. Besides, it trains and compares five numerical machine learning classification models and three text-mining methods to find the best-fitted numerical and textual analysis approaches. According to the results, the proposed SVM model outperforms the numerical benchmarks in terms of Accuracy, Precision, Recall, F1 score, and model Training latency. Moreover, BERT gives the best results if the dataset is complex, while Word2vec works better with simple features in textual analysis. Saeid Pourroostaei Ardakani, Kaifeng Jin, Tianhong Cai, Anthony Bellotti, Xiuping Hua |
J. Supercomput. | 6 |
| 2024 | SFIDMT-ART: A metamorphic group generation method based on Adaptive Random Testing applied to source and follow-up input domainsabstractThe performance of metamorphic testing relates strongly to the quality of test cases. However, most related research has only focused on source test cases, ignoring follow-up test cases to some extent. In this paper, we identify a potential problem that may be encountered with existing metamorphic group generation algorithms. We then propose a possible solution to address this problem. Based on this solution, we design a new algorithm for generating effective source and follow-up test cases. To improve the performance (test effectiveness and efficiency) of metamorphic testing. We introduce the concept of the input-domain difference problem, which is likely to affect the performance of metamorphic group generation algorithms. We propose a new test-case distribution criterion for metamorphic testing to address this problem. Based on our proposed criterion, we further present a new metamorphic group generation algorithm, from a black-box perspective, with new distance metrics to facilitate this algorithm. Our algorithm performs significantly better than existing algorithms, in terms of test effectiveness, efficiency and test-case diversity. Through experiments, we find that the input-domain difference problem is likely to affect the performance of metamorphic group generation algorithms. The experimental results demonstrate that our algorithm can achieve good test efficiency, effectiveness, and test-case diversity. Zhihao Ying, Dave Towey, Anthony Bellotti, Tsong Yueh Chen, Zhiquan Zhou 0001 |
Inf. Softw. Technol. | 3 |
| 2023 | Reliable prediction intervals with directly optimized inductive conformal regression for deep learningabstractBy generating prediction intervals (PIs) to quantify the uncertainty of each prediction in deep learning regression, the risk of wrong predictions can be effectively controlled. High-quality PIs need to be as narrow as possible, whilst covering a preset proportion of real labels. At present, many approaches to improve the quality of PIs can effectively reduce the width of PIs, but they do not ensure that enough real labels are captured. Inductive Conformal Predictor (ICP) is an algorithm that can generate effective PIs which is theoretically guaranteed to cover a preset proportion of data. However, typically ICP is not directly optimized to yield minimal PI width. In this study, we propose Directly Optimized Inductive Conformal Regression (DOICR) for neural networks that takes only the average width of PIs as the loss function and increases the quality of PIs through an optimized scheme, under the validity condition that sufficient real labels are captured in the PIs. Benchmark experiments show that DOICR outperforms current state-of-the-art algorithms for regression problems using underlying Deep Neural Network structures for both tabular and image data. Haocheng Lei, Anthony Bellotti |
Neural Networks | 2 |
| 2022 | Using Metamorphic Relation Violation Regions to Support a Simulation Framework for the Process of Metamorphic TestingabstractMetamorphic testing (MT) has been growing in pop-ularity, but it can still be quite challenging and time-consuming to assess its performance. Typical approaches to performance assessment can require a series of steps, and depend on a variety of factors, often requiring serendipity. This can be a bottleneck for some aspects of MT research. Central to MT, metamorphic relations (MRs) represent necessary properties of the system under test (SUT). In traditional software testing, simulations are often employed to examine and compare the performance of dif-ferent testing strategies. However, these simulations are typically designed based on the assumed availability (and applicability) of a test oracle - a mechanism to decide the correctness of the SUT output or behaviour. A key reason for the popularity of MT is its proven record of effective software testing, without the need for a test oracle. This strength, however, also means that traditional ways of using simulations to analyse software testing approaches are not applicable for MT. This lack of cheap and fast ways to conduct simulation analyses of MT is a hurdle for many aspects of MT research, and may be an obstacle to its more widespread adoption. To address this, in this paper we introduce the concept of MR-violation regions (MRVRs), and show how they can be used for a certain category of MRs, Deterministic MRs (DMRs), to build simulation tools for MT. We analyse the differences between MRVRs and traditional, oracle-defined failure regions; and report on a preliminary case study exploring MRVRs in numerical-input-domain systems from previous MT studies. We anticipate that the proposed MT simulation framework may facilitate more research into MT, and may help lead to its more widespread adoption. Zhihao Ying, Anthony Bellotti, Dave Towey, Tsong Yueh Chen, Zhiquan Zhou 0001 |
COMPSAC | 2 |
| 2021 | PBMC Cell Classification from Single Cell mRNA Expression by Artificial Neural Networks, Profiles, Gene Markers, and Protein MarkersabstractWe performed classification of healthy Peripheral Blood Mononuclear Cells cell types using four methods Artificial Neural Network (ANN), Profiles, Protein Markers (PMs), and RNA markers (RNAMs). Profiles represent patterns of gene expressions characteristic of the subtypes of cells. PMs are protein found exclusively in certain types or subtypes of cells, or represent particular cell states, RNAMs are genes which demonstrate significant differential expressions between cell types. A total of 109 datasets from four different sources containing $\sim$ 120,000 single cells gene expression were used to train and test prediction models. We combined the methods which perform prediction using the whole set of gene features (ANN and Profiles), and those that used specific gene features (PMs and RNAMs) to predict the cell type. The overall classification accuracy was 94.8% for ANN, 94.5% for Profiles, 90.7% for PMs, 67.9% for RNAMs. The combination of four methods showed accuracy of 90.9% with high confidence of positive predictions. The combination of four methods allowed identification of mislabeled cell types in test data sets. Minjie Lyu, Yihan Zhang 0003, Luning Yang, Huan Jin, Anthony Bellotti, Nenad S. Mitic, Vladimir Brusic |
BIBM | 7 |
| 2021 | Normalized nonconformity measures for automated valuation models
Zhe Lim, Anthony Bellotti |
Expert Syst. Appl. | 2 |
| 2019 | Aggregating Algorithm for prediction of packsabstractThis paper formulates a protocol for prediction of packs, which is a special case of on-line prediction under delayed feedback. Under the prediction of packs protocol, the learner must make a few predictions without seeing the respective outcomes and then the outcomes are revealed in one go. The paper develops the theory of prediction with expert advice for packs by generalising the concept of mixability. We propose a number of merging algorithms for prediction of packs with tight worst case loss upper bounds similar to those for Vovk’s Aggregating Algorithm. Unlike existing algorithms for delayed feedback settings, our algorithms do not depend on the order of outcomes in a pack. Empirical experiments on sports and house price datasets are carried out to study the performance of the new algorithms and compare them against an existing method. Dmitry Adamskiy, Anthony Bellotti, Raisa Dzhamtyrova, Yuri Kalnishkan |
Mach. Learn. | 2 |
| 2016 | Improving clustering performance by incorporating uncertainty
Maha Bakoben, Anthony Bellotti, Niall M. Adams |
Pattern Recognit. Lett. | 2 |