VLDB 2026 Research / reviewers in the wild / expert
Breno Dantas Cruz
dblp:193/4176
· DBLP profile ↗
9ranked-venue papers
3as first author
6since 2021 · last 2023
0000-0002-0243-6591ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Decomposing a Recurrent Neural Network into Modules for Enabling Reusability and ReplacementabstractCan we take a recurrent neural network (RNN) trained to translate between languages and augment it to support a new natural language without retraining the model from scratch? Can we fix the faulty behavior of the RNN by replacing portions associated with the faulty behavior? Recent works on decomposing a fully connected neural network (FCNN) and convolutional neural network (CNN) into modules have shown the value of engineering deep models in this manner, which is standard in traditional SE but foreign for deep learning models. However, prior works focus on the image-based multi-class classification problems and cannot be applied to RNN due to (a) different layer structures, (b) loop structures, (c) different types of input-output architectures, and (d) usage of both non-linear and logistic activation functions. In this work, we propose the first approach to decompose an RNN into modules. We study different types of RNNs, i.e., Vanilla, LSTM, and GRU. Further, we show how such RNN modules can be reused and replaced in various scenarios. We evaluate our approach against 5 canonical datasets (i.e., Math QA, Brown Corpus, Wiki-toxicity, Cline OOS, and Tatoeba) and 4 model variants for each dataset. We found that decomposing a trained model has a small cost (Accuracy: -0.6%, BLEU score: +0.10%). Also, the decomposed modules can be reused and replaced without needing to retrain. Sayem Mohammad Imtiaz, Fraol Batole, Astha Singh, Rangeet Pan, Breno Dantas Cruz, Hridesh Rajan |
ICSE | 5 |
| 2023 | Design by Contract for Deep Learning APIsabstractDeep Learning (DL) techniques are increasingly being incorporated in critical software systems today. DL software is buggy too. Recent work in SE has characterized these bugs, studied fix patterns, and proposed detection and localization strategies. In this work, we introduce a preventative measure. We propose design by contract for DL libraries, DL Contract for short, to document the properties of DL libraries and provide developers with a mechanism to identify bugs during development. While DL Contract builds on the traditional design by contract techniques, we need to address unique challenges. In particular, we need to document properties of the training process that are not visible at the functional interface of the DL libraries. To solve these problems, we have introduced mechanisms that allow developers to specify properties of the model architecture, data, and training process. We have designed and implemented DL Contract for Python-based DL libraries and used it to document the properties of Keras, a well-known DL library. We evaluate DL Contract in terms of effectiveness, runtime overhead, and usability. To evaluate the utility of DL Contract, we have developed 15 sample contracts specifically for training problems and structural bugs. We have adopted four well-vetted benchmarks from prior works on DL bug detection and repair. For the effectiveness, DL Contract correctly detects 259 bugs in 272 real-world buggy programs, from well-vetted benchmarks provided in prior work on DL bug detection and repair. We found that the DL Contract overhead is fairly minimal for the used benchmarks. Lastly, to evaluate the usability, we conducted a survey of twenty participants who have used DL Contract to find and fix bugs. The results reveal that DL Contract can be very helpful to DL application developers when debugging their code. Shibbir Ahmed, Sayem Mohammad Imtiaz, Syeda Khairunnesa Samantha, Breno Dantas Cruz, Hridesh Rajan |
ESEC/SIGSOFT FSE | 4 |
| 2023 | Trusted and privacy-preserving sensor data onloading
Breno Dantas Cruz, Eli Tilevich |
Comput. Commun. | 2 |
| 2022 | DeepDiagnosis: Automatically Diagnosing Faults and Recommending Actionable Fixes in Deep Learning ProgramsabstractDeep Neural Networks (DNNs) are used in a wide variety of applications. However, as in any software application, DNN-based apps are afflicted with bugs. Previous work observed that DNN bug fix patterns are different from traditional bug fix patterns. Furthermore, those buggy models are non-trivial to diagnose and fix due to inexplicit errors with several options to fix them. To support developers in locating and fixing bugs, we propose DeepDiagnosis, a novel debugging approach that localizes the faults, reports error symptoms and suggests fixes for DNN programs. In the first phase, our technique monitors a training model, periodically checking for eight types of error conditions. Then, in case of problems, it reports messages containing sufficient information to perform actionable repairs to the model. In the evaluation, we thoroughly examine 444 models - 53 real-world from GitHub and Stack Overflow, and 391 curated by AUTOTRAINER. DeepDiagnosis provides superior accuracy when compared to UMLUAT and DeepLocalize. Our technique is faster than AUTOTRAINER for fault localization. The results show that our approach can support additional types of models, while state-of-the-art was only able to handle classification ones. Our technique was able to report bugs that do not manifest as numerical errors during training. Also, it can provide actionable insights for fix whereas DeepLocalize can only report faults that lead to numerical errors during training. DeepDiagnosis manifests the best capabilities of fault detection, bug localization, and symptoms identification when compared to other approaches. Mohammad Wardat, Breno Dantas Cruz, Wei Le, Hridesh Rajan |
ICSE | 2 |
| 2022 | Secure and flexible message-based communication for mobile apps within and across devices
Breno Dantas Cruz, Eli Tilevich |
J. Syst. Softw. | 2 |
| 2021 | HTPD: Secure and Flexible Message-Based Communication for Mobile Apps
Breno Dantas Cruz, Eli Tilevich |
SecureComm (2) | 2 |
| 2020 | Understanding the Potential of Edge-Based Participatory Sensing: an Experimental StudyabstractParticipatory sensing uses both local devices for data collection and cloud-based servers for processing. However, transferring the collected data to the cloud can lead to draining device battery power and cause network bandwidth bottlenecks, especially for large multimedia files. In this paper, we investigate how the processing resources at the edge of the network can be leveraged to enable efficient participatory sensing that avoids heavy network traffic. In particular, we report on the experiences of designing, implementing, and evaluating a sensing system that constructs indoor maps by recognizing door signs. A distinguishing characteristic of our system is an almost exclusive use of edge-based processing for tasks that include ML-based image recognition, human-assisted data verification, data model retraining, and administrative data flow aggregation. Our evaluation shows that our system architecture effectively leverages the available edge resources, while greatly reducing network traffic. Based on our experiences of implementing and evaluating our system prototype, we identify several open research directions for further advancing edge-based participatory sensing. Breno Dantas Cruz, Junjie Cheng, Zheng Song 0001, Eli Tilevich |
VTC Spring | 1 |
| 2017 | Detecting Vague Words & Phrases in Requirements Documents in a Multilingual EnvironmentabstractVagueness in software requirements documents can lead to several maintenance problems, especially when the customer and development team do not share the same language. Currently, companies rely on human translators to maintain communication and limit vagueness by translating the requirement documents by hand. In this paper, we describe two approaches that automatically identify vagueness in requirements documents in a multilingual environment. We perform two studies for calibration purposes under strict industrial limitations, and describe the tool that we ultimately deploy. In the first study, six participants, two native Portuguese speakers and four native Spanish speakers, evaluated both approaches. Then, we conducted a field study to test the performance of the best approach in real-world environments at two companies. We describe several lessons learned for research and industrial deployment. Breno Dantas Cruz, Bargav Jayaraman, Anurag Dwarakanath, Collin McMillan |
RE | 1 |
| 2016 | TraceLab Components for Reproducing Source Code Summarization ExperimentsabstractThis artifact is a reproducibility package for experiments in source code summarization. The artifact is implemented as a set of components for the TraceLab research infrastructure. We have converted two implementations of state-of-the-art source code summarization into prepackaged and easily-reusable TraceLab components. Prior to this conversion, the implementations were accessible but difficult to use, being scattered across numerous scripts in various languages with many dependencies. We provide the components, detailed tutorials, and two example virtual machine images via our online appendix. Breno Dantas Cruz, Paul W. McBurney, Collin McMillan |
ICSME | 1 |