VLDB 2026 Research / reviewers in the wild / expert
Sina Gholamian
dblp:131/0424
· DBLP profile ↗
7ranked-venue papers
5as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Handwritten and Printed Text Segmentation: A Signature Case StudyabstractWhile analyzing scanned documents, handwritten text can overlap with printed text. This overlap causes difficulties during the optical character recognition (OCR) and digitization process of documents, and subsequently, hurts downstream NLP tasks. Prior research either focuses solely on the binary classification of handwritten text or performs a three-class segmentation of the document, i.e., recognition of handwritten, printed, and background pixels. This approach results in the assignment of overlapping handwritten and printed pixels to only one of the classes, and thus, they are not accounted for in the other class. Thus, in this research, we develop novel approaches to address the challenges of handwritten and printed text segmentation. Our objective is to recover text from different classes in their entirety, especially enhancing the segmentation performance on overlapping sections. To support this task, we introduce a new dataset, SignaTR6K, collected from real legal documents, as well as a new model architecture for the handwritten and printed text segmentation task. Our best configuration outperforms prior work on two different datasets by 17.9% and 7.3% on IoU scores. The SignaTR6K dataset is accessible for download via the following link: https://forms.office.com/r/2a5RDg7cAY. Sina Gholamian, Ali Vahdat |
ICCV | 1 |
| 2021 | Leveraging Code Clones and Natural Language Processing for Log Statement PredictionabstractSoftware developers embed logging statements inside the source code as an imperative duty in modern software development as log files are necessary for tracking down runtime system issues and troubleshooting system management tasks. Prior research has emphasized the importance of logging statements in the operation and debugging of software systems. However, the current logging process is mostly manual and ad hoc, and thus, proper placement and content of logging statements remain as challenges. To overcome these challenges, methods that aim to automate log placement and log content, i.e., ‘where, what, and how to log’, are of high interest. Thus, we propose to accomplish the goal of this research, that is “to predict the log statements by utilizing source code clones and natural language processing (NLP)”, as these approaches provide additional context and advantage for log prediction. We pursue the following four research objectives: (RO1) investigate whether source code clones can be leveraged for log statement location prediction, (RO2) propose a clone-based approach for log statement prediction, (RO3) predict log statement’s description with code-clone and NLP models, and (RO4) examine approaches to automatically predict additional details of the log statement, such as its verbosity level and variables. For this purpose, we perform an experimental analysis on seven open-source java projects, extract their method-level code clones, investigate their attributes, and utilize them for log location and description prediction. Our work demonstrates the effectiveness of log-aware clone detection for automated log location and description prediction and outperforms the prior work. Sina Gholamian |
ASE | 1 |
| 2021 | On the Naturalness and Localness of Software LogsabstractLogs are an essential part of the development and maintenance of large and complex software systems as they contain rich information pertaining to the dynamic content and state of the system. As such, developers and practitioners rely heavily on the logs to monitor their systems. In parallel, the increasing volume and scale of the logs, due to the growing complexity of modern software systems, renders the traditional way of manual log inspection insurmountable. Consequently, to handle large volumes of logs efficiently and effectively, various prior research aims to automate the analysis of log files. Thus, in this paper, we begin with the hypothesis that log files are natural and local and these attributes can be applied for automating log analysis tasks. We guide our research with six research questions with regards to the naturalness and localness of the log files, and present a case study on anomaly detection and introduce a tool for anomaly detection, called ANALOG, to demonstrate how our new findings facilitate the automated analysis of logs. Sina Gholamian, Paul A. S. Ward |
MSR | 1 |
| 2021 | What Distributed Systems Say: A Study of Seven Spark Application LogsabstractExecution logs are a crucial medium as they record runtime information of software systems. Although extensive logs are helpful to provide valuable details to identify the root cause in postmortem analysis in case of a failure, this may also incur performance overhead and storage cost. Therefore, in this research, we present the result of our experimental study on seven Spark benchmarks to illustrate the impact of different logging verbosity levels on the execution time and storage cost of distributed software systems. We also evaluate the log effectiveness and the information gain values, and study the changes in performance and the generated logs for each benchmark with various types of distributed system failures. Our research draws insightful findings for developers and practitioners on how to set up and utilize their distributed systems to benefit from the execution logs. Sina Gholamian, Paul A. S. Ward |
SRDS | 1 |
| 2017 | Efficient incremental data analytics with apache sparkabstractAs smart electricity meters are becoming more popular and starting to replace conventional meters worldwide, new area of research for meter data analytics has emerged. Wide spectrum of computations in this context has been applied, ranging from computationally inexpensive tasks such as calculating monthly bills and peak usage, to elaborate computations to provide energy saving feedback to consumers in order to reduce peak energy demand. Examples include model building approaches for usage predictions and recommendations. Although research efforts in this field are progressing, majority of research in this domain still has overlooked the incremental aspects of energy data analytics, or in best cases, researches have not been able to properly utilize the incremental nature of the energy data. We have noticed that incremental approaches can significantly improve performance of smart meter analytics. For example, per-hour readings of a smart meter can efficiently become integrated with the previous readings and result in an incremental re-computation of a particular smart meter task. In this paper, we introduce UW Incremental Spark Analytics (UWISA), our incremental smart meter data platform, which applies efficient incremental techniques for calculating “energy-temperature” model (also called three-line model) [9]. Our platform can achieve better multi-core scalability and speedup of 4.5× (on average) compared to non-incremental implementation and speedup of higher than 2× when compared to previous incremental research for smart meter datasets up to tens of GBs. We also investigate the reasons behind better performance of incremental method when compared to the non-incremental and Spark Streaming approaches. Sina Gholamian, Wojciech M. Golab, Paul A. S. Ward |
IEEE BigData | 1 |
| 2015 | SLA: A Stage-Level Latency Analysisfor Real-Time Communicationin a Pipelined Resource ModelabstractWe present a communication analysis for hard real-time systems interconnects. The objective is to provide tight estimates on the worst-case communication latency between communicating processing elements that use a priority-aware communication medium for data transmission. The communication model consists of communication tasks transmitting data across a series of pipelined resources. The analysis incorporates interferences caused by multiple communication tasks requesting the pipelined resources, and it captures parallel transmission of data between multiple pipeline stages. We call this analysis a stage-level analysis. We evaluate the proposed analysis through simulation of synthetic benchmarks, and we apply the analysis to an instantiation of a platform proposed by Shi and Burns. Our experiments confirm that stage-level analysis provides tight upper-bounds when compared to previous work and improves schedulability by 34 percent. Hany Kashif, Sina Gholamian, Hiren D. Patel |
IEEE Trans. Computers | 2 |
| 2013 | ORTAP: An Offset-based response time analysis for a pipelined communication resource modelabstractThis work addresses the challenge of computing worst-case response times of hard real-time applications deployed on multiprocessor systems. In particular, the worst-case response time analysis (WCRTA) focuses on the communication between distributed tasks of hard real-time applications. The proposed WCRTA models the communication as a pipelined communication resource model. This model incorporates the effect of pipelining, and the parallel transmission of data. Applications of such a model include multiprocessor systems that use complex interconnects such as network-on-chips (NoC)s with priorities. In this paper, we present an exponential analysis, and a polynomial analysis, and prove its correctness. As an application, we apply the pipelined communication resource model to priority-aware NoCs, and we compare the proposed analyses against prior analysis techniques. Our experimental evaluation on two instances of 4 × 4 and 8 × 8 NoCs with 512,000 synthetic benchmarks shows 48.3% and 66.7% improvement in schedulability for the two NoC sizes over prior work. Hany Kashif, Sina Gholamian, Rodolfo Pellizzoni, Hiren D. Patel, Sebastian Fischmeister |
IEEE Real-Time and Embedded Technology and Applications Symposium | 2 |