EDBT 2026 Demo / reviewers in the wild / expert
Roman Vaculín
dblp:47/4468
· DBLP profile ↗
27ranked-venue papers
5as first author
5since 2021 · last 2024
0009-0003-0867-2311ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 3 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Software engineering, systems software and programming languages · 8 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Identifying Homogeneous and Interpretable Groups for Conformal PredictionabstractConformal prediction methods are a tool for uncertainty quantification of a model’s prediction, providing a model-agnostic and distribution-free statistical wrapper that generates prediction intervals/sets for a given model with finite sample generalization guarantees. However, these guarantees hold only on average, or conditioned on the output values of the predictor or on a set of predefined groups, which a-priori may not relate to the prediction task at hand. We propose a method to learn a generalizable partition function of the input space (or representation mapping) into interpretable groups of varying sizes where the non-conformity scores - a measure of discrepancy between prediction and target - are as homogeneous as possible when conditioned to the group. The learned partition can be integrated with any of the group conditional conformal approaches to produce conformal sets with group conditional guarantees on the discovered regions. Since these learned groups are expressed as strictly a function of the input, they can be used for downstream tasks such as data collection or model selection. We show the effectiveness of our method in reducing worst case group coverage outcomes in a variety of datasets. Natalia Martinez Gil, Dhaval Patel 0002, Chandra Reddy, Venkata Sitaramagiridharganesh Ganapavarapu, Roman Vaculín, Jayant Kalagnanam |
UAI | 5 |
| 2023 | Efficient Pruning for Machine Learning Under Homomorphic Encryption
Ehud Aharoni, Moran Baruch, Pradip Bose, Alper Buyuktosunoglu, Nir Drucker, Subhankar Pal, Tomer Pelleg, Kanthi K. Sarpatwar, Hayim Shaul, Omri Soceanu, Roman Vaculín |
ESORICS (4) | 11 |
| 2023 | AI Explainability 360 Toolkit for Time-Series and Industrial Use CasesabstractWith the growing adoption of AI, trust and explainability have become critical which has attracted a lot of research attention over the past decade and has led to the development of many popular AI explainability libraries such as AIX360, Alibi, OmniXAI, etc. Despite that, applying explainability techniques in practice often poses challenges such as lack of consistency between explainers, semantically incorrect explanations, or scalability. Furthermore, one of the key modalities that has been less explored, both from the algorithmic and practice point of view, is time-series. Several application domains involve time-series including Industry 4.0, asset monitoring, supply chain or finance to name a few. Venkata Sitaramagiridharganesh Ganapavarapu, Sumanta Mukherjee, Natalia Martinez Gil, Kanthi K. Sarpatwar, Amaresh Rajasekharan, Amit Dhurandhar, Vijay Arya, Roman Vaculín |
KDD | 8 |
| 2022 | Physics-based multiple time-series univariate forecastingabstractThe Koopman operator theory provides a recipe for data-driven analysis of dynamical systems. The forecasting problem addresses the future value estimation of an observation from its recent past. One common modeling approach is producing forecasts via modeling the data-generating process, where observations are a function of the state variables of the data-generating process. One major challenge to this is the lack of knowledge of the data-generating process. Delay embedding is a common approach in dynamical system analysis to approximate the state-space manifold from a few state value measurements. In this work, we propose a deep learning framework that employs delay embedding, and Koopman operator theory in the context of univariate forecasting to produce a long-term stable forecast. In this work, we empirically show the correctness of the proposed framework. Our study shows the proposed model can generalize across data-set and is capable of producing a stable long-range forecast. Sumanta Mukherjee, Arindam Jati, Kanthi K. Sarpatwar, Roman Vaculín |
IEEE Big Data | 4 |
| 2021 | AutoAI-TS: AutoAI for Time Series ForecastingabstractA large number of time series forecasting models including traditional statistical models, machine learning models and more recently deep learning have been proposed in the literature. However, choosing the right model along with good parameter values that performs well on a given data is still challenging. Automatically providing a good set of models to users for a given dataset saves both time and effort from using trial-and-error approaches with a wide variety of available models along with parameter optimization. We present AutoAI for Time Series Forecasting (AutoAI-TS) that provides users with a zero configuration (zero-conf) system to efficiently train, optimize and choose best forecasting model among various classes of models for the given dataset. With its flexible zero-conf design, AutoAI-TS automatically performs all the data preparation, model creation, parameter optimization, training and model selection for users and provides a trained model that is ready to use. For given data, AutoAI-TS utilizes a wide variety of models including classical statistical models, Machine Learning (ML) models, statistical-ML hybrid models and deep learning models along with various transformations to create forecasting pipelines. It then evaluates and ranks pipelines using the proposed T-Daub mechanism to choose the best pipeline. The paper describe in detail all the technical aspects of AutoAI-TS along with extensive benchmarking on a variety of real world data sets for various use-cases. Benchmark results show that AutoAI-TS, with no manual configuration from the user, automatically trains and selects pipelines that on average outperform existing state-of-the-art time series forecasting toolkits. Syed Yousaf Shah, Dhaval Patel 0002, Long Vu, Xuan-Hong Dang, Peter Kirchner, Horst Samulowitz, Gregory Bramble, Wesley M. Gifford, Venkata Sitaramagiridharganesh Ganapavarapu, Roman Vaculín, Petros Zerfos |
SIGMOD Conference | 12 |
| 2019 | Constructing and Compressing Frames in Blockchain-based Verifiable Multi-party ComputationabstractIn previous work, we proposed a scalable multi-party verification scheme for expensive iterative computations on a Blockchain substrate by appropriate storage and endorsement of frames of iterates. In this work, we extend the framework to verify sets of complete computations with different unordered hyperparameters and develop frame ordering and compression algorithms to enable scalability in the system. We illustrate the efficacy of the proposed approach by verifying the OpenMalaria epidemiological simulation. Ravi Kiran Raman, Kush R. Varshney, Roman Vaculín, Nelson Bore, Sekou L. Remy, Eleftheria Kyriaki Pissadaki, Michael Hind |
ICASSP | 3 |
| 2019 | Similarity Preserving Representation Learning for Time Series ClusteringabstractA considerable amount of clustering algorithms take instance-feature matrices as their inputs. As such, they cannot directly analyze time series data due to its temporal nature, usually unequal lengths, and complex properties. This is a great pity since many of these algorithms are effective, robust, efficient, and easy to use. In this paper, we bridge this gap by proposing an efficient representation learning framework that is able to convert a set of time series with various lengths to an instance-feature matrix. In particular, we guarantee that the pairwise similarities between time series are well preserved after the transformation , thus the learned feature representation is particularly suitable for the time series clustering task. Given a set of $n$ time series, we first construct an $n\times n$ partially-observed similarity matrix by randomly sampling $\mathcal{O}(n \log n)$ pairs of time series and computing their pairwise similarities. We then propose an efficient algorithm that solves a non-convex and NP-hard problem to learn new features based on the partially-observed similarity matrix. By conducting extensive empirical studies, we demonstrate that the proposed framework is much more effective, efficient, and flexible compared to other state-of-the-art clustering methods. Jinfeng Yi, Roman Vaculín, Lingfei Wu 0001, Inderjit S. Dhillon |
IJCAI | 3 |
| 2019 | Differentially Private Distributed Data Summarization under Covariate ShiftabstractWe envision Artificial Intelligence marketplaces to be platforms where consumers, with very less data for a target task, can obtain a relevant model by accessing many private data sources with vast number of data samples. One of the key challenges is to construct a training dataset that matches a target task without compromising on privacy of the data sources. To this end, we consider the following distributed data summarizataion problem. Given K private source datasets denoted by $[D_i]_{i\in [K]}$ and a small target validation set $D_v$, which may involve a considerable covariate shift with respect to the sources, compute a summary dataset $D_s\subseteq \bigcup_{i\in [K]} D_i$ such that its statistical distance from the validation dataset $D_v$ is minimized. We use the popular Maximum Mean Discrepancy as the measure of statistical distance. The non-private problem has received considerable attention in prior art, for example in prototype selection (Kim et al., NIPS 2016). Our work is the first to obtain strong differential privacy guarantees while ensuring the quality guarantees of the non-private version. We study this problem in a Parsimonious Curator Privacy Model, where a trusted curator coordinates the summarization process while minimizing the amount of private information accessed. Our central result is a novel protocol that (a) ensures the curator does not access more than $O(K^{\frac{1}{3}}|D_s| + |D_v|)$ points (b) has formal privacy guarantees on the leakage of information between the data owners and (c) closely matches the best known non-private greedy algorithm. Our protocol uses two hash functions, one inspired by the Rahimi-Recht random features method and the second leverages state of the art differential privacy mechanisms. We introduce a novel ``noiseless'' differentially private auctioning protocol, which may be of independent interest. Apart from theoretical guarantees, we demonstrate the efficacy of our protocol using real-world datasets. Kanthi K. Sarpatwar, Karthikeyan Shanmugam 0001, Venkata Sitaramagiridharganesh Ganapavarapu, Ashish Jagmohan, Roman Vaculín |
NeurIPS | 5 |
| 2016 | Safe distribution and parallel execution of data-centric workflows over the publish/subscribe abstractionabstractWe present a unique representation of data-centric workflows, designed to exploit the loosely coupled nature of publish/subscribe systems to enable their safe distribution and parallel execution. We argue for the practicality of our approach by mapping a standard and industry-strength data-centric workflow model, namely, IBM Business Artifacts with Guard-Stage-Milestone (GSM), into the publish/subscribe abstraction. Martin Jergler, Hans-Arno Jacobsen, Mohammad Sadoghi, Richard Hull 0001, Roman Vaculín |
ICDE | 5 |
| 2015 | Online Topic-based Social Influence Analysis for the Wimbledon ChampionshipsabstractVarious industries are turning to social media to identify key influencers on topics of interest. Following this trend, the All England Lawn Tennis and Croquet Club (AELTC) is keen to analyze the `social pulse' around the famous Wimbledon Championships. IBM developed and deployed social influence analysis capability for AELTC during the 2014 edition of the Championship. The design and implementation of influence analysis technology in the real world involves several challenges. In this paper, we define various functional and usability criteria that social influence scores should satisfy, and propose a multi-dimensional definition of influence that satisfies these criteria. We highlight the need to identify both all-time influencers and recent influencers, and track user influences over multiple time-scales for this purpose. We also stress the importance of aspect-specific influence analysis, and investigate an approach that uses an aspect hierarchy that annotates tweets with topics or aspects before analyzing them for influence. We also describe interesting insights discovered by our tool and the lessons that we learnt from this engagement. Varun Embar, Indrajit Bhattacharya, Vinayaka Pandit, Roman Vaculín |
KDD | 4 |
| 2015 | Safe Distribution and Parallel Execution of Data-Centric Workflows over the Publish/Subscribe AbstractionabstractIn this work, we develop an approach for the safe distribution and parallel execution of data-centric workflows over the publish/subscribe abstraction. In essence, we design a unique representation of data-centric workflows, specifically designed to exploit the loosely coupled and distributed nature of publish/subscribe systems. Furthermore, we argue for the practicality and expressiveness of our approach by mapping a standard and industry-strength data-centric workflow model, namely, IBM Business Artifacts with Guard-Stage-Milestone (GSM), into the publish/subscribe abstraction. In short, the contributions of this work are three-fold: (1) mapping of data-centric workflows into publish/subscribe to achieve distributed and parallel execution; (2) detailed theoretical analysis of the mapping; and (3) formulation of the complexity of the optimal workflow distribution over the publish/subscribe abstraction as an NP-hard problem. Mohammad Sadoghi, Martin Jergler, Hans-Arno Jacobsen, Richard Hull 0001, Roman Vaculín |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2014 | Splitting GSM schemas: A framework for outsourcing of declarative artifact systems
Rik Eshuis, Richard Hull 0001, Yutian Sun, Roman Vaculín |
Inf. Syst. | 4 |
| 2013 | Splitting GSM Schemas: A Framework for Outsourcing of Declarative Artifact Systems
Rik Eshuis, Richard Hull 0001, Yutian Sun, Roman Vaculín |
BPM | 4 |
| 2013 | Barcelona: A Design and Runtime Environment for Declarative Artifact-Centric BPM
Terry Heath, David Boaz, Manmohan Gupta, Roman Vaculín, Yutian Sun, Richard Hull 0001, Lior Limonad |
ICSOC | 4 |
| 2013 | Data management perspectives on business process management: tutorial overviewabstractTraditional approaches to Business Process Management (BPM) focus primarily on the process aspects, and treat the persistent data accessed and manipulated by the business processes as second class citizens. A recent approach to BPM, based on "business artifacts", is centered on a modeling framework that places data and process on an equal footing. The approach has been shown useful in various application domains, and one variant of business artifacts forms the basis of the emerging OMG Case Management Model and Notation (CMMN) standard. Research results have been developed around conceptual models, enterprise interoperation, business intelligence, and verification. This data-centric approach has the potential to provide the basis for a new generation of BPM technology in support of diverse application, and fueled by the insights into abstraction and data management that have been the hallmark of database research since the 70's. Richard Hull 0001, Jianwen Su, Roman Vaculín |
SIGMOD Conference | 3 |
| 2013 | On the equivalence of incremental and fixpoint semantics for business artifacts with Guard-Stage-Milestone lifecycles
Elio Damaggio, Richard Hull 0001, Roman Vaculín |
Inf. Syst. | 3 |
| 2012 | Ontology of Dynamic Entities
Lior Limonad, Pieter De Leenheer, Mark H. Linehan, Richard Hull 0001, Roman Vaculín |
ER | 5 |
| 2012 | Data-centric Web Services Based on Business ArtifactsabstractExisting Web services standards consider data primarily on the level of inputs and outputs specifications, with the major focus on functional aspects of interactions. Majority of applications rely on data sources, but such data sources are not part of the Web service specifications and cannot be accessed directly by clients. The fact that data are treated independently or as second-class citizens severely limits re-use, flexibility, customization and integration options of current Web services. In this paper we suggest to extend the WS specifications by introducing a data-centric Web services model that integrates functional and data perspectives in one coherent framework. The approach is based on Business Artifacts and in particular on the declarative modular Guard-Stage-Milestone (GSM) model. We introduce a Web Data- and Artifact- centric Service (W-DAS) model using GSM in its core which in addition to usual application specific WS operations defines a set of data access interfaces including CRUD operations, artifacts retrieval interface for querying, filtering and sorting data, and operations for arbitrary custom defined ad hoc run-time queries. We discuss W-DAS publish-subscribe mechanisms and implementation. Roman Vaculín, Terry Heath, Richard Hull 0001 |
ICWS | 1 |
| 2011 | On the Equivalence of Incremental and Fixpoint Semantics for Business Artifacts with Guard-Stage-Milestone Lifecycles
Elio Damaggio, Richard Hull 0001, Roman Vaculín |
BPM | 3 |
| 2011 | Business Artifact-Centric Modeling for Real-Time Performance Monitoring
Roman Vaculín, Zhe Shan 0001, Anil Nigam, Frederick Y. Wu |
BPM | 2 |
| 2011 | Declarative business artifact centric modeling of decision and knowledge intensive business processesabstractIn this paper we address the problem of modeling collaborative decision and knowledge intensive business processes (sometimes referred to as Decision Intensive Processes, or DIP processes). DIP processes assist users in performing decision intensive tasks, and provide users with a guidance relevant to process execution context. DIP processes are by nature collaborative, data-driven, need to support various kinds of flexibility at design and run time, and need to integrate with external services and information sources. Such a combination presents significant challenges for contemporary business processes technologies. We present a solution based on a business artifacts paradigm (a.k.a. business entities with lifecycles) using a Guard-Stage-Milestone (GSM) model for declarative lifecycles specification. We introduce a CoreControl - MicroProcess process design pattern, which allows a natural blending of a business functional process structure (usual for most business processes), with a decision & knowledge driven structure providing domain specific decision guidance to users. The proposed design pattern along with the declarative GSM BA approach provide suitable design primitives for DIP process, as demonstrated on a real problem from the supply chain solutions enablement domain. Roman Vaculín, Richard Hull 0001, Terry Heath, Craig Cochran, Anil Nigam, Noi Sukaviriya |
EDOC | 1 |
| 2011 | Intelligent Content-Based Privacy Assistant for FacebookabstractAlthough most online social networks now offer fine-grained controls of information sharing, these are rarely used, both because their use imposes additional burden on the user and because there are too many control settings for an average user to handle. To mitigate this problem, we have developed an Intelligent Privacy Assistant for Face book that partially automates the assignment of sharing permissions, taking into account the content of the information published and user's high-level sharing policies. The Assistant uses a novel social web privacy language, employs named entity recognition algorithms to annotate sensitive parts of published information and an answer set programming system to evaluate user's privacy policies and determine the list of safe recipients. On a test scenario, the Assistant reached 73.8% and 95.2% performance in correctly determining safe and unsafe recipients, respectively. Michal Jakob, Zbynek Moler, Michal Pechoucek, Roman Vaculín |
Web Intelligence | 4 |
| 2010 | Learning Task Specific Web Services Compositions with Loops and Conditional Branches from Example ExecutionsabstractMajority of the existing approaches to service composition, including the widely popular planning based techniques, are not able to automatically compose practical workflows that include complex repetitive behaviors (loops), taking into account possibility of failures and non-determinism of web service execution results. In this work, we present a learning based approach for composing task specific workflows. We present an approach for learning task specific web service compositions from a very small number of observations (one or more) of example service execution sequences (traces) that solve a given goal. The workflows learned by this approach generalize to the tasks justified by the observed execution trace. The generalization captures the repetitive executions of service sequences, conditional branching executions, and repetitions and branching resulting from failures. We evaluate the approach on a complex web services application involving arbitrary number of repetitive executions and failed executions. Harini Veeraraghavan, Roman Vaculín, Manuela M. Veloso |
Web Intelligence | 2 |
| 2009 | Efficient Discovery of Collision-Free Service CombinationsabstractMajority of service discovery research considers only primitive services as a suitable match for a given query while service combinations are not allowed. However, many realistic queries cannot be matched by individual services and only a combination of several services can satisfy such queries. Allowing service combinations or proper compositions of primitive services as a valid match introduces problems such as unwanted side-effects (i.e., producing an effect that is not requested), effect duplicities (i.e., producing some effect more than once) and contradictory effects (i.e., producing both an effect and its negation). Also the ranking of matched services has to be reconsidered for service combinations. In this paper, we address all the mentioned issues and present a matchmaking algorithm for retrieval of the best top k collision-free service combinations satisfying a given query. Roman Vaculín, Katia P. Sycara |
ICWS | 1 |
| 2008 | Recovery Mechanisms for Semantic Web Services
Kevin Wiesner, Roman Vaculín, Martin J. Kollingbaum, Katia P. Sycara |
DAIS | 2 |
| 2008 | Modeling and Discovery of Data Providing ServicesabstractAbstract Web Services providing access to datasources with structured data have an important place in the SOA. In this paper we focus on modeling and discovery of generic data providing services (DPS), with the goal of making data providing services available for interactions with service requesters in contexts such as service composition and mediation. In our model RDF Views are used to represent the content provided by the DPS. A characterization of match between description of DPS as RDF Views and the OWL-S service request is specified, based on which we developed a flexible matchmaking algorithm for discovery of data providing services. Finally, we propose a realization of the DPS using a SOAP version of the SPARQL protocol and a dynamic configuration interface allowing easy interactions of service requesters with data providing services. Roman Vaculín, Huajun Chen, Roman Neruda, Katia P. Sycara |
ICWS | 1 |
| 2007 | Towards automatic mediation of OWL-S process modelsabstractThe framework for automatic mediation of two process models composed of semantically annotated Web services is presented. Process mediation is hard because of many possible mismatches between process models. We introduce algorithms for the process models analysis to find possible mappings between provider's and requester's process models, or to identify incompatibilities that cannot be reconciled with given set of available data mediators and external services. Results of the analysis phase are used in the mediator runtime component. In particular, we show how the workflow and dataflow mismatches can be resolved. Roman Vaculín, Katia P. Sycara |
ICWS | 1 |