EDBT 2026 Demo / reviewers in the wild / expert
Ekta Gujral
dblp:206/7395
· DBLP profile ↗
12ranked-venue papers
8as first author
3since 2021 · last 2022
0000-0001-7255-3374ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 12 · 8 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Aptera: Automatic PARAFAC2 Tensor AnalysisabstractIn data mining, PARAFAC2 is a powerful and a multi-layer tensor decomposition method that is ideally suited for unsupervised modeling of data which forms “irregular” tensors, e.g., patient's diagnostic profiles, where each patient's recovery timeline does not necessarily align with other patients. In real-world applications, where no ground truth is available, how can we automatically choose how many components to analyze? Although extremely trivial, finding the number of components is very hard. So far, under traditional settings, to determine a reasonable number of components, when using PARAFAC2 data, is to compute decomposition with a different number of components and then analyze the outcome manually. This is an inefficient and time-consuming path, first, due to large data volume and second, the human evaluation makes the selection biased. In this paper, we introduce Aptera, a novel automatic PARAFAC2 tensor mining that is based on locating the L-curve corner. The automation of the PARAFAC2 model quality assessment helps both novice and qualified researchers to conduct detailed and advanced analysis. We extensively evaluate Aptera 's performance on synthetic data, outperforming existing state-of-the-art methods on this very hard problem. Finally, we apply Aptera to a variety of real-world datasets and demonstrate its robustness, scalability, and estimation reliability. Ekta Gujral, Evangelos E. Papalexakis |
ASONAM | 1 |
| 2021 | Mining Bursty Groups from Interaction DataabstractEmpirical studies and theoretical models both highlight burstinessas a common temporal pattern in online behavior. A key driver for burstiness is the self-exciting nature of online interactions. For example, posts in online groups often incite posts in response. Such temporal dependencies are easily lost when interaction data is aggregated in snapshots which are subsequently analyzed independently. An alternative is to model individual interactions as a multi-dimensional self-exciting process, thus, enforcing both temporal and network dependencies. Point processes, however, are challenging to employ for large real-world datasets as fitting them incurs super-linear cost in the number of events. How can we efficiently detect online groups exhibiting bursty self-exciting temporal behavior in large real-world datasets? Alexander Gorovits, Ekta Gujral, Evangelos E. Papalexakis, Petko Bogdanov |
CIKM | 3 |
| 2021 | NED: Niche Detection in User Content Consumption DataabstractExplainable machine learning methods have attracted increased interest in recent years. In this work, we pose and study the niche detection problem, which imposes an explainable lens on the classical problem of co-clustering interactions across two modes. In the niche detection problem, our goal is to identify niches, or co-clusters with node-attribute oriented explanations. Niche detection is applicable to many social content consumption scenarios, where an end goal is to describe and distill high-level insights about user-content associations: not only that certain users like certain types of content, but rather the types of users and content, explained via node attributes. Some examples are an e-commerce platform with who-buys-what interactions and user and product attributes, or a mobile call platform with who-calls-whom interactions and user attributes. Discovering and characterizing niches has powerful implications for user behavior understanding, as well as marketing and targeted content production. Unlike prior works, ours focuses on the intersection of explainable methods and co-clustering. First, we formalize the niche detection problem and discuss preliminaries. Next, we design an end-to-end framework, NED, which operates in two steps: discovering co-clusters of user behaviors based on interaction densities, and explaining them using attributes of involved nodes. Finally, we show experimental results on several public datasets, as well as a large-scale industrial dataset from Snapchat, demonstrating that NED improves in both co-clustering (20% accuracy) and explanation-related objectives (12% average precision) compared to state-of-the-art methods. Ekta Gujral, Leonardo Neves, Evangelos E. Papalexakis, Neil Shah |
CIKM | 1 |
| 2020 | C3 APTION: Constrainted Coupled CP and PARAFAC2 Tensor DecompositionabstractGiven data from a variety of sources that share a number of dimensions, how can we effectively decompose them jointly into interpretable latent factors? The coupled tensor decomposition framework captures this idea by jointly supporting the decomposition of several CP tensors. However, coupling tends to suffer when one dimension of data is irregular, i.e., one of the dimensions of the tensor is uneven, such as in the case of PARAFAC2. In this work, we provide a scalable method for decomposing coupled CP and PARAFAC2 tensor datasets through non-negativity-constrained least squares optimization on a variety of objective functions. We offer the following contributions: (1) Our algorithm can perform coupled factorization with an active-set, block principal pivoting and least square optimization method including the Frobenius norm induced non-negative factorization. (2) C3APTION scales to billions of non-zero elements in both the data and model. Comprehensive experiments on large data confirmed that C3APTION is up to 5× faster and 70 - 80% accurate than several baselines. We present results showing the scalability of this novel implementation on a billion elements as well as demonstrate the high level of interpretability in the latent factors produced, implying that coupling is indeed a promising framework for large-scale, unsupervised pattern exploration and cluster discovery. Ekta Gujral, Georgios Theocharous, Evangelos E. Papalexakis |
ASONAM | 1 |
| 2020 | OnlineBTD: Streaming Algorithms to Track the Block Term Decomposition of Large TensorsabstractIn data mining, block term tensor decomposition (BTD) is a relatively under-explored but very powerful multilayer factor analysis method that is ideally suited for modeling for batch processing of data which is either low or multi-linear rank, e.g., EEG/ECG signals, that extract "rich" structures (> rank – 1) from tensor data while still maintaining a lot of the desirable properties of popular tensor decompositions methods such as the interpretability, uniqueness, and etc. These days data, however, is constantly changing which hinders its use for large data. The tracking of the BTD decomposition for the dynamic tensors is a very pivotal and challenging task due to the variability of incoming data and lack of efficient online algorithms in terms of accuracy, time and space.In this paper, we fill this gap by proposing an efficient method OnlineBTD to compute the BTD decomposition of streaming tensor datasets containing millions of entries. In terms of effectiveness, our proposed method shows comparable results with the prior work, BTD, while being computationally much more efficient. We evaluate OnlineBTD on six synthetic and three diverse real datasets, indicatively, our proposed method shows 10 – 60% speedup and saves 40 – 70% memory usage over the traditional baseline methods and is capable of handling larger tensor streams for which the classic BTD fails to run. To the best of our knowledge, OnlineBTD is the first approach to track streaming block term decomposition while not only being able to provide stable decompositions but also provides better performance in terms of efficiency and scalability. Ekta Gujral, Evangelos E. Papalexakis |
DSAA | 1 |
| 2020 | SPADE: Streaming PARAFAC2 DEcomposition for Large DatasetsabstractIn tensor mining, PARAFAC2 is a powerful and a multi-modal factor analysis method that is ideally suited for modeling for batch processing of data which forms “irregular” tensors, e.g., user movie viewing profiles, where each user's timeline does not necessarily align with other users. However, these days data is dynamically changing which hinders the use of this model for large data. The tracking of the PARAFAC2 decomposition for the dynamic tensors is very pivotal and challenging task due to the variability of incoming data and lack of online efficient algorithm in terms of time and memory. In this paper, we fill this gap by proposing an efficient method to compute the PARAFAC2 decomposition of streaming large tensor datasets containing millions of entries, called SPADE. In terms of effectiveness, our proposed method shows comparable results with the prior work, PARAFAC2, while being computationally much more efficient. We evaluate SPADE on both synthetic and real datasets, indicatively, our proposed method shows 10–23× speedup and saves 17–150× memory usage over the baseline methods and is also capable of handling larger tensor streams (≍ 7 million users) for which the batch baseline was not able to operate. To the best of our knowledge, SPADE is the first approach to online PARAFAC2 decomposition while not only being able to provide on par accuracy but also provide better performance in terms of scalability and efficiency. Ekta Gujral, Georgios Theocharous, Evangelos E. Papalexakis |
SDM | 1 |
| 2020 | Beyond Rank-1: Discovering Rich Community Structure in Multi-Aspect GraphsabstractHow are communities in real multi-aspect or multi-view graphs structured? How we can effectively and concisely summarize and explore those communities in a high-dimensional, multi-aspect graph without losing important information? State-of-the-art studies focused on patterns in single graphs, identifying structures in a single snapshot of a large network or in time evolving graphs and stitch them over time. Ekta Gujral, Ravdeep Pasricha, Evangelos E. Papalexakis |
WWW | 1 |
| 2018 | t-PNE: Tensor-Based Predictable Node EmbeddingsabstractGraph representations have increasingly grown in popularity during the last years. Existing embedding approaches explicitly encode network structure. Despite their good performance in downstream processes (e.g., node classification), there is still room for improvement in different aspects, like effectiveness. In this paper, we propose, t-PNE, a method that addresses this limitation. Contrary to baseline methods, which generally learn explicit node representations by solely using an adjacency matrix, t-PNE avails a multi-view information graph-the adjacency matrix represents the first view, and a nearest neighbor adjacency, computed over the node features, is the second view-in order to learn explicit and implicit node representations, using the Canonical Polyadic (a.k.a. CP) decomposition. We argue that the implicit and the explicit mapping from a higher-dimensional to a lower-dimensional vector space is the key to learn more useful and highly predictable representations. Extensive experiments show that t-PNE drastically outperforms baseline methods by up to 158.6% with respect to Micro-Fl, in several multi-label classification problems. Saba A. Al-Sayouri, Ekta Gujral, Danai Koutra, Evangelos E. Papalexakis, Sarah S. Lam |
ASONAM | 2 |
| 2018 | LARC: Learning Activity-Regularized Overlapping Communities Across TimeabstractCommunities are essential building blocks of complex networks enjoying significant research attention in terms of modeling and detection algorithms. Common across models is the premise that node pairs that share communities are likely to interact more strongly. Moreover, in the most general setting a node may be a member of multiple communities, and thus, interact with more than one cohesive group of other nodes. If node interactions are observed over a long period and aggregated into a single static network, the communities may be hard to discern due to their in-network overlap. Alternatively, if interactions are observed over short time periods, the communities may be only partially observable. How can we detect communities at an appropriate temporal resolution that resonates with their natural periods of activity? We propose LARC, a general framework for joint learning of the overlapping community structure and the periods of activity of communities, directly from temporal interaction data. We formulate the problem as an optimization task coupling community fit and smooth temporal activation over time. To the best of our knowledge, the tensor version of LARC is the first tensor-based community detection method to introduce such smoothness constraints. We propose efficient algorithms for the problem, achieving a $2.6x$ quality improvement over all baselines for high temporal resolution datasets, and consistently detecting better-quality communities for different levels of data aggregation and varying community overlap. In addition, LARC elucidates interpretable temporal patterns of community activity corresponding to botnet attacks, transportation change points and public forum interaction trends, while being computationally practical---few minutes on large real datasets. Finally, LARC provides a comprehensive \em unsupervised parameter estimation methodology yielding high accuracy and rendering it easy-to-use for practitioners. Alexander Gorovits, Ekta Gujral, Evangelos E. Papalexakis, Petko Bogdanov |
KDD | 2 |
| 2018 | Identifying and Alleviating Concept Drift in Streaming Tensor Decomposition
Ravdeep Pasricha, Ekta Gujral, Evangelos E. Papalexakis |
ECML/PKDD (2) | 2 |
| 2018 | SMACD: Semi-supervised Multi-Aspect Community DetectionabstractCommunity detection in real-world graphs has been shown to benefit from using multi-aspect information, e.g., in the form of “means of communication” between nodes in the network. An orthogonal line of work, broadly construed as semi-supervised learning, approaches the problem by introducing a small percentage of node assignments to communities and propagates that knowledge throughout the graph. In this paper we introduce SMACD, a novel semi-supervised multi-aspect community detection method along with an automated parameter tuning algorithm which essentially renders SMACD parameter-free. To the best of our knowledge, SMACD is the first approach to incorporate multi-aspect graph information and semi-supervision, while being able to discover overlapping and non-overlapping communities. We extensively evaluate SMACD's performance in comparison to state-of-the-art approaches across eight real and two synthetic datasets, and demonstrate that SMACD, through combining semi-supervision and multi-aspect edge information, outperforms the baselines. Ekta Gujral, Evangelos E. Papalexakis |
SDM | 1 |
| 2018 | SamBaTen: Sampling-based Batch Incremental Tensor DecompositionabstractTensor decompositions are invaluable tools in analyzing multimodal datasets. In many real-world scenarios, such datasets are far from being static, to the contrary they tend to grow over time. For instance, in an online social network setting, as we observe new interactions over time, our dataset gets updated in its “time” mode. How can we maintain a valid and accurate tensor decomposition of such a dynamically evolving multimodal dataset, without having to re-compute the entire decomposition after every single update? In this paper we introduce SamBaTen, a Sampling-based Batch Incremental Tensor Decomposition algorithm, which incrementally maintains the decomposition given new updates to the tensor dataset. SamBaTen is able to scale to datasets that the state-of-the-art in incremental tensor decomposition is unable to operate on, due to its ability to effectively summarize the existing tensor and the incoming updates, and perform all computations in the reduced summary space. We extensively evaluate SamBaTen using synthetic and real datasets. Indicatively, SamBaTen achieves comparable accuracy to state-of-the-art incremental and non-incremental techniques, while being up to 25–30 times faster. Furthermore, SamBaTen scales to very large sparse and dense dynamically evolving tensors of dimensions up to 100K × 100K × 100K where state-of-the-art incremental approaches were not able to operate. Ekta Gujral, Ravdeep Pasricha, Evangelos E. Papalexakis |
SDM | 1 |