Zubin Abraham

dblp:27/209 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 6 · 3 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Environmental and earth informatics · 100%

Topics — the 2 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
spatiotemporal data mining
0.312018
Distribution Preserving Multi-task Regression for Spatio-Temporal Data · ICDM 2018
Environmental and earth informatics › climate science
climate data analysis
0.112018
Distribution Preserving Multi-task Regression for Spatio-Temporal Data · ICDM 2018

Methods — techniques the papers use, named apart from their topics

l2-distance divergence · 0.7nonparametric density estimation · 0.3non-parametric density estimation · 0.3
YearPublicationVenuePosition
2025 Attention-Driven Causal Discovery: From Transformer Matrices to Granger Causal Graphs for Non-Stationary Time-series Data
abstract
Causal discovery in non-stationary time series data is crucial for understanding complex systems but remains challenging due to evolving relationships over time. This paper presents a novel two-stage approach for causal discovery in non-stationary multivariate time series data. The first stage employs a Temporal Attention Forecasting Network (TAFNet), a modified Transformer architecture, to capture complex temporal dependencies and generate informative attention matrices. The second stage utilizes these matrices in an iterative process for Granger causality discovery, refining the predicted causal graph while improving forecasting accuracy. The proposed method addresses the limitations of existing approaches and provides a more complete understanding of causal relationships in non-stationary systems. Extensive experiments demonstrate the method’s superior performance compared to state-of-the-art approaches, particularly in handling non-linear relationships and scaling to high-dimensional data.
Jiageng Zhu, Kehao Li, Zheda Mai, Hanchen Xie, Wael Abd-Almageed, Zubin Abraham
ICASSP6
2018 Distribution Preserving Multi-task Regression for Spatio-Temporal Data
abstract
For many spatio-temporal applications, building regression models that can reproduce the true data distribution is often as important as building models with high prediction accuracy. For example, knowing the future distribution of daily temperature and precipitation can help scientists determine their long-term trends and assess their potential impact on human and natural systems. As conventional methods are designed to minimize residual errors, the shape of their predicted distribution may not be consistent with their actual distribution. To overcome this challenge, this paper presents a novel, distribution-preserving multi-task learning framework for multi-location prediction of spatio-temporal data. The framework employs a non-parametric density estimation approach with L2-distance to measure the divergence between the predicted and true distribution of the data. Experimental results using climate data from more than 1500 weather stations in the United States show that the proposed framework reduces the distribution error for more than 78% of the stations without degrading the prediction accuracy significantly.
Pang-Ning Tan, Zubin Abraham, Lifeng Luo, Pouyan Hatami
ICDM3
2017 Concept Drift Detection with Hierarchical Hypothesis Testing
abstract
When using statistical models (such as a classifier) in a streaming environment, there is often a need to detect and adapt to concept drifts to mitigate any deterioration in the model's predictive performance over time. Unfortunately, the ability of popular concept drift approaches in detecting these drifts in the relationship of the response and predictor variable is often dependent on the distribution characteristics of the data streams, as well as its sensitivity on parameter tuning. This paper presents Hierarchical Linear Four Rates (HLFR), a framework that detects concept drifts for different data stream distributions (including imbalanced data) by leveraging a hierarchical set of hypothesis tests in an online setting. The performance of HLFR is compared to benchmark approaches using both simulated and real-world datasets spanning the breadth of concept drift types. HLFR significantly outperforms benchmark approaches in terms of accuracy, G-mean, recall, delay in detection and adaptability across the various datasets.
Shujian Yu, Zubin Abraham
SDM2
2015 Concept drift detection for streaming data
abstract
Common statistical prediction models often require and assume stationarity in the data. However, in many practical applications, changes in the relationship of the response and predictor variables are regularly observed over time, resulting in the deterioration of the predictive performance of these models. This paper presents Linear Four Rates (LFR), a framework for detecting these concept drifts and subsequently identifying the data points that belong to the new concept (for relearning the model). Unlike conventional concept drift detection approaches, LFR can be applied to both batch and stream data; is not limited by the distribution properties of the response variable (e.g., datasets with imbalanced labels); is independent of the underlying statistical-model; and uses user-specified parameters that are intuitively comprehensible. The performance of LFR is compared to benchmark approaches using both simulated and commonly used public datasets that span the gamut of concept drift types. The results show LFR significantly outperforms benchmark approaches in terms of recall, accuracy and delay in detection of concept drifts across datasets.
Zubin Abraham
IJCNN2
2013 Position Preserving Multi-Output Prediction
Zubin Abraham, Pang-Ning Tan, Perdinan, Julie Winkler, Shiyuan Zhong, Malgorzata Liszewska
ECML/PKDD (2)1
2013 Distribution Regularized Regression Framework for Climate Modeling
abstract
Regression-based approaches are widely used in climate modeling to capture the relationship between a climate variable of interest and a set of predictor variables. These approaches are often designed to minimize the overall prediction errors. However, some climate modeling applications emphasize more on fitting the distribution properties of the observed data. For example, histogram equalization techniques such as quantile mapping have been successfully used to debias outputs from computer-simulated climate models to obtain more realistic projections of future climate scenarios. In this paper, we show the limitations of current regression-based approaches in terms of preserving the distribution of observed climate data and present a multi-objective regression framework that simultaneously fits the distribution properties and minimizes the prediction error. The framework is highly flexible and can be applied to linear, nonlinear, and conditional quantile models. The paper demonstrates the effectiveness of the framework in modeling the daily minimum and maximum temperature as well as precipitation for climate stations in the Great Lakes region. The framework showed marked improvement over traditional regression-based approaches in all 14 climate stations evaluated.
Zubin Abraham, Malgorzata Liszewska, Perdinan, Pang-Ning Tan, Julie Winkler, Shiyuan Zhong
SDM1
2012 Extreme Value Prediction for Zero-Inflated Data
Fan Xin, Zubin Abraham
PAKDD (1)2
2010 An Integrated Framework for Simultaneous Classification and Regression of Time-Series Data
abstract
Zero-inflated time series data are commonly encountered in many applications, including climate and ecological modeling, disease monitoring, manufacturing defect detection, and traffic monitoring. Such data often leads to poor model fitting using standard regression methods because they tend to underestimate the frequency of zeros and the magnitude of non-zero values. This paper presents an integrated framework that simultaneously performs classification and regression to accurately predict future values of a zero-inflated time series. A regression model is initially applied to predict the value of the time series. The regression output is then fed into a classification model to determine whether the predicted value should be adjusted to zero. Our regression and classification models are trained to optimize a joint objective function that considers both classification errors on the time series and regression errors on data points that have non-zero values. We demonstrate the effectiveness of our framework in the context of its application to a precipitation downscaling problem for climate impact assessment studies.
Zubin Abraham, Pang-Ning Tan
SDM1