David C. Anastasiu

dblp:79/7457 · also Dragos C. Anastasiu · DBLP profile ↗
← Back
23ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-8604-9248ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 PFformer: A Position-Free Transformer Variant for Extreme-Adaptive Multivariate Time Series Forecasting
David C. Anastasiu
PAKDD (6)2
2025 Explainable AI for Real-Time Video Anomaly Anticipation
abstract
With computational power expanding on the edge and deep learning models now capable of real-time inference, it is time to rethink our approach to anomalies. Instead of waiting for anomalies to happen and then detecting them, why not predict them before they occur—and prevent them altogether? Imagine a system that can look at the data at time t and warn us that something might go wrong at t + h. If h is long enough, we can act—automatically or semi-automatically—to stop the anomaly in its tracks. Of course, that might mean changing some people’s plans, and when the anomaly does not happen (because it was prevented), they might wonder why those changes were necessary. That is why these systems need to explain themselves—showing us, visually or descriptively, what they thought was about to go wrong. In this paper, we explore the challenges and opportunities of building real-time video anomaly anticipation systems and share a vision for how these tools could make a real-world impact.
David C. Anastasiu
SDM1
2025 MC-ANN: A Mixture Clustering-Based Attention Neural Network for Time Series Forecasting
abstract
Time Series Forecasting (TSF) has been researched extensively, yet predicting time series with big variances and extreme events remains a challenging problem. Extreme events in reservoirs occur rarely but tend to cause huge problems, e.g., flooding entire towns or neighborhoods, which makes accurate reservoir water level prediction exceedingly important. In this work, we develop a novel extreme-adaptive forecasting approach to accommodate the big variance in hydrologic datasets. We model the time series data distribution as a mixture of both point-wise and segment-wise Gaussian distributions. In particular, we develop a novel End-To-End Mixture Clustering Attention Neural Network (MC-ANN) model for univariate time series forecasting, which we show is able to predict future reservoir water levels effectively. MC-ANN consists of two modules: 1) a grouped Auto-Encoder-based Forecaster (AEF) and 2) a mixture clustering-based learnable Weights Attention Network (WAN) with an attention mechanism. The WAN component is crucial, skillfully adjusting weights to distinguish data with varying distributions, enabling each AEF to concentrate on clusters of data with similar characteristics. Through extensive experiments on real-world datasets, we show MC-ANN's effectiveness (10-45% root mean square error reductions over state-of-the-art methods), underlining its notable potential for practical applications in univariate, skewed, long-term time series prediction tasks.
David C. Anastasiu
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Learning from Polar Representation: An Extreme-Adaptive Model for Long-Term Time Series Forecasting
abstract
In the hydrology field, time series forecasting is crucial for efficient water resource management, improving flood and drought control and increasing the safety and quality of life for the general population. However, predicting long-term streamflow is a complex task due to the presence of extreme events. It requires the capture of long-range dependencies and the modeling of rare but important extreme values. Existing approaches often struggle to tackle these dual challenges simultaneously. In this paper, we specifically delve into these issues and propose Distance-weighted Auto-regularized Neural network (DAN), a novel extreme-adaptive model for long-range forecasting of stremflow enhanced by polar representation learning. DAN utilizes a distance-weighted multi-loss mechanism and stackable blocks to dynamically refine indicator sequences from exogenous data, while also being able to handle uni-variate time-series by employing Gaussian Mixture probability modeling to improve robustness to severe events. We also introduce Kruskal-Wallis sampling and gate control vectors to handle imbalanced extreme data. On four real-life hydrologic streamflow datasets, we demonstrate that DAN significantly outperforms both state-of-the-art hydrologic time series prediction methods and general methods designed for long-term time series prediction.
Jack Xu, David C. Anastasiu
AAAI3
2024 Long-Term Hydrologic Time Series Prediction with LSPM
abstract
Predicting multivariate time series has been a topic of interest among researchers for a long time, especially in hydrological prediction. Due to the presence of extreme events, hydrological prediction requires capturing long-range dependencies and modeling rare but significant extreme values. Accurate prediction of these dependencies is often accomplished using complex models, such as stacked RNNs or transformer-based models, which can be computationally expensive and challenging to train. In addition, existing studies have identified a strong correlation between streamflow and rainfall data. However, the use of additional input data in these studies has often been insufficient, resulting in predictions with low accuracy. In this paper, we address these issues and propose LSPM, a Long Short-term Polar-Learning time series forecasting Model. LSPM learns polar representations through a feature reuse method called EDDU (Encoder Double-Decoder Unit). EDDU creatively incorporates exogenous input to generate long-term predictions based on these learned representations. To maximize the use of indicator sequences from exogenous data, LSPM enhances short-term predictions by a carefully designed loss function and integrates them into the overall forecast, improving robustness to short-term severe events. Experiments on four real-life hydrologic streamflow datasets demonstrate that LSPM significantly outperforms both state-of-the-art hydrologic time series prediction methods and general methods designed for long-term time series prediction.
David C. Anastasiu
CIKM2
2023 An Extreme-Adaptive Time Series Prediction Model Based on Probability-Enhanced LSTM Neural Networks
abstract
Forecasting time series with extreme events has been a challenging and prevalent research topic, especially when the time series data are affected by complicated uncertain factors, such as is the case in hydrologic prediction. Diverse traditional and deep learning models have been applied to discover the nonlinear relationships and recognize the complex patterns in these types of data. However, existing methods usually ignore the negative influence of imbalanced data, or severe events, on model training. Moreover, methods are usually evaluated on a small number of generally well-behaved time series, which does not show their ability to generalize. To tackle these issues, we propose a novel probability-enhanced neural network model, called NEC+, which concurrently learns extreme and normal prediction functions and a way to choose among them via selective back propagation. We evaluate the proposed model on the difficult 3-day ahead hourly water level prediction task applied to 9 reservoirs in California. Experimental results demonstrate that the proposed model significantly outperforms state-of-the-art baselines and exhibits superior generalization ability on data with diverse distributions.
Jack Xu, David C. Anastasiu
AAAI3
2023 SEED: An Effective Model for Highly-Skewed Streamflow Time Series Data Forecasting
abstract
Accurate time series forecasting is crucial in various domains, but predicting highly-skewed and heavy-tailed univariate series poses challenges. We introduce the Segment-Expandable Encoder-Decoder (SEED) model, designed for such time series. SEED incorporates segment representation learning, Kullback-Leibler divergence regularization, and an importance-enhanced sampling policy. We tested our model on the 3-day ahead single-shot prediction task on four hydrologic datasets. Experimental results demonstrate SEED’s effectiveness in optimizing the forecasting process (10-30% of root mean square error reductions over state-of-the-art methods), underlining its notable potential for practical applications in univariate, skewed, long-term time series prediction tasks.
Jack Xu, David C. Anastasiu
IEEE Big Data3
2023 CosTaL: an accurate and scalable graph-based clustering algorithm for high-dimensional single-cell data analysis
abstract
With the aim of analyzing large-sized multidimensional single-cell datasets, we are describing a method for Cosine-based Tanimoto similarity-refined graph for community detection using Leiden's algorithm (CosTaL). As a graph-based clustering method, CosTaL transforms the cells with high-dimensional features into a weighted k-nearest-neighbor (kNN) graph. The cells are represented by the vertices of the graph, while an edge between two vertices in the graph represents the close relatedness between the two cells. Specifically, CosTaL builds an exact kNN graph using cosine similarity and uses the Tanimoto coefficient as the refining strategy to re-weight the edges in order to improve the effectiveness of clustering. We demonstrate that CosTaL generally achieves equivalent or higher effectiveness scores on seven benchmark cytometry datasets and six single-cell RNA-sequencing datasets using six different evaluation metrics, compared with other state-of-the-art graph-based clustering methods, including PhenoGraph, Scanpy and PARC. As indicated by the combined evaluation metrics, Costal has high efficiency with small datasets and acceptable scalability for large datasets, which is beneficial for large-scale analysis.
Jonathan Nguyen, David C. Anastasiu, Edgar A. Arriaga
Briefings Bioinform.3
2022 CLP: A Platform for Competitive Learning
Arpita Vats, Gheorghi Guzun, David C. Anastasiu
EC-TEL3
2022 "Managing, Mining and Learning in the Legal Data Domain"
Andrea Tagarelli, Ester Zumpano, David C. Anastasiu, Andrea Calì, Gottfried Vossen
Inf. Syst.3
2019 CityFlow: A City-Scale Benchmark for Multi-Target Multi-Camera Vehicle Tracking and Re-Identification
abstract
Urban traffic optimization using traffic cameras as sensors is driving the need to advance state-of-the-art multi-target multi-camera (MTMC) tracking. This work introduces CityFlow, a city-scale traffic camera dataset consisting of more than 3 hours of synchronized HD videos from 40 cameras across 10 intersections, with the longest distance between two simultaneous cameras being 2.5 km. To the best of our knowledge, CityFlow is the largest-scale dataset in terms of spatial coverage and the number of cameras/videos in an urban environment. The dataset contains more than 200K annotated bounding boxes covering a wide range of scenes, viewing angles, vehicle models, and urban traffic flow conditions. Camera geometry and calibration information are provided to aid spatio-temporal analysis. In addition, a subset of the benchmark is made available for the task of image-based vehicle re-identification (ReID). We conducted an extensive experimental evaluation of baselines/state-of-the-art approaches in MTMC tracking, multi-target single-camera (MTSC) tracking, object detection, and image-based ReID on this dataset, analyzing the impact of different network architectures, loss functions, spatio-temporal models and their combinations on task effectiveness. An evaluation server is launched with the release of our benchmark at the 2019 AI City Challenge (https://www.aicitychallenge.org/) that allows researchers to compare the performance of their newest techniques. We expect this dataset to catalyze research in this field, propel the state-of-the-art forward, and lead to deployed traffic optimization(s) in the real world.
Milind Naphade, Ming-Yu Liu 0001, Xiaodong Yang 0001, Stanley T. Birchfield, Ratnesh Kumar 0004, David C. Anastasiu, Jenq-Neng Hwang
CVPR8
2019 Tutorial: Are You My Neighbor?: Bringing Order to Neighbor Computing Problems
abstract
Finding nearest neighbors is an important topic that has attracted much attention over the years and has applications in many fields, such as market basket analysis, plagiarism and anomaly detection, community detection, ligand-based virtual screening, etc. As data are easier and easier to collect, finding neighbors has become a potential bottleneck in analysis pipelines. Performing pairwise comparisons given the massive datasets of today is no longer feasible. The high computational complexity of the task has led researchers to develop approximate methods, which find many but not all of the nearest neighbors. Yet, for some types of data, efficient exact solutions have been found by carefully partitioning or filtering the search space in a way that avoids most unnecessary comparisons.
David C. Anastasiu, Huzefa Rangwala, Andrea Tagarelli
KDD1
2019 Parallel cosine nearest neighbor graph construction
David C. Anastasiu, George Karypis
J. Parallel Distributed Comput.1
2018 Data Structure for Efficient Line of Sight Queries
abstract
Given the great amounts of data being transmitted between devices in the 21st century, existing channels of wireless communication are getting congested. In the wireless space, the focus up to now has been on the microwave frequency range. An alternative for high-speed medium- and long-range communication is the millimeter wave spectrum, which is most effectively used through point-to-point links. In this paper, we develop and compare methods for verifying the Line of Sight (LOS) constraint between two points in a city. To be useful for online wireless network planning systems, the methods must be able to process terabytes of 3D city geolocation data and provide answers in milliseconds. We evaluate our methods using data for the city of San Jose, a major metropolitan area in Silicon Valley, California. Our results indicate that our Hierarchical Polygon Aggregation (HPA) method is able to achieve millisecond-level query times with very little loss of precision.
Swapnil Gaikwad, Melody Moh, David C. Anastasiu
CIKM3
2018 Improving Student Motivation through Competitive Active Learning
abstract
This Research Work in Progress paper introduces the Competitive Learning Platform (CLP), an online tool that provides automatic partial performance feedback to students or groups of students on individual or collaborative assignments. CLP motivates students to think outside-the-box and come up with novel solutions that can lead to improved assignment results before the assignment deadline. We developed the CLP system as an Active Learning tool for encouraging and motivating student engagement in a STEM course. In this paper, we describe the CLP system and present the results of a set of analyses aimed at gauging the impact of competitive Active Learning activities using the CLP system on student motivation, engagement, and performance. The analyses are based on CLP submission, student outcome, and student feedback data obtained from 5 STEM undergraduate and graduate course offerings over 3 semesters. Results indicate that competitive active learning is beneficial in this setting, leading to active student participation and improved motivation.
Manika Kapoor, Shuai Hua, David C. Anastasiu
FIE3
2016 Efficient Identification of Tanimoto Nearest Neighbors
abstract
Tanimoto, or (extended) Jaccard, is an important similarity measure which has seen prominent use in fields such as data mining and chemoinformatics. Many of the existing state-of-the-art methods for market-basket analysis, plagiarism and anomaly detection, compound database search, and ligand-based virtual screening rely heavily on identifying Tanimoto nearest neighbors. Given the rapidly increasing size of data that must be analyzed, new algorithms are needed that can speed up nearest neighbor search, yet provide reliable results. While many search algorithms address the complexity of the task by retrieving only some of the nearest neighbors, we propose a method that finds all of the exact nearest neighbors efficiently by leveraging recent advances in similarity search filtering. We provide tighter filtering bounds for the Tanimoto coefficient and show that our method, TAPNN, greatly outperforms existing baselines across a variety of real-world datasets and similarity thresholds.
David C. Anastasiu, George Karypis
DSAA1
2015 L2Knng: Fast Exact K-Nearest Neighbor Graph Construction with L2-Norm Pruning
abstract
The k-nearest neighbor graph is often used as a building block in information retrieval, clustering, online advertising, and recommender systems algorithms. The complexity of constructing the exact k-nearest neighbor graph is quadratic on the number of objects that are compared, and most existing methods solve the problem approximately. We present L2Knng, an efficient algorithm that finds the exact cosine similarity k-nearest neighbor graph for a set of sparse high-dimensional objects. Our algorithm quickly builds an approximate solution to the problem, identifying many of the most similar neighbors, and then uses theoretic bounds on the similarity of two vectors, based on the L2-norm of part of the vectors, to find each object's exact k-neighborhood. We perform an extensive evaluation of our algorithm, comparing against both exact and approximate baselines, and demonstrate the efficiency of our method across a variety of real-world datasets and neighborhood sizes. Our approximate and exact L2Knng variants compute the k-nearest neighbor graph up to an order of magnitude faster than their respective baselines.
David C. Anastasiu, George Karypis
CIKM1
2015 Understanding computer usage evolution
abstract
The proliferation of computing devices in recent years has dramatically changed the way people work, play, communicate, and access information. The personal computer (PC) now has to compete with smartphones, tablets, and other devices for tasks it used to be the default device for. Understanding how PC usage evolves over time can help provide the best overall user experience for current customers, can help determine when they need brand new systems vs. upgraded components, and can inform future product design to better anticipate user needs.
David C. Anastasiu, Al Mamunur Rashid, Andrea Tagarelli, George Karypis
ICDE1
2014 L2AP: Fast cosine similarity search with prefix L-2 norm bounds
abstract
The All-Pairs similarity search, or self-similarity join problem, finds all pairs of vectors in a high dimensional sparse dataset with a similarity value higher than a given threshold. The problem has been classically solved using a dynamically built inverted index. The search time is reduced by early pruning of candidates using size and value-based bounds on the similarity. In the context of cosine similarity and weighted vectors, leveraging the Cauchy-Schwarz inequality, we propose new ℓ2-norm bounds for reducing the inverted index size, candidate pool size, and the number of full dot-product computations. We tighten previous candidate generation and verification bounds and introduce several new ones to further improve our algorithm's performance. Our new pruning strategies enable significant speedups over baseline approaches, most times outperforming even approximate solutions. We perform an extensive evaluation of our algorithm, L2AP, and compare against state-of-the-art exact and approximate methods, AllPairs, MMJoin, and BayesLSH, across a variety of real-world datasets and similarity thresholds.
David C. Anastasiu, George Karypis
ICDE1
2013 A novel two-box search paradigm for query disambiguation
David C. Anastasiu, Byron J. Gao, George Karypis
World Wide Web1
2011 A framework for personalized and collaborative clustering of search results
abstract
How to organize and present search results plays a critical role in the utility of search engines. Due to the unprecedented scale of the Web and diversity of search results, the common strategy of ranked lists has become increasingly inadequate, and clustering has been considered as a promising alternative. Clustering divides a long list of disparate search results into a few topic-coherent clusters, allowing the user to quickly locate relevant results by topic navigation. While many clustering algorithms have been proposed that innovate on the automatic clustering procedure, we introduce ClusteringWiki, the first prototype and framework for personalized clustering that allows direct user editing of the clustering results. Through a Wiki interface, the user can edit and annotate the membership, structure and labels of clusters for a personalized presentation. In addition, the edits and annotations can be shared among users as a mass-collaborative way of improving search result organization and search engine utility.
David C. Anastasiu, Byron J. Gao, David Buttler
CIKM1
2011 ClusteringWiki: personalized and collaborative clustering of search results
abstract
How to organize and present search results plays a critical role in the utility of search engines. Due to the unprecedented scale of the Web and diversity of search results, the common strategy of ranked lists has become increasingly inadequate, and clustering has been considered as a promising alternative. Clustering divides a long list of disparate search results into a few topic-coherent clusters, allowing the user to quickly locate relevant results by topic navigation. While many clustering algorithms have been proposed that innovate on the automatic clustering procedure, we introduce ClusteringWiki, the first prototype and framework for personalized clustering that allows direct user editing of clustering results. Through a Wiki interface, the user can edit and annotate the membership, structure and labels of clusters for a personalized presentation. In addition, the edits and annotations can be shared among users as a mass collaborative way of improving search result organization and search engine utility.
David C. Anastasiu, Byron J. Gao, David Buttler
SIGIR1
2009 The gardener's problem for web information monitoring
abstract
We introduce and theoretically study the Gardener's problem that well models many web information monitoring scenarios, where numerous dynamically changing web sources are monitored and local information needs to be periodically updated under communication and computation capacity constraints. Typical such examples include maintenance of inverted indexes for search engines and maintenance of extracted structures for unstructured data management systems. We formulate a corresponding multicriteria optimization problem and propose heuristic solutions.
Byron J. Gao, Mingji Xia, Walter Cai, David C. Anastasiu
CIKM4