EDBT 2026 Demo / reviewers in the wild / expert
Aparajita Haldar
dblp:208/2848
· DBLP profile ↗
6ranked-venue papers
1as first author
4since 2021 · last 2024
0000-0002-8485-0330ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Graph learning · 40% Efficient and distributed learning · 40% Language models and text generation · 10% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 87% Distributed systems · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 100% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
0.8 | 1 | 2024 | D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural Networks · Proc. VLDB Endow. 2024 |
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training |
0.8 | 1 | 2024 | D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural Networks · Proc. VLDB Endow. 2024 |
Graph data management › graph processing
streaming graph processing |
0.8 | 1 | 2024 | D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural Networks · Proc. VLDB Endow. 2024 |
Machine learning › Efficient and distributed learning
distributed training |
0.6 | 1 | 2022 | Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022 |
Machine learning › Graph learning › graph neural network
graph convolutional network |
0.6 | 1 | 2022 | Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022 |
Parallel and multicore computing › parallel algorithms
distributed-memory parallel algorithms |
0.6 | 1 | 2022 | Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022 |
Parallel and multicore computing
graph partitioning |
0.6 | 1 | 2022 | Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022 |
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference |
0.3 | 1 | 2018 | Collecting Diverse Natural Language Inference Problems for Sentence Representation Evaluation · EMNLP 2018 |
Distributed systems
communication optimization |
0.2 | 1 | 2022 | Scalable Graph Convolutional Network Training on Distributed-Memory Systems · Proc. VLDB Endow. 2022 |
Methods — techniques the papers use, named apart from their topics
windowed forward pass · 1.5unrolled computation graph · 1.5non-blocking point-to-point communication · 1.1hypergraph partitioning · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural NetworksabstractGraph Neural Network (GNN) models on streaming graphs entail algorithmic challenges to continuously capture its dynamic state, as well as systems challenges to optimize latency, memory, and throughput during both inference and training. We present D3-GNN, the first distributed, hybrid-parallel, streaming GNN system designed to handle real-time graph updates under online query setting. Our system addresses data management, algorithmic, and systems challenges, enabling continuous capturing of the dynamic state of the graph and updating node representations with fault-tolerance and optimal latency, load-balance, and throughput. D3-GNN utilizes streaming GNN aggregators and an unrolled, distributed computation graph architecture to handle cascading graph updates. To counteract data skew and neighborhood explosion issues, we introduce inter-layer and intra-layer windowed forward pass solutions. Experiments on large-scale graph streams demonstrate that D3-GNN achieves high efficiency and scalability. Compared to DGL, D3-GNN achieves a significant throughput improvement of about 76x for streaming workloads. The windowed enhancement further reduces running times by around 10x and message volumes by up to 15x at higher parallelism. Rustam Guliyev, Aparajita Haldar, Hakan Ferhatosmanoglu |
Proc. VLDB Endow. | 2 |
| 2022 | RAGUEL: Recourse-Aware Group Unfairness EliminationabstractWhile machine learning and ranking-based systems are in widespread use for sensitive decision-making processes (e.g., determining job candidates, assigning credit scores), they are rife with concerns over unintended biases in their outcomes, which makes algorithmic fairness (e.g., demographic parity, equal opportunity) an objective of interest. 'Algorithmic recourse' offers feasible recovery actions to change unwanted outcomes through the modification of attributes. We introduce the notion of ranked group-level recourse fairness, and develop a 'recourse-aware ranking' solution that satisfies ranked recourse fairness constraints while minimizing the cost of suggested modifications. Our solution suggests interventions that can reorder the ranked list of database records and mitigate group-level unfairness; specifically, disproportionate representation of sub-groups and recourse cost imbalance. This re-ranking identifies the minimum modifications to data points, with these attribute modifications weighted according to their ease of recourse. We then present an efficient block-based extension that enables re-ranking at any granularity (e.g., multiple brackets of bank loan interest rates, multiple pages of search engine results). Evaluation on real datasets shows that, while existing methods may even exacerbate recourse unfairness, our solution – RAGUEL – significantly improves recourse-aware fairness. RAGUEL outperforms alternatives at improving recourse fairness, through a combined process of counterfactual generation and re-ranking, whilst remaining efficient for large-scale datasets. Aparajita Haldar, Teddy Cunningham, Hakan Ferhatosmanoglu |
CIKM | 1 |
| 2022 | Scalable Graph Convolutional Network Training on Distributed-Memory SystemsabstractGraph Convolutional Networks (GCNs) are extensively utilized for deep learning on graphs. The large data sizes of graphs and their vertex features make scalable training algorithms and distributed memory systems necessary. Since the convolution operation on graphs induces irregular memory access patterns, designing a memory- and communication-efficient parallel algorithm for GCN training poses unique challenges. We propose a highly parallel training algorithm that scales to large processor counts. In our solution, the large adjacency and vertex-feature matrices are partitioned among processors. We exploit the vertex-partitioning of the graph to use non-blocking point-to-point communication operations between processors for better scalability. To further minimize the parallelization overheads, we introduce a sparse matrix partitioning scheme based on a hypergraph partitioning model for full-batch training. We also propose a novel stochastic hypergraph model to encode the expected communication volume in mini-batch training. We show the merits of the hypergraph model, previously unexplored for GCN training, over the standard graph partitioning model which does not accurately encode the communication costs. Experiments performed on real-world graph datasets demonstrate that the proposed algorithms achieve considerable speedups over alternative solutions. The optimizations achieved on communication costs become even more pronounced at high scalability with many processors. The performance benefits are preserved in deeper GCNs having more layers as well as on billion-scale graphs. Gunduz Vehbi Demirci, Aparajita Haldar, Hakan Ferhatosmanoglu |
Proc. VLDB Endow. | 2 |
| 2022 | RoleSim*: Scaling axiomatic role-based similarity ranking on large graphsabstractAbstract RoleSim and SimRank are among the popular graph-theoretic similarity measures with many applications in, e.g., web search, collaborative filtering, and sociometry. While RoleSim addresses the automorphic (role) equivalence of pairwise similarity which SimRank lacks, it ignores the neighboring similarity information out of the automorphically equivalent set. Consequently, two pairs of nodes, which are not automorphically equivalent by nature, cannot be well distinguished by RoleSim if the averages of their neighboring similarities over the automorphically equivalent set are the same. To alleviate this problem: 1) We propose a novel similarity model, namely RoleSim*, which accurately evaluates pairwise role similarities in a more comprehensive manner. RoleSim* not only guarantees the automorphic equivalence that SimRank lacks, but also takes into account the neighboring similarity information outside the automorphically equivalent sets that are overlooked by RoleSim. 2) We prove the existence and uniqueness of the RoleSim* solution, and show its three axiomatic properties (i.e., symmetry, boundedness, and non-increasing monotonicity). 3) We provide a concise bound for iteratively computing RoleSim* formula, and estimate the number of iterations required to attain a desired accuracy. 4) We induce a distance metric based on RoleSim* similarity, and show that the RoleSim* metric fulfills the triangular inequality, which implies the sum-transitivity of its similarity scores. 5) We present a threshold-based RoleSim* model that reduces the computational time further with provable accuracy guarantee. 6) We propose a single-source RoleSim* model, which scales well for sizable graphs. 7) We also devise methods to scale RoleSim* based search by incorporating its triangular inequality property with partitioning techniques. Our experimental results on real datasets demonstrate that RoleSim* achieves higher accuracy than its competitors while scaling well on sizable graphs with billions of edges. Weiren Yu, Sima Iranmanesh, Aparajita Haldar, Maoyin Zhang, Hakan Ferhatosmanoglu |
World Wide Web | 3 |
| 2020 | Variational Recurrent Sequence-to-Sequence Retrieval for Stepwise Illustration
Vishwash Batra, Aparajita Haldar, Yulan He 0001, Hakan Ferhatosmanoglu, George Vogiatzis, Tanaya Guha |
ECIR (1) | 2 |
| 2018 | Collecting Diverse Natural Language Inference Problems for Sentence Representation EvaluationabstractAdam Poliak, Aparajita Haldar, Rachel Rudinger, J. Edward Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2018. Adam Poliak, Aparajita Haldar, Rachel Rudinger, Edward J. Hu, Ellie Pavlick, Aaron Steven White, Benjamin Van Durme |
EMNLP | 2 |