VLDB 2026 Research / reviewers in the wild / expert
Timothy Castiglia
dblp:169/2472 · also Timothy J. Castiglia
· DBLP profile ↗
6ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0003-3432-935XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 73% Trustworthy machine learning · 13% Representation and self-supervised learning · 13% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Distributed systems · 93% Interconnection networks and networks-on-chip · 7% |
Topics — the 11 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
federated learning |
1.2 | 2 | 2023 | LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023 Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022 |
Machine learning › Efficient and distributed learning › federated learning › federated learning architecture
vertical federated learning |
1.2 | 2 | 2023 | LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023 Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.7 | 1 | 2023 | LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023 |
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training |
0.6 | 1 | 2022 | Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022 |
Machine learning › Efficient and distributed learning › distributed training
gradient compression |
0.6 | 1 | 2022 | Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022 |
Distributed systems
distributed machine learning |
0.5 | 1 | 2021 | Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021 |
Distributed systems › distributed machine learning
distributed stochastic gradient descent |
0.5 | 1 | 2021 | Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021 |
Distributed systems › distributed machine learning
distributed training |
0.5 | 1 | 2021 | Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021 |
Distributed systems
heterogeneous networks |
0.5 | 1 | 2021 | Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021 |
Interconnection networks and networks-on-chip › network topology
hierarchical interconnection network |
0.1 | 1 | 2021 | Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021 |
Methods — techniques the papers use, named apart from their topics
spurious feature removal · 0.7pre-training · 0.7top-k sparsification · 0.6quantization · 0.6convergence analysis · 0.6stochastic gradient descent · 0.5local SGD · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Flexible Vertical Federated Learning With Heterogeneous PartiesabstractWe propose flexible vertical federated learning (Flex-VFL), a distributed machine algorithm that trains a smooth, nonconvex function in a distributed system with vertically partitioned data. We consider a system with several parties that wish to collaboratively learn a global function. Each party holds a local dataset; the datasets have different features but share the same sample ID space. The parties are heterogeneous in nature: the parties' operating speeds, local model architectures, and optimizers may be different from one another and, further, they may change over time. To train a global model in such a system, Flex-VFL utilizes a form of parallel block coordinate descent (P-BCD), where parties train a partition of the global model via stochastic coordinate descent. We provide theoretical convergence analysis for Flex-VFL and show that the convergence rate is constrained by the party speeds and local optimizer parameters. We apply this analysis and extend our algorithm to adapt party learning rates in response to changing speeds and local optimizer parameters. Finally, we compare the convergence time of Flex-VFL against synchronous and asynchronous VFL algorithms, as well as illustrate the effectiveness of our adaptive extension. Timothy Castiglia, Shiqiang Wang 0001, Stacy Patterson |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated LearningabstractWe propose LESS-VFL, a communication-efficient feature selection method for distributed systems with vertically partitioned data. We consider a system of a server and several parties with local datasets that share a sample ID space but have different feature sets. The parties wish to collaboratively train a model for a prediction task. As part of the training, the parties wish to remove unimportant features in the system to improve generalization, efficiency, and explainability. In LESS-VFL, after a short pre-training period, the server optimizes its part of the global model to determine the relevant outputs from party models. This information is shared with the parties to then allow local feature selection without communication. We analytically prove that LESS-VFL removes spurious features from model training. We provide extensive empirical evidence that LESS-VFL can achieve high accuracy and remove spurious features at a fraction of the communication cost of other feature selection approaches. Timothy Castiglia, Yi Zhou 0015, Shiqiang Wang 0001, Swanand Kadhe, Nathalie Baracaldo, Stacy Patterson |
ICML | 1 |
| 2022 | Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned DataabstractWe propose Compressed Vertical Federated Learning (C-VFL) for communication-efficient training on vertically partitioned data. In C-VFL, a server and multiple parties collaboratively train a model on their respective features utilizing several local iterations and sharing compressed intermediate results periodically. Our work provides the first theoretical analysis of the effect message compression has on distributed training over vertically partitioned data. We prove convergence of non-convex objectives at a rate of $O(\frac{1}{\sqrt{T}})$ when the compression error is bounded over the course of training. We provide specific requirements for convergence with common compression techniques, such as quantization and top-$k$ sparsification. Finally, we experimentally show compression can reduce communication by over $90%$ without a significant decrease in accuracy over VFL without compression. Timothy Castiglia, Anirban Das 0004, Shiqiang Wang 0001, Stacy Patterson |
ICML | 1 |
| 2022 | Cross-Silo Federated Learning for Multi-Tier Networks with Vertical and Horizontal Data PartitioningabstractWe consider federated learning in tiered communication networks. Our network model consists of a set of silos, each holding a vertical partition of the data. Each silo contains a hub and a set of clients, with the silo’s vertical data shard partitioned horizontally across its clients. We propose Tiered Decentralized Coordinate Descent (TDCD), a communication-efficient decentralized training algorithm for such two-tiered networks. The clients in each silo perform multiple local gradient steps before sharing updates with their hub to reduce communication overhead. Each hub adjusts its coordinates by averaging its workers’ updates, and then hubs exchange intermediate updates with one another. We present a theoretical analysis of our algorithm and show the dependence of the convergence rate on the number of vertical partitions and the number of local updates. We further validate our approach empirically via simulation-based experiments using a variety of datasets and objectives. Anirban Das 0004, Timothy Castiglia, Shiqiang Wang 0001, Stacy Patterson |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2021 | Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks
Timothy Castiglia, Anirban Das 0004, Stacy Patterson |
ICLR | 1 |
| 2020 | A Hierarchical Model for Fast Distributed Consensus in Dynamic NetworksabstractState machine replication is a foundational tool that is used to provide availability and fault tolerance in distributed systems. Safe replication requires a consensus algorithm as a method to achieve agreement on the order of system updates. Designing algorithms that are safe while retaining high throughput is imperative for today's systems. Two of the most widely adopted consensus algorithms, Paxos [1] and Raft [2] , have received much attention and use in industry. While both Paxos and Raft provide safe specification for maintaining a replicated log of system updates, Raft aims for ease of understandability and implementation. Timothy Castiglia, Colin Goldberg, Stacy Patterson |
ICDCS | 1 |