Timothy Castiglia

dblp:169/2472 · also Timothy J. Castiglia · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0003-3432-935XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Efficient and distributed learning · 73% Trustworthy machine learning · 13% Representation and self-supervised learning · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 93% Interconnection networks and networks-on-chip · 7%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
1.222023
LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023
Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022
Machine learning › Efficient and distributed learning › federated learning › federated learning architecture
vertical federated learning
1.222023
LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023
Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection
0.712023
LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023
Machine learning › Trustworthy machine learning
interpretability
0.712023
LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning · ICML 2023
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.612022
Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.612022
Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data · ICML 2022
Distributed systems
distributed machine learning
0.512021
Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021
Distributed systems › distributed machine learning
distributed stochastic gradient descent
0.512021
Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021
Distributed systems › distributed machine learning
distributed training
0.512021
Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021
Distributed systems
heterogeneous networks
0.512021
Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021
Interconnection networks and networks-on-chip › network topology
hierarchical interconnection network
0.112021
Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks · ICLR 2021

Methods — techniques the papers use, named apart from their topics

spurious feature removal · 0.7pre-training · 0.7top-k sparsification · 0.6quantization · 0.6convergence analysis · 0.6stochastic gradient descent · 0.5local SGD · 0.5
YearPublicationVenuePosition
2024 Flexible Vertical Federated Learning With Heterogeneous Parties
abstract
We propose flexible vertical federated learning (Flex-VFL), a distributed machine algorithm that trains a smooth, nonconvex function in a distributed system with vertically partitioned data. We consider a system with several parties that wish to collaboratively learn a global function. Each party holds a local dataset; the datasets have different features but share the same sample ID space. The parties are heterogeneous in nature: the parties' operating speeds, local model architectures, and optimizers may be different from one another and, further, they may change over time. To train a global model in such a system, Flex-VFL utilizes a form of parallel block coordinate descent (P-BCD), where parties train a partition of the global model via stochastic coordinate descent. We provide theoretical convergence analysis for Flex-VFL and show that the convergence rate is constrained by the party speeds and local optimizer parameters. We apply this analysis and extend our algorithm to adapt party learning rates in response to changing speeds and local optimizer parameters. Finally, we compare the convergence time of Flex-VFL against synchronous and asynchronous VFL algorithms, as well as illustrate the effectiveness of our adaptive extension.
Timothy Castiglia, Shiqiang Wang 0001, Stacy Patterson
IEEE Trans. Neural Networks Learn. Syst.1
2023 LESS-VFL: Communication-Efficient Feature Selection for Vertical Federated Learning
abstract
We propose LESS-VFL, a communication-efficient feature selection method for distributed systems with vertically partitioned data. We consider a system of a server and several parties with local datasets that share a sample ID space but have different feature sets. The parties wish to collaboratively train a model for a prediction task. As part of the training, the parties wish to remove unimportant features in the system to improve generalization, efficiency, and explainability. In LESS-VFL, after a short pre-training period, the server optimizes its part of the global model to determine the relevant outputs from party models. This information is shared with the parties to then allow local feature selection without communication. We analytically prove that LESS-VFL removes spurious features from model training. We provide extensive empirical evidence that LESS-VFL can achieve high accuracy and remove spurious features at a fraction of the communication cost of other feature selection approaches.
Timothy Castiglia, Yi Zhou 0015, Shiqiang Wang 0001, Swanand Kadhe, Nathalie Baracaldo, Stacy Patterson
ICML1
2022 Compressed-VFL: Communication-Efficient Learning with Vertically Partitioned Data
abstract
We propose Compressed Vertical Federated Learning (C-VFL) for communication-efficient training on vertically partitioned data. In C-VFL, a server and multiple parties collaboratively train a model on their respective features utilizing several local iterations and sharing compressed intermediate results periodically. Our work provides the first theoretical analysis of the effect message compression has on distributed training over vertically partitioned data. We prove convergence of non-convex objectives at a rate of $O(\frac{1}{\sqrt{T}})$ when the compression error is bounded over the course of training. We provide specific requirements for convergence with common compression techniques, such as quantization and top-$k$ sparsification. Finally, we experimentally show compression can reduce communication by over $90%$ without a significant decrease in accuracy over VFL without compression.
Timothy Castiglia, Anirban Das 0004, Shiqiang Wang 0001, Stacy Patterson
ICML1
2022 Cross-Silo Federated Learning for Multi-Tier Networks with Vertical and Horizontal Data Partitioning
abstract
We consider federated learning in tiered communication networks. Our network model consists of a set of silos, each holding a vertical partition of the data. Each silo contains a hub and a set of clients, with the silo’s vertical data shard partitioned horizontally across its clients. We propose Tiered Decentralized Coordinate Descent (TDCD), a communication-efficient decentralized training algorithm for such two-tiered networks. The clients in each silo perform multiple local gradient steps before sharing updates with their hub to reduce communication overhead. Each hub adjusts its coordinates by averaging its workers’ updates, and then hubs exchange intermediate updates with one another. We present a theoretical analysis of our algorithm and show the dependence of the convergence rate on the number of vertical partitions and the number of local updates. We further validate our approach empirically via simulation-based experiments using a variety of datasets and objectives.
Anirban Das 0004, Timothy Castiglia, Shiqiang Wang 0001, Stacy Patterson
ACM Trans. Intell. Syst. Technol.2
2021 Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks
Timothy Castiglia, Anirban Das 0004, Stacy Patterson
ICLR1
2020 A Hierarchical Model for Fast Distributed Consensus in Dynamic Networks
abstract
State machine replication is a foundational tool that is used to provide availability and fault tolerance in distributed systems. Safe replication requires a consensus algorithm as a method to achieve agreement on the order of system updates. Designing algorithms that are safe while retaining high throughput is imperative for today's systems. Two of the most widely adopted consensus algorithms, Paxos [1] and Raft [2] , have received much attention and use in industry. While both Paxos and Raft provide safe specification for maintaining a replicated log of system updates, Raft aims for ease of understandability and implementation.
Timothy Castiglia, Colin Goldberg, Stacy Patterson
ICDCS1