VLDB 2026 Research / reviewers in the wild / expert
Baturalp Buyukates
dblp:230/4023
· DBLP profile ↗
17ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-5941-0667ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aggressive, Imperceptible, or Both: Architecture-Aware Hybrid Byzantines in Federated Learning
Emre Ozfatura, Kerem Ozfatura, Baturalp Buyukates, Mert Coskuner, Alptekin Küpçü, Deniz Gündüz |
EuroS&P | 3 |
| 2026 | Balancing Information Accuracy and Response Timeliness in Networked LLMsabstractRecent advancements in Large Language Models (LLMs) have transformed many fields including scientific discovery, content generation, biomedical text mining, and educational technology. However, the substantial requirements for training data, computational resources, and energy consumption pose significant challenges for their practical deployment. A promising alternative is to leverage smaller, specialized language models and aggregate their outputs to improve overall response quality. In this work, we investigate a networked LLM system composed of multiple users, a central task processor, and clusters of topic-specialized LLMs. Each user submits categorical binary (true/false) queries, which are routed by the task processor to a selected cluster of $m$ LLMs. After gathering individual responses, the processor returns a final aggregated answer to the user. We characterize both the information accuracy and response timeliness in this setting, and formulate a joint optimization problem to balance these two competing objectives. Our extensive simulations demonstrate that the aggregated responses consistently achieve higher accuracy than those of individual LLMs. Notably, this improvement is more significant when the participating LLMs exhibit similar standalone performance. Yigit Turkmen, Baturalp Buyukates, Melih Bastopcu |
INFOCOM | 2 |
| 2025 | Reconsidering LLM Uncertainty Estimation Methods in the WildabstractYavuz Faruk Bakman, Duygu Nur Yaldiz, Sungmin Kang, Tuo Zhang, Baturalp Buyukates, Salman Avestimehr, Sai Praneeth Karimireddy. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yavuz Faruk Bakman, Duygu Nur Yaldiz, Sungmin Kang, Baturalp Buyukates, Amir Salman Avestimehr, Sai Praneeth Karimireddy |
ACL (1) | 5 |
| 2025 | FedGrAINS: Personalized SubGraph Federated Learning with AdaptIve Neighbor SamplingabstractGraphs are crucial for modeling relational and biological data. As datasets grow larger in real-world scenarios, the risk of exposing sensitive information increases, making privacy-preserving training methods like federated learning (FL) essential to ensure data security and compliance with privacy regulations. Recently proposed personalized subgraph FL methods have become the de-facto standard for training personalized Graph Neural Networks (GNNs) in a federated manner while dealing with the missing links across clients’ subgraphs due to privacy restrictions. However, personalized subgraph FL faces significant challenges due to the heterogeneity in client subgraphs, such as degree distributions among the nodes, which complicate federated training of graph models. To address these challenges, we propose FedGrAINS, a novel data-adaptive and sampling-based regularization method for subgraph FL. FedGrAINS leverages generative flow networks (GFlowNets) to evaluate node importance concerning clients’ tasks, dynamically adjusting the message-passing step in clients’ GNNs. This adaptation reflects task-optimized sampling aligned with a trajectory balance objective. Experimental results demonstrate that the inclusion of FedGrAINS as a regularizer consistently improves the FL performance compared to baselines that do not leverage such regularization. Emir Ceyani, Baturalp Buyukates, Carl Yang 0001, Amir Salman Avestimehr |
SDM | 3 |
| 2024 | MARS: Meaning-Aware Response Scoring for Uncertainty Estimation in Generative LLMsabstractYavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Salman Avestimehr. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Chenyang Tao, Dimitrios Dimitriadis, Amir Salman Avestimehr |
ACL (1) | 3 |
| 2024 | Predicting Uncertainty of Generative LLMs with MARS: Meaning-Aware Response ScoringabstractGenerative Large Language Models (LLMs) have recently been widely utilized for their unprecedented capabil-ities across many tasks. Considering their use in high-stakes environments and for mission-critical applications, the fact that LLMs often can generate inaccurate or misleading results can be potentially harmful, which motivates us to study the correctness of generative LLM outputs. Uncertainty Estimation (UE) in generative LLMs is a developing area, with state-of-the-art probability-based techniques frequently using length-normalized scoring. As an alternative to length-normalized scoring in UE, in this work, we propose Meaning-Aware Response Scoring (MARS). The key idea of MARS is to consider the semantic contribution of each token of the generated sequence to the context of the question during UE. Through extensive experiments on three question-answering datasets across five pretrained LLMs, we show that utilizing MARS during UE results in a universal and significant improvement in UE performance. Yavuz Faruk Bakman, Duygu Nur Yaldiz, Baturalp Buyukates, Amir Salman Avestimehr, Chenyang Tao, Dimitrios Dimitriadis |
ISIT | 3 |
| 2024 | Frequency Domain Diffusion Model with Scale-Dependent Noise ScheduleabstractDiffusion models have played a crucial role in the recent advancements in generative image modeling. These models are characterized by a forward process that incrementally corrupts images. The modeling objective is to develop a reverse process capable of reconstructing the original image from degraded inputs so that the trained model can then be leveraged to generate natural images from pure noise. In this work, we introduce a novel diffusion process that operates in the frequency domain. Typically, the frequency domain representation of an image exhibits a sparse structure, with energy predominantly concentrated in low frequency components. This inherent sparsity aids us in the effective separation of signal and noise during the reverse process. We utilize this property to introduce a scale-dependent noise schedule, offering precise control over various image scales. Working in the frequency domain allows us to modify the training protocol, resulting in significant computation enhancements, achieving a speedup of 2.7-8.5 x without a significant drop in generated image quality, compared to the image domain models, which operate with fixed noise schedules. Amir Ziashahabi, Baturalp Buyukates, Artan Sheshmani, Yi-Zhuang You, Amir Salman Avestimehr |
ISIT | 2 |
| 2024 | FedSecurity: A Benchmark for Attacks and Defenses in Federated Learning and Federated LLMsabstractThis paper introduces FedSecurity, an end-to-end benchmark that serves as a supplementary component of the FedML library for simulating adversarial attacks and corresponding defense mechanisms in Federated Learning (FL). FedSecurity eliminates the need for implementing the fundamental FL procedures, e.g., FL training and data loading, from scratch, thus enables users to focus on developing their own attack and defense strategies. It contains two key components, including FedAttacker that conducts a variety of attacks during FL training, and FedDefender that implements defensive mechanisms to counteract these attacks. FedSecurity has the following features: i) It offers extensive customization options to accommodate a broad range of machine learning models (e.g., Logistic Regression, ResNet, and GAN) and FL optimizers (e.g., FedAVG, FedOPT, and FedNOVA); ii) it enables exploring the effectiveness of attacks and defenses across different datasets and models; and iii) it supports flexible configuration and customization through a configuration file and some APIs. We further demonstrate FedSecurity's utility and adaptability through federated training of Large Language Models (LLMs) to showcase its potential on a wide range of complex applications. Baturalp Buyukates, Zijian Hu 0001, Weizhao Jin, Lichao Sun 0001, Chulin Xie, Yuhang Yao 0003, Kai Zhang 0039, Qifan Zhang 0002, Carlee Joe-Wong, Amir Salman Avestimehr, Chaoyang He 0001 |
KDD | 2 |
| 2023 | Gradient Coding With Dynamic Clustering for Straggler-Tolerant Distributed LearningabstractDistributed implementations are crucial in speeding up large scale machine learning applications. Distributed gradient descent (GD) is widely employed to parallelize the learning task by distributing the dataset across multiple workers. A significant performance bottleneck for the per-iteration completion time in distributed synchronous GD is straggling workers. Coded distributed computation techniques have been introduced recently to mitigate stragglers and to speed up GD iterations by assigning redundant computations to workers. In this paper, we introduce a novel paradigm of dynamic coded computation, which assigns redundant data to workers to acquire the flexibility to dynamically choose from among a set of possible codes depending on the past straggling behavior. In particular, we propose gradient coding (GC) with dynamic clustering, called GC-DC, and regulate the number of stragglers in each cluster by dynamically forming the clusters at each iteration. With time-correlated straggling behavior, GC-DC adapts to the straggling behavior over time; in particular, at each iteration, GC-DC aims at distributing the stragglers across clusters as uniformly as possible based on the past straggler behavior. For both homogeneous and heterogeneous worker models, we numerically show that GC-DC provides significant improvements in the average per-iteration completion time without an increase in the communication load compared to the original GC scheme. Baturalp Buyukates, Emre Ozfatura, Sennur Ulukus, Deniz Gündüz |
IEEE Trans. Commun. | 1 |
| 2021 | Gradient Coding with Dynamic Clustering for Straggler MitigationabstractIn distributed synchronous gradient descent (GD) the main performance bottleneck for the per-iteration completion time is the slowest straggling workers. To speed up GD iterations in the presence of stragglers, coded distributed computation techniques are implemented by assigning redundant computations to workers. In this paper, we propose a novel gradient coding (GC) scheme that utilizes dynamic clustering, denoted by GC-DC, to speed up gradient calculations. Under time-correlated straggling behavior, GC-DC aims at regulating the number of straggling workers in each cluster based on the straggler behavior in the previous iteration. We numerically show that GC-DC provides significant improvements in the average completion time (of each iteration) with no increase in the communication load compared to the original GC scheme. Baturalp Buyukates, Emre Ozfatura, Sennur Ulukus, Deniz Gündüz |
ICC | 1 |
| 2021 | Selective Encoding Policies for Maximizing Information FreshnessabstractAn information source generates independent and identically distributed status update messages from an observed random phenomenon which takes n distinct values based on a given probability mass function (PMF). These update packets are encoded at the transmitter node to be sent to a receiver node which wants to track the observed random variable with as little age as possible. The transmitter node implements a selective k encoding policy such that rather than encoding all possible n realizations, the transmitter node encodes the most probable k realizations. We consider three different policies regarding the remaining n-k less probable realizations: highest k selective encoding which disregards whenever a realization from the remaining n-k values occurs; randomized selective encoding which encodes and sends the remaining n-k realizations with a certain probability to further inform the receiver node at the expense of longer codewords for the selected k realizations; and highest k selective encoding with an empty symbol which sends a designated empty symbol when one of the remaining n-k realizations occurs. For all of these three encoding schemes, we find the average age and determine the age-optimal real codeword lengths, including the codeword length for the empty symbol in the case of the latter scheme, such that the average age at the receiver node is minimized. Through numerical evaluations for arbitrary PMFs, we show that these selective encoding policies result in a lower average age than encoding every realization, and find the corresponding age-optimal k values. Since we focus on real-valued codeword lengths in this paper, the resulting age value obtained in each case studied here serves as a lower bound to what can be attained by integer-valued codeword lengths in that case. Melih Bastopcu, Baturalp Buyukates, Sennur Ulukus |
IEEE Trans. Commun. | 2 |
| 2021 | Scaling Laws for Age of Information in Wireless NetworksabstractWe study age of information in a multiple source-multiple destination setting with a focus on its scaling in large wireless networks. There are n nodes uniformly and independently distributed on a fixed area that are randomly paired with each other to form n source-destination (S-D) pairs. Each source node wants to keep its destination node as up-to-date as possible. To accommodate successful communication between all n S-D pairs, we first propose a three-phase transmission scheme which utilizes local cooperation between the nodes along with what we call mega update packets to serve multiple S-D pairs at once. We show that under the proposed scheme average age of an S-D pair scales as O(n1/4logn) as the number of users, n, in the network grows. Next, we observe that communications that take place in Phases I and III of the proposed scheme are scaled-down versions of network-level communications. With this along with scale-invariance of the system, we introduce hierarchy to improve this scaling result and show that when hierarchical cooperation between users is utilized, an average age scaling of O(nα(h)logn) per-user is achievable, where h denotes the number of hierarchy levels and α(h) = 1/3·2h+1. We note that α(h) tends to 0 as h increases, and asymptotically, the average age scaling of the proposed hierarchical scheme is O(logn). To the best of our knowledge, this is the best average age scaling result in a status update system with multiple S-D pairs. Baturalp Buyukates, Alkan Soysal, Sennur Ulukus |
IEEE Trans. Wirel. Commun. | 1 |
| 2020 | Age-Based Coded Computation for Bias Reduction in Distributed LearningabstractCoded computation can speed up distributed learning in the presence of straggling workers. Partial recovery of the gradient vector can further reduce the computation time at each iteration; however, this can result in biased estimators, which may slow down convergence, or even cause divergence. Estimator bias is particularly prevalent when the straggling behavior is correlated over time, which results in the gradient estimators being dominated by a few fast servers. To mitigate biased estimators, we design a timely dynamic encoding framework for partial recovery that includes an ordering operator that changes the codewords and computation orders at workers over time. To regulate the recovery frequencies, we adopt an age metric in the design of the dynamic encoding scheme. The proposed age-based scheme prioritizes the recovery of computations with relatively large age. We show through numerical results that the proposed dynamic encoding strategy increases the timeliness of the recovered computations, which, as a result, reduces the bias in model updates, and accelerates the convergence compared to conventional static partial recovery schemes. Emre Ozfatura, Baturalp Buyukates, Deniz Gündüz, Sennur Ulukus |
GLOBECOM | 2 |
| 2020 | Optimal Selective Encoding for Timely Updates with Empty SymbolabstractAn information source generates independent and identically distributed status update messages from an observed random phenomenon which takes n distinct values based on a given pmf. These update packets are encoded at the transmitter to be sent to a receiver which wants to track the observed random variable with as little age as possible. The transmitter implements a selective k encoding policy such that rather than encoding all possible n realizations, the transmitter encodes the most probable k realizations and sends a designated empty symbol when one of the remaining n-k realizations occurs. We consider two scenarios: when the empty symbol does not reset the age and when the empty symbol resets the age. We find the time average age of information and the age-optimal real codeword lengths, including the codeword length for the empty symbol, for both of these scenarios. Through numerical evaluations for arbitrary pmfs, we show that this selective encoding policy yields a lower age at the receiver than encoding every realization and find the corresponding age-optimal k values. Baturalp Buyukates, Melih Bastopcu, Sennur Ulukus |
ISIT | 1 |
| 2020 | Timely Distributed Computation With StragglersabstractWe consider a status update system in which the update packets need to be processed to extract the embedded useful information. The source node sends the acquired information to a computation unit (CU) which consists of a master node and n worker nodes. The master node distributes the received computation task to the worker nodes. Upon computation, the master node aggregates the results and sends them back to the source node to keep it updated. We investigate the age performance of uncoded and coded (repetition coded, MDS coded, and multi-message MDS (MM-MDS) coded) schemes in the presence of stragglers under i.i.d. exponential transmission delays and i.i.d shifted exponential computation times. We show that asymptotically MM-MDS coded scheme outperforms the other schemes. Furthermore, we characterize the optimal codes such that the average age is minimized. Baturalp Buyukates, Sennur Ulukus |
IEEE Trans. Commun. | 1 |
| 2019 | Age of Information Scaling in Large Networks with Hierarchical CooperationabstractGiven n randomly located source-destination (S- D) pairs on a fixed area network that want to communicate with each other, we study the age of information with a particular focus on its scaling as the network size n grows. We propose a three- phase transmission scheme that utilizes hierarchical cooperation between users along with mega update packets and show that an average age scaling of O(nα(h)log n) per-user is achievable where h denotes the number of hierarchy levels and α(h) = 1 / (3.2h+1) which tends to 0 as h increases such that asymptotically average age scaling of the proposed scheme is O(log n). To the best of our knowledge, this is the best average age scaling result in a status update system with multiple S-D pairs. Baturalp Buyukates, Alkan Soysal, Sennur Ulukus |
GLOBECOM | 1 |
| 2019 | Age of Information Scaling in Large NetworksabstractWe study age of information in a multiple source-multiple destination setting with a focus on its scaling in large wireless networks. There are n nodes that are randomly paired with each other on a fixed area to form n source-destination (SD) pairs. We propose a three-phase transmission scheme which utilizes local cooperation between the nodes by forming what we call mega update packets to serve multiple S-D pairs at once. We show that under the proposed scheme average age of an S-D pair scales as O(n1/4log n) as the number of users, n, in the network grows. To the best of our knowledge, this is the best age scaling result for a multiple source-multiple destination setting. Baturalp Buyukates, Alkan Soysal, Sennur Ulukus |
ICC | 1 |