Caner Gocmen

dblp:223/3236 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2024
0009-0003-7725-8309ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Performance modeling and evaluation · 55% Cloud and datacenter computing · 32% Distributed systems · 14%
Theoretical computer science
1 paper
Algorithmic game theory and mechanism design · 100%
Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.812024
ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing · OSDI 2024
Performance modeling and evaluation › benchmarking
performance regression testing
0.812024
ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing · OSDI 2024
Algorithmic game theory and mechanism design › market equilibrium
competitive equilibrium
0.712023
Fair Allocation Over Time, with Applications to Content Moderation · KDD 2023
Algorithmic game theory and mechanism design
fair division
0.712023
Fair Allocation Over Time, with Applications to Content Moderation · KDD 2023
Algorithmic game theory and mechanism design
market equilibrium
0.712023
Fair Allocation Over Time, with Applications to Content Moderation · KDD 2023
Distributed systems
anomaly detection
0.312018
A Real-time Framework for Detecting Efficiency Regressions in a Globally Distributed Codebase · KDD 2018
Performance modeling and evaluation
workload characterization
0.312018
A Real-time Framework for Detecting Efficiency Regressions in a Globally Distributed Codebase · KDD 2018
Performance modeling and evaluation
benchmarking
0.212024
ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing · OSDI 2024
Computational social science and digital humanities › platform governance
content moderation
0.212023
Fair Allocation Over Time, with Applications to Content Moderation · KDD 2023

Methods — techniques the papers use, named apart from their topics

leximin allocation · 1.3eisenberg-gale program · 1.3signal processing · 0.7sequential statistics · 0.7
YearPublicationVenuePosition
2024 ServiceLab: Preventing Tiny Performance Regressions at Hyperscale through Pre-Production Testing
Mike Chow, Yang Wang 0009, Ayichew Hailu, Rohan Bopardikar, Jialiang Qu, David Meisner, Santosh Sonawane, Rodrigo Paim, Mack Ward, Ivor Huang, Matt McNally, Daniel Hodges, Zoltan Farkas, Caner Gocmen, Elvis Huang, Chunqiang Tang
OSDI17
2023 Fair Allocation Over Time, with Applications to Content Moderation
abstract
In today's digital world, interaction with online platforms is ubiquitous, and thus content moderation is important for protecting users from content that do not comply with pre-established community guidelines. Given the vast volume of content generated online daily, having an efficient content moderation system throughout every stage of planning is particularly important. We study the short-term planning problem of allocating human content reviewers to different harmful content categories. We use tools from fair division and study the application of competitive equilibrium and leximin allocation rules for addressing this problem. On top of the traditional Fisher market setup, we additionally incorporate novel aspects that are of practical importance. The first aspect is the forecasted workload of different content categories, which puts constraints on the allocation chosen by the planner. We show how a formulation that is inspired by the celebrated Eisenberg-Gale program allows us to find an allocation that not only satisfies the forecasted workload, but also fairly allocates the remaining working hours from the content reviewers among all content categories. A fair allocation of oversupply provides a guardrail in cases where the actual workload deviates from the predicted workload. The second practical consideration is time dependent allocation that is motivated by the fact that partners need scheduling guidance for the reviewers across days to achieve efficiency. To address the time component, we introduce new extensions of the various fair allocation approaches for the single-time period setting, and we show that many properties extend in essence, albeit with some modifications. Lastly, related to the time component, we additionally investigate how to satisfy markets' desire for smooth allocation (i.e, an allocation that does not vary much from time to time) so that the switch in staffing is minimized. We demonstrate the performance of our proposed approaches through real-world data obtained from Meta.
Amine Allouah, Christian Kroer, Vashist Avadhanula, Nona Bohanon, Anil Dania, Caner Gocmen, Sergey Pupyrev, Parikshit Shah, Nicolás E. Stier Moses, Ken Rodríguez Taarup
KDD7
2018 A Real-time Framework for Detecting Efficiency Regressions in a Globally Distributed Codebase
abstract
Multiple teams at Facebook are tasked with monitoring compute and memory utilization metrics that are important for managing the efficiency of the codebase. An efficiency regression is characterized by instances where the CPU utilization or query per second (QPS) patterns of a function or endpoint experience an unexpected increase over its prior baseline. If the code changes responsible for these regressions get propagated to Facebook's fleet of web servers, the impact of the inefficient code will get compounded over billions of executions per day, carrying potential ramifications to Facebook's scaling efforts and the quality of the user experience. With a codebase ingesting in excess of 1,000 diffs across multiple pushes per day, it is important to have a real-time solution for detecting regressions that is not only scalable and high in recall, but also highly precise in order to avoid overrunning the remediation queue with thousands of false positives. This paper describes the end-to-end regression detection system designed and used at Facebook. The main detection algorithm is based on sequential statistics supplemented by signal processing transformations, and the performance of the algorithm was assessed with a mixture of online and offline tests across different use cases. We compare the performance of our algorithm against a simple benchmark as well as a commercial anomaly detection software solution.
Martin Valdez-Vivas, Caner Gocmen, Andrii Korotkov, Ethan Fang, Kapil Goenka, Sherry Chen
KDD2