Balaji Varadarajan

dblp:34/2065 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 61% Empirical software engineering · 30% Software testing · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 50% Storage systems · 33% Cloud and datacenter computing · 17%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
change management
0.412019
Keeping Master Green at Scale · EuroSys 2019
Software maintenance and evolution › release engineering
continuous integration
0.412019
Keeping Master Green at Scale · EuroSys 2019
Empirical software engineering
mining software repositories
0.412019
Keeping Master Green at Scale · EuroSys 2019
Cloud and datacenter computing
datacenter infrastructure
0.112012
Data Infrastructure at LinkedIn · ICDE 2012
Storage systems › key-value storage
distributed key-value store
0.112012
Data Infrastructure at LinkedIn · ICDE 2012
Distributed systems
fault tolerance
0.112012
Data Infrastructure at LinkedIn · ICDE 2012
Storage systems
key-value storage
0.112012
Data Infrastructure at LinkedIn · ICDE 2012
Distributed systems › middleware
message systems
0.112012
Data Infrastructure at LinkedIn · ICDE 2012
Distributed systems
replication
0.112012
Data Infrastructure at LinkedIn · ICDE 2012

Methods — techniques the papers use, named apart from their topics

distributed storage · 0.1change data capture · 0.1
YearPublicationVenuePosition
2019 Keeping Master Green at Scale
abstract
Giant monolithic source-code repositories are one of the fundamental pillars of the back end infrastructure in large and fast-paced software companies. The sheer volume of everyday code changes demands a reliable and efficient change management system with three uncompromisable key requirements --- always green master, high throughput, and low commit turnaround time. Green refers to a master branch that always successfully compiles and passes all build steps, the opposite being red. A broken master (red) leads to delayed feature rollouts because a faulty code commit needs to be detected and rolled backed. Additionally, a red master has a cascading effect that hampers developer productivity--- developers might face local test/build failures, or might end up working on a codebase that will eventually be rolled back.
Sundaram Ananthanarayanan, Masoud Saeida Ardekani, Denis Haenikel, Balaji Varadarajan, Simon Soriano, Ali-Reza Adl-Tabatabai
EuroSys4
2012 All aboard the Databus!: Linkedin's scalable consistent change data capture platform
abstract
In Internet architectures, data systems are typically categorized into source-of-truth systems that serve as primary stores for the user-generated writes, and derived data stores or indexes which serve reads and other complex queries. The data in these secondary stores is often derived from the primary data through custom transformations, sometimes involving complex processing driven by business logic. Similarly data in caching tiers is derived from reads against the primary data store, but needs to get invalidated or refreshed when the primary data gets mutated. A fundamental requirement emerging from these kinds of data architectures is the need to reliably capture, flow and process primary data changes.
Shirshanka Das, Chavdar Botev, Kapil Surlaker, Bhaskar Ghosh, Balaji Varadarajan, Sunil Nagaraj, David Zhang 0001, Jemiah Westerman, Phanindra Ganti, Boris Shkolnik, Sajid Topiwala, Alexander Pachev, Naveen Somasundaram, Subbu Subramaniam
SoCC5
2012 Data Infrastructure at LinkedIn
abstract
Linked In is among the largest social networking sites in the world. As the company has grown, our core data sets and request processing requirements have grown as well. In this paper, we describe a few selected data infrastructure projects at Linked In that have helped us accommodate this increasing scale. Most of those projects build on existing open source projects and are themselves available as open source. The projects covered in this paper include: (1) Voldemort: a scalable and fault tolerant key-value store, (2) Data bus: a framework for delivering database changes to downstream applications, (3) Espresso: a distributed data store that supports flexible schemas and secondary indexing, (4) Kafka: a scalable and efficient messaging system for collecting various user activity events and log data.
Aditya Auradkar, Chavdar Botev, Shirshanka Das, Dave De Maagd, Alex Feinberg, Phanindra Ganti, Bhaskar Ghosh, Kishore Gopalakrishna, Brendan Harris, Joel Koshy, Kevin Krawez, Jay Kreps, Shi Lu, Sunil Nagaraj, Neha Narkhede, Sasha Pachev, Igor Perisic, Lin Qiao, Tom Quiggle, Jun Rao, Bob Schulman, Abraham Sebastian, Oliver Seeliger, Adam Silberstein, Boris Shkolnik, Chinmay Soman, Roshan Sumbaly, Kapil Surlaker, Sajid Topiwala, Cuong Tran 0003, Balaji Varadarajan, Jemiah Westerman, Zach White, David Zhang 0001
ICDE32
2006 Utility scoring of product reviews
abstract
We identify a new task in the ongoing research in text sentiment analysis: predicting utility of product reviews, which is orthogonal to polarity classification and opinion extraction. We build regression models by incorporating a diverse set of features, and achieve highly competitive performance for utility scoring on three real-world data sets.
Balaji Varadarajan
CIKM2