Reza Bosagh Zadeh

dblp:61/10451 · also Reza Zadeh · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Recommender systems · 38% Graph data management · 38% Data mining · 19%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 69% Cloud and datacenter computing · 21% Distributed systems · 10%
Theoretical computer science
2 papers
Mathematical optimization · 68% Algorithms and data structures · 32%
Artificial intelligence
2 papers
Efficient and distributed learning · 54% Learning theory · 23% Information extraction and text analysis · 23%
Human-computer interaction and pervasive computing
2 papers
Collaborative and social computing · 87% Design research and methods · 13%

Topics — the 18 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
machine learning libraries
0.212016
MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016
High-performance computing › numerical linear algebra
distributed matrix computation
0.212016
Matrix Computations and Optimization in Apache Spark · KDD 2016
High-performance computing › numerical linear algebra
singular value decomposition
0.212016
Matrix Computations and Optimization in Apache Spark · KDD 2016
Mathematical optimization › continuous optimization
convex optimization
0.212016
Matrix Computations and Optimization in Apache Spark · KDD 2016
Mathematical optimization › continuous optimization › matrix optimization › matrix recovery
matrix completion
0.212015
Matrix completion and low-rank SVD via fast alternating least squares · J. Mach. Learn. Res. 2015
Algorithms and data structures
numerical linear algebra
0.212015
Matrix completion and low-rank SVD via fast alternating least squares · J. Mach. Learn. Res. 2015
Recommender systems
graph-based recommendation
0.212013
WTF: the who to follow service at Twitter · WWW 2013
Graph data management › graph processing
graph processing systems
0.212013
WTF: the who to follow service at Twitter · WWW 2013
Graph data management › graph processing › graph processing systems
in-memory graph processing
0.212013
WTF: the who to follow service at Twitter · WWW 2013
Data mining
similarity computation
0.212013
Dimension independent similarity computation · J. Mach. Learn. Res. 2013
Recommender systems
user recommendation
0.212013
WTF: the who to follow service at Twitter · WWW 2013
Collaborative and social computing
computer-supported cooperative work
0.112011
Research team integration: what it is and why it matters · CSCW 2011
Machine learning › Learning theory
clustering theory
0.112010
Supervised Clustering · NIPS 2010
Natural language and speech › Information extraction and text analysis
supervised clustering
0.112010
Supervised Clustering · NIPS 2010
Cloud and datacenter computing
cluster computing framework
0.112016
Matrix Computations and Optimization in Apache Spark · KDD 2016
Cloud and datacenter computing
cluster resource management and scheduling
0.112016
MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016
Distributed systems
distributed data processing
0.112016
MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016
Web and social media mining
social network analysis
0.012013
WTF: the who to follow service at Twitter · WWW 2013

Methods — techniques the papers use, named apart from their topics

linear algebra primitives · 0.5distributed optimization · 0.5distributed matrix operations · 0.5convex optimization · 0.5qualitative analysis · 0.2alternating least squares · 0.2sketching · 0.2random walk · 0.2hashing · 0.2SALSA · 0.2retrospective interviews · 0.1retrospective interview · 0.1single-linkage clustering · 0.1query-efficient algorithms · 0.1
YearPublicationVenuePosition
2016 Matrix Computations and Optimization in Apache Spark
abstract
We describe matrix computations available in the cluster programming framework, Apache Spark. Out of the box, Spark provides abstractions and implementations for distributed matrices and optimization routines using these matrices. When translating single-node algorithms to run on a distributed cluster, we observe that often a simple idea is enough: separating matrix operations from vector operations and shipping the matrix operations to be ran on the cluster, while keeping vector operations local to the driver. In the case of the Singular Value Decomposition, by taking this idea to an extreme, we are able to exploit the computational power of a cluster, while running code written decades ago for a single core. Another example is our Spark port of the popular TFOCS optimization package, originally built for MATLAB, which allows for solving Linear programs as well as a variety of other convex programs. We conclude with a comprehensive set of benchmarks for hardware accelerated matrix computations from the JVM, which is interesting in its own right, as many cluster programming frameworks use the JVM. The contributions described in this paper are already merged into Apache Spark and available on Spark installations by default, and commercially supported by a slew of companies which provide further services.
Reza Bosagh Zadeh, Alexander Ulanov, Burak Yavuz, Li Pu, Shivaram Venkataraman, Evan Randall Sparks, Aaron Staple, Matei Zaharia
KDD1
2016 MLlib: Machine Learning in Apache Spark
abstract
Apache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open- source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes several underlying statistical, optimization, and linear algebra primitives. Shipped with Spark, MLlib supports several languages and provides a high-level API that leverages Spark's rich ecosystem to simplify the development of end-to-end machine learning pipelines. MLlib has experienced a rapid growth due to its vibrant open-source community of over 140 contributors, and includes extensive documentation to support further growth and to let users quickly get up to speed.
Joseph K. Bradley, Burak Yavuz, Evan Randall Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, D. B. Tsai, Manish Amde, Sean Owen, Doris Xin, Reynold Xin, Michael J. Franklin, Reza Bosagh Zadeh, Matei Zaharia, Ameet Talwalkar
J. Mach. Learn. Res.14
2015 Matrix completion and low-rank SVD via fast alternating least squares
Trevor J. Hastie, Rahul Mazumder, Jason D. Lee, Reza Bosagh Zadeh
J. Mach. Learn. Res.4
2013 WTF: the who to follow service at Twitter
abstract
WTF ("Who to Follow") is Twitter's user recommendation service, which is responsible for creating millions of connections daily between users based on shared interests, common connections, and other related factors. This paper provides an architectural overview and shares lessons we learned in building and running the service over the past few years. Particularly noteworthy was our design decision to process the entire Twitter graph in memory on a single server, which significantly reduced architectural complexity and allowed us to develop and deploy the service in only a few months. At the core of our architecture is Cassovary, an open-source in-memory graph processing engine we built from scratch for WTF. Besides powering Twitter's user recommendations, Cassovary is also used for search, discovery, promoted products, and other services as well. We describe and evaluate a few graph recommendation algorithms implemented in Cassovary, including a novel approach based on a combination of random walks and SALSA. Looking into the future, we revisit the design of our architecture and comment on its limitations, which are presently being addressed in a second-generation system under development.
Pankaj Gupta 0002, Ashish Goel, Jimmy Lin, Aneesh Sharma, Reza Bosagh Zadeh
WWW6
2013 Dimension independent similarity computation
Reza Bosagh Zadeh, Ashish Goel
J. Mach. Learn. Res.1
2011 What's in a move?: normal disruption and a design challenge
abstract
The CHI community has led efforts to support teamwork, but has neglected team disruption, as may occur if team members relocate to another institution. We studied moves in 548 interdisciplinary research projects with 2691 researchers (PIs). Moves, and thus disruptions, were not rare, especially in large distributed projects. Overall, one-third of all projects experienced at least one member relocating but most moves reflected churn across high-ranking institutions. When collaborators moved, the project was disrupted. Our data suggest that moves exemplify normal disruptions. A design challenge is to help projects adapt to disruption.
Reza Bosagh Zadeh, Aruna D. Balakrishnan, Sara B. Kiesler, Jonathon N. Cummings
CHI1
2011 Research team integration: what it is and why it matters
abstract
Science policy across the world emphasizes the desirability of research teams that can integrate diverse perspectives and expertise into new knowledge, methods, and products. However, integration in research work is not well understood. Based on retrospective interviews with 55 researchers from 52 diverse research projects, we categorized teams as co-acting (50%), coordinated (15%), and integrated (35%). Integration, when it existed, usually began when PIs chose collaborators and pursued integration throughout the project. We describe researchers' experiences and research climates that discouraged or encouraged integration. Implications for policy choices and design include changes in team structuring and technology support.
Aruna D. Balakrishnan, Sara B. Kiesler, Jonathon N. Cummings, Reza Bosagh Zadeh
CSCW4
2010 Supervised Clustering
abstract
Despite the ubiquity of clustering as a tool in unsupervised learning, there is not yet a consensus on a formal theory, and the vast majority of work in this direction has focused on unsupervised clustering. We study a recently proposed framework for supervised clustering where there is access to a teacher. We give an improved generic algorithm to cluster any concept class in that model. Our algorithm is query-efficient in the sense that it involves only a small amount of interaction with the teacher. We also present and study two natural generalizations of the model. The model assumes that the teacher response to the algorithm is perfect. We eliminate this limitation by proposing a noisy model and give an algorithm for clustering the class of intervals in this noisy model. We also propose a dynamic model where the teacher sees a random subset of the points. Finally, for datasets satisfying a spectrum of weak to strong properties, we give query bounds, and show that a class of clustering functions containing Single-Linkage will find the target clustering under the strongest property.
Pranjal Awasthi, Reza Bosagh Zadeh
NIPS2
2009 Building Strong Multilingual Aligned Corpora
Reza Bosagh Zadeh
EAMT1
2009 A Uniqueness Theorem for Clustering
Reza Bosagh Zadeh, Shai Ben-David
UAI1