Suhrid Balakrishnan

dblp:88/3442 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
0since 2021 · last 2013
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 6 first-authorDatabases, data management, data science and information retrieval · 6 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Computer networks · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Recommender systems · 70% Information retrieval · 30%
Network and information security
1 paper
Web and mobile security · 50% Hardware security and side channels · 50%
Theoretical computer science
2 papers
Mathematical optimization · 67% Algorithms and data structures · 33%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational social science and digital humanities · 100%
Human-computer interaction and pervasive computing
1 paper
Ubiquitous computing and smart environments · 100%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Recommender systems
collaborative filtering
0.112012
Collaborative ranking · WSDM 2012
Information retrieval › evaluation › effectiveness metrics
discounted cumulative gain
0.112012
Collaborative ranking · WSDM 2012
Recommender systems
ranking-based recommendation
0.112012
Collaborative ranking · WSDM 2012
Web and mobile security
mobile security
0.112012
Tapprints: your finger taps have fingerprints · MobiSys 2012
Hardware security and side channels
side-channel attack
0.112012
Tapprints: your finger taps have fingerprints · MobiSys 2012
Recommender systems › collaborative filtering
matrix factorization
0.112010
Two of a Kind or the Ratings Game? Adaptive Pairwise Preferences and Latent Factor Models · ICDM 2010
Machine learning › Efficient and distributed learning
large-scale learning
0.112008
Algorithms for Sparse Linear Classifiers in the Massive Data Setting · J. Mach. Learn. Res. 2008
Mathematical optimization › statistical estimation › regression › sparse regression
solution path algorithm
0.112007
Finding Predictive Runs with LAPS · ICDM 2007
Mathematical optimization › statistical estimation › regression
sparse regression
0.112007
Finding Predictive Runs with LAPS · ICDM 2007
Computational social science and digital humanities
marketing
0.012012
Computational Television Advertising · ICDM 2012
Information retrieval › ranking
learning to rank
0.012012
Collaborative ranking · WSDM 2012
Ubiquitous computing and smart environments
mobile sensing
0.012012
Tapprints: your finger taps have fingerprints · MobiSys 2012
Recommender systems
pairwise preference learning
0.012010
Two of a Kind or the Ratings Game? Adaptive Pairwise Preferences and Latent Factor Models · ICDM 2010

Methods — techniques the papers use, named apart from their topics

mathematical optimization · 0.3machine learning · 0.3gyroscope · 0.3accelerometer · 0.3sparse linear classifiers · 0.2large-scale optimization · 0.2matrix factorization · 0.1learning to rank · 0.1information gain · 0.1bayesian framework · 0.1lasso · 0.1group lasso · 0.1fused lasso · 0.1
YearPublicationVenuePosition
2013 Detecting hidden enemy lines in IP address space
abstract
If an outbound flow is observed at the boundary of a protected network, destined to an IP address within a few addresses of a known malicious IP address, should it be considered a suspicious flow? Conventional blacklisting is not going to cut it in this situation, and the established fact that malicious IP addresses tend to be highly clustered in certain portions of IP address space, should indeed raise suspicions. We present a new approach for perimeter defense that addresses this concern. At the heart of our approach, we attempt to infer internal, hidden boundaries in IP address space, that lie within publicly known boundaries of registered IP netblocks. Our hypothesis is that given a known bad IP address, other IP address in the same internal contiguous block are likely to share similar security properties, and may therefore be vulnerable to being similarly hacked and used by attackers in the future. In this paper, we describe how we infer hidden internal boundaries in IPv4 netblocks, and what effect this has on being able to predict malicious IP addresses.
Suhas Mathur, Baris Coskun, Suhrid Balakrishnan
NSPW3
2012 Computational Television Advertising
abstract
Ever wonder why that Kia Ad ran during Iron Chef? Traditional advertising methodology on television is a fascinating mix of marketing, branding, measurement, and predictive modeling. While still a robust business, it is at risk with the recent growth of online and time-shifted (recorded) television. A particular issue is that traditional methods for television advertising are far less efficient than their counterparts in the online world which employ highly sophisticated computational techniques. This paper formalizes an approach to eliminate some of these inefficiencies by recasting the process of television advertising media campaign generation in a computational framework. We describe efficient mathematical approaches to solve for the task of finding optimal campaigns for specific target audiences. In two case studies, our campaigns report gains in key operational metrics of up to 56% compared to campaigns generated by traditional methods.
Suhrid Balakrishnan, Sumit Chopra, David L. Applegate, Simon Urbanek
ICDM1
2012 Tapprints: your finger taps have fingerprints
abstract
This paper shows that the location of screen taps on modern smartphones and tablets can be identified from accelerometer and gyroscope readings. Our findings have serious implications, as we demonstrate that an attacker can launch a background process on commodity smartphones and tablets, and silently monitor the user's inputs, such as keyboard presses and icon taps. While precise tap detection is nontrivial, requiring machine learning algorithms to identify fingerprints of closely spaced keys, sensitive sensors on modern devices aid the process. We present TapPrints, a framework for inferring the location of taps on mobile device touch-screens using motion sensor data combined with machine learning analysis. By running tests on two different off-the-shelf smartphones and a tablet computer we show that identifying tap locations on the screen and inferring English letters could be done with up to 90% and 80% accuracy, respectively. By optimizing the core tap detection capability with additional information, such as contextual priors, we are able to further magnify the core threat.
Emiliano Miluzzo, Alexander Varshavsky, Suhrid Balakrishnan, Romit Roy Choudhury
MobiSys3
2012 Collaborative ranking
abstract
Typical recommender systems use the root mean squared error (RMSE) between the predicted and actual ratings as the evaluation metric. We argue that RMSE is not an optimal choice for this task, especially when we will only recommend a few (top) items to any user. Instead, we propose using a ranking metric, namely normalized discounted cumulative gain (NDCG), as a better evaluation metric for this task. Borrowing ideas from the learning to rank community for web search, we propose novel models which approximately optimize NDCG for the recommendation task. Our models are essentially variations on matrix factorization models where we also additionally learn the features associated with the users and the items for the ranking task. Experimental results on a number of standard collaborative filtering data sets validate our claims. The results also show the accuracy and efficiency of our models and the benefits of learning features for ranking.
Suhrid Balakrishnan, Sumit Chopra
WSDM1
2012 Two of a kind or the ratings game? Adaptive pairwise preferences and latent factor models
Suhrid Balakrishnan, Sumit Chopra
Frontiers Comput. Sci.1
2010 Two of a Kind or the Ratings Game? Adaptive Pairwise Preferences and Latent Factor Models
abstract
While latent factor models are built using ratings data, which is typically assumed static, the ability to incorporate different kinds of subsequent user feedback is an important asset. For instance, the user might want to provide additional information to the system in order to improve his personal recommendations. To this end, we examine a novel scheme for efficiently learning (or refining) user parameters from such feedback. We propose a scheme where users are presented with a sequence of pair wise preference questions: "Do you prefer item A over B?". User parameters are updated based on their response, and subsequent questions are chosen adaptively after incorporating the feedback. We operate in a Bayesian framework and the choice of questions is based on an information gain criterion. We validate the scheme on the Netflix movie ratings data set. A user study and automated experiments validate our findings.
Suhrid Balakrishnan, Sumit Chopra
ICDM1
2010 On-demand set-based recommendations
abstract
This paper investigates the problem of generating on-demand recommendations over a dataset of items where the input is a selection of a few of the items. As an example in the context of a movie dataset, the user may wish to see a list of movies related to the three animation movies 'Finding Nemo', 'Up' and 'Spirited Away'. In this case, it would be expected that the list returned would contain other animation movies like 'Wall-E', 'Princess Mononoke' etc. Thus, this problem can be viewed as a type of "clustering on demand" problem [1]. It is the set form of input that distinguishes this problem from a standard information retrieval problem where the query is usually a single item or an abstraction of a single item in the dataset. In this paper, we present several new approaches to dealing with this problem. We also show some representative results on a movie text dataset.
Suhrid Balakrishnan
RecSys1
2010 Feature-rich continuous language models for speech recognition
abstract
State-of-the-art probabilistic models of text such as n-grams require an exponential number of examples as the size of the context grows, a problem that is due to the discrete word representation. We propose to solve this problem by learning a continuous-valued and low-dimensional mapping of words, and base our predictions for the probabilities of the target word on non-linear dynamics of the latent space representation of the words in context window. We build on neural networks-based language models; by expressing them as energy-based models, we can further enrich the models with additional inputs such as part-of-speech tags, topic information and graphs of word similarity. We demonstrate a significantly lower perplexity on different text corpora, as well as improved word accuracy rate on speech recognition tasks, as compared to Kneser-Ney back-off n-gram-based language models.
Piotr Mirowski, Sumit Chopra, Suhrid Balakrishnan, Srinivas Bangalore
SLT3
2009 Reinforcement learning for dialog management using least-squares Policy iteration and fast feature selection
abstract
Reinforcement learning (RL) is a promising technique for creating a dialog manager. RL accepts features of the current dialog state and seeks to find the best action given those features. Although it is often easy to posit a large set of potentially useful features, in practice, it is difficult to find the subset which is large enough to contain useful information yet compact enough to reliably learn a good policy. In this paper, we propose a method for RL optimization which automatically performs feature selection. The algorithm is based on least-squares policy iteration, a state-of-the-art RL algorithm which is highly sampleefficient and can learn from a static corpus or on-line. Experiments in dialog simulation show it is more stable than a baseline RL algorithm taken from a working dialog system.
Lihong Li 0001, Jason D. Williams, Suhrid Balakrishnan
INTERSPEECH3
2009 Estimating Probability of Correctness for ASR N-Best Lists
Jason D. Williams, Suhrid Balakrishnan
SIGDIAL Conference2
2008 Algorithms for Sparse Linear Classifiers in the Massive Data Setting
Suhrid Balakrishnan, David Madigan
J. Mach. Learn. Res.1
2007 Finding Predictive Runs with LAPS
abstract
We present an extension to the Lasso [6] for binary classification problems with ordered attributes. Inspired by the Fused Lasso [5] and the Group Lasso [7, 3] models, we aim to both discover and model runs (contiguous subgroups of the variables) that are highly predictive. We call the extended model LAPS (the Lasso with Attribute Partition Search). Such problems commonly arise in financial and medical domains, where predictors are time series variables, for example. This paper outlines the formulation of the problem, an algorithm to obtain the model coefficients and experiments showing applicability to practical problems of this type.
Suhrid Balakrishnan, David Madigan
ICDM1
2006 Decision Trees for Functional Variables
abstract
Classification problems with functionally structured input variables arise naturally in many applications. In a clinical domain, for example, input variables could include a time series of blood pressure measurements. In a financial setting, different time series of stock returns might serve as predictors. In an archaeological application, the 2D profile of an artifact may serve as a key input variable. In such domains, accuracy of the classifier is not the only reasonable goal to strive for; classifiers that provide easily interpretable results are also of value. In this work, we present an intuitive scheme for extending decision trees to handle functional input variables. Our results show that such decision trees are both accurate and readily interpretable.
Suhrid Balakrishnan, David Madigan
ICDM1