Brandon Norick

dblp:36/8557 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9Artificial intelligence and machine learning · 7 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Efficient and distributed learning · 44% Graph learning · 31% Representation and self-supervised learning · 25%
Databases, data mining, and information retrieval
5 papers
Data mining · 82% Recommender systems · 10% Information retrieval · 8%

Topics — the 18 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
data curation
0.912025
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset · ACL (1) 2025
Machine learning › Representation and self-supervised learning › pre-training
pretraining data
0.912025
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset · ACL (1) 2025
Data mining › representation learning
hyperedge-based embedding
0.522017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016
Machine learning › Graph learning
heterogeneous graph
0.522018
Active Learning on Heterogeneous Information Networks: A Multi-armed Bandit Approach · ICDM 2018
Personalized entity recommendation: a heterogeneous information network approach · WSDM 2014
Machine learning › Efficient and distributed learning
active learning
0.312018
Active Learning on Heterogeneous Information Networks: A Multi-armed Bandit Approach · ICDM 2018
Machine learning › Efficient and distributed learning › data-efficient learning
label-efficient learning
0.312018
Active Learning on Heterogeneous Information Networks: A Multi-armed Bandit Approach · ICDM 2018
Machine learning › Graph learning › graph neural network
node classification
0.312018
Active Learning on Heterogeneous Information Networks: A Multi-armed Bandit Approach · ICDM 2018
Data mining › structured data mining › graph mining
heterogeneous information network
0.312017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Data mining › structured data mining › graph mining
network embedding
0.312017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Data mining
pattern mining
0.312017
Embedding Learning with Events in Heterogeneous Information Networks · IEEE Trans. Knowl. Data Eng. 2017
Machine learning › Graph learning
network embedding
0.212016
Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016
Data mining › structured data mining › graph mining › heterogeneous information network
heterogeneous information network mining
0.212016
Large-Scale Embedding Learning in Heterogeneous Event Data · ICDM 2016
Recommender systems › knowledge-aware recommendation
entity recommendation
0.212014
Personalized entity recommendation: a heterogeneous information network approach · WSDM 2014
Data mining › structured data mining › graph mining
information network analysis
0.212014
NewsNetExplorer: automatic construction and exploration of news information networks · SIGMOD Conference 2014
Information retrieval › document retrieval › domain-specific retrieval
news retrieval
0.212014
NewsNetExplorer: automatic construction and exploration of news information networks · SIGMOD Conference 2014
Data mining › clustering › graph clustering
heterogeneous information network clustering
0.112012
Integrating meta-path selection with user-guided object clustering in heterogeneous information networks · KDD 2012
Recommender systems › collaborative filtering
implicit feedback
0.112014
Personalized entity recommendation: a heterogeneous information network approach · WSDM 2014
Data mining › text mining
information extraction
0.112014
NewsNetExplorer: automatic construction and exploration of news information networks · SIGMOD Conference 2014

Methods — techniques the papers use, named apart from their topics

deduplication · 0.9data filtering · 0.9hyperedge prediction · 0.5embedding learning · 0.5network embedding · 0.4heterogeneous information network · 0.4multi-armed bandit · 0.3proximity prediction · 0.3hyperedge modeling · 0.3similarity search · 0.2ranking-based clustering · 0.2OLAP · 0.2iterative algorithm · 0.1
YearPublicationVenuePosition
2025 Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
abstract
Dan Su, Kezhi Kong, Ying Lin, Joseph Jennings, Brandon Norick, Markus Kliegl, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Dan Su 0003, Kezhi Kong, Joseph Jennings, Brandon Norick, Markus Kliegl, Mostofa Patwary, Mohammad Shoeybi, Bryan Catanzaro
ACL (1)5
2018 Active Learning on Heterogeneous Information Networks: A Multi-armed Bandit Approach
abstract
Active learning exploits inherent structures in the unlabeled data to minimize the number of labels required to train an accurate model. It enables effective machine learning in applications with high labeling cost, such as document classification and drug response prediction. We investigate active learning on heterogeneous information networks, with the objective of obtaining accurate node classifications while minimizing the number of labeled nodes. Our proposed algorithm harnesses a multi-armed bandit (MAB) algorithm to determine network structures that identify the most important nodes to the classification task, accounting for node types and without assuming label assortativity. Evaluations on real-world network classification tasks demonstrate that our algorithm outperforms existing methods independent of the underlying classification model.
Doris Xin, Ahmed El-Kishky, De Liao, Brandon Norick, Jiawei Han 0001
ICDM4
2017 Embedding Learning with Events in Heterogeneous Information Networks
abstract
In real-world applications, objects of multiple types are interconnected, formingHeterogeneous Information Networks. In such heterogeneous information networks, we make the key observation that many interactions happen due to someeventand the objects in each event form a complete semantic unit. By taking advantage of such a property, we propose a generic framework calledHyperEdge-BasedEmbedding(Hebe) to learn object embeddings with events in heterogeneous information networks, where ahyperedgeencompasses the objects participating in one event. TheHebeframework models the proximity among objects in each event with two methods: (1) predicting a target object given other participating objects in the event, and (2) predicting if the event can be observed given all the participating objects. Since each hyperedge encapsulates more information of a given event,Hebeis robust to data sparseness and noise. In addition,Hebeis scalable when the data size spirals. Extensive experiments on large-scale real-world datasets show the efficacy and robustness of the proposed framework.
Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Lance M. Kaplan, Jiawei Han 0001
IEEE Trans. Knowl. Data Eng.5
2016 Large-Scale Embedding Learning in Heterogeneous Event Data
abstract
Heterogeneous events, which are defined as events connecting strongly-typed objects, are ubiquitous in the real world. We propose a HyperEdge-Based Embedding (Hebe) framework for heterogeneous event data, where a hyperedge represents the interaction among a set of involving objects in an event. The Hebe framework models the proximity among objects in an event by predicting a target object given the other participating objects in the event (hyperedge). Since each hyperedge encapsulates more information on a given event, Hebe is robust to data sparseness. In addition, Hebe is scalable when the data size spirals. Extensive experiments on large-scale real-world datasets demonstrate the efficacy and robustness of Hebe.
Huan Gui, Fangbo Tao, Meng Jiang 0001, Brandon Norick, Jiawei Han 0001
ICDM5
2014 NewsNetExplorer: automatic construction and exploration of news information networks
abstract
News data is one of the most abundant and familiar data sources. News data can be systematically utilized and ex- plored by database, data mining, NLP and information re- trieval researchers to demonstrate to the general public the power of advanced information technology. In our view, news data contains rich, inter-related and multi-typed data objects, forming one or a set of gigantic, interconnected, het- erogeneous information networks. Much knowledge can be derived and explored with such an information network if we systematically develop effective and scalable data-intensive information network analysis technologies. By further developing a set of information extraction, in- formation network construction, and information network mining methods, we extract types, topical hierarchies and other semantic structures from news data, construct a semi- structured news information network NewsNet. Further, we develop a set of news information network exploration and mining mechanisms that explore news in multi-dimensional space, which include (i) OLAP-based operations on the hierarchical dimensional and topical structures and rich-text, such as cell summary, single dimension analysis, and promo- tion analysis, (ii) a set of network-based operations, such as similarity search and ranking-based clustering, and (iii) a set of hybrid operations or network-OLAP operations, such as entity ranking at different granularity levels. These form the basis of our proposed NewsNetExplorer system. Although some of these functions have been studied in recent research, effective and scalable realization of such functions in large networks still poses multiple challenging research problems. Moreover, some functions are our on-going research tasks. By integrating these functions, NewsNetExplorer not only provides with us insightful recommendations in NewsNet exploration system but also helps us gain insight on how to perform effective information extraction, integration and mining in large unstructured datasets.
Fangbo Tao, George Brova, Jiawei Han 0001, Heng Ji 0001, Chi Wang 0001, Brandon Norick, Ahmed El-Kishky, Xiang Ren 0001, Yizhou Sun
SIGMOD Conference6
2014 Personalized entity recommendation: a heterogeneous information network approach
abstract
Among different hybrid recommendation techniques, network-based entity recommendation methods, which utilize user or item relationship information, are beginning to attract increasing attention recently. Most of the previous studies in this category only consider a single relationship type, such as friendships in a social network. In many scenarios, the entity recommendation problem exists in a heterogeneous information network environment. Different types of relationships can be potentially used to improve the recommendation quality. In this paper, we study the entity recommendation problem in heterogeneous information networks. Specifically, we propose to combine heterogeneous relationship information for each user differently and aim to provide high-quality personalized recommendation results using user implicit feedback data and personalized recommendation models.
Xiao Yu 0007, Xiang Ren 0001, Yizhou Sun, Quanquan Gu, Bradley Sturt, Urvashi Khandelwal, Brandon Norick, Jiawei Han 0001
WSDM7
2013 Recommendation in heterogeneous information networks with implicit user feedback
abstract
Recent studies suggest that by using additional user or item relationship information when building hybrid recommender systems, the recommendation quality can be largely improved. However, most such studies only consider a single type of relationship, e.g., social network. Notice that in many applications, the recommendation problem exists in an attribute-rich heterogeneous information network environment. In this paper, we study the entity recommendation problem in heterogeneous information networks. We propose to combine various relationship information from the network with user feedback to provide high quality recommendation results.
Xiao Yu 0007, Xiang Ren 0001, Yizhou Sun, Bradley Sturt, Urvashi Khandelwal, Quanquan Gu, Brandon Norick, Jiawei Han 0001
RecSys7
2013 PathSelClus: Integrating Meta-Path Selection with User-Guided Object Clustering in Heterogeneous Information Networks
Yizhou Sun, Brandon Norick, Jiawei Han 0001, Xifeng Yan, Philip S. Yu, Xiao Yu 0007
ACM Trans. Knowl. Discov. Data2
2012 User guided entity similarity search using meta-path selection in heterogeneous information networks
abstract
With the emergence of web-based social and information applications, entity similarity search in information networks, aiming to find entities with high similarity to a given query entity, has gained wide attention. However, due to the diverse semantic meanings in heterogeneous information networks, which contain multi-typed entities and relationships, similarity measurement can be ambiguous without context. In this paper, we investigate entity similarity search and the resulting ambiguity problems in heterogeneous information networks. We propose to use a meta-path-based ranking model ensemble to represent semantic meanings for similarity queries, exploit the possibility of using using user-guidance to understand users query. Experiments on real-world datasets show that our framework significantly outperforms competitor methods.
Xiao Yu 0007, Yizhou Sun, Brandon Norick, Tiancheng Mao, Jiawei Han 0001
CIKM3
2012 Integrating meta-path selection with user-guided object clustering in heterogeneous information networks
abstract
Real-world, multiple-typed objects are often interconnected, forming heterogeneous information networks. A major challenge for link-based clustering in such networks is its potential to generate many different results, carrying rather diverse semantic meanings. In order to generate desired clustering, we propose to use meta-path, a path that connects object types via a sequence of relations, to control clustering with distinct semantics. Nevertheless, it is easier for a user to provide a few examples ("seeds") than a weighted combination of sophisticated meta-paths to specify her clustering preference. Thus, we propose to integrate meta-path selection with user-guided clustering to cluster objects in networks, where a user first provides a small set of object seeds for each cluster as guidance. Then the system learns the weights for each meta-path that are consistent with the clustering result implied by the guidance, and generates clusters under the learned weights of meta-paths. A probabilistic approach is proposed to solve the problem, and an effective and efficient iterative algorithm, PathSelClus, is proposed to learn the model, where the clustering quality and the meta-path weights are mutually enhancing each other. Our experiments with several clustering tasks in two real networks demonstrate the power of the algorithm in comparison with the baselines.
Yizhou Sun, Brandon Norick, Jiawei Han 0001, Xifeng Yan, Philip S. Yu, Xiao Yu 0007
KDD2
2011 Semantically-Guided Clustering of Text Documents via Frequent Subgraphs Discovery
Rafal A. Angryk, Mahmud Shahriar Hossain, Brandon Norick
ISMIS3
2010 Effects of the number of developers on code quality in open source software: a case study
abstract
Eleven open source software projects were analyzed to determine if the number of committing developers impacts code quality. We use cyclomatic complexity, lines of code per function, comment density, and maximum nesting as surrogate measures of code quality. We find no significant evidence to suggest that the number of committing developers affects the quality of software.
Brandon Norick, Justin Krohn, Eben Howard, Ben Welna, Clemente Izurieta
ESEM1