EDBT 2026 Demo / reviewers in the wild / expert
Furkan Kocayusufoglu
dblp:194/4848
· DBLP profile ↗
6ranked-venue papers
4as first author
3since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 57% Recommender systems · 35% Data mining · 8% | |
| Artificial intelligence
2 papers |
Information extraction and text analysis · 33% Graph learning · 25% Generative modeling · 25% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% |
Topics — the 14 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Graph learning
graph neural network |
0.6 | 1 | 2022 | FlowGEN: A Generative Model for Flow Graphs · KDD 2022 |
Machine learning › Generative modeling
implicit generative model |
0.6 | 1 | 2022 | FlowGEN: A Generative Model for Flow Graphs · KDD 2022 |
Recommender systems › neural recommendation
attention-based recommendation |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval › retrieval models › neural retrieval
neural ranking model |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval › e-commerce search
personalized product search |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval
personalized search |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Information retrieval
ranking |
0.6 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Machine learning › Representation and self-supervised learning › text embedding › text representation learning
document representation |
0.4 | 1 | 2019 | RiSER: Learning Better Representations for Richly Structured Emails · WWW 2019 |
Natural language and speech › Information extraction and text analysis › document analysis
email analysis |
0.4 | 1 | 2019 | RiSER: Learning Better Representations for Richly Structured Emails · WWW 2019 |
Natural language and speech › Information extraction and text analysis › text classification
email classification |
0.4 | 1 | 2019 | RiSER: Learning Better Representations for Richly Structured Emails · WWW 2019 |
Recommender systems › collaborative filtering › matrix factorization
boolean matrix factorization |
0.3 | 1 | 2018 | Summarizing Network Processes with Network-Constrained Boolean Matrix Factorization · ICDM 2018 |
Data mining › pattern mining
graph pattern mining |
0.3 | 1 | 2018 | Summarizing Network Processes with Network-Constrained Boolean Matrix Factorization · ICDM 2018 |
Recommender systems › collaborative filtering
matrix factorization |
0.3 | 1 | 2018 | Summarizing Network Processes with Network-Constrained Boolean Matrix Factorization · ICDM 2018 |
Recommender systems
user modeling |
0.2 | 1 | 2022 | Multi-Resolution Attention for Personalized Item Search · WSDM 2022 |
Methods — techniques the papers use, named apart from their topics
soft-thresholding · 0.6multi-head attention · 0.6flow graph neural network · 0.6physics-informed machine learning · 0.5attention-based LSTM · 0.4HTML structure encoding · 0.4subgraph search · 0.3sampling · 0.3monte carlo markov chain · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | FlowGEN: A Generative Model for Flow GraphsabstractFlow graphs capture the directed flow of a quantity of interest (e.g., water, power, vehicles) being transported through an underlying network. Modeling and generating realistic flow graphs is key in many applications in infrastructure design, transportation, and biomedical and social sciences. However, they pose a great challenge to existing generative models due to a complex dynamics that is often governed by domain-specific physical laws or patterns. We introduce FlowGEN, an implicit generative model for flow graphs, that learns how to jointly generate graph topologies and flows with diverse dynamics directly from data using a novel (flow) graph neural network. Experiments show that our approach is able to effectively reproduce relevant local and global properties of flow graphs, including flow conservation, cyclic trends, and congestion around hotspots. Furkan Kocayusufoglu, Arlei Silva, Ambuj K. Singh |
KDD | 1 |
| 2022 | Multi-Resolution Attention for Personalized Item SearchabstractPersonalized item search has become an essential tool for online platforms---where users interact with a large corpus of items (e.g., click, purchase, like) via a search query---to provide their users with a more satisfactory search experience. The record (or history) of users' past interactions serves as a valuable asset to achieve personalization. While user history data can span over a long period of time, only a part of the history is relevant to a user's current search intent. Moreover, since historical interactions take place at aperiodic points in time, modeling their relevance to the current search query entangles complex temporal dependencies. We propose multi-resolution attention to address these challenges for personalized item search. Our approach captures higher-order temporal relations between user queries and their history across several temporal subspaces (i.e., resolutions), each corresponding to distinct temporal ranges with adaptive time boundaries that are also learned directly from data. We achieve this by coupling the conventional multi-head attention module with a differentiable soft-thresholding mechanism, which essentially operates as a masking function in the temporal domain. Comparisons with strong baselines on an open-source benchmark dataset confirm the efficacy of our approach. Furkan Kocayusufoglu, Anima Singh, George Roumpos, Heng-Tze Cheng, Sagar Jain, Ed H. Chi, Ambuj K. Singh |
WSDM | 1 |
| 2021 | Combining Physics and Machine Learning for Network Flow Estimation
Arlei Silva, Furkan Kocayusufoglu, Saber Jafarpour, Francesco Bullo, Ananthram Swami, Ambuj K. Singh |
ICLR | 2 |
| 2020 | DANR: Discrepancy-aware Network RegularizationabstractNetwork regularization is an effective tool for incorporating structural prior knowledge to learn coherent models over networks, and has yielded provably accurate estimates in applications ranging from spatial economics to neuroimaging studies. Recently, there has been an increasing interest in extending network regularization to the spatio-temporal case to accommodate the evolution of networks. However, in both static and spatio-temporal cases, missing or corrupted edge weights can compromise the ability of network regularization to discover desired solutions. To address these gaps, we propose a novel approach—discrepancy-aware network regularization (DANR)—that is robust to inadequate regularizations and effectively captures model evolution and structural changes over spatio-temporal networks. We develop a distributed and scalable algorithm based on alternating direction method of multipliers (ADMM) to solve the proposed problem with guaranteed convergence to global optimum solutions. Experimental results on both synthetic and real-world networks demonstrate that our approach achieves improved performance on various tasks, and enables interpretation of model changes in evolving networks. Hongyuan You, Furkan Kocayusufoglu, Ambuj K. Singh |
SDM | 2 |
| 2019 | RiSER: Learning Better Representations for Richly Structured EmailsabstractRecent studies show that an overwhelming majority of emails are machine-generated and sent by businesses to consumers. Many large email services are interested in extracting structured data from such emails to enable intelligent assistants. This allows experiences like being able to answer questions such as “What is the address of my hotel in New York?” or “When does my flight leave?”. A high-quality email classifier is a critical piece in such a system. In this paper, we argue that the rich formatting used in business-to-consumer emails contains valuable information that can be used to learn better representations. Most existing methods focus only on textual content and ignore the rich HTML structure of emails. We introduce RiSER (Richly Structured Email Representation) - an approach for incorporating both the structure and content of emails. RiSER projects the email into a vector representation by jointly encoding the HTML structure and the words in the email. We then use this representation to train a classifier. To our knowledge, this is the first description of a neural technique for combining formatting information along with the content to learn improved representations for richly formatted emails. Experimenting with a large corpus of emails received by users of Gmail, we show that RiSER outperforms strong attention-based LSTM baselines. We expect that these benefits will extend to other corpora with richly formatted documents. We also demonstrate with examples where leveraging HTML structure leads to better predictions. Furkan Kocayusufoglu, Ying Sheng 0002, Nguyen Vo, James B. Wendt, Sandeep Tata, Marc Najork |
WWW | 1 |
| 2018 | Summarizing Network Processes with Network-Constrained Boolean Matrix FactorizationabstractUnderstanding and modeling complex network processes is an important task in many real-world applications. The first challenge is to discover patterns in such complex data. In this work, our goal is to summarize different processes in a network by a small yet interpretable set of network patterns, each of which represents a local community of connected nodes frequently participating in the same network processes. We formulate this problem as a Boolean Matrix Factorization with a network constraint, which we prove to be NP-hard. We then propose an efficient algorithm that incrementally adds the best patterns and achieve scalability with two further improvements. First, to decide which network processes contain which network patterns, we introduce two mapping algorithms with linear costs. Second, to systematically mine the exponential subgraph search space for good patterns, we devise two sampling algorithms based on Monte Carlo Markov Chain. Experimental results on both synthetic and real-world datasets show that our solutions are scalable and find network patterns that effectively summarize network processes. Furkan Kocayusufoglu, Minh X. Hoang, Ambuj K. Singh |
ICDM | 1 |