VLDB 2026 Research / reviewers in the wild / expert
Toyotaro Suzumura
dblp:99/844
· DBLP profile ↗
22ranked-venue papers in the field
3as first author
7since 2021 · last 2025
0000-0001-6412-8386ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 8 (2 first)Information Retrieval & Web Search · 7 (1 first)Database Systems & Data Management · 5Data Mining & Knowledge Discovery · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Look Into News Avoidance Through AWRS: An Avoidance-Aware Recommender SystemabstractIn recent years, journalists have expressed concerns about the increasing trend of news article avoidance, especially within specific domains. This issue has been exacerbated by the rise of recommender systems. Our research indicates that recommender systems should consider avoidance as a fundamental factor. We argue that news articles can be characterized by three principal elements: exposure, relevance, and avoidance, all of which are closely interconnected. To address these challenges, we introduce AWRS, an Avoidance-Aware Recommender System. This framework incorporates avoidance awareness when recommending news, based on the premise that news article avoidance conveys significant information about user preferences. Evaluation results on three news datasets in different languages (English, Norwegian, and Japanese) demonstrate that our method outperforms existing approaches. Igor L. R. Azevedo, Toyotaro Suzumura, Yuichiro Yasui |
SDM | 2 |
| 2024 | ARIM-mdx Data System: Towards a Nationwide Data Platform for Materials ScienceabstractIn modern materials science, effective and high-volume data management across leading-edge experimental facilities and world-class supercomputers is indispensable for cutting-edge research. However, existing integrated systems that handle data from these resources have primarily focused just on smaller-scale cross-institutional or single-domain operations. As a result, they often lack the scalability, efficiency, agility, and interdisciplinarity, needed for handling substantial volumes of data from various researchersIn this paper, we introduce ARIM-mdx data system1, aiming at a nationwide data platform for materials science in Japan. Currently in its trial phase, the platform has been involving 11 universities and institutes all over Japan, and it is utilized by over 800 researchers from around 140 organizations in academia and industry, being intended to gradually expand its reach. The ARIM-mdx data system, as a pioneering nationwide data platform, has the potential to contribute to the creation of new research communities and accelerate innovations. Masatoshi Hanai, Ryo Ishikawa, Mitsuaki Kawamura, Masato Ohnishi, Norio Takenaka, Kou Nakamura, Daiju Matsumura, Seiji Fujikawa, Hiroki Sakamoto, Yukinori Ochiai, Tetsuo Okane, Shin-Ichiro Kuroki, Atsuo Yamada, Toyotaro Suzumura, Kenjiro Taura, Yoshio Mita, Naoya Shibata, Yuichi Ikuhara |
IEEE Big Data | 14 |
| 2024 | Graph-Based Audience Expansion Model for Marketing CampaignsabstractAudience Expansion, a technique for identifying new audiences with similar behaviors to the original target or seed users. The major challenges include a heterogeneous user base, intricate marketing campaigns, constraints imposed by sparsity, and limited seed users, which lead to overfitting. In this context, we propose a novel solution named AudienceLinkNet, specifically designed to address the challenges associated with audience expansion in the context of Rakuten's diverse services and its clients. Our approach formulates the audience expansion problem as a graph problem and explores the combination of a Pre-trained Knowledge Graph Embedding Model and a Graph Convolutional Networks (GCNs). It emphasizes the structural retention properties of GCNs, enabling the model to overcome challenges related to cross-service data usage, sparsity and limited seed data. AudienceLinkNet simplifies the targeting process for small and large marketing campaigns and better utilizes demographics and behavioral attributes for targeting. Extensive experiments on our advertising platform, Rakuten AIris Target Prospecting, demonstrate the effectiveness of our audience expansion model. Additionally, we present the limitations of AudienceLinkNet. Daisuke Kikuta, Yu Hirate, Toyotaro Suzumura |
SIGIR | 4 |
| 2023 | Revisiting Mobility Modeling with Graph: A Graph Transformer Model for Next Point-of-Interest RecommendationabstractNext Point-of-Interest (POI) recommendation plays a crucial role in urban mobility applications. Recently, POI recommendation models based on Graph Neural Networks (GNN) have been extensively studied and achieved, however, the effective incorporation of both spatial and temporal information into such GNN-based models remains challenging. Temporal information is extracted from users' trajectories, while spatial information is obtained from POIs. Extracting distinct fine-grained features unique to each piece of information is difficult since temporal information often includes spatial information, as users tend to visit nearby POIs. To address the challenge, we propose Mobility Graph Transformer (MobGT) that enables us to fully leverage graphs to capture both the spatial and temporal features in users' mobility patterns. MobGT combines individual spatial and temporal graph encoders to capture unique features and global user-location relations. Additionally, it incorporates a mobility encoder based on Graph Transformer to extract higher-order information between POIs. To address the long-tailed problem in spatial-temporal data, MobGT introduces a novel loss function, Tail Loss. Experimental results demonstrate that MobGT outperforms state-of-the-art models on various datasets and metrics, achieving 24% improvement on average. Our codes are available at https://github.com/Yukayo/MobGT. Xiaohang Xu 0002, Toyotaro Suzumura, Jiawei Yong, Masatoshi Hanai, Chuang Yang 0002, Hiroki Kanezashi, Renhe Jiang, Shintaro Fukushima |
SIGSPATIAL/GIS | 2 |
| 2023 | ✨ Going Beyond Local: Global Graph-Enhanced Personalized News RecommendationsabstractPrecisely recommending candidate news articles to users has always been a core challenge for personalized news recommendation systems. Most recent works primarily focus on using advanced natural language processing techniques to extract semantic information from rich textual data, employing content-based methods derived from local historical news. However, this approach lacks a global perspective, failing to account for users’ hidden motivations and behaviors beyond semantic information. To address this challenge, we propose a novel model called GLORY (Global-LOcal news Recommendation sYstem), which combines global representations learned from other users with local representations to enhance personalized recommendation systems. We accomplish this by constructing a Global-aware Historical News Encoder, which includes a global news graph and employs gated graph neural networks to enrich news representations, thereby fusing historical news representations by a historical news aggregator. Similarly, we extend this approach to a Global Candidate News Encoder, utilizing a global entity graph and a candidate news aggregator to enhance candidate news representation. Evaluation results on two public news datasets demonstrate that our method outperforms existing approaches. Furthermore, our model offers more diverse recommendations1. Boming Yang, Dairui Liu, Toyotaro Suzumura, Ruihai Dong, Irene Li |
RecSys | 3 |
| 2023 | Exploring 360-Degree View of Customers for Lookalike ModelingabstractLookalike models are based on the assumption that user similarity plays an important role towards product selling and enhancing the existing advertising campaigns from a very large user base. Challenges associated to these models reside on the heterogeneity of the user base and its sparsity. In this work, we propose a novel framework that unifies the customers' different behaviors or features such as demographics, buying behaviors on different platforms, customer loyalty behaviors and build a lookalike model to improve customer targeting for Rakuten Group, Inc. Extensive experiments on real e-commerce and travel datasets demonstrate the effectiveness of our proposed lookalike model for user targeting task. Daisuke Kikuta, Satyen Abrol, Yu Hirate, Toyotaro Suzumura, Pablo Loyola, Takuma Ebisu, Manoj Kondapaka |
SIGIR | 5 |
| 2023 | Can Persistent Homology provide an efficient alternative for Evaluation of Knowledge Graph Completion Methods?abstractIn this paper we present a novel method, Knowledge Persistence (), for faster evaluation of Knowledge Graph (KG) completion approaches. Current ranking-based evaluation is quadratic in the size of the KG, leading to long evaluation times and consequently a high carbon footprint. addresses this by representing the topology of the KG completion methods through the lens of topological data analysis, concretely using persistent homology. The characteristics of persistent homology allow to evaluate the quality of the KG completion looking only at a fraction of the data. Experimental results on standard datasets show that the proposed metric is highly correlated with ranking metrics (Hits@N, MR, MRR). Performance evaluation shows that is computationally efficient: In some cases, the evaluation time (validation+test) of a KG completion method has been reduced from 18 hours (using Hits@10) to 27 seconds (using ), and on average (across methods & data) reduces the evaluation time (validation+test) by ≈ 99.96%. Anson Bastos, Kuldeep Singh 0001, Abhishek Nadgeri, Johannes Hoffart, Manish Singh 0002, Toyotaro Suzumura |
WWW | 6 |
| 2020 | Memory Efficient Graph Convolutional Network based Distributed Link PredictionabstractGraph Convolutional Networks (GCN) have found multiple applications of graph-based machine learning. However, training GCNs on large graphs of billions of nodes and edges with rich node attributes consume significant amount of time and memory resources. This makes it impossible to train such GCNs on general purpose commodity hardware. Such use cases demand high-end servers with accelerators and ample amounts of memory. In this paper we implement a memory efficient GCN based link prediction on top of a distributed graph database server called JasmineGraph1. Our approach is based on federated training on partitioned graphs with multiple parallel workers. We conduct experiments with three real world graph datasets called DBLP-V11, Reddit, and Twitter. We demonstrate that our approach produces optimal performance for a given hardware setting. JasmineGraph was able to train a GCN on the largest dataset DBLP-V11(>10GB) in 20 hours and 24 minutes for 5 training rounds and 3 epochs by partitioning it into 16 partitions with 2 workers on a single server while the conventional training method could not process it at all due to lack of memory. The second largest dataset Reddit took 9 hours 8 minutes to train with conventional training while JasmineGraph took only 3 hours and 11 minutes with 8 partitions-4 workers in the same hardware giving 3 times improved performance. In case of Twitter dataset JasmineGraph was able to give 5 times improved performance. (10 hours 31 minutes vs 2 hours 6 minutes;16 partitions-16 workers). Damitha Seneviratne, Isuru Wijesiri, Suchitha Dehigaspitiya, Miyuru Dayarathna, Sanath Jayasena, Toyotaro Suzumura |
IEEE BigData | 6 |
| 2020 | The Impact of COVID-19 on Flight NetworksabstractAs COVID-19 transmissions spread worldwide, governments have announced and enforced travel restrictions to prevent further infections. Such restrictions have a direct effect on the volume of international flights among these countries, resulting in extensive social and economic costs. To better understand the situation in a quantitative manner, we analyzed the OpenSky Network data to clarify flight patterns and flight densities around the world. Then we observed relationships between flight numbers with new infection cases and the economy (the unemployment rate) in Barcelona. We found that the number of daily flights gradually decreased and then suddenly dropped 64% during the second half of March in 2020 after the United States and Europe enacted travel restrictions. We also observed a 51% decrease in the global flight network density decreased during this period. Regarding new COVID-19 cases, the United States had an unexpected surge regardless of travel restrictions. Finally, the layoffs for temporary workers in the tourism and airplane business increased by 4.3 fold in the weeks following Spain's decision to close its borders. Toyotaro Suzumura, Hiroki Kanezashi, Mishal Dholakia, Euma Ishii, Sergio Álvarez-Napagao, Raquel Pérez-Arnal, Dario Garcia-Gasulla |
IEEE BigData | 1 |
| 2019 | Distributed Edge Partitioning for Trillion-edge GraphsabstractWe propose Distributed Neighbor Expansion (Distributed NE), a parallel and distributed graph partitioning method that can scale to trillion-edge graphs while providing high partitioning quality. Distributed NE is based on a new heuristic, called parallel expansion, where each partition is constructed in parallel by greedily expanding its edge set from a single vertex in such a way that the increase of the vertex cuts becomes local minimal. We theoretically prove that the proposed method has the upper bound in the partitioning quality. The empirical evaluation with various graphs shows that the proposed method produces higher-quality partitions than the state-of-the-art distributed graph partitioning algorithms. The performance evaluation shows that the space efficiency of the proposed method is an order-of-magnitude better than the existing algorithms, keeping its time efficiency comparable. As a result, Distributed NE can partition a trillion-edge graph using only 256 machines within 70 minutes. Masatoshi Hanai, Toyotaro Suzumura, Wen Jun Tan, Elvis S. Liu, Georgios Theodoropoulos 0001, Wentong Cai 0001 |
Proc. VLDB Endow. | 2 |
| 2017 | A generalized incremental bottom-up community detection framework for highly dynamic graphsabstractHow to efficiently detect communities for dynamic graphs has attracted significant attention due to its widespread applications, such as “Recommended System” in social networking or “Anti-Money Laundry” for banks. However, with the growing size and the increasing of changing frequency for dynamic graphs, it is better to detect communities incrementally than applying community detection algorithm for each of the graph. In this paper, we propose a generalized bottom-up community detection framework to help standard community detection algorithms to detect communities in highly dynamic graphs incrementally. The evaluations on real data sets show that our generalized framework does have the ability to accelerate standard community detection methods for highly dynamic graphs (up to 96%). In addition, for super hubs in large scale graph, we also propose an approximate process to meet the requirement of real-time processing. By giving “mathematical proof”, “efficiency” and “accuracy” evaluations from real data sets, we demonstrate this approximate process can not only accelerate the community detection process but also can preserve the accuracy, as the normalized mutual information for regular communities and approximate communities is larger than 92%. Toyotaro Suzumura, Lingli Chen, Guangmin Hu |
IEEE BigData | 2 |
| 2017 | Scalable time-versioning support for property graph databasesabstractWhen graphs change over time, it is important to make the changes trackable for many graph-based applications. We propose an implementation of OLTP-oriented graph database that supports time-versioning. There has been a few snapshot-based approaches for supporting time-versions, but they usually require the full-restoration of the graph, and lack the resolution of the time space. Using a B-tree as the datastructure for the backend storage, our database allow fast and scalable support for restoring the arbitrary part of the graph, without slowing down the normal accesses to the current graph. Experimental results show that our scheme is much efficient than the straightforward solutions, in terms of space and performance. Warut D. Vijitbenjaronk, Jinho Lee 0001, Toyotaro Suzumura, Ilie Gabriel Tanase |
IEEE BigData | 3 |
| 2017 | Efficient Breadth-First Search on Massively Parallel and Distributed-Memory MachinesabstractThere are many large-scale graphs in real world such as Web graphs and social graphs. The interest in large-scale graph analysis is growing in recent years. Breadth-First Search (BFS) is one of the most fundamental graph algorithms used as a component of many graph algorithms. Our new method for distributed parallel BFS can compute BFS for one trillion vertices graph within half a second, using large supercomputers such as the K-Computer. By the use of our proposed algorithm, the K-Computer was ranked 1st in Graph500 using all the 82,944 nodes available on June and November 2015 and June 2016 38,621.4 GTEPS. Based on the hybrid BFS algorithm by Beamer (Proceedings of the 2013 IEEE 27th International Symposium on Parallel and Distributed Processing Workshops and PhD Forum, IPDPSW ’13, IEEE Computer Society, Washington, 2013 ), we devise sets of optimizations for scaling to extreme number of nodes, including a new efficient graph data structure and several optimization techniques such as vertex reordering and load balancing. Our performance evaluation on K-Computer shows that our new BFS is 3.19 times faster on 30,720 nodes than the base version using the previously known best techniques. Koji Ueno, Toyotaro Suzumura, Naoya Maruyama, Katsuki Fujisawa, Satoshi Matsuoka |
Data Sci. Eng. | 2 |
| 2016 | An incremental local-first community detection method for dynamic graphsabstractCommunity detections for large-scale real world networks have been more popular in social analytics. In particular, dynamically growing network analyses become important to find long-term trends and detect anomalies. In order to analyze such networks, we need to obtain many snapshots and apply same analytic methods to them. However, it is inefficient to extract communities from these whole newly generated networks with little differences every time, and then it is impossible to follow the network growths in the real time. We proposed an incremental community detection algorithm for high-volume graph streams. It is based on the top of a well-known batch-oriented algorithm named DEMON [1]. We also evaluated performance and precisions of our proposed incremental algorithm with real-world big networks with up to 410,236 vertices and 2,439,437 edges and computed in less than one second to detect communities in an incremental fashion - which achieves up to 107 times faster than the original algorithm without sacrificing accuracies. Hiroki Kanezashi, Toyotaro Suzumura |
IEEE BigData | 2 |
| 2016 | Extreme scale breadth-first search on supercomputersabstractBreadth-First Search(BFS) is one of the most fundamental graph algorithms used as a component of many graph algorithms. Our new method for distributed parallel BFS can compute BFS for one trillion vertices graph within half a second, using large supercomputers such as the K-Computer. By the use of our proposed algorithm, the K-Computer was ranked 1st in Graph500 using all the 82,944 nodes available on June and November 2015 and June 2016 38,621.4 GTEPS. Based on the hybrid-BFS algorithm by Beamer[3], we devise sets of optimizations for scaling to extreme number of nodes, including a new efficient graph data structure and optimization techniques such as vertex reordering and load balancing. Performance evaluation on the K shows our new BFS is 3.19 times faster on 30,720 nodes than the base version using the previously-known best techniques. Koji Ueno, Toyotaro Suzumura, Naoya Maruyama, Katsuki Fujisawa, Satoshi Matsuoka |
IEEE BigData | 2 |
| 2015 | ScaleGraph: A high-performance library for billion-scale graph analyticsabstractRecently, large-scale graph analytics has become a very popular topic owing to the emergence of gigantic graphs whose number of vertices and edges is in millions, billions or even trillions. Many graph analytics libraries and frameworks have been proposed with various computational models and programming languages to deal with such graphs. X10 programming language is a PGAS language that aims at both software performance and programmer's productivity. We introduce ScaleGraph library developed using X10 programming to illustrate the use of X10 for large-scale graph analytics. ScaleGraph library provides XPregel framework that is inspired by Google's Pregel computation model, serving as a building block for implementing graph kernels. We also optimized X10 runtime in some parts such as collective communication and memory management. We evaluated the performance and scalability of ScaleGraph libraries. The result shows that most graph kernels have good performance and scalability. ScaleGraph library is 9.4 times faster than Giraph in the experiment of PageRank with 16 machine nodes. To the best of our knowledge, ScaleGraph is the first X10-based library to address performance, scalability and productivity issues in dealing with large-scale graph analytics. Toyotaro Suzumura, Koji Ueno |
IEEE BigData | 1 |
| 2013 | A Mechanism for Stream Program Performance Recovery in Resource Limited Compute Clusters
Miyuru Dayarathna, Toyotaro Suzumura |
DASFAA (2) | 2 |
| 2013 | Automatic optimization of stream programs via source program operator graph transformations
Miyuru Dayarathna, Toyotaro Suzumura |
Distributed Parallel Databases | 2 |
| 2012 | Highly Scalable Speech Processing on Data Stream Management System
Shunsuke Nishii, Toyotaro Suzumura |
DASFAA (2) | 2 |
| 2009 | Highly scalable web applications with zero-copy data transferabstractThe performance of server-side applications is becoming increasingly important as more applications exploit the Web application model. Extensive work has been done to improve the performance of individual software components such as Web servers and programming language runtimes. This paper describes a novel approach to boost Web application performance by improving inter-process communication between a programming language runtime and Web server runtime. The approach reduces redundant processing for memory copying and the context switch overhead between user space and kernel space by exploiting the zero-copy data transfer methodology, such as the sendfile system call. In order to transparently utilize this optimization feature with existing Web applications, we propose enhancements of the PHP runtime, FastCGI protocol, and Web server. Our proposed approach achieves a 126% performance improvement with micro-benchmarks and a 44% performance improvement for a standard Web benchmark, SPECweb2005. Toyotaro Suzumura, Michiaki Tatsubori, Scott Trent, Akihiko Tozawa, Tamiya Onodera |
WWW | 1 |
| 2009 | HTML templates that fly: a template engine approach to automated offloading from server to clientabstractWeb applications often use HTML templates to separate the webpage presentation from its underlying business logic and objects. This is now the de facto standard programming model for Web application development. This paper proposes a novel implementation for existing server-side template engines, FlyingTemplate, for (a) reduced bandwidth consumption in Web application servers, and (b) off-loading HTML generation tasks to Web clients. Instead of producing a fully-generated HTML page, the proposed template engine produces a skeletal script which includes only the dynamic values of the template parameters and the bootstrap code that runs on a Web browser at the client side. It retrieves a client-side template engine and the payload templates separately. With the goals of efficiency, implementation transparency, security, and standards compliance in mind, we developed FlyingTemplate with two design principles: effective browser cache usage, and reasonable compromises which restrict the template usage patterns and relax the security policies slightly but in a controllable way. This approach allows typical template-based Web applications to run effectively with FlyingTemplate. As an experiment, we tested the SPECweb2005 banking application using FlyingTemplate without any other modifications and saw throughput improvements from 1.6x to 2.0x in its best mode. In addition, FlyingTemplate can enforce compliance with a simple security policy, thus addressing the security problems of client-server partitioning in the Web environment. Michiaki Tatsubori, Toyotaro Suzumura |
WWW | 2 |
| 2005 | An adaptive, fast, and safe XML parser based on byte sequences memorizationabstractXML (Extensible Markup Language) processing can incur significant runtime overhead in XML-based infrastructural middleware such as Web service application servers. This paper proposes a novel mechanism for efficiently processing similar XML documents. Given a new XML document as a byte sequence, the XML parser proposed in this paper normally avoids syntactic analysis but simply matches the document with previously processed ones, reusing those results. Our parser is adaptive since it partially parses and then remembers XML document fragments that it has not met before. Moreover, it processes safely since its partial parsing correctly checks the well-formedness of documents. Our implementation of the proposed parser complies with the JSR 63 standard of the Java API for XML Processing (JAXP) 1.1 specification. We evaluated Deltarser performance with messages using Google Web services. Comparing to Piccolo (and Apache Xerces), it effectively parses 35 % (106%) faster in a server-side use-case scenario, and 73 % (126%) faster in a client-side use-case scenario. Toshiro Takase, Hisashi Miyashita, Toyotaro Suzumura, Michiaki Tatsubori |
WWW | 3 |