VLDB 2026 Research / reviewers in the wild / expert
Shuichiro Haruta
dblp:173/8348
· DBLP profile ↗
21ranked-venue papers
4as first author
13since 2021 · last 2024
0000-0002-0695-9963ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hypergraph Contrastive Learning with Graph Structure Learning for RecommendationabstractIn this paper, to address the insufficient capturing of high-order correlations and the vulnerability to noise in conventional collaborative filtering models, we propose HyperGraph Contrastive Learning with graph structure learning for recommendation (HGCL). HGCL employs user and item hypergraphs as contrastive views in contrastive learning, where edges connect the item and user's k-order reachable neighbors. This approach allows HGCL to model relationships with distant nodes and explicitly capture high-order correlations, alleviating the issue of over-smoothing. Subsequently, HGCL performs contrastive learning between representations obtained from the user-item interaction graph and hypergraphs. This integrates the user-item relationship features from the interaction graph with the high-order correlations for each user and item from the hypergraphs, resulting in more effective representations. Furthermore, to address the weakness of hypergraphs against noise, we modify the hypergraph structure learning method for the recommendation task and incorporate it into HGCL. Based on the user and item representations, HGCL detects potential node relationships and noise in the hypergraph for the recommendation task. By adding or removing these nodes from the hypergraph, HGCL acquires a denoised hypergraph. By applying these processes to user and item hypergraphs, HGCL obtains improved hypergraphs with reduced noise effects and achieves more effective recommendations. Experimental results on real-world datasets demonstrate that HGCL outperforms baseline models, achieving up to 5.25% improvement in NDCG@20. Yuma Dose, Shuichiro Haruta, Yihong Zhang 0001, Takahiro Hara |
ICMLA | 2 |
| 2024 | Domain Adaptation Utilizing Texts and Visions for Cross-domain Recommendations with No Shared UsersabstractIn recent years, many researchers focus on cross-domain recommendations (CDRs). CDRs leverage data across multiple services (domains) to enable the transfer of user preferences toward sparse domains. Although some approaches transfer knowledge through shared users across domains, there are privacy concerns. Thus, several CDRs use text features like product reviews, instead of shared users. However, real-world datasets often have noisy text features, complicating effective domain adaptation. In this paper, our objective is to enhance domain adaptation performance while ensuring effective recom-mendation by integrating visual features, alongside text features. First, to mitigate the effect by noisy item descriptions, we incorporate visual features that exclusively capture item-specific information. We obtain effective multimodal representations to improve domain adaptation by employing contrastive learning that aligns visual features with fixed text features. On the other hand, relying solely on this approach could potentially undermine the distinctiveness inherent in both text and images. As a second proposal to address this concern, we minimize mutual information between embeddings of text and image, which encourages the preservation of distinctive features within both modalities. These two proposals lead to a balanced approach to domain adaptation while keeping the effectiveness of recommendation systems. Experimental results show the domain adaptation by the proposed method is effective. Kentaro Shiga, Shuichiro Haruta, Takahiro Hara |
ICMLA | 2 |
| 2024 | QWalkVec: Node Embedding by Quantum Walk
Rei Sato, Shuichiro Haruta, Kazuhiro Saito, Mori Kurokawa |
PAKDD (1) | 2 |
| 2024 | Mutual Information-based Preference Disentangling and Transferring for Non-overlapped Multi-target Cross-domain RecommendationsabstractBuilding high-quality recommender systems is challenging for new services and small companies, because of their sparse interactions. Cross-domain recommendations (CDRs) alleviate this issue by transferring knowledge from data in external domains. However, most existing CDRs leverage data from only a single external domain and serve only two domains. CDRs serving multiple domains require domain-shared entities (i.e., users and items) to transfer knowledge, which significantly limits their applications due to the hardness and privacy concerns of finding such entities. We therefore focus on a more general scenario, non-overlapped multi-target CDRs (NO-MTCDRs), which require no domain-shared entities and serve multiple domains. Existing methods require domain-shared users to learn user preferences and cannot work on NO-MTCDRs. We hence propose MITrans, a novel mutual information-based (MI-based) preference disentangling and transferring framework to improve recommendations for all domains. MITrans effectively leverages knowledge from multiple domains as well as learning both domain-shared and domain-specific preferences without using domain-shared users. In MITrans, we devise two novel MI constraints to disentangle domain-shared and domain-specific preferences. Moreover, we introduce a module that fuses domain-shared preferences in different domains and combines them with domain-specific preferences to improve recommendations. Our experimental results on two real-world datasets demonstrate the superiority of MITrans in terms of recommendation quality and application range against state-of-the-art overlapped and non-overlapped CDRs. Zhi Li 0084, Daichi Amagata, Yihong Zhang 0001, Takahiro Hara, Shuichiro Haruta, Kei Yonekawa, Mori Kurokawa |
SIGIR | 5 |
| 2023 | Unvisited Out-Of-Town POI Recommendation with Simultaneous Learning of Multiple RegionsabstractIn recent years, Point Of Interest (POI) recommendations have been actively studied because of the widespread use of location-based social network services. Out-Of-Town POI recommendation methods recommend POIs outside a user’s residence, such as travel and business trip destinations. Most existing methods require a certain number of past visit sequences at destination regions or can be used only between two specific unidirectional regions. In this study, we propose an Unvisited Out-Of-Town POI Recommendation framework (UOPR), which can recommend POIs for Out-Of-Town regions that have not yet been visited by users. In addition, UOPR can learn users’ visit patterns for multiple regions simultaneously. UOPR takes into account the fact that user’s Out-Of-Town visiting tendencies vary depending on the geographical distance from their residences. It therefore adopts an approach that calculates two scores, one determined by user’s Home-Town visiting tendencies and the other by their Out-Of-Town visiting tendencies, and uses them differently. Furthermore, we propose a mask-input learning and user-embedding creation method that enables learning with enhanced interaction between POI embeddings in an Out-Of-Town POI recommendation environment, where the data are often sparse. UOPR achieves higher recommendation accuracy than existing methods on a dataset collected on a real-world service. Rikuto Tsubouchi, Takahiro Hara, Kei Yonekawa, Shuichiro Haruta |
IEEE Big Data | 4 |
| 2023 | A Graph-Based Recommendation Model Using Contrastive Learning for Inductive ScenarioabstractGraph-based recommendation models, which utilize a user-item interaction graph whose edges represent users' preferences, are well known because of their effectiveness. However, many existing models are intrinsically transductive, meaning they can make recommendations only for users and items that exist in the training data. To solve this challenge, some researchers focus on recommendations in the inductive scenario where new users and new items emerge in the inference phase. In this paper, we explore the potential of contrastive learning for the inductive scenario to further improve recommendation accuracy. The situation where new users and new items emerge can be interpreted as a change in the interaction graph. Therefore, it is crucial to capture the change in the interaction graph for inductive recommendations. Based on the above perspective, we propose the contrastive learning framework, INductive Contrastive Learning (INCL). INCL creates an augmented graph by adding or removing edges of the original interaction graph and simulates changes in the interaction graph. By performing contrastive learning, INCL fosters each representation generated from the original interaction graph and the augmented graph to be similar. As a result, INCL is trained to generate robust representations that can adapt to changes in the interaction graph. Experimental results on real-world datasets demonstrate that INCL outperforms existing models in the inductive scenario, achieving up to 2.13% improvement in Recall@20. Yuma Dose, Shuichiro Haruta, Takahiro Hara |
ICMLA | 2 |
| 2023 | Semantic Relation Transfer for Non-overlapped Cross-domain Recommendations
Zhi Li 0084, Daichi Amagata, Yihong Zhang 0001, Takahiro Hara, Shuichiro Haruta, Kei Yonekawa, Mori Kurokawa |
PAKDD (3) | 5 |
| 2022 | A Novel Graph Aggregation Method Based on Feature Distribution Around Each Ego-node for Heterophily
Shuichiro Haruta, Tatsuya Konishi, Mori Kurokawa |
ACML | 1 |
| 2022 | Debiasing Graph Transfer Learning via Item Semantic Clustering for Cross-Domain RecommendationsabstractDeep learning-based recommender systems may lead to over-fitting when lacking training interaction data. This over-fitting significantly degrades recommendation performances. To address this data sparsity problem, cross-domain recommender systems (CDRSs) exploit the data from an auxiliary source domain to facilitate the recommendation on the sparse target domain. Most existing CDRSs rely on overlapping users or items to connect domains and transfer knowledge. However, matching users is an arduous task and may involve privacy issues when data comes from different companies, resulting in a limited application for the above CDRSs. Some studies develop CDRSs that require no overlapping users and items by transferring learned user interaction patterns. However, they ignore the bias in user interaction patterns between domains and hence suffer from an inferior performance compared with single-domain recommender systems. In this paper, based on the above findings, we propose a novel CDRS, namely semantic clustering enhanced debiasing graph neural recommender system (SCDGN), that requires no overlapping users and items and can handle the domain bias. More precisely, SCDGN semantically clusters items from both domains and constructs a cross-domain bipartite graph generated from item clusters and users. Then, the knowledge is transferred via this cross-domain user-cluster graph from source to the target. Furthermore, we design a debiasing graph convolutional layer for SCDGN to extract unbiased structural knowledge from the cross-domain user-cluster graph. Our Experimental results on three public datasets and a pair of proprietary datasets verify the effectiveness of SCDGN over stateof-the-art models in terms of cross-domain recommendations. Zhi Li 0084, Daichi Amagata, Yihong Zhang 0001, Takahiro Hara, Shuichiro Haruta, Kei Yonekawa, Mori Kurokawa |
IEEE Big Data | 5 |
| 2022 | Multi-view Contrastive Multiple Knowledge Graph Embedding for Knowledge CompletionabstractKnowledge graphs (KGs) are useful information sources to make machine learning efficient with human knowledge. Since KGs are often incomplete, KG completion has become an important problem to complete missing facts in KGs. Whereas most of the KG completion methods are conducted on a single KG, multiple KGs can be effective to enrich embedding space for KG completion. However, most of the recent studies have concentrated on entity alignment prediction and ignored KG-invariant semantics in multiple KGs that can improve the completion performance. In this paper, we propose a new multiple KG embedding method composed of intra-KG and inter-KG regularization to introduce KG-invariant semantics into KG embedding space using aligned entities between related KGs. The intra-KG regularization adjusts local distance between aligned and not-aligned entities using contrastive loss, while the inter-KG regularization globally correlates aligned entity embeddings between KGs using multi-view loss. Our experimental results demonstrate that our proposed method combining both regularization terms largely outperforms existing baselines in the KG completion task. Mori Kurokawa, Kei Yonekawa, Shuichiro Haruta, Tatsuya Konishi, Hideki Asoh, Chihiro Ono, Masafumi Hagiwara |
ICMLA | 3 |
| 2021 | A Website Fingerprinting Attack based on the Virtual Memory of the Process on Android DevicesabstractWebsite Fingerprinting Attack (WFA) which identifies websites browsed on Android devices is extremely dangerous because it creates an opportunity for stealing private information. As the most feasible WFA method, we focus on a method that can identify a website by using the power consumption model restored from CPU data. However, that is not effective in a real situation where multiple background tasks run because CPU data which are unrelated to browsing are confused. Furthermore, the previous method cannot accurately identify simple websites that are subject to background tasks. Thus, a more feasible method is required to indicate the dangers. In this paper, we propose a website fingerprinting attack based on virtual memory of process on Android device. We focus on the fact that a specific process about browsing websites works when a website is browsed. Because each process has its virtual memory which is independent of each other, the useful feature of a task can be extracted from the virtual memory without noise. Therefore, the proposed method can precisely identify a browsed website by using the virtual memory-based features even if background tasks work. Furthermore, the proposed method can obtain effective information even for a simple website. By computer simulation with a real dataset, we demonstrate that the proposed method can improve up to 86%, 89%, and 82% in precision, recall, and F-measure, respectively for websites which the previous scheme cannot identify at all. Tatsuya Okazaki, Hiroya Kato, Shuichiro Haruta, Iwao Sasase |
APCC | 3 |
| 2021 | An Empirical Study on News Recommendation in Multiple Domain Settings
Shuichiro Haruta, Mori Kurokawa |
MobiQuitous | 1 |
| 2021 | A Study on Metrics for Concept Drift Detection Based on Predictions and Parameters of Ensemble Model
Kei Yonekawa, Shuichiro Haruta, Tatsuya Konishi, Kazuhiro Saito, Hideki Asoh, Mori Kurokawa |
MobiQuitous | 2 |
| 2020 | A Preprocessing Methodology by Using Additional Steganography on CNN-based SteganalysisabstractThere exists a need of “image steganalysis” which reveals whether steganographic signals are embedded in an image to improve information security. Among various steganalysis, Convolutional Neural Networks (CNN) based steganalysis is promising since it can automatically learn the features of diverse steganographic algorithms. However, we discover the detection performance of CNN is degraded when an image is intentionally reduced by the nearest-neighbor interpolation before steganography. This is because spatial frequency in a reduced image gets high, which disturbs the training. In order to overcome this shortcoming, in this paper, we propose a preprocessing methodology by using additional steganography on CNN-based steganalysis. In the proposed preprocessing, steganographic signals are additionally embedded into both reduced original images and reduced steganographic ones since a difference of spatial frequencies between them gets obvious, which helps CNN learn features. Whenever reduced images are trained in CNN or inspected whether they are steganographic ones or not, steganography is applied to them once by the proposed preprocessing. Thus, an image is regarded as a steganographic one if the trained model judges steganography is applied to it twice; otherwise it is an original one. Since the proposed methodology is very simple, its computational cost is low. Our evaluation shows accuracy in a model with the proposed preprocessing is 10.6% higher than that in the conventional one. Besides, even in the situation where another steganography is additionally embedded, the proposed preprocessing yields 7% higher accuracy compared with the conventional one. Hiroya Kato, Kyohei Osuge, Shuichiro Haruta, Iwao Sasase |
GLOBECOM | 3 |
| 2019 | A Novel Visual Similarity-based Phishing Detection Scheme using Hue Information with Auto Updating DatabaseabstractIn this paper, we propose a novel visual similarity-based phishing detection scheme using hue information with auto updating database. Since a PWS (Phishing Website) is created based on targeted legitimate website or other subspecies whose hue information is similar each other, many PWSs can be exhaustively detected by tracing similar colored subspecies. Based on this notion, the proposed scheme detects a new PWS which has similar hue information to already detected PWSs. By repeating this procedure, the detection scope can be effectively expanded. In order to avoid the misdetection of legitimate websites which have similar hue information to database's ones, the proposed scheme utilizes the fact that the combination of used colors is hard to be similar among legitimate websites and PWSs. By the computer simulation with real dataset, we demonstrate that the proposed scheme improves the detection performance as the number of detected PWSs increases. Shuichiro Haruta, Fumitaka Yamazaki, Hiromu Asahina, Iwao Sasase |
APCC | 1 |
| 2019 | Trust-based Verification Attack Prevention Scheme using Tendency of Contents Request on NDNabstractTo realize content distribution, NDN (Named Data Networking) is gathering attention. Since NDN is vulnerable to spreading fake contents, router based verification schemes are proposed to solve this problem. However, routers are vulnerable to the attack which puts a burden to them by verification of contents (verification attack). In order to detect it, the scheme leveraging the fact that the number of the request of unverified contents and the verification of them increase under the attack is proposed. While verification attack can be detected by that scheme, the attack has already occurred. In order to detect the attack before it occurs, in this paper, we propose a trust-based verification attack prevention scheme using tendency of contents request on NDN. We focus on the fact that the access interval to unverified contents tends to be short dramatically just before verification attack occurs. By leveraging this fact, the router determines that verification attack has occurred and restricts requests of all users temporarily. However, in this case, it is impossible to identify attackers, and the requests of legitimate users are also restricted. Therefore, we focus on the fact that legitimate users tend not to request contents in a cache in many cases. Meanwhile, in order to conduct verification attack, attackers need to request such contents for a short time. By giving low trust value to users requesting these contents, a router can identify attackers and restrict only attackers' requests. Our evaluation results show our scheme can detect verification attack before the attack. Furthermore, we clearly demonstrate that our scheme can restrict only attackers' requests. Hironori Nakano, Hiroya Kato, Shuichiro Haruta, Masashi Yoshida, Iwao Sasase |
APCC | 3 |
| 2019 | Android Malware Detection Scheme Based on Level of SSL Server CertificateabstractDetecting Android malware is imperative. As a promising Android malware detection scheme, we focus on the scheme leveraging the differences of traffic patterns between benign apps and malware. Those differences can be captured even if the packet is encrypted. However, since such features are just statistic based ones, they cannot identify whether each traffic is malicious. Thus, it is necessary to design the scheme which is applicable to encrypted traffic data and supports identification of malicious traffic. In this paper, we propose an Android malware detection scheme based on the level of SSL server certificate. Attackers tend to use an untrusted certificate to encrypt malicious payloads in many cases because passing rigorous examination is required to get a trusted certificate. Thus, we utilize SSL server certificate based features for detection since their certificates tend to be untrusted. Furthermore, in order to obtain the more exact features, we introduce required permission based weight values because malware inevitably require permissions regarding malicious actions. By computer simulation with real dataset, we show our scheme achieves an accuracy of 92.7 %. True positive rate and false positive rate are 5.6% higher and 3.3% lower than the previous scheme, respectively. Our scheme can cope with encrypted malicious payloads and 89 malware which are not detected by the previous scheme. Hiroya Kato, Shuichiro Haruta, Iwao Sasase |
GLOBECOM | 2 |
| 2018 | Encounter Record Reduction Scheme based on Theoretical Contact Probability for Flooding Attack Mitigation in DTNabstractDelay Tolerant Network (DTN) is characterized by a lack of end-to-end connectivity. Due to this, detecting flooding attack in DTN is a challenging and important task. Among several schemes against flooding attack in DTN, the scheme using Encounter Record (ER) that consists of past transmission history of each node is gathering attention. Since ER entries are exchanged between nodes, a node can detect an attacker whose transmission rate is too much. Although an attacker may falsify entry to pretend that his/her transmission rate is less than the actual value, it can also be detected through the contradiction between a falsified entry and another entry. However, since a node sends all ER entries in its own buffer regardless of whether an entry is helpful to detect an attacker or not, the energy consumption increases as the number of entries increases. In this paper, we propose an ER reduction scheme based on theoretical contact probability for flooding attack mitigation. We focus on the fact that if the falsified entry does not exist, corresponding entries are not helpful to detect an attacker. Since the falsified entries are propagated over the network with the lapse of time, the probability that there is no contradicting entries over the network gets higher if a node has not received any contradicting entries for a sufficient time. By removing such entries, the energy consumption can be reduced while the effectiveness of ER is kept. By computer simulation, we demonstrate our scheme successfully reduce the energy consumption while the same level of performance is achieved. Keisuke Arai, Shuichiro Haruta, Hiromu Asahina, Iwao Sasase |
APCC | 2 |
| 2017 | Obfuscated malicious javascript detection scheme using the feature based on divided URLabstractOn web application services, detecting obfuscated malicious JavaScript utilized for the attacks such as Drive-by-Download is an urgent demand. Obfuscation is a technique that modifies some elements of program codes and is used to evade the pattern matching of traditional anti-virus softwares. In particular, encode obfuscation is adopted in almost all malicious JavaScript codes as the most effective technique to hide their malicious intents. Therefore, many approaches focus on encode obfuscation to detect malicious JavaScript. However, we point out that malicious JavaScript obfuscated by the techniques except for encode obfuscation can easily evade those approaches. Motivated by the above, in this paper, we first investigated the malicious files that previous schemes cannot detect, and found that some files contain divided URL in their codes. In order to detect such JavaScript codes as malicious, we propose obfuscated malicious JavaScript detection scheme using the feature based on divided URL. We focus on the fact that the segments of URL are declared as variables and connected later. Our scheme stores variables and their contents in the dictionary type object and in the connection parts, verifies that malicious URL can be reconstructed. By the computer simulation with real dataset, we show that our scheme improves the detection effectiveness of the conventional scheme. Shoya Morishige, Shuichiro Haruta, Hiromu Asahina, Iwao Sasase |
APCC | 2 |
| 2017 | Traceroute-based target link flooding attack detection scheme by analyzing hop count to the destinationabstractRecently, the detection of target link flooding attack which is a new type of DDoS (Distributed Denial of Service) is required. Target link flooding attack is used for disconnecting a specific area from the Internet. It is more difficult to detect and mitigate this attack than legacy DDoS since attacking flows do not reach the target region. Among several schemes for target link flooding attack, the scheme focusing on traceroute is gathering attention. The idea behind that is the attacker needs to send traceroute to investigate the topology around targeted region before attack starts. That scheme detects the attack by finding rapid increase of traceroute. However, it cannot work when attacker's traceroute ratio is low. In this paper, we propose traceroute-based target link flooding attack detection scheme by analyzing hop count to the destination. Since the attacker must choose the link flooded to disconnect the target area, the destinations of attacker's traceroutes are concentrated within several hops from the target link while legitimate user's ones are distributed uniformly. By analyzing the number of traceroutes as per hop counts, the change can be emphasized and the attack symptom might be more easily captured. By computer simulations, we first prove the above hypotheses and show that our scheme has more robustness compared with the conventional scheme. Kei Sakuma, Hiromu Asahina, Shuichiro Haruta, Iwao Sasase |
APCC | 3 |
| 2017 | Visual Similarity-Based Phishing Detection Scheme Using Image and CSS with Target Website FinderabstractThe detection of phishing websites and identifying their target are imperative. Among several phishing detection schemes, the scheme using visual similarity is gathering attention. It takes a screenshot of website and stores it to the database. If the inputted website''s screenshot is similar to database''s one, it is judged as phishing. However, if multiple similar websites exist, the first inputted website is regarded as legitimate. As a result, it cannot correctly detect legitimate website and identifying phishing target becomes difficult. As a second shortcoming, if the screenshot of phishing website is locally different from ones in the database, false negative occurs. In this paper, we propose visual similarity-based phishing detection scheme using image and CSS with target website finder. To remedy first shortcoming, we focus on the fact that legitimate websites are often linked by other websites and regard such website as legitimate and store the screenshot and CSS in the database. Since CSS is a file which defines the websites visual contents, attackers often steal legitimate CSS to mimic the legitimate website. Thus, by detecting the website which plagiarizes appearance or CSS of legitimate website, we detect phishing website and its target simultaneously. Moreover, we can alleviate the second shortcoming by using CSS because it is probable that the websites which have locally different appearance use identical CSS. By computer simulation with real dataset, we demonstrate our scheme improves detection accuracy while finding phishing target. Shuichiro Haruta, Hiromu Asahina, Iwao Sasase |
GLOBECOM | 1 |