Haojun Hou

dblp:315/9025 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Mining Long Short-Term Evolution Patterns for Temporal Knowledge Graph Reasoning
Haojun Hou, Chunfeng Yuan, Yihua Huang 0001
ICPR (4)3
2023 STSD: Modeling Spatial Temporal Staticity and Dynamicity in Traffic Forecasting
abstract
As one of the typical tasks of spatial temporal forecasting, traffic prediction has attracted extensive research attention in recent studies. Recent works usually combine time series modeling methods and graph neural networks to capture the temporal dynamic trends and spatial static dependencies together. However, we find that the temporal static characteristic and spatial dynamic correlations do also exist in the traffic flow, which are neglected or partially considered in previous methods. To capture such characteristics, we propose a novel method for traffic prediction, which models the Spatial Temporal Staticity and Dynamicity (STSD) together. Specifically, besides the traditional spatial and temporal encoder, we propose a pre-training module based on self-supervised contrastive learning to learn the location representation and thus capture the temporal static characteristic. Then, we propose a flow-aware dynamic graph learning module to capture the spatial dynamic correlations. Besides, to correct the original graph built from node distance, a feature reconstruction-based static graph learning module is introduced. The original, static, and dynamic graph are combined together in STSD to model the spatial staticity and dynamicity. Extensive experimental results on real-world datasets reveal that STSD can outperform existing traffic forecasting methods.
Haojun Hou, Peiliang Wang, Chunfeng Yuan, Yihua Huang 0001
ICDM2
2023 High-Level Data Abstraction and Elastic Data Caching for Data-Intensive AI Applications on Cloud-Native Platforms
abstract
Nowdays, it is prevalent to train deep learning models in cloud-native platforms that actively leverage containerization and orchestration technologies for high elasticity, low and flexible operation cost, and many other benefits. However, it also faces new challenges and our work is focusing on those related to I/O throughput for training, including complex data access, lack of matching dynamic I/O requirement, and inefficient I/O resource scheduling across different jobs. We proposeFluid, a cloud-native platform that provides DL training jobs with high-level data abstraction calledFluid Datasetto access training data from heterogeneous sources with elastic data acceleration. In addition, it comes with an on-the-fly cache system autoscaler that can match the online training speed and increase the number of cache replicas adaptively to alleviate I/O bottlenecks. To improve the overall performance of multiple DL jobs, Fluid co-orchestrate the data cache and DL jobs by arranging job scheduling in an appropriate order and can also schedule data cache and DL jobs on the same node to realize cache affinity. Experimental results show significant performance improvement of each individual DL job which uses dynamic computing resources with Fluid. For scheduling multiple DL jobs with same datasets, Fluid achieves around 2x performance speedup when integrated with existing widely-used and cutting-edge scheduling solutions through the appropriate job scheduling order. Besides, the cache affinity scheduling policy also improves job execution performance significantly. Fluid is now an open source project hosted by Cloud Native Computing Foundation (CNCF) with many production adopters.
Rong Gu 0001, Yang Che, Haipeng Dai 0001, Haojun Hou, Li Yi 0003, Yihua Huang 0001, Guihai Chen
IEEE Trans. Parallel Distributed Syst.8
2022 Fluid: Dataset Abstraction and Elastic Acceleration for Cloud-native Deep Learning Training Jobs
abstract
Nowdays, it is prevalent to train deep learning (DL) models in cloud-native platforms that actively leverage containerization and orchestration technologies for high elasticity, low and flexible operation cost, and many other benefits. However, it also faces new challenges and our work is focusing on those related to I/O throughput for training, including complex data access with complicated performance tuning, lack of cache capacity with specialized hardware to match its high and dynamic I/O requirement, and inefficient I/O resource sharing across different training jobs. We propose Fluid, a cloud-native platform that provides DL training jobs with a data abstraction called Fluid Dataset to access training data from heterogeneous sources in a unified manner with transparent and elastic data acceleration powered by auto-tuned cache runtimes. In addition, it comes with an on-the-fly cache system autoscaler that can intelligently scale up and down the cache capacity to match the online training speed of each individual DL job. To improve the overall performance of multiple DL jobs, Fluid can co-orchestrate the data cache and DL jobs by arranging job scheduling in an appropriate order. Our experimental results show significant performance improvement of each individual DL job which uses dynamic computing resources with Fluid. In addition, for scheduling multiple DL jobs with same datasets, Fluid gives around 2x performance speedup when integrated with existing widely-used and cutting-edge scheduling solutions. Fluid is now an open source project hosted by Cloud Native Computing Foundation (CNCF) with adopters in production including Alibaba Cloud, Tencent Cloud, Weibo.com, China Telecom, etc.
Rong Gu 0001, Yang Che, Haojun Hou, Haipeng Dai 0001, Li Yi 0003, Guihai Chen, Yihua Huang 0001
ICDE6
2022 AutoGSR: Neural Architecture Search for Graph-based Session Recommendation
abstract
Session-based recommendation aims to predict next click action (e.g., item) of anonymous users based on a fixed number of previous actions. Recently, Graph Neural Networks (GNNs) have shown superior performance in various applications. Inspired by the success of GNNs, tremendous endeavors have been devoted to introduce GNNs into session-based recommendation and have achieved significant results. Nevertheless, due to the highly diverse types of potential information in sessions, existing GNNs-based methods perform differently on different session datasets, leading to the need for efficient design of neural networks adapted to various session recommendation scenarios. To address this problem, we propose Automated neural architecture search for Graph-based Session Recommendation, namely AutoGSR, a framework that provides a practical and general solution to automatically find the optimal GNNs-based session recommendation model. In AutoGSR, we propose two novel GNN operations to build an expressive and compact search space. Building upon the search space, we employ a differentiable search algorithm to search for the optimal graph neural architecture. Furthermore, to consider all types of session information together, we propose to learn the item meta knowledge, which acts as a priori knowledge for guiding the optimization of final session representations. Comprehensive experiments on three real-world datasets demonstrate that AutoGSR is able to find effective neural architectures and achieve state-of-the-art results. To the best of our knowledge, we are the first to study the neural architecture search for the session-based recommendation.
Jingfan Chen, Haojun Hou, Chunfeng Yuan, Yihua Huang 0001
SIGIR3