VLDB 2026 Research / reviewers in the wild / expert
Jiahui Jin 0001
dblp:06/8559-1
· DBLP profile ↗
64ranked-venue papers
11as first author
37since 2021 · last 2026
0000-0001-9570-1456ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 15 · 8 since 2021Systems, architecture and hardware · 14 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 13 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Pricing Online LLM Services with Data-Calibrated Stackelberg Routing GameabstractThe proliferation of Large Language Models (LLMs) has established LLM routing as a standard service delivery mechanism, where users select models based on cost, Quality of Service (QoS), among other things. However, optimal pricing in LLM routing platforms requires precise modeling for dynamic service markets, and solving this problem in real time at scale is computationally intractable. In this paper, we propose PriLLM, a novel practical and scalable solution for real-time dynamic pricing in competitive LLM routing. PriLLM models the service market as a Stackelberg game, where providers set prices and users select services based on multiple criteria. To capture real-world market dynamics, we incorporate both objective factors (eg cost, QoS) and subjective user preferences into the model. For scalability, we employ a deep aggregation network to learn provider abstraction that preserve user-side equilibrium behavior across pricing strategies. Moreover, PriLLM offers interpretability by explaining its pricing decisions. Empirical evaluation on real-world data shows that PriLLM achieves over 95% of the optimal profit while only requiring less than 5% of the optimal solution's computation time. Zhendong Guo, Wenchao Bai, Jiahui Jin 0001 |
AAAI | 3 |
| 2026 | Don't Be Misled by Style: A Style-Adaptive Reranker for Capturing Effective Knowledge in Retrieval-Augmented GenerationabstractRuwen Zhang, Bo Liu, Zhang Sheng Xiang, Yida Chen, Hantao Zhao, Ding Ding, Jiahui Jin, Jiuxin Cao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ruwen Zhang, Bo Liu 0004, Zhang Sheng Xiang, Hantao Zhao, Ding Ding 0002, Jiahui Jin 0001, Jiuxin Cao |
ACL (1) | 7 |
| 2026 | Integrating Heterogeneous Spatio-Temporal Interactions for Traffic Speed PredictionabstractPredicting traffic speed is a crucial task in intelligent transportation systems, as it helps analyze traffic congestion and improve road flow. The complex spatio-temporal interactions present in traffic data make accurate predictions challenging. In recent years, many studies have focused on extracting and learning spatio-temporal features. Deep learning methods, particularly spatio-temporal graph learning models, show promising performance and become the mainstream approach in this area of research. However, existing methods cannot exploit and unify the heterogeneous spatio-temporal interactions hidden in traffic data to achieve multi-correlation modeling. As a result, they cannot effectively model the complex evolving patterns in traffic dynamics. To this end, we propose a heterogeneous spatio-temporal traffic graph learning framework (HSTGL) to capture these diverse spatio-temporal interactions comprehensively. In terms of design, HSTGL consists of three modules: spatio-temporal heterogeneous graph construction, spatio-temporal heterogeneous graph attention learning, and heterogeneous information supplementation. The first two modules utilize similar temporal pattern clustering and heterogeneous spatio-temporal graph attention mechanisms (HSTGAT) to learn heterogeneous spatio-temporal interactions in traffic data. The latter feature fusion module (FFM) is developed to complement potential heterogeneous information in the global spatio-temporal context. Our HSTGL conducts extensive experiments on three real-world public traffic datasets: METR-LA, PEMS-BAY, and PEMSD7M. The results demonstrate that HSTGL achieves superior predictive performance compared to representative benchmark methods with average improvements of 3.3% in MAE, 1.8% in RMSE, and 5.2% in MAPE compared to the optimal baseline. Xigang Sun, Jiahui Jin 0001, Haojia Zhu, Wenchao Bai |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceabstractIn few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their application. Recent Mamba demonstrates efficiency in modeling long sequences, but directly applying Mamba to FSAR overlooks the importance of local feature modeling and alignment. Moreover, long sub-sequences within the same class accumulate intra-class variance, which adversely impacts FSAR performance. To solve these challenges, we propose a Matryoshka MAmba and CoNtrasTive LeArning framework (Manta). Firstly, the Matryoshka Mamba introduces multiple Inner Modules to enhance local feature representation, rather than directly modeling global features. An Outer Module captures dependencies of timeline between these local features for implicit temporal alignment. Secondly, a hybrid contrastive learning paradigm, combining both supervised and unsupervised methods, is designed to mitigate the negative effects of intra-class variance accumulation. The Matryoshka Mamba and the hybrid contrastive learning paradigm operate in two parallel branches within Manta, enhancing Mamba for FSAR of long sub-sequence. Manta achieves new state-of-the-art performance on prominent benchmarks, including SSv2, Kinetics, UCF101, and HMDB51. Extensive empirical studies prove that Manta significantly improves FSAR of long sub-sequence from multiple perspectives. Wenbo Huang 0001, Jinghui Zhang 0001, Guang Li 0008, Lei Zhang 0130, Shuoyuan Wang, Fang Dong 0001, Jiahui Jin 0001, Takahiro Ogawa 0001, Miki Haseyama |
AAAI | 7 |
| 2025 | Exploring Multimodal Relation Extraction of Hierarchical Tabular Data with Multi-task LearningabstractRelation Extraction (RE) is a key task in table understanding, aiming to extract semantic relations between columns.However, complex tables with hierarchical headers are hard to obtain high-quality textual formats (e.g., Markdown) for input under practical scenarios like webpage screenshots and scanned documents, while table images are more accessible and intuitive.Besides, existing works overlook the need of mining relations among multiple columns rather than just the semantic relation between two specific columns in real-world practice.In this work, we explore utilizing Multimodal Large Language Models (MLLMs) to address RE in tables with complex structures.We creatively extend the concept of RE to include calculational relations, enabling multi-task learning of both semantic and calculational RE for mutual reinforcement.Specifically, we reconstruct table images into graph structure based on neighboring nodes to extract graph-level visual features.Such feature enhancement alleviates the insensitivity of MLLMs to the positional information within table images.We then propose a Chain-of-Thought distillation framework with self-correction mechanism to enhance MLLMs' reasoning capabilities without increasing parameter scale.Our method significantly outperforms most baselines on wide datasets.Additionally, we release a benchmark dataset for calculational RE in complex tables. Aibo Song, Jingyi Qiu, Jiahui Jin 0001, Tianbo Zhang, Xiaolin Fang 0001 |
ACL (1) | 4 |
| 2025 | Bilateral Virtual Companions: The Impact of Virtual Humans' Movement and Voice Realism on User Perception and Experience in Multi-user VR CinemasabstractDespite the increasing prevalence of online social interaction, challenges such as insufficient immersion and lack of interactivity still persist. To overcome these limitations, this study developed a multi-user virtual reality (VR) cinema system with motion capture (Mocap) and multi-user VR technology. The system demonstrated strengths in overcoming physical space restrictions and saving travel costs, providing a more enriched interactive experience and fulfilling social needs under special circumstances, and facilitating metaverse applications. Furthermore, to give insight into the further design of multi-user VR cinemas, this study investigated the impact of virtual humans' (VHs') characteristics on user perception and experience of bilateral virtual companions, by evaluating the impact of movement and voice realism on immersion, social presence and intimacy. The results show that while neither movement nor voice realism significantly influences immersion, voice realism rather than movement realism significantly affects social presence and intimacy. Jingfeng Hu, Ding Ding 0002, Xiangyu Xu 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001 |
CSCWD | 5 |
| 2025 | GER-LLM: Efficient and Effective Geospatial Entity Resolution with Large Language ModelabstractGeospatial Entity Resolution (GER) plays a central role in integrating spatial data from diverse sources.However, existing methods are limited by their reliance on large amounts of training data and their inability to incorporate commonsense knowledge.While recent advances in Large Language Models (LLMs) offer strong semantic reasoning and zero-shot capabilities, directly applying them to GER remains inadequate due to their limited spatial understanding and high inference cost.In this work, we present GER-LLM, a framework that integrates LLMs into the GER pipeline.To address the challenge of spatial understanding, we design a spatially informed blocking strategy based on adaptive quadtree partitioning and Area of Interest (AOI) detection, preserving both spatial proximity and functional relationships.To mitigate inference overhead, we introduce a group prompting mechanism with graph-based conflict resolution, enabling joint evaluation of diverse candidate pairs and enforcing global consistency across alignment decisions.Extensive experiments on real-world datasets demonstrate the effectiveness of our approach, yielding significant improvements over state-of-the-art methods.The data and code is available in https Haojia Zhu, Jiahui Jin 0001 |
EMNLP | 3 |
| 2025 | Riding the Wave: Multi-Scale Spatial-Temporal Graph Learning for Highway Traffic Flow Prediction Under Overload ScenariosabstractHighway traffic flow prediction under overload scenarios (HIPO) is a critical problem in intelligent transportation systems, which aims to forecast future traffic patterns on highway segments during periods of exceptionally high demand. Despite its importance, this problem has rarely been explored in recent research due to the unique challenges posed by irregular flow patterns, complex traffic behaviors, and sparse contextual data. In this paper, we propose a Heterogeneous Spatial-Temporal graph network With Adaptive contrastiVE learning (HST-WAVE) to address the HIPO problem. Specifically, we first construct a heterogeneous traffic graph according to the physical highway structure. Then, we develop a multi-scale temporal weaving Transformer and a coupled heterogeneous graph attention network to capture the irregular traffic flow patterns and complex transition behaviors. Furthermore, we introduce an adaptive temporal enhancement contrastive learning strategy to bridge the gap between divergent temporal patterns and mitigate data sparsity. We conduct extensive experiments on two real-world highway network datasets (No. G56 and G60 in Hangzhou, China), showing that our model can effectively handle the HIPO problem and achieve state-of-the-art performance. The source code is available at https://github.com/luck-seu/HST-WAVE. Xigang Sun, Jiahui Jin 0001, Hancheng Wang, Xiangguo Sun |
IJCAI | 2 |
| 2025 | Urban Region Pre-training and Prompting: A Graph-based ApproachabstractUrban region representation is crucial for various urban downstream tasks. However, despite the proliferation of methods and their success, acquiring general urban region knowledge and adapting to different tasks remains challenging. Existing work pays limited attention to the fine-grained functional layout semantics in urban regions, limiting their ability to capture transferable knowledge across regions. Further, inadequate handling of the unique features and relationships required for different downstream tasks may also hinder effective task adaptation. In this paper, we propose a Graph-based Urban Region Pre-training and Prompting framework (GURPP) for region representation learning. Specifically, we first construct an urban region graph and develop a subgraph-centric urban region pre-training model to capture the heterogeneous and transferable patterns of entity interactions. This model pre-trains knowledge-rich region embeddings using contrastive learning and multi-view learning methods. To further refine these representations, we design two graph-based prompting methods: a manually-defined prompt to incorporate explicit task knowledge and a task-learnable prompt to discover hidden knowledge, which enhances the adaptability of these embeddings to different tasks. Extensive experiments on various urban region prediction tasks and different cities demonstrate the superior performance of our framework. Jiahui Jin 0001, Yifan Song 0003, Dong Kan, Haojia Zhu, Xiangguo Sun, Xigang Sun, Jinghui Zhang 0001 |
KDD (2) | 1 |
| 2025 | IM-POI: Bridging ID and Multi-modal Gaps in Next POI RecommendationabstractNext Point-of-Interest (POI) recommendation aims to predict user's subsequent destinations based on historical check-in sequences, thereby enhancing travel experiences. While traditional methods primarily rely on unique identifiers (IDs) to represent POIs, they face data scarcity challenges. Recent multi-modal approaches offer alternatives but struggle with two key issues: inadequate handling of heterogeneity between ID and multi-modal features, and difficulties in unified framework integration, limiting their potential benefits. To address these limitations, we propose IM-POI, a novel framework that leverages the complementary strengths of both ID embeddings and multi-modal representations for next POI recommendation. In our framework, a global POI weighted transition graph inspired by TF-IDF captures sequential dependencies and enhances memorization capabilities, while a geographical graph incorporates spatial information into multi-modal features to be consistent with real-world visitation patterns. To address representation integration, we introduce an IM-Aligner module to prevent representation collapse during distribution matching. Extensive experiments on three real-world datasets demonstrate that IM-POI significantly outperforms state-of-the-art baselines. Jiahui Jin 0001, Xigang Sun, Yukun Ban |
ACM Multimedia | 2 |
| 2025 | Rule-Based Graph Cleaning with GPUs on a Single MachineabstractThis paper studies cost-effective graph cleaning with a single machine. We adopt a rule-based method that may embed machine learning models as predicates in the rules. Graph cleaning with the rules involves rule discovery, error detection and correction. These tasks are both computation-heavy and I/O-intensive as they repeatedly invoke costly graph pattern matching, and produce a large amount of a large volume of intermediate results, among other things. In light of these, no existing single-machine system is able to carry out these tasks even on not-too-large graphs, even using GPUs. Thus we develop MiniClean, a single-machine system for cleaning large graphs. It proposes (1) a workflow that better fits a single machine by pipelining CPU, GPU and I/O operations; (2) memory footprint reduction with bundled processing and data compression; and (3) a multi-mode parallel model for SIMD, pipelined and independent parallelism, and their scheduling to maximize CPU--GPU synergy. Using real-life graphs, we empirically verify that MiniClean outperforms the SOTA single-machine systems by at least 65.34× and multi-machine systems with 32 nodes by at least 8.09×. Wenchao Bai, Wenfei Fan, Shuhao Liu 0001, Kehan Pang, Xiaoke Zhu, Jiahui Jin 0001 |
Proc. ACM Manag. Data | 6 |
| 2025 | An Event-Centric Framework for Predicting Crime Hotspots With Flexible Time Intervals
Jiahui Jin 0001, Yi Hong 0003, Guandong Xu, Jinghui Zhang 0001, Hancheng Wang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | DiG-In-GNN: Discriminative Feature Guided GNN-Based Fraud Detector against Inconsistencies in Multi-Relation Fraud GraphabstractFraud detection on multi-relation graphs aims to identify fraudsters in graphs. Graph Neural Network (GNN) models leverage graph structures to pass messages from neighbors to the target nodes, thereby enriching the representations of those target nodes. However, feature and structural inconsistency in the graph, owing to fraudsters' camouflage behaviors, diminish the suspiciousness of fraud nodes which hinders the effectiveness of GNN-based models. In this work, we propose DiG-In-GNN, Discriminative Feature Guided GNN against Inconsistency, to dig into graphs for fraudsters. Specifically, we use multi-scale contrastive learning from the perspective of the neighborhood subgraph where the target node is located to generate guidance nodes to cope with the feature inconsistency. Then, guided by the guidance nodes, we conduct fine-grained neighbor selection through reinforcement learning for each neighbor node to precisely filter nodes that can enhance the message passing and therefore alleviate structural inconsistency. Finally, the two modules are integrated together to obtain discriminable representations of the nodes. Experiments on three fraud detection datasets demonstrate the superiority of the proposed method DiG-In-GNN, which obtains up to 20.73% improvement over previous state-of-the-art methods. Our code can be found at https://github.com/GraphBerry/DiG-In-GNN. Jinghui Zhang 0001, Zhengjia Xu, Dingyang Lyu 0001, Dian Shen, Jiahui Jin 0001, Fang Dong 0001 |
AAAI | 6 |
| 2024 | Congestion-aware Stackelberg pricing game in urban Internet-of-Things networks: A case study
Jiahui Jin 0001, Zhendong Guo, Wenchao Bai, Biwei Wu, Xiang Liu 0014, Weiwei Wu 0001 |
Comput. Networks | 1 |
| 2024 | Relation-oriented few-shot knowledge graph prototype networks
Yingying Xue, Aibo Song, Jiahui Jin 0001, Jingyi Qiu, Xiaolin Fang 0001, Xiaorui Zhai |
Neurocomputing | 3 |
| 2024 | Learning context-aware region similarity with effective spatial normalization over Point-of-Interest data
Jiahui Jin 0001, Yifan Song 0003, Dong Kan, Binjie Zhang, Jinghui Zhang 0001, Hongru Lu |
Inf. Process. Manag. | 1 |
| 2024 | Federated variational generative learning for heterogeneous data in distributed environments
Wei Xie 0011, Runqun Xiong, Jinghui Zhang 0001, Jiahui Jin 0001, Junzhou Luo |
J. Parallel Distributed Comput. | 4 |
| 2024 | Real-Time Network-Level Traffic Signal Control: An Explicit Multiagent Coordination MethodabstractTraffic signal control (TSC) has been one of the most useful ways for reducing urban road congestion. The challenge of TSC includes 1) real-time signal decision, 2) the complexity in traffic dynamics, and 3) the network-level coordination. Reinforcement learning (RL) methods can query policies by mapping the traffic state to the signal decision in real-time, however, are inadequate for different traffic flow environment. By observing real traffic information, online planning methods can compute the signal decisions in a responsive manner. Unfortunately, existing online planning methods either require high computation complexity or get stuck in local coordination. Against this background, we propose an explicit multiagent coordination (EMC)-based online planning methods that can satisfy adaptive, real-time and network-level TSC. By multiagent, we model each intersection as an autonomous agent, and the coordination efficiency is modeled by a cost function between neighbor intersections. By network-level coordination, each agent exchanges messages of cost function with its neighbors in a fully decentralized manner. By real-time, the message-passing procedure can interrupt at any time when the real time limit is reached and agents select the optimal signal decisions according to current message. Finally, we test our EMC method in both synthetic and real road network datasets. Experimental results are encouraging: compared to RL and conventional transportation baselines, our EMC method performs reasonably well in terms of adapting to real-time traffic dynamics, minimizing vehicle travel time and scalability to city-scale road networks. Wanyuan Wang, Haipeng Zhang 0005, Tianchi Qiao, Jiahui Jin 0001, Zhibin Li 0003, Weiwei Wu 0001, Yichuan Jiang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | Matching Tabular Data to Knowledge Graph with Effective Core Column Set DiscoveryabstractMatching tabular data to a knowledge graph (KG) is critical for understanding the semantic column types, column relationships, and entities of a table. Existing matching approaches rely heavily on core columns that represent primary subject entities on which other columns in the table depend. However, discovering these core columns before understanding the table’s semantics is challenging. Most prior works use heuristic rules, such as the leftmost column, to discover a single core column, while an insightful discovery of the core column set that accurately captures the dependencies between columns is often overlooked. To address these challenges, we introduce Dependency-aware Core Column Set Discovery ( DaCo ), an iterative method that uses a novel rough matching strategy to identify both inter-column dependencies and the core column set. Additionally, DaCo can be seamlessly integrated with pre-trained language models, as proposed in the optimization module. Unlike other methods, DaCo does not require labeled data or contextual information, making it suitable for real-world scenarios. In addition, it can identify multiple core columns within a table, which is common in real-world tables. We conduct experiments on six datasets, including five datasets with single core columns and one dataset with multiple core columns. Our experimental results show that DaCo outperforms existing core column set detection methods, further improving the effectiveness of table understanding tasks. Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Jiaoyan Chen 0001, Xiaolin Fang 0001, Tianbo Zhang |
ACM Trans. Web | 3 |
| 2023 | Time-Aware POI Recommendation Based on Multi-Grained Location GroupingabstractThe task of point-of-interest (POI) recommendation aims to recommend locations to users in location-based applications. Among them, the task of time-aware POI recommendation aims to capture the user’s preferences that change dynamically over time, so as to make more accurate recommendations to users at a specific time. While existing works take into account the spatial, temporal and category context of POIs, they cannot capture user preferences that are more fine-grained than the category granularity. Additionally, RNN-based methods suffer from the problem of long-term dependency when capturing a user’s check-in patterns. To address these challenges, we propose a novel model with POI multi-grained grouping method which captures the user’s co-visit patterns and weekly patterns, to obtain finer-grained POI groups. The model also utilizes the transformer model to capture the user’s check-in preference patterns. We evaluate our model on two real-world datasets, and the experimental results demonstrate the effectiveness of our proposed model. Haoxiang Zhang 0003, Wenchao Bai, Jingyi Ding, Jiahui Jin 0001 |
CSCWD | 4 |
| 2023 | Dependency-Aware Core Column Discovery for Table Understanding
Jingyi Qiu, Aibo Song, Jiahui Jin 0001, Tianbo Zhang, Jingyi Ding, Xiaolin Fang 0001, Jianguo Qian |
ISWC | 3 |
| 2023 | Label Information Enhanced Fraud Detection against Low Homophily in GraphsabstractNode classification is a substantial problem in graph-based fraud detection. Many existing works adopt Graph Neural Networks (GNNs) to enhance fraud detectors. While promising, currently most GNN-based fraud detectors fail to generalize to the low homophily setting. Besides, label utilization has been proved to be significant factor for node classification problem. But we find they are less effective in fraud detection tasks due to the low homophily in graphs. In this work, we propose GAGA, a novel Group AGgregation enhanced TrAnsformer, to tackle the above challenges. Specifically, the group aggregation provides a portable method to cope with the low homophily issue. Such an aggregation explicitly integrates the label information to generate distinguishable neighborhood information. Along with group aggregation, an attempt towards end-to-end trainable group encoding is proposed which augments the original feature space with the class labels. Meanwhile, we devise two additional learnable encodings to recognize the structural and relational context. Then, we combine the group aggregation and the learnable encodings into a Transformer encoder to capture the semantic information. Experimental results clearly show that GAGA outperforms other competitive graph-based fraud detectors by up to 24.39% on two trending public datasets and a real-world industrial dataset from Baidu. Even more, the group aggregation is demonstrated to outperform other label utilization methods (e.g., C&S, BoT/UniMP) in the low homophily setting. Jinghui Zhang 0001, Zhengjie Huang, Weibin Li 0004, Shikun Feng, Ziheng Ma, Yu Sun 0029, Dianhai Yu, Fang Dong 0001, Jiahui Jin 0001, Beilun Wang, Junzhou Luo |
WWW | 10 |
| 2023 | Intra- and inter-semantic with multi-scale evolving patterns for dynamic graph learning
Yingying Xue, Aibo Song, Xiaolin Fang 0001, Jiahui Jin 0001, Xiangguo Sun, Yingxue Zhang 0008 |
Knowl. Based Syst. | 4 |
| 2022 | AirBC: A Lightweight Reputation-based Blockchain Scheme for Resource-constrained UANETabstractUAV Ad-hoc Network (UANET) has been widely used in many fields. However, the collaborative communication and data sharing among multiple UAVs in UANET are often attacked and threatened, due to the limited software and hardware capability of UAVs and the openness of wireless network environment. In this work, we introduce blockchain technology into UANET to enhance its security. Instead of directly adopting traditional blockchain which require huge storage, computation and communication resources, we propose a lightweight reputation-based blockchain scheme for resource-constrained UANET, named AirBC. Firstly, we present a lightweight storage strategy by elimination and compression to reduce storage overhead for UAV nodes. Secondly, we propose an improved reputation-enhanced Practical Byzantine Fault Tolerance (PBFT) consensus, as well as a reputation evaluation scheme based on reliable recording of UAV behaviors. In our scheme, UAVs with high reputation are selected into a miner committee to perform the consensus, thus improving efficiency. Meanwhile, the committee is updated at regular intervals to ensure scalability of UANET. Thirdly, we adopt a weighted proposal voting scheme to enhance the ability of group decision-making for UANET. Finally, to evaluate our approach, simulations are conducted and their results demonstrate that AirBC can reduce 63% storage overhead and 69% consensus latency on average for different scale UANET. Zhoujie Wang, Runqun Xiong, Jiahui Jin 0001, Chuan Liang |
CSCWD | 3 |
| 2022 | Generating Dynamic Urban Traffic Based on Stochastic Origin-Destination MatrixabstractUrban traffic data plays an important role in urban transportation planning. Due to the scarcity of real-life urban traffic data, many transportation planning applications need to generate synthesized traffic flows based on the real-life trajectory datasets. However, those synthesized traffic flows can only fit the input trajectories, which are static and does not reflect the real traffic distributions. In this paper, we use a stochastic origin-destination (OD) matrix to represent the density of the dynamic traffic flows and then develop a dynamic traffic flow generator. We extract the stochastic OD matrix from the trajectory data, design an efficient neural network to the predict successive stochastic OD matrices, and deploy our model on a real-world road network. The proposed model surpasses the existing generative model in RMSE, MAE, VAR, KL indicators, and is significantly better than the existing model in the MAE indicator. Our traffic generator is able to dynamically adjust urban traffics to generate different simulation environments. Zhichang Wang, Binjie Zhang, Nu Xia, Jiahui Jin 0001 |
CSCWD | 4 |
| 2022 | MetaSync: Relaxing Synchronization Frequency in Distributed Urban Traffic SimulationabstractRecently, city-scale microscopic simulators are developed to analyze data-driven algorithms and strategies. Since the simulators usually require intensive computing and memory resources, distributed approaches are applied to these simulators. However, distributed simulators need to frequently synchronize workers and transmit detailed vehicle information between workers. This process consumes an unacceptable amount of time. In this paper, we speed up the simulation through lowering the synchronization frequency by MetaSync. MetaSync transmits the distributions of vehicles on roads instead of the detailed vehicle information. This allows a worker to self-generate vehicles according to the distributions without frequently interacting with other workers. We also develop a simulator using MetaSync. Results show it is 50% faster than normal synchronization and provides competitive simulation accuracy. Our strategy also enables distributed simulators to run microscopic simulations on a city-scale map with dynamic route planning in an acceptable amount of time. Haojia Zhu, Jiahui Jin 0001 |
CSCWD | 2 |
| 2022 | Aggregate Queries on Knowledge Graphs: Fast Approximation with Semantic-aware SamplingabstractA knowledge graph (KG) manages large-scale and real-world facts as a big graph in a schema-flexible manner. Aggregate query is a fundamental query over KGs, e.g., “what is the average price of cars produced in Germany?”. Despite its importance, answering aggregate queries on KGs has received little attention in the literature. Aggregate queries can be supported based on factoid queries, e.g., “find all cars produced in Germany”, by applying an additional aggregate operation on factoid queries' answers. However, this straightforward method is challenging because both the accuracy and efficiency of factoid query processing will seriously impact the performance of aggregate queries. In this paper, we propose a “sampling-estimation” model to answer aggregate queries over KGs, which is the first work to provide an approximate aggregate result with an effective accuracy guarantee, and without relying on factoid queries. Specifically, we first present a semantic-aware sampling to collect a high-quality random sample through a random walk based on knowledge graph embedding. Then, we propose unbiased estimators for COUNT, SUM, and a consistent estimator for AVG to compute the approximate aggregate results based on the random sample, with an accuracy guarantee in the form of confidence interval. We extend our approach to support iterative improvement of accuracy, and more complex queries with filter, GROUP-BY, and different graph shapes, e.g., chain, cycle, star, flower. Extensive experiments over real-world KGs demonstrate the effectiveness and efficiency of our approach. Yuxiang Wang 0001, Arijit Khan 0001, Jiahui Jin 0001, Qifan Hong |
ICDE | 4 |
| 2022 | Multi-Step Occurrence Prediction of Urban Crimes with Enhanced GRU-ODE-Bayes ModelabstractCrime occurrence prediction is an important task for public security and urban governance. In real world, some crimes are quite harmful, thus it’s necessary to predict these crimes in the next few days (weeks). However, these harmful crime events may have relatively lower frequency, and their records are discontinuous and sporadic, making the multistep occurrence prediction difficult. Due to the sparsity and irregularity of the crime sequences, most of the existing crime prediction techniques do not distinguish crime types and have difficulties to predict low-frequency crimes’ occurrences. In addition, they cannot accurately predict the occurrence of crimes over multiple recent time slots due to sporadic crime records. In this paper, we propose Multi-Time-Slot Crime Occurrence Prediction (MCOP) model, which aims to (1) predict crime types separately and (2) predict crime occurrences in the next multiple time slots. Specifically, MCOP uses a GRU-ODE-Bayes model to handle the discontinuous and sporadic crime events sequences. MCOP is enhanced by a data augmentation to address the data sparsity problem, and exploits the near repeat phenomenon to enable continuous crime predictions. We evaluated our model using real-world urban data, and the results showed our model outperforms the baseline techniques. Binjie Zhang, Zijia Miao, Jiahui Jin 0001 |
SMC | 5 |
| 2022 | A random walk sampling on knowledge graphs for semantic-oriented statistical tasks
Qifan Hong, Yuxiang Wang 0001, Jiahui Jin 0001, Xinle Xuan |
Data Knowl. Eng. | 4 |
| 2022 | Identifying critical nodes in power networks: A group-driven framework
Aibo Song, Yingying Xue, Jiahui Jin 0001 |
Expert Syst. Appl. | 5 |
| 2022 | FedAda: Fast-convergent adaptive federated learning in heterogeneous mobile edge computing environment
Jinghui Zhang 0001, Jiahui Jin 0001, Aibo Song, Wei Zhao 0023, Liangsheng Wen |
World Wide Web | 6 |
| 2021 | PipePar: A Pipelined Hybrid Parallel Approach for Accelerating Distributed DNN TrainingabstractLarge scale DNN training tasks are exceedingly compute-intensive and time-consuming, which are usually executed on highly-parallel platforms. Data and model parallelization is a common way to speed up the training progress across devices. However, they tend to achieve sub-optimal performance due to the communication overheads and unbalanced load among servers. Recent emerging pipelining solutions mitigate the above issues, incorporating the advantages of data and model parallelism. In this paper, we make a step further towards optimizing the execution of pipelining. We introduce PipePar, a pipeline-parallel DNN training method that provides optimized execution strategies of layer-stacked DNNs. PipePar considers the entire tensor partition space of pipelining and explores potential hybrid parallel configurations of each stage in the pipeline. Additionally, we notice the network heterogeneity between different GPU servers and it is inevitable to transfer tensors with different bandwidths and latency. So, taking into account both computation and communication capacity of different GPU servers, PipePar is intended to find a elastic load distribution strategy at different levels. We evaluate PipePar with a set of real-world DNNs on 4 GPU servers. Our experimental results show that PipePar is able to find an efficient strategy that are up to 2.16× faster than state-of-the-art hybrid parallelization approaches. Jiange Li, Jinghui Zhang 0001, Jiahui Jin 0001, Fang Dong 0001 |
CSCWD | 4 |
| 2021 | Region-aware POI Recommendation with Semantic Spatial GraphabstractThe development of Location-Based Social Networks (LBSNs) offers an opportunity for understanding user preferences and promoting Point-of-Interest (POI) recommendation. The user preferences usually change when the environment changes, which will affect the performance of POI recommendation. Many existing methods extract the environmental features from static regions, but they cannot capture user preferences in real time with the fine-grained changes in user locations. Meanwhile, the similarity of POI categories, which is significant to capture user preferences, is usually ignored. To address these issues, we propose RegDM, a region-aware POI recommendation model that employs a semantic spatial graph to model the relations among POIs. With the semantic spatial graph, RegDM uses a Graph Neural Network (GNN) to extract fine-grained region features and user preferences for personalized recommendation. We evaluate RegDM on two datasets, and the experiment results demonstrate the effectiveness of our model. Jiakai Tang, Jiahui Jin 0001, Zijia Miao, Binjie Zhang, Jinghui Zhang 0001 |
CSCWD | 2 |
| 2021 | Building Portable ECG Classification Model with Cross-Dimension Knowledge Distillation
Renjie Tang, Junbo Qian, Jiahui Jin 0001, Junzhou Luo |
ICA3PP (2) | 3 |
| 2021 | Revenue Maximization of Electric Vehicle Charging Services with Hierarchical Game
Biwei Wu, Xiaoxuan Zhu, Xiang Liu 0014, Jiahui Jin 0001, Runqun Xiong, Weiwei Wu 0001 |
WASA (2) | 4 |
| 2021 | Relation-based multi-type aware knowledge graph embedding
Yingying Xue, Jiahui Jin 0001, Aibo Song, Yingxue Zhang 0008 |
Neurocomputing | 2 |
| 2021 | Top-k star queries on knowledge graphs through semantic-aware bounding match scores
Yuxiang Wang 0001, Xiaoliang Xu 0001, Qifan Hong, Jiahui Jin 0001, Tianxing Wu 0001 |
Knowl. Based Syst. | 4 |
| 2020 | Semantic Guided and Response Times Bounded Top-k Similarity Search over Knowledge GraphsabstractRecently, graph query is widely adopted for querying knowledge graphs. Given a query graph GQ, the graph query finds subgraphs in a knowledge graph G that exactly or approximately match GQ. We face two challenges on graph query: (1) the structural gap between GQand the predefined schema in G causes mismatch with query graph, (2) users cannot view the answers until the graph query terminates, leading to a longer system response time (SRT). In this paper, we propose a semantic-guided and response-time-bounded graph query to return the top-k answers effectively and efficiently. We leverage a knowledge graph embedding model to build the semantic graph SGQ, and we define the path semantic similarity (pss) over SGQas the metric to evaluate the answer's quality. Then, we propose an A* semantic search on SGQto find the top-k answers with the greatest pss via a heuristic pss estimation. Furthermore, we make an approximate optimization on A* semantic search to allow users to trade off the effectiveness for SRT within a user- specific time bound. Extensive experiments over real datasets confirm the effectiveness and efficiency of our solution. Yuxiang Wang 0001, Arijit Khan 0001, Tianxing Wu 0001, Jiahui Jin 0001, Haijiang Yan |
ICDE | 4 |
| 2020 | A Topology-Adaptive Deep Model for Power System Multifault DiagnosisabstractQuickly identifying faulty sections is tremendously important for power systems, yet challenging due to handling the variations of complex alarm patterns. Existing works have focused on finding fault section clues solely from alarm information (and ignoring power system topology information). So they are only sensitive to alarms from power systems with pre-assumed topology structures, and encounter difficulties when a system's topology changes. To adapt to unknown or varying system topologies, here we present a Topology-Adaptive Deep Model (TADM) for power system multifault diagnosis. TADM mines the underlying mapping from alarm and topology information to each section's fault status. It consists of a deep iterative network (DIN), a one-layer fully connected network (FCN), and section-wise multifault diagnosis (SWMD) subnetwork. TADM first models a fault power system as a graph, from which DIN iteratively integrates the alarm and topology information in the region from each node to its T -hop neighbors, and learns their local correlation. Limited to T 's size, FCN then combines all local correlations to determine the global correlation between alarm and topology information across the entire power system. To implement multifault diagnosis, learned local and global correlations serve as topology-related fault representations for input as an SWMD (to predict all sections' fault states one by one). A comprehensive experimental study demonstrates that TADM outperforms state-of-the-art models in both multifault diagnosis and adapting to system topologies. The source code of the TADM is available onlline1. Aibo Song, Jiahui Jin 0001, Mingyu Zhai, Yingying Xue |
SMC | 3 |
| 2020 | Optimizing execution for pipelined-based distributed deep learning in a heterogeneously networked GPU clusterabstractSummary Exorbitant resources (computing and memory) are required to train a deep neural network (DNN). Often researchers deploy an approach that uses distributed parallel training to acquire larger models faster on GPUs. This approach has its detriments, though; on one hand, a GPU's expanded capacity to compute also produces bigger bottlenecks in inter‐GPU's communications during model training, and multi‐GPU systems lead to complex connectivity. Workload schedulers then end up having to consider hardware topology and requirements for workload communication, in hopes of allocating GPU resources to optimize execution time and improve usage in a heterogeneous environment. On the other hand, the high memory requirements to train a DNN model make running the training processes on GPUs onerous. To contend with this, we introduce two execution optimization methods based on pipeline‐hybrid parallelism (using both data and model parallelism) in a GPU cluster with heterogeneous networking. First, we propose a model partition algorithm that accelerates pipeline‐hybrid parallelism training between heterogeneously network‐connected GPUs. Second, we introduce a cost‐balanced recomputing algorithm to reduce memory usage in the pipeline mode. Experiments show that our solution (Pipe‐Torch) averages a speedup of 1.4× compared with data parallelism, and reduces the memory footprint while maintaining pipelined load‐balanced training. Jinghui Zhang 0001, Jiange Li, Jiahui Jin 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | Intelligent online catastrophe assessment and preventive control via a stacked denoising autoencoder
Mingyu Zhai, Jiahui Jin 0001, Aibo Song, Jikeng Lin, Zhiang Wu 0001, Yixin Zhao |
Neurocomputing | 3 |
| 2020 | Accelerating Skycube Computation with Partial and Parallel Processing for Service SelectionabstractRecently researchers use skyline techniques to optimize service selection procedure, where they can filter those low-quality web services from the large amount of candidates and return a much smaller high-quality service set. The skycube concept is adopted for quickly responding to the skyline queries with different combinations of Quality of Web Service (QoWS) parameters. As the skycube computation is quite time-consuming, it is a compelling challenge to accelerate this procedure. However, the current solutions usually have a number of redundant computations which will significantly affect the efficiency. To address such drawbacks, after an in-depth analysis of skycube computation procedure, we introduce a partial skycube, which only consists of the skylines with frequently used combinations of QoWS. Then the computational relationships between the skyline on one subspace and its parent-space are studied. Based on the relationships, we develop ParCube algorithm to speedup partial skycube computation by reusing the intermediate comparison results. Meanwhile, at the execution phase, ParCube can be further optimized with parallel execution mode and optimized scheduling strategy. Finally, we evaluate the efficiency and scalability of ParCube on both single machine and cluster environment. The results show that ParCube can efficiently compute partial skycube and scale well in cluster environment. Fang Dong 0001, Junzhou Luo, Jiahui Jin 0001, Jiyuan Shi, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2020 | Facilitating Application-Aware Bandwidth Allocation in the Cloud with One-Step-Ahead Traffic InformationabstractBandwidth allocation to virtual machines (VMs) has a significant impact on the performance of communication-intensive big data applications hosted in VMs. It is crucial to accurately determine how much bandwidth to be reserved for VMs and when to adjust it. Past approaches typically resort to predicting the long-term network demands of applications for bandwidth allocation. However, lacking of prediction accuracy, these methods lead to the unpredictable application performance. Recently, it is conceded that the network demands of applications can only be accurately derived right before each of their execution phases. Hence, it is challenging to timely allocate the bandwidth to VMs with limited information. In this paper, we design and implement AppBag, an Application-aware Bandwidth guarantee framework, which allocates the accurate bandwidth to VMs with one-step-ahead traffic information. We propose an algorithm to allocate the bandwidth to VMs and map them onto feasible hosts. To reduce the overhead when adjusting the allocation, an efficient Lazy Migration (LM) algorithm is proposed with bounded performance. We conduct extensive evaluations using real-world applications, showing that AppBag can handle the bandwidth requests at run-time, while reducing the execution time of applications by 47.3 percent and the global traffic by 36.7 percent, compared to the state-of-the-art methods. Dian Shen, Junzhou Luo, Fang Dong 0001, Jiahui Jin 0001, Junxue Zhang 0001, Jun Shen 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2019 | MBECN: Enabling ECN with Micro-burst Traffic in Multi-queue Data CenterabstractModern multi-queue data centers often use the standard Explicit Congestion Notification (ECN) scheme to achieve high network performance. However, one substantial drawback of this approach is that micro-burst traffic can cause the instantaneous queue length to exceed the ECN's threshold, resulting in numerous mismarkings. After enduring too many mismarkings, senders may overreact, leading to severe throughput loss. As a solution to this dilemma, we propose our own adaptationthe Micro-burst ECN (MBECN) scheme-to mitigate mismarking. MBECN finds a more appropriate threshold baseline for each queue to absorb micro-bursts, based on steady-state analysis and an ideal generalized processor sharing (GPS) model. By adopting a queue-occupation-based dynamically adjusting algorithm, MBECN effectively handles packet backlog without hurting latency. Through testbed experiments, we find that MBECN improves throughput by ~20% and reduces flow completion time (FCT) by ~40%. Using large scale simulations, we find that throughput can be improved by 1.5~2.4× with DCTCP and 1.26~1.35× with ECN*. We also measure network delay and find that latency only increases by 7.36%. Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Junzhou Luo, Zhiang Wu 0001 |
CLUSTER | 3 |
| 2019 | QAECN: Dynamically Tuning ECN Threshold with Micro-burst in Multi-queue Data CentersabstractPacket loss is a common problem in data center networks. The factors causing packet loss are various. Among them, micro-burst is the most important reason. Some previous works have studied the causes and influence o f micro-burst in single queue data center. However, through simulations and experiments, we find that micro-burst could bring m ore serious performance degradation in multi-queue data centers. The micro-burst traffic could cause E CN marking ratio rising from 4% to 22%, and cause throughput loss by up to 40%. Through observing queue length, we find that the standard E CN, which adopts immutable threshold, is not suitable for micro-burst traffic because micro-burst could trigger spurious congestion signals frequently, especially in DCTCP. In this paper, we not only show how much influence the micro-burst brings, but also propose Queue-length Aware ECN (QA-ECN) scheme to mitigate micro-burst. Finally, the simulations and experiments show that QAECN could reduce ECN marking ratio to 2.5%. In addition, the throughput and flow completion time could be improved by up to 22.9% and 34.1%, respectively. Kexi Kang, Jinghui Zhang 0001, Jiahui Jin 0001, Dian Shen, Runqun Xiong, Junzhou Luo |
CSCWD | 3 |
| 2019 | Graph partition-based data and task co-scheduling of scientific workflow in geo-distributed datacentersabstractSummary Most large‐scale scientific workflows take place in multiple collaborative datacenters for access to community‐wide resources, while adhering to each datacenter's non‐uniform resource limits. However, moving both initial input datasets with predetermined locations and intermediate datasets needing placement decisions across geo‐distributed datacenters hinders efficient execution of large‐scale data‐intensive scientific workflows. Thus, scientific workflow's data and task co‐scheduling deal with situations such as pre‐placed initial input datasets, placement of intermediate datasets and each datacenter's non‐uniform computation and storage constraint, while minimizing the cross‐datacenter data transfer. Since this scheduling problem is known to be NP‐hard, here, we propose a novel approach, based on the multilevel graph coarsening and uncoarsening framework, together with a specialized hybrid genetic algorithm having distinctive graph partition driven features of repair and local improvement, for scheduling data‐intensive scientific workflows in geo‐distributed datacenters and optimizing the cross‐datacenter data transfer volume. Extensive simulations, based on four real‐world workflow traces, show that our algorithm significantly reduces the overall geo‐distributed data transfer and demonstrate its effectiveness. Jinghui Zhang 0001, Jiahui Jin 0001, Aibo Song |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | A data-locality-aware task scheduler for distributed social graph queries
Jiahui Jin 0001, Junzhou Luo, Mingyang Du, Yongcheng Dang, Jinghui Zhang 0001, Aibo Song |
Future Gener. Comput. Syst. | 1 |
| 2019 | Offloading Delay Constrained Transparent Computing Tasks With Energy-Efficient Transmission Power Scheduling in Wireless IoT EnvironmentabstractBillions of lightweight Internet of Things (IoT) devices have been deployed for various applications nowadays. Most of them first collect interested data and then process them in some degree according to application requirements. Transparent computing (TC) is a promising technique that makes such lightweight devices suitable to process even large-size applications. The advantage of TC is to separate code storage from its execution, allowing IoT devices to load code blocks from nearby TC storage server on demand. Distinct from existing work, this paper allows the TC IoT devices to offload some tasks to servers, since wireless IoT devices are usually powered by batteries, having limited energy resources. If a task is offloaded, a challenging problem is that its input data collected by the IoT device must be transferred as well, which incurs additional transmission time and energy. This paper proposes a two-step approach aiming at minimizing the energy consumption of the IoT device while satisfies the delay constraint. This approach first studies the offloading decision problem that determines for each task whether to offload task data or load task code blocks, while loading code indicates code receiving and executing energy cost. Second, the transmission power scheduling problem is investigated to further reduce offloading energy for a given delay constrained offloading task set. Heuristic decision making algorithms and optimal power scheduling algorithm are proposed, respectively. Such two-step approach is shown by extensive simulation to be near optimal for the original problem thanks to the optimal design of the power scheduling algorithm. Feng Shan, Junzhou Luo, Jiahui Jin 0001, Weiwei Wu 0001 |
IEEE Internet Things J. | 3 |
| 2019 | GStar: an efficient framework for answering top-k star queries on billion-node knowledge graphs
Jiahui Jin 0001, Junzhou Luo, Samamon Khemmarat, Fang Dong 0001, Lixin Gao 0001 |
World Wide Web | 1 |
| 2018 | An Effective Model for Edge-Side Collaborative Storage in Data-Intensive Edge ComputingabstractEdge Computing is a new computing paradigm that performs data processing at the edge of the network (i.e., edge servers) to lower data processing latency. Existing research works have paid lots of attention to how to offload computation tasks from terminals to edge servers, but most of them ignored how to store tasks' necessary data like pretrained models or databases on edge servers. Recently, the data-intensive tasks like deep learning and augmented reality are becoming common, which need large data storages and powerful computation resources. This leads to a cumbersome challenge, since many lightweight edge servers have limited resources. If an edge server does not have a task's necessary data, it needs to offload the task to cloud data centers or download the necessary data from the cloud. Both cases could increase the data processing latency. To address this problem, this paper proposes an edge-side collaborative storage framework (ECS). In ECS, the edge servers collaboratively store and process data-intensive tasks' necessary data. Particularly, if an edge server does not have the necessary data, it will forward the task to the nearest servers that contain the data. An effective iterative data placement algorithm is also proposed to improve ECS's performance. The experimental results show that ECS is 2× better than the traditional non-shared storage framework in terms of the cache hit rate. Junzhou Luo, Jiahui Jin 0001, Runqun Xiong, Fang Dong 0001 |
CSCWD | 3 |
| 2018 | Cooperative storage by exploiting graph-based data placement algorithm for edge computing environmentabstractSummary Edge computing is a new computing paradigm that performs data processing at the edge of the network (ie, edge servers) to lower data processing latency. Prior research significantly focused on offloading tasks from terminals to edge servers, yet most ignored how to store task's necessary data (such as databases and pretrained machine‐learning models) on edge servers. Today, as data‐intensive tasks such as deep learning and augmented reality become common, large data storage and powerful computation resources are needed. This is a cumbersome challenge, because many lightweight edge servers have limited resources. If an edge server does not have a task's necessary data, then it needs to offload the task to cloud datacenters or download the necessary data from the cloud. Either case could increase data processing latency. To address this problem, this paper proposes an edge‐side collaborative storage framework called Edge‐side Cooperative Storage (ECS). In ECS, edge servers collaboratively store and process data‐intensive tasks's necessary data. Here, we particularly focus on how to effectively place data on ECS, using an approach that differs from existing works (that model data placement problems as linear/integer programming problems). Our work models cooperative storage as a graph and solves the data placement problem by using a graph‐based iterative algorithm. This algorithm easily extends to a distributed version, so distributed ECSworks efficiently without a centralized scheduler. We also evaluate ECS's effectiveness and convergence through simulations. Simulation results show that ECSis 2× better than a traditional nonshared storage framework in terms of the cache hit rate. Jiahui Jin 0001, Junzhou Luo |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Skew-aware online aggregation over joins through guided samplingabstractSummary Online aggregation is a query processing technique that returns approximate answers with error guarantees (in the form of confidence intervals) continuously during the query execution process. This approach offers users a suitable tradeoff between query efficiency and accuracy. The key issue of online aggregation is how to ensure a random sample collection's efficiency and effectiveness. However, the often‐used “blind” sampling method does not adequately consider dataset statistics and other useful information, leading to inefficient sampling and poor sample quality. This becomes a glaring performance issue for skewed data distribution over joins. To alleviate this problem, we utilize dataset statistics to propose a new “guided” sampling approach, which consists of a logic‐partition‐based weighted Gaussian sampling method tailored for the skewed join key, as well as a two‐level sample allocation method that applies to the skewed measured value. Extensive experiments using the TPC‐H benchmark for skewed data distribution demonstrate our solution's superior performance. Yuxiang Wang 0001, Jiahui Jin 0001, Longbin Zhang |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | HaDaap: A hotness-aware data placement strategy for improving storage efficiency in heterogeneous Hadoop clustersabstractSummary Enterprises increasingly use the Hadoop Distributed File System (HDFS) to manage and store big data for many applications. However, HDFS uses triple replication, leading to staggering data center storage costs. As big data increases in volume and its heat levels becomes more sensitive, there comes a point where storing so much cold data actually makes it less accessible and more expensive. Meanwhile, as data centers expand, the heterogeneity of nodes also becomes an issue. Rack‐aware data placement adopted by HDFS results in an unbalanced load and uneven resource allocation because it ignores the data nodes' heterogeneity. Here, we attempt to resolve these problems by proposing a hotness‐aware data placement strategy (named HaDaap). In HaDaap, the first step is to use a hotness‐aware data clustering algorithm to set the data's degree of heat. Then, cold data (with a redundancy of erasure code) are placed through a Double Sort Exchange algorithm to reduce storage costs and increase data availability. Finally, hot data are placed via a dynamic replication placement mechanism that comprehensively factors availability, load, and storage costs. Experimental results show that with these enhancements, HaDaap uses resources rationally and substantially reduces storage costs by considering the difference of data hotness in heterogeneous Hadoop clusters. Runqun Xiong, Jiahui Jin 0001, Junzhou Luo |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Virtual network fault diagnosis mechanism based on fault injectionabstractDiagnosing faults in virtual networks is always a popular research area. Existing researches primarily focus on diagnosing faults in physical networks, while they could not identify the faults introduced by virtual networks. Besides, the high complexity of algorithms and the requirement for modifying hardware may limit their scope of use. To address these drawbacks, in this paper, we propose a novel approach to diagnose faults in virtual networks. The rational of our approach is that the faults can be identified when located in the packet traces, with the knowledge that the possible known faults that can happen in that location. To achieve this goal, we apply packet marking, fault injection and machine learning techniques to provide precise fault diagnosis. Experimental results show that our approach can efficiently identify 73% of the faults while for virtual network-specific faults, our approach can diagnose 86% of them. Our system can also support real-time or near real-time fault analysis. Fang Dong 0001, Dian Shen, Runqun Xiong, Jiahui Jin 0001 |
CSCWD | 5 |
| 2017 | GScheduler: Optimizing resource provision by using GPU usage pattern extraction in cloud environmentsabstractGPU-based clusters are widely chosen for accelerating a variety of scientific applications in high-end cloud environments. With their growing popularity, there is a necessity for improving the system throughput and decreasing the turnaround time for co-executing applications on the same GPU device. However, resource contention among multiple applications on a multi-tasked GPU leads to the performance degradation of applications. Previous works are not accurate enough to learn the characteristics of GPU application before execution, or cannot get such information timely, which may lead to misleading scheduling decisions. In this paper, we present GScheduler, a framework to detect and reduce interference for co-executing applications on the GPU-based cloud. The most important feature of GScheduler is to utilize GPU usage pattern extractor for detecting interference between applications. It is composed of key function-call graph extractor and key GPU resource usage vector extractor, the former is used to detect the similarity of GPU usage mode between applications, while the latter is used to calculate the similarity of GPU resource requirements in-between. In addition, an interference aware scheduler is proposed to minimize the interference. We evaluated our framework with 26 diverse, real-world CUDA applications. When compared with state-of the-art interference-oblivious schedulers, our framework improves system throughput by 36% on average, and achieves a 30.5% reduction of turnaround time on average. Zhuqing Xu, Fang Dong 0001, Jiahui Jin 0001, Junzhou Luo, Jun Shen 0001 |
SMC | 3 |
| 2017 | Enabling application-aware flexible graph partition mechanism for parallel graph processing systemsabstractSummary With the emerging of the large‐scale graph data,Pregel‐like graph parallel processing systems have been an essential tool to efficiently process the graph data. The first step to use thePregel‐like systems is to partition the graph into multiple blocks and distribute them on multiple machines. The partition strategy plays a significant role in determining the performance because a good partition could both ensure load balance and optimize network communication overhead, and vice versa. However, existing partition strategies fail to meet the requirements because they suffer from the following drawbacks: (1) they ignore the application features and (2) they ignore the multi‐application feature in productive environment. To overcome those drawbacks, we proposed thesuperblockpartition strategy, which utilizes theatomic blocksgenerated by pre‐processing of the original graph and could be constructed and re‐constructed dynamically according to the submitted applications in real time. The hash‐based and clustering‐based pre‐partition methods are covered in details. The application feature extraction method and heuristicsuperblockpartition algorithm are proposed to construct the superblocks. Experimental results show that thesuperblockpartition strategy could boost the graph processing performance and its partition efficiency also outperforms the hash‐based and topology optimal partition strategy. Copyright © 2016 John Wiley & Sons, Ltd. Fang Dong 0001, Junxue Zhang 0001, Junzhou Luo, Dian Shen, Jiahui Jin 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Querying Web-Scale Knowledge Graphs Through Effective Pruning of Search SpaceabstractWeb-scale knowledge graphs containing billions of entities are common nowadays. Querying these graphs can be modeled as a subgraph matching problem. Since knowledge graphs are incomplete and noisy in nature, it is important to discover answers matching exactly as well as answers similar to queries. Existing graph matching algorithms usually use graph indices to accelerate query processing. For billion-node graphs, it may be infeasible to build the graph indices due to the amount of work and the memory/ storage required. In this paper, we propose an efficient algorithm for finding the best k answers for a given query without precomputing graph indices. An answer's quality is measured by a matching score that is computed online. To accelerate query processing, we propose a novel technique for bounding the matching scores during the computation. By using bounds, the low quality answers can be efficiently pruned. The bounding technique can be implemented in a distributed environment, allowing our approach to efficiently query web-scale knowledge graphs. We evaluate the effectiveness and the efficiency of our approach on real-world datasets. The result shows that our bounding technique can reduce the running time up to two orders of magnitude comparing to an approach without using bounds. Jiahui Jin 0001, Junzhou Luo, Samamon Khemmarat, Lixin Gao 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | Resource provisioning optimization for service hosting on cloud platformabstractWith the popularity of cloud computing technology, service hosting is used as a typical model to deploy different kinds of services on cloud platform. In recent years, how to effectively provide resources for service hosting has attracted more and more attention. However, most of the existing works only focused on how to effectively provide virtual machines for service hosting. They ignored how to efficiently place these virtual machines into physical servers, when considering multidimensional resource requirements. This may result in unreasonable virtual machine placement in servers, thereby causing the underutilization of resource. To address this problem, we propose a novel resource provisioning method including virtual machine provisioning for hosting service and virtual machine placement in servers. The proposed method decides how many virtual machines should be provided for each service by utilizing queuing theory. Then based on the virtual machines to be provided, the proposed method models the virtual machine placement problem as a variant of cutting stock problem, and decides how many servers should be provided by solving this problem. The proposed method is evaluated by simulations. Experimental results show the proposed method achieves a better performance than these baseline methods. Jiyuan Shi, Fang Dong 0001, Jinghui Zhang 0001, Jiahui Jin 0001, Junzhou Luo |
CSCWD | 4 |
| 2015 | Querying Web-Scale Information Networks Through Bounding Matching ScoresabstractWeb-scale information networks containing billions of entities are common nowadays. Querying these networks can be modeled as a subgraph matching problem. Since information networks are incomplete and noisy in nature, it is important to discover answers that match exactly as well as answers that are similar to queries. Existing graph matching algorithms usually use graph indices to improve the efficiency of query processing. For web-scale information networks, it may not be feasible to build the graph indices due to the amount of work and the memory/storage required. In this paper, we propose an efficient algorithm for finding the best k answers for a given query without precomputing graph indices. The quality of an answer is measured by a matching score that is computed online. To speed up query processing, we propose a novel technique for bounding the matching scores during the computation. By using bounds, we can efficiently prune the answers that have low qualities without having to evaluate all possible answers. The bounding technique can be implemented in a distributed environment, allowing our approach to efficiently answer the queries on web-scale information networks. We demonstrate the effectiveness and the efficiency of our approach through a series of experiments on real-world information networks. The result shows that our bounding technique can reduce the running time up to two orders of magnitude comparing to an approach that does not use bounds. Jiahui Jin 0001, Samamon Khemmarat, Lixin Gao 0001, Junzhou Luo |
WWW | 1 |
| 2014 | A distributed approach for top-k star queries on massive information networksabstractMassive information networks, such as the knowledge graph by Google, contain billions of labeled entities. Star queries, which aim to identify an entity, given a set of related entities, are common on such networks. Answering star queries can be modeled as a graph pattern matching problem. Traditional approaches apply graph indices to accelerate the query processing. Unfortunately, it is so costly that it is nearly infeasible to build indices on billion node graphs since the time or storage complexity of most indexing techniques is super-linear to the graph size. In this paper, we propose an algorithm to identify the top-k best answers for a star query. Instead of using expensive indices, our algorithm utilizes novel bounding techniques to derive the top-k best answers efficiently. Further, the algorithm can be implemented in a distributed manner scaling to billions of entities and hundreds of machines. We demonstrate the effectiveness and the efficiency of our approach through a series of experiments on real-world information networks. Jiahui Jin 0001, Samamon Khemmarat, Lixin Gao 0001, Junzhou Luo |
ICPADS | 1 |
| 2012 | Improving Online Aggregation Performance for Skewed Data Distribution
Yuxiang Wang 0001, Junzhou Luo, Aibo Song, Jiahui Jin 0001, Fang Dong 0001 |
DASFAA (1) | 4 |
| 2012 | Performance evaluation and analysis of SEU Cloud Computing Platform - Using general benchmarks and real world AMS applicationabstractCloud computing, as a popular technique to support and achieve CSCW, is gaining increasing importance in recent years, where the virtualization becomes the key technique. However, although utilizing virtualization can implement more efficient and flexible resource allocation, it may also come at the cost of increased system complexity and dynamics. In order to effectively adapt to performance fluctuations for ensuring high-performance, a generic approach to predict the performance influences of cloud platforms is highly desirable. To address this request, in this paper, the major factors that affect the performance of cloud and the relevant variation discipline are evaluated and analyzed thoroughly using a series of benchmarks in SEU (Southeast University) Cloud Computing Platform, where not only a general methodology on quantifying the performance influence but also the most important impact factors are proposed. Moreover, we use a real world application, as AMS experiment, to further evaluate the relevant performance. Fang Dong 0001, Junzhou Luo, Jiahui Jin 0001 |
SMC | 3 |
| 2011 | BAR: An Efficient Data Locality Driven Task Scheduling Algorithm for Cloud ComputingabstractLarge scale data processing is increasingly common in cloud computing systems like MapReduce, Hadoop, and Dryad in recent years. In these systems, files are split into many small blocks and all blocks are replicated over several servers. To process files efficiently, each job is divided into many tasks and each task is allocated to a server to deals with a file block. Because network bandwidth is a scarce resource in these systems, enhancing task data locality(placing tasks on servers that contain their input blocks) is crucial for the job completion time. Although there have been many approaches on improving data locality, most of them either are greedy and ignore global optimization, or suffer from high computation complexity. To address these problems, we propose a heuristic task scheduling algorithm called Balance-Reduce(BAR), in which an initial task allocation will be produced at first, then the job completion time can be reduced gradually by tuning the initial task allocation. By taking a global view, BAR can adjust data locality dynamically according to network state and cluster workload. The simulation results show that BAR is able to deal with large problem instances in a few seconds and outperforms previous related algorithms in term of the job completion time. Jiahui Jin 0001, Junzhou Luo, Aibo Song, Fang Dong 0001, Runqun Xiong |
CCGRID | 1 |
| 2010 | Resource Load Based Stochastic DAGs Scheduling Mechanism for Grid EnvironmentabstractThe dynamic feature is one of the most important differences between Grid and traditional heterogeneous distributed systems, thus the most significant challenge for task scheduling in Grid environment is how to relieve the resource performance dynamism effectively. However, the existing schedule algorithms usually suppose that computation or communication times are deterministic and static, thus they will lead to bad performance in the practical Grid environment. To address this problem, a mechanism which is used to estimate the probability distribution of task execution time based on resource load is proposed. And then a Resource Load based Stochastic DAGs Scheduling algorithm for Grid environments is introduced. The simulation results show that our mechanism can achieve a significant improvement in several metrics (such as normalized real schedule length) and can relieve the influence brought by the dynamic nature of Grid effectively. Fang Dong 0001, Junzhou Luo, Aibo Song, Jiahui Jin 0001 |
HPCC | 4 |