VLDB 2026 Research / reviewers in the wild / expert
Jian Wan 0001
dblp:87/2888-1
· DBLP profile ↗
83ranked-venue papers
4as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 12 since 2021Artificial intelligence and machine learning · 14 · 11 since 2021Systems, architecture and hardware · 14 · 5 since 2021Computer networks · 9 · 4 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Security and privacy · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One-Shot Federated Learning With Lightweight Intermediate Models in Wireless Sensor Network Setting
Yexin Dou, Haijiang Wang 0003, Jian Wan 0001, Lei Zhang 0196, Jie Huang 0014 |
IEEE Internet Things J. | 3 |
| 2026 | Collaborative knowledge and personalized preference alignment for sequential recommendation
Weiqi Yue, Tingting Liang, Xixi Sun, Xin Zhang 0079, Leilei Zheng, Yuyu Yin, Jian Wan 0001 |
Knowl. Based Syst. | 7 |
| 2026 | FLUXLog: A Federated Mixture-of-Experts Framework for Unified Log Anomaly DetectionabstractTraditional log anomaly detection systems are centralized, which poses the risk of privacy leakage during data transmission. Previous research mainly focuses on single-domain logs, requiring domain-specific models and retraining, which limits flexibility and scalability. In this paper, we propose a unified federated cross-domain log anomaly detection approach, FLUXLog, which is based on MoE (Mixture of Experts) to handle heterogeneous log data. Based on our insights, we establish a two phase training process: pre-training the gating network to assign expert weights based on data distribution, followed by expert driven top-down feature fusion. The following training of the gating network is based on fine-tuning the adapters, providing the necessary flexibility for the model to adapt across domains while maintaining expert specialization. This training paradigm enables a Hybrid Specialization Strategy, fostering both domain-specific expertise and cross-domain generalization. The Cross Gated Experts Module (CGEM) then fuses expert weights and dual channel outputs. Experiments on public datasets demonstrate that our model outperforms baseline models in handling unified cross-domain log data. Yixiao Xia, Yinghui Zhao, Jian Wan 0001, Congfeng Jiang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2025 | CoT4Rec: Revealing User Preferences Through Chain of Thought for Recommender SystemsabstractLarge Language Models (LLMs) offer groundbreaking advancements in recommender systems through superior text analysis and decision-making support. However, integrating LLMs into recommender systems still suffers from the problems of identifier uninterpretability and lack of transparency. To address these issues and fully leverage the capabilities of LLMs, we propose a chain of thought (CoT) based recommendation framework called CoT4Rec which employs LLMs as data enhancers for user preference analysis. Initially, we design a CoT reasoning strategy that can derive more behaviorally-aligned user preference features by clustering users’ historical interactions. Subsequently, we propose a two-stage recommendation model that not only makes full use of the world knowledge embedded in LLMs but also generates a logically transparent reasoning path. By integrating a user preference analyzer early in the recommendation pipeline, the model deeply analyzes users' historical interactions, helping to enhance the personalization and transparency of the recommender system. CoT4Rec demonstrates superior performance over existing state-of-the-art models in recommendation tasks across four public datasets, achieving improvements ranging from 2.2% to 12.2%. Weiqi Yue, Yuyu Yin, Xin Zhang 0079, Binbin Shi, Tingting Liang, Jian Wan 0001 |
AAAI | 6 |
| 2025 | A Novel Negative Sample Generation Method for Contrastive Learning in Hierarchical Text ClassificationabstractHierarchical text classification (HTC) is an important task in natural language processing (NLP). Existing methods typically utilize both text features and the hierarchical structure of labels to categorize text effectively. However, these approaches often struggle with fine-grained labels, which are closely similar, leading to difficulties in accurate classification. At the same time, contrastive learning has significant advantages in strengthening fine-grained label features and discrimination. However, the performance of contrastive learning strongly depends on the construction of negative samples. In this paper, we design a hierarchical sequence ranking (HiSR) method for generating diverse negative samples. These samples maximize the effectiveness of contrastive learning to enhance the ability of the model to distinguish between fine-grained labels and improve the performance of the model in HTC. Specifically, we transform the entire label set into linear sequences based on the hierarchical structure and rank these sequences according to their quality. During model training, the most suitable negative samples were dynamically selected from the ranked sequences. Then contrastive learning amplifies the differences between similar fine-grained labels by emphasizing the distinction between the ground truth and the generated negative samples, thereby enhancing the discriminative ability of the model. Our method has been tested on three public datasets and achieves state-of-art (SOTA) on two of them, demonstrating its effectiveness. Juncheng Zhou, Lijuan Zhang 0004, Yachen He, Rongli Fan, Lei Zhang 0196, Jian Wan 0001 |
COLING | 6 |
| 2025 | Federated Learning with Dual-View Feature Fusion and Composite Aggregation
Wenjian Xu, Yuyang Ji, Xiaohong Qian, Jie Huang 0014, Lei Zhang 0196, Jian Wan 0001 |
ICA3PP (8) | 8 |
| 2025 | MSPFT: Multivariate Time Series Prediction Transformer with Multi-Scale Patch Fusion Mechanism
Wenhao Fang, Junfeng Yuan, Jian Wan 0001, Yuyu Yin |
ICIC (19) | 4 |
| 2025 | HRSTORY: Historical News Review Based Online Story Discovery
Haoran Ye, Jian Wan 0001, Yong Liao 0003 |
KDD (1) | 3 |
| 2025 | FedAMM: Federated Learning for Brain Tumor Segmentation with Arbitrary Missing Modalities
Yukun Shi, Meiting Xue, Jian Wan 0001 |
MICCAI (8) | 5 |
| 2025 | Device placement using Laplacian PCA and graph attention networksabstractAbstract The exponential growth in data and parameters in modern neural networks has created the need to distribute these models across multiple devices for efficient training, resulting in the device placement problem. Existing graph encoding approaches for device placement suffer from low efficiency when searching for optimal parallel strategies, primarily due to suboptimal positional information retrieval. To address these challenges, we propose the Laplacian Principal Component Analysis-graph attention networks (LPCA-GAT) model. Firstly, we employ GAT to capture complex relationships between nodes and generate node encodings. Secondly, we leverage LPCA on the graph Laplacian matrix to extract crucial low-dimensional positional information. Finally, by integrating these two components, we obtain the final node encodings. This enhances the representation capability of nodes within the graph, enabling efficient device placement. The experimental results demonstrate that LPCA-GAT achieves superior device placement results, specifically accelerating execution and computation time by 13.24% and 96.38%, respectively, leading to significant improvements in both operational efficiency and performance. Lupeng Yue, Jian Wan 0001 |
Comput. J. | 6 |
| 2025 | An improved two-stage zero-shot relation triplet extraction model with hybrid cross-entropy loss and discriminative reranking
Diyou Li, Lijuan Zhang 0004, Juncheng Zhou, Jie Huang 0014, Naixue Xiong, Lei Zhang 0196, Jian Wan 0001 |
Expert Syst. Appl. | 7 |
| 2025 | Two-layer tensor decomposition for temporal kge graph completion
Lupeng Yue, Kaisheng Zeng, Mingyao Zhou, Jian Wan 0001 |
Expert Syst. Appl. | 7 |
| 2025 | Sequences and nodes: Probability-guided contrastive learning for hierarchical text classification
Lijuan Zhang 0004, Juncheng Zhou, Rongli Fan, Naixue Xiong, Lei Zhang 0196, Jian Wan 0001 |
Knowl. Based Syst. | 6 |
| 2025 | scGCRC: Graph and Contrastive-Based Representation Learning for Single-Cell RNA-Seq Data ClusteringabstractThe advent of single-cell RNA-sequencing (scRNA-seq) technology promotes biological analysis at the cellular level. Clustering cells to identify the type of cell is an important step in scRNA-seq analysis. Most of the existing clustering methods based on deep learning technology first adopt an autoencoder-decoder module to learn the low-dimensional features of cells and then apply other modules to learn the clustering relationship features of cells. However, the two-stage learning process makes the model training more difficult. Here we propose a novel cell representation learning method that is based on a local self-attention network and contrastive learning for scRNA-seq clustering. In particular, a local self-attention network automatically aggregates potential information of cells based on a cell relationship graph, and a dual contrastive learning module simultaneously optimizes the cell representation in cell- and cluster-level. The cell-level module makes related cells similar at the feature level, whereas the cluster-level module enables cells to form clusters at the cluster level. Finally, the powerful Leiden community discovery algorithm is used for clustering based on learned representation. In brief, we construct cell pairs through cell relationships and utilize contrastive learning to directly learn cell representations in a low-dimensional space while preserving their local structural relationships without pretraining an autoencoder-decoder module. Three benchmark experiments on 160 subsample datasets with different numbers of cell types, 3 datasets of different protocols, and 9 real public datasets demonstrate the superior performance of the proposed method compared with baseline methods. Jian Wan 0001, Xin Zhang 0079, Yuyu Yin |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Multitask-Based Self-Supervised Learning for Recommendation in Social SystemsabstractIn computational social systems, recommendation functionality plays a pivotal role in influencing user behavior, enhancing user experience, and driving engagement. To help recommendation functionality to better suggest relevant content or items, the social platforms usually utilize large-scale knowledge discovery techniques to analyze trends in user interactions and extract patterns from large datasets. Click-through rate (CTR) prediction is crucial in recommendation systems for measuring effectiveness, understanding user behavior, training and optimizing models, impacting business outcomes, enhancing personalization, and identifying issues. It provides actionable insights that assist in continuously refining and improving the recommendation process. Traditional deep learning-based CTR prediction models cannot work well for recommendation in social systems due to the data sparsity and the long-tail data problems since the representation learned from the user behavior is basically dominated by the major part of the data. In this article, we propose a multitask-based self-supervised learning model (MTSSL) that can better deal with sparse and long-tail user interaction data. Specifically, we first transform the CTR prediction task into the multitask joint learning framework with a set of shared subnetworks. Each subnetwork learns a representation of the entire user data, and hence, the sparse and long-tail data would have opportunity to fall into the best matched representation space of historical user behavior. Moreover, two kinds of self-supervision signals are employed to guide the learning of the representations. Extensive experiments over four user interaction datasets demonstrate the superiority of our proposed MTSSL over state-of-art models for recommendations. In terms of online A/B test, our model achieves around 3% better performance than the counterparts. Wenjian Xu, Fanxiang Zeng, Nan Zhang 0036, Honghao Gao, Yuyu Yin, Zulong Chen, Maolei Huang, Jian Wan 0001 |
IEEE Trans. Comput. Soc. Syst. | 8 |
| 2025 | DT-CTFP: 6G-Enabled Digital Twin Collaborative Traffic Flow PredictionabstractIn the era of big data, intelligent transportation systems are crucial for the development of smart cities, significantly impacting urban economic growth and planning. The integration of 6G networks and digital twin technology presents unprecedented opportunities to enhance urban traffic management through real-time data synchronization and high-fidelity simulations. Accurate traffic flow prediction is vital for congestion control, intelligent route planning, and effective urban traffic management. However, existing deep learning models often struggle to capture the complex spatio-temporal dependencies and dynamic spatial relationships inherent in urban traffic data, particularly in data-scarce environments. Given the spatial heterogeneity of urban data, where dense and sparse regions coexist, improving prediction accuracy in sparse areas is critical to ensuring overall forecasting performance. To address these challenges, we propose a novel framework called 6G-Enabled Digital Twin Collaborative Traffic Flow Prediction (DT-CTFP), which integrates advanced deep learning models within a 6G-supported digital twin environment. The framework leverages real-time data processing capabilities and ultra-low latency of 6G networks to capture complex traffic features and dynamic spatial dependencies. In data-rich regions, the Dynamic Graph Multi-Attention (DGMA) model is used to learn fine-grained spatio-temporal patterns, while for data-scarce regions, the Cross-Area Transfer Prediction (CATP) model utilizes meta-learning techniques to transfer knowledge from data-rich urban areas, improving prediction accuracy in areas with limited data. Experimental results demonstrate the superiority of the DT-CTFP framework, achieving up to 6% reductions in RMSE and 4% reductions in MAE across multiple datasets, highlighting its enhanced prediction accuracy and efficiency. These results emphasize the framework’s capacity to improve traffic management and vehicle-road cooperation within a digital twin smart city. Baofu Wu, Junfeng Yuan, Peng Zhan, Yuyu Yin, Jian Wan 0001, Honghao Gao |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | scCRT: a contrastive-based dimensionality reduction model for scRNA-seq trajectory inferenceabstractTrajectory inference is a crucial task in single-cell RNA-sequencing downstream analysis, which can reveal the dynamic processes of biological development, including cell differentiation. Dimensionality reduction is an important step in the trajectory inference process. However, most existing trajectory methods rely on cell features derived from traditional dimensionality reduction methods, such as principal component analysis and uniform manifold approximation and projection. These methods are not specifically designed for trajectory inference and fail to fully leverage prior information from upstream analysis, limiting their performance. Here, we introduce scCRT, a novel dimensionality reduction model for trajectory inference. In order to utilize prior information to learn accurate cells representation, scCRT integrates two feature learning components: a cell-level pairwise module and a cluster-level contrastive module. The cell-level module focuses on learning accurate cell representations in a reduced-dimensionality space while maintaining the cell-cell positional relationships in the original space. The cluster-level contrastive module uses prior cell state information to aggregate similar cells, preventing excessive dispersion in the low-dimensional space. Experimental findings from 54 real and 81 synthetic datasets, totaling 135 datasets, highlighted the superior performance of scCRT compared with commonly used trajectory inference methods. Additionally, an ablation study revealed that both cell-level and cluster-level modules enhance the model's ability to learn accurate cell features, facilitating cell lineage inference. The source code of scCRT is available at https://github.com/yuchen21-web/scCRT-for-scRNA-seq. Jian Wan 0001, Xin Zhang 0079, Tingting Liang, Yuyu Yin |
Briefings Bioinform. | 2 |
| 2024 | A novel joint extraction model based on cross-attention mechanism and global pointer using context shield windowabstractRelational triple extraction is a critical step in knowledge graph construction. Compared to pipeline-based extraction, joint extraction is gaining more attention because it can better utilize entity and relation information without causing error propagation issues. Yet, the challenge with joint extraction lies in handling overlapping triples. Existing approaches adopt sequential steps or multiple modules, which often accumulate errors and interfere with redundant data. In this study, we propose an innovative joint extraction model with cross-attention mechanism and global pointers with context shield window. Specifically, our methodology begins by inputting text data into a pre-trained RoBERTa model to generate word vector representations. Subsequently, these embeddings are passed through a modified cross-attention layer along with entity type embeddings to address missing entity type information. Next, we employ the global pointer to transform the extraction problem into a quintuple extraction problem, which skillfully solves the issue of overlapping triples. It is worth mentioning that we design a context shield window on the global pointer, which facilitates the identification of correct entities within a limited range during the entity extraction process. Finally, the capability of our model against malicious samples is improved by adding adversarial training during the training process. Demonstrating superiority over mainstream models, our approach achieves impressive results on three publicly available datasets. Zhengwei Zhai, Rongli Fan, Jie Huang 0014, Naixue Xiong, Lijuan Zhang 0004, Jian Wan 0001, Lei Zhang 0196 |
Comput. Speech Lang. | 6 |
| 2024 | VPI: Vehicle Programming Interface for Vehicle Computing
Baofu Wu, Ren Zhong, Jian Wan 0001, Ji-Lin Zhang, Weisong Shi |
J. Comput. Sci. Technol. | 4 |
| 2024 | Complex expressional characterizations learning based on block decomposition for temporal knowledge graph completion
Lupeng Yue, Kaisheng Zeng, Jian Wan 0001, Mingyao Zhou |
Knowl. Based Syst. | 6 |
| 2023 | MD-TransUNet: TransUNet with Multi-attention and Dilated Convolution for Brain Stroke Lesion Segmentation
Jian Wan 0001, Xin Zhang 0079 |
CollaborateCom (2) | 2 |
| 2023 | Block Decomposition with Multi-granularity Embedding for Temporal Knowledge Graph Completion
Lupeng Yue, Kaisheng Zeng, Jian Wan 0001 |
DASFAA (2) | 6 |
| 2023 | Evolutionary Interest Representation Network for Click-Through Rate PredictionabstractThe rapid development of the Internet has revo-lutionized our lives, providing us with an array of convenient services. However, this revolution has also led to the problem of data overload. Personalized recommendation services have been widely recognized as an effective solution to this problem by both the industrial and academic communities. Accurate prediction of click-through rate (CTR) is a critical task in personalized recommendation systems since it can improve the user’s shopping experience, ultimately increasing revenue for the platform. To achieve accurate CTR prediction, it is crucial to capture user interests. While several methods for interest modeling exist, challenges related to user interest evolution and user behavior noise need further attention. This paper proposes a novel CTR prediction model named the Evolutionary Interest Representation Network (EIRN), which learns the evolving interest representation of users from their behaviors, profiles, and environmental attributes. In the model, we introduce an advertisement correlation graph decomposition layer and a noise filter layer to enhance and purify the user behavior representation. An interest extraction layer is designed to capture sequential correlations and common characteristics between user behaviors, representing their interests. Additionally, as user interests evolve over time and their importance changes, an attention-based interest evolution layer is designed to track the evolution of user interests. Our proposed method is validated using four public datasets and an industrial dataset, showing its superiority over existing CTR prediction methods. Zhenming Jin, Jian Wan 0001, Xingyu Guo |
ICWS | 3 |
| 2023 | A Scalable Hybrid Total FETI Method for Massively Parallel FEM SimulationsabstractThe Hybrid Total Finite Element Tearing and Interconnecting (HTFETI) method plays an important role in solving large-scale and complex engineering problems. This method needs to handle numerous matrix-vector multiplications. Directly calling the vendor-optimized library for general matrix-vector multiplication (gemv) on GPU leads to low performance, since it does not consider optimizations for different matrix sizes in HTFETI, i.e. different row and column sizes. In addition, state-of-the-art graph partitioning methods cannot guarantee load balancing for HTFETI, since the matrix size is determined by the length of the subdomain boundary. To solve the problems above, we first port gemv to the multi-stream pipeline scheme and develop a new batched kernel function on GPU, which brings 15%~30% throughput improvement and 37% average GFLOPs improvement, respectively. We also propose a multi-grained load-balancing scheme based on graph repartitioning and work-stealing, and the load imbalance ratio is down to 1.05~1.09 from 1.5. We have successfully applied the scalable HTFETI method to simulate the whole core assembly of China Experimental Fast Reactor (CEFR) for steady-state analysis, and the efficiencies of weak scalability and strong scalability reach 78% and 72% on 12,288 GPUs, respectively. As far as we know, this is the first time that HTFETI has been used in large-scale and high-fidelity whole core assembly simulation. Kehao Lin, Chunbao Zhou, Ningming Nie, Jue Wang 0013, Shigang Li 0002, Yangde Feng, Yangang Wang 0002, Kehan Yao, Tiechui Yao, Jian Wan 0001 |
PPoPP | 12 |
| 2023 | Adaptive Federated Learning With Non-IID DataabstractAbstract With the widespread use of Internet of things(IoT) devices, it generates an enormous volume of data, and it is a challenge to mine the IoT data value while ensuring security and privacy. Federated learning is a decentralized approach for training data located on edge devices, such as mobile phones and IoT devices, while keeping privacy, efficiency, and security. However, the Non-IID (non-independent and identically distributed) data, always greatly impacts the performance of the global model. In this paper, we propose a FedDynamic algorithm to solve the statistical challenge of federated learning caused by Non-IID. As Non-IID data can lead to significant differences in model parameters between edge devices, we set different weights for different devices during model aggregation to get a high-performance global model. We analyze and exact key indices (local model accuracy, local data quality, and model difference between local models and the global model), which can reflect the quality of the model, and calculate the aggregation weight for edge devices based on the key indices. Furthermore, we dynamically adjust aggregation weight based on accuracy’s variety to solve weight staleness during the training process. Experiments on the MNIST, FMNIST, EMNIST, CINIC-10 and CIFAR-10 datasets show that the FedDynamic algorithm has better accuracy and convergence performance, compared to the FedAvg, FedProx and Scaffold algorithms. Yuankai Mu, Junfeng Yuan, Siyuan Teng, Jian Wan 0001, Yunquan Zhang |
Comput. J. | 6 |
| 2023 | Time-Aware Smart City Services Based on QoS Prediction: A Contrastive Learning ApproachabstractSmart cities are designed to satisfy the needs of residents and improve their quality of life by providing a wide range of smart city services. One of the keys to the efficient operation of smart city services is the accurate forecast of the missing Quality of Service (QoS). Presently, many approaches utilize the context information of users and services, such as geographic location and network location, to somewhat increase the prediction accuracy and forecast the missing QoS values. However, because the network conditions and server status are unpredictable, time is also considered as one of the important factors affecting QoS prediction, which brings more challenges as follows: higher data dimension, more complex data characteristics, and higher data sparsity. To overcome these challenges, we propose an approach for time-aware Web service QoS prediction based on contrastive learning (named CLpred). CLpred utilizes a sequential data input format for QoS data and models these QoS sequences through transformer encoder with CLpred framework. Therefore, it can downscale QoS data and extract a more efficient representation in complex QoS data. Furthermore, it makes it possible to apply data augmentation methods to address the problems of data sparsity. In order to prove the superiority of the proposed approach, particularly inside the presence of extremely high-data sparsity, extensive experiments are conducted on the well-known service QoS data set WSDREAM. Yuyu Yin, Qianhui Di, Jian Wan 0001, Tingting Liang |
IEEE Internet Things J. | 3 |
| 2023 | PGDENet: Progressive Guided Fusion and Depth Enhancement Network for RGB-D Indoor Scene ParsingabstractScene parsing is a fundamental task in computer vision. Various RGB-D (color and depth) scene parsing methods based on fully convolutional networks have achieved excellent performance. However, color and depth information are different in nature and existing methods cannot optimize the cooperation of high-level and low-level information when aggregating modal information, which introduces noise or loss of key information in the aggregated features and generates inaccurate segmentation maps. The features extracted from the depth branch are weak because of the low quality of the depth map, which results in unsatisfactory feature representation. To address these drawbacks, we propose a progressive guided fusion and depth enhancement network (PGDENet) for RGB-D indoor scene parsing. First, high-quality RGB images are used to improve depth data through a depth enhancement module, in which the depth maps are strengthened in terms of channel and spatial correlations. Then, we integrate information from the RGB and enhance depth modalities using a progressive complementary fusion module, in which we start with high-level semantic information and move down layerwise to guide the fusion of adjacent layers while reducing hierarchy-based differences. Extensive experiments are conducted on two public indoor scene datasets, and the results show that the proposed PGDENet outperforms state-of-the-art methods in RGB-D scene parsing. Wujie Zhou, Enquan Yang, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003 |
IEEE Trans. Multim. | 4 |
| 2022 | Attention and Edge-Label Guided Graph Convolutional Networks for Named Entity RecognitionabstractIt has been shown that named entity recognition (NER) could benefit from incorporating the long-distance structured information captured by dependency trees.However, dependency trees built by tools usually have a certain percentage of errors.Under such circumstances, how to better use relevant structured information while ignoring irrelevant or wrong structured information from the dependency trees to improve NER performance is still a challenging research problem.In this paper, we propose the Attention and Edge-Label guided Graph Convolution Network (AELGCN) model.Then, we integrate it into BiLSTM-CRF to form BiLSTM-AELGCN-CRF model.We design an edge-aware node joint update module and introduce a node-aware edge update module to explore hidden structured information entirely and solve the wrong dependency label information to some extent.After two modules, we apply attention-guided GCN, which automatically learns how to attend to the relevant structured information selectively.We conduct extensive experiments on several standard datasets across four languages and achieve better results than previous approaches.Through experimental analysis, it is found that our proposed model can better exploit the structured information on the dependency tree to improve the recognition of long entities. Zhongyi Xie, Jian Wan 0001, Yong Liao 0003 |
EMNLP | 3 |
| 2022 | To Turn or Not To Turn, SafeCross is the AnswerabstractBlind area has plagued drivers’ safety ever since the dawn of automobiles. Thanks to the fast-growing vision-based perception technologies, autonomous driving systems can monitor the driving circumstance through a 360-degree view, and hence most blind areas can be avoided. However, in the left turn scenario at an intersection, the opposite road may be blocked by another vehicle parking at the same intersection (see Fig. 1), and in this case, the blind area cannot be observed by the onboard perception module of the autonomous vehicle. A potential fatal collision may occur if the autonomous vehicle turns left while a vehicle is running through the blind area. In this paper, we propose Safecross, a framework that oversees an intersection and delivers blind area warnings to the left-turn vehicles at the intersection if running vehicles are detected in the blind area. In order to provide accurate and reliable real-time warnings in all possible weather conditions, the architecture of Safecross has four major components: video pre-processing (VP) module, video classification (VC) module, few-shot learning (FL) module, and model switching (MS) module. Especially, the VP and VC modules will train a basic model to identify the blind area when a blocking vehicle appears at the intersection. Since the range of the blind area varies in different weather conditions, the FL and MS modules can adapt the basic model to the new condition in real-time to make the blind area identification more accurate. Intuitively, if the blind area is identified timely and accurately, the left-turn throughput of the intersection can be maximized. We have conducted extensive experiments to evaluate our proposed framework. The experiments are performed on a total of 2855 video segments with a time span of 180 days, including sunny, rainy, and snowy weather conditions. Experimental results show how Safecross can guarantee the vehicle’s safety while increasing the left-turn traffic throughput by 50%. Baofu Wu, Yuankai He, Zheng Dong 0002, Jian Wan 0001, Weisong Shi |
ICDCS | 4 |
| 2022 | A blockchain-based fine-grained data sharing scheme for e-healthcare systemabstractThe cloud-aided sharing of e-healthcare data has great positive significance for research. However, due to the privacy consideration, these data are usually encrypted before uploading to the cloud server which impedes data sharing between different medical institutions. Conditional proxy re-encryption (CPRE) allows the proxy to converse ciphertext, especially by specifying a condition embed in the re-encryption key to achieve fine-grained access control over the ciphertext. Unfortunately, existing CPRE schemes cannot ensure the privacy of the condition, which may contain some sensitive private information. Furthermore, a malicious proxy server may return part of the results and even false data to save its computation or bandwidth. To solve these problems, we propose a blockchain-based condition invisible proxy re-encryption scheme for the e-healthcare system. The proposed scheme guarantees the confidentiality of the data by hiding the condition in the re-encryption key so that the proxy cannot learn anything about the condition. Moreover, the ciphertext-searching algorithm is leveraged by executing the smart contract in the blockchain which ensures the results are correct and complete. Finally, experiment results demonstrate the practicability of the proposed scheme in applications. Gaofan Lin, Haijiang Wang 0003, Jian Wan 0001, Lei Zhang 0196, Jie Huang 0014 |
J. Syst. Archit. | 3 |
| 2022 | Conditional Embedding Pre-Training Language Model for Image Captioning
Pengfei Li 0009, Min Zhang 0029, Peijie Lin, Jian Wan 0001, Ming Jiang 0009 |
Neural Process. Lett. | 4 |
| 2022 | A Secure Dynamic Mix Zone Pseudonym Changing Scheme Based on Traffic Context PredictionabstractTraffic context plays an important role in supporting automated driving and intelligent transportation systems. Smart vehicles explore surrounding environments by analyzing sensor data and periodically communicating with neighbors and road infrastructures. The context can be well learned in this way to support driving, but the vehicle trajectory can be also easily exposed under eavesdropping attacks. The pseudonym is proposed to hide the real identity of the vehicles. However, the effectiveness of anonymity, the safety of driving, the convenience of implementation and the utilization of resources in previous approaches have not been well-balanced. Therefore, focusing on efficiently replacing pseudonyms with the premise of ensuring driving safety, we propose a secure dynamic silent mix zone pseudonym changing scheme (TLAS) based on the real-time traffic context prediction for urban regions. It naturally takes the area in front of the red traffic light as a silent mix zone, which avoids the driving security issue caused by signal silence. Besides, the area length is dynamically configured according to the traffic context predicted in the last green light cycle, so the anonymous effect can be improved. In addition, considering the resource utilization and accuracy requirement, the adaptive prediction algorithm is applied. We conduct simulation experiments with real-world traffic history using SUMO and OMNET++, the results show that TLAS strategy can indeed achieve a better anonymous effect (reducing standardized traceability rate by 8.2%) with lower driving speed for safety concern. Youhuizi Li, Yuyu Yin, Xu Chen 0048, Jian Wan 0001, Gangyong Jia, Kewei Sha |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | CCAFNet: Crossflow and Cross-Scale Adaptive Fusion Network for Detecting Salient Objects in RGB-D ImagesabstractOwing to the widespread adoption of depth sensors, salient object detection (SOD) supported by depth maps for reliable complementary information is being increasingly investigated. Existing SOD models mainly exploit the relation between an RGB image and its corresponding depth information across three fusion domains: input RGB-D images, extracted feature maps, and output salient object. However, these models do not leverage the crossflows between high- and low-level information well. Moreover, the decoder in these models uses conventional convolution that involves several calculations. To further improve RGB-D SOD, we propose a crossflow and cross-scale adaptive fusion network (CCAFNet) to detect salient objects in RGB-D images. First, a channel fusion module allows for effective fusing depth and high-level RGB features. This module extracts accurate semantic information features from high-level RGB features. Meanwhile, a spatial fusion module combines low-level RGB and depth features with accurate boundaries and subsequently extracts detailed spatial information from low-level depth features. Finally, a purification loss is proposed to precisely learn the boundaries of salient objects and obtain additional details of the objects. The results of comprehensive experiments on seven common RGB-D SOD datasets indicate that the performance of the proposed CCAFNet is comparable to those of state-of-the-art RGB-D SOD models. Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003 |
IEEE Trans. Multim. | 4 |
| 2022 | Local Epochs Inefficiency Caused by Device Heterogeneity in Federated LearningabstractFederated learning is a new framework of machine learning, it trains models locally on multiple clients and then uploads local models to the server for model aggregation iteratively until the model converges. In most cases, the local epochs of all clients are set to the same value in federated learning. In practice, the clients are usually heterogeneous, which leads to the inconsistent training speed of clients. The faster clients will remain idle for a long time to wait for the slower clients, which prolongs the model training time. As the time cost of clients’ local training can reflect the clients’ training speed, and it can be used to guide the dynamic setting of local epochs, we propose a method based on deep learning to predict the training time of models on heterogeneous clients. First, a neural network is designed to extract the influence of different model features on training time. Second, we propose a dimensionality reduction rule to extract the key features which have a great impact on training time based on the influence of model features. Finally, we use the key features extracted by the dimensionality reduction rule to train the time prediction model. Our experiments show that, compared with the current prediction method, our method reduces 30% of model features and 25% of training data for the convolutional layer, 20% of model features and 20% of training data for the dense layer, while maintaining the same level of prediction error. Xin Wang 0157, Junfeng Yuan, Jian Wan 0001 |
Wirel. Commun. Mob. Comput. | 5 |
| 2021 | Federated Learning Model Training Method Based on Data Features Perception AggregationabstractThe rapidly expanding number of Internet of Things (IoT) devices is generating huge quantities of data, but public concern over data privacy means users are apprehensive to send data to a central server for machine learning purposes. Federated learning is an emerging concept, which allows edge devices to collaboratively learn and share models, while keeping training data on devices. Federated learning decouples “model training” and “direct access to original training data”. But, in the IoT where the wireless network resource is constrained, the key problem of federated learning is the communication overhead for parameter synchronization, which wastes bandwidth, increases training time, and even impacts the model accuracy. Moreover, the IoT devices collect data from different users, so the distribution over devices can be highly non-independent identically distributed (non-IID), which results in the variation of feature distribution and label distribution. As a result, the test accuracy of the federated model is reduced, and the communication cost of training the federated model is increased. In this paper, we propose the FedCC algorithm, to improve the accuracy of the federated model in the non-IID scenario. The FedCC method constructs client groups by mining data similarity, and selects one model of every client group to upload to the cloud server for model aggregation. Our experiments show that FedCC not only outperforms popular state-of-the-art federated learning algorithms on CNN and MLP architectures trained on MNIST and CIFAR-10 datasets, but also reduces the overall communication cost. Zeng Yan, Yan Zhong Yi, Nailiang Zhao, Jian Wan 0001, Jun Yu 0002 |
VTC Fall | 6 |
| 2021 | EDM-Fuzzy: An Euclidean Distance Based Multiscale Fuzzy Entropy Technology for Diagnosing Faults of Industrial SystemsabstractSample entropy (SampEn) technologies have been widely applied in diagnosing the faults of industrial systems. However, there are two disadvantages of these technologies. First, all of these technologies measure the distance of two vectors solely based on the maximum distance between the corresponding elements in the two vectors, which is not able to fully reflect the distance of the two vectors. Second, these methodologies measure the similarity of two vectors with either zero or one, which may cause sudden changes in entropy values. Therefore, we proposed a Euclidean distance based multiscale fuzzy entropy (EDM-Fuzzy), which measures the similarity of two vectors with continuous values from zero to one based on the Euclidean distance of the two vectors. The results from the synthetic and real signals demonstrated that EDM-Fuzzy has higher accuracy in measuring the complexity of signals. As a result, EDM-Fuzzy obtains a higher accuracy in detecting the bearing faults than the state-of-the-art entropy technologies. Xiao Wang 0047, Jian Wan 0001, Naixue Xiong |
IEEE Trans. Ind. Informatics | 3 |
| 2021 | SDABS: A Flexible and Efficient Multi-Authority Hybrid Attribute-Based Signature Scheme in Edge EnvironmentabstractThe explosive growth of the Internet of Things and modern networking technologies lay the foundation for the development of intelligent transportation systems and smart cities. To analyzing massive data under the required time for transportation issues, the edge computing paradigm is applied, which pre-processing large amounts of data at the network edge to save bandwidth and improve response time. However, data reliability and security are still facing many challenges in the edge environment. In this article, we propose a multi-authority hybrid attribute-based signature scheme (SDABS). It is composed of four phases: system initialization, signature generation, signature verification, and attribute revocation phases. To better describe frequently changing features in the transportation systems like location, the dynamic attribute is introduced in building the signature. The multi-layer policy tree is applied to support flexible and various access policies, which also naturally form user groups and help data searching. Besides, the multi-authority structure is more suitable for the distributed edge environment. We evaluate SDABS from both theoretical analysis and practical analysis. Compared with two classical signature schemes (MABS and ODMA-ABS), experimental results demonstrate that the proposed SDABS can achieve better performance at an acceptable cost in the terms of attributes and attribute authorities. Youhuizi Li, Xu Chen 0048, Yuyu Yin, Jian Wan 0001, Li Kuang, Zeyong Dong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Research on Named Entity Recognition of Electronic Medical Records Based on RoBERTa and Radical-Level FeatureabstractClinical named entity recognition (CNER) identifies entities from unstructured medical records and classifies them into predefined categories. It is of great significance for follow‐up clinical studies. Most of the existing CNER methods fail to give enough thought to Chinese radical‐level characteristics and the specialty of the Chinese field. This paper proposes the Ra‐RC model, which combines radical features and a deep learning structure to fix this problem. A bidirectional encoder representation of transformer (RoBERTa) is utilized to learn medical features thoroughly. Simultaneously, we use the bidirectional long short‐term memory (BiLSTM) network to extract radical‐level information to capture the internal relevance of characteristics and stitch the eigenvectors generated by RoBERTa. In addition, the relationship between labels is considered to obtain the optimal tag sequence by applying conditional random field (CRF). The experimental results demonstrate that the proposed Ra‐RC model achieves F1 score 93.26% and 82.87% on the CCKS2017 and CCKS2019 datasets, respectively. Jie Huang 0014, Caie Xu, Huilin Zheng, Lei Zhang 0196, Jian Wan 0001 |
Wirel. Commun. Mob. Comput. | 6 |
| 2020 | Energy aware edge computing: A survey
Congfeng Jiang, Tiantian Fan, Honghao Gao, Weisong Shi, Liangkai Liu, Christophe Cérin, Jian Wan 0001 |
Comput. Commun. | 7 |
| 2020 | QoS Prediction for Service Recommendation with Deep Feature Learning in Edge Computing Environment
Yuyu Yin, Yueshen Xu, Jian Wan 0001, Zhida Mai |
Mob. Networks Appl. | 4 |
| 2020 | An Industrial Analysis Technology About Occupational Adaptability and Association Rules in Social NetworksabstractWith the enormous growth of the content and users of social network, user portrait based on social network have been widely used in industrial applications such as artificial intelligence, recommendation systems, etc. The issue has been extensively studied, and most of them have focused on user attribute mining and behavior prediction. However, there is little research on the intrinsic relationship between big data-based interests and skills and occupational adaptability. To clarify this issue, in this article, we first collect a large amount of user data from LinkedIn. Then, we filter high-frequency interests and skills, and explore relevant characteristics of interests and skills through relevant analytical methods, found that there are many association rules. Next, the users are grouped according to the characteristics of occupational adaptability, and the relationship between the association rules of interests and skills and user adaptability is studied. Finally, designed an occupational adaptive classifier based on association rules. This article reveals the connections and dependencies between human interests and professional skills, as well as the impact of these association rules on user career adaptability, also discusses the prospects of these results in industrial applications. We hope that our research results will provide a solid theoretical foundation for other areas of research and industrial applications, such as career adaptability judgment, interest training, career development, career recommendation, etc. Huayou Si, Haopeng Wu, Li Zhou 0008, Jian Wan 0001, Naixue Xiong |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Distributed machine learning load balancing strategy in cloud computing services
Jian Wan 0001, Li Zhou 0008, Baofu Wu, Jue Wang 0013 |
Wirel. Networks | 3 |
| 2019 | A Collaborative Anomaly Detection Approach of Marine Vessel Trajectory (Short Paper)
Jian Wan 0001, Jie Huang 0014, Gangyong Jia, Wei Zhang 0138 |
CollaborateCom | 2 |
| 2019 | Priority-Based Optimization of I/O Isolation for Hybrid Deployed Services
Youhuizi Li, Li Zhou 0008, Zujie Ren, Jian Wan 0001 |
CollaborateCom | 5 |
| 2019 | An Edge Computing-Based Framework for Marine Fishery Vessels Monitoring Systems
Fengwei Zhu, Jie Huang 0014, Jian Wan 0001 |
CollaborateCom | 4 |
| 2019 | A member recognition approach for specific organizations based on relationships among users in social networking Twitter
Huayou Si, Wei Zhang 0138, Jian Wan 0001, Naixue Xiong |
Future Gener. Comput. Syst. | 4 |
| 2018 | Towards Building a Scalable Data Analytics System on Clouds: An Early Experience on AliCloudabstractWith the development of big data, big data processing systems, such as Hadoop and Spark, are widely used to handle large-scale data. To avoid the complexity and expensiveness of building a self-owned big data processing system, cloud providers tend to deploy big data processing tools as cloud services. Typical examples include Amazon EMR, Azure HDInsight and AliCloud E-MapReduce. However, how to build a cost-efficient system and scale the system is still challenging. In this paper, we have conducted a case study on AliCloud E-MapReduce, and analyzed the system performance upon local and remote file systems. We compared the scalability of Hadoop and Spark by using scaleout and scale-up strategies respectively. Based on the analysis results, we derive several observations and implications, which will contribute to guide the performance optimization. Congfeng Jiang, Zujie Ren, Youhuizi Li, Jian Wan 0001, Jiangbin Lin |
IEEE CLOUD | 5 |
| 2018 | PARDA: A Dataset for Scholarly PDF Document Metadata Extraction Evaluation
Tiantian Fan, Yeliang Qiu, Congfeng Jiang, Wei Zhang 0138, Jian Wan 0001 |
CollaborateCom | 7 |
| 2018 | How Good is Query Optimizer in Spark?
Zujie Ren, Na Yun, Youhuizi Li, Jian Wan 0001, Lihua Yu, Xinxin Fan |
CollaborateCom | 4 |
| 2018 | Two-Phase Web Service QoS Prediction with Restricted Boltzmann Machine
Yuyu Yin, Yueshen Xu, Liang Chen 0001, Jian Wan 0001 |
ICSOC | 5 |
| 2018 | EASE: Energy Efficiency and Proportionality Aware Virtual Machine SchedulingabstractServers have different energy efficiency and energy proportionality (EP) due to their hardware configuration (i.e., CPU generation and memory installation) and workload. However, current virtual machine (VM) scheduling in virtualized environments will saturate servers without considering their energy efficiency and EP differences. This article will discuss EASE, the energy efficiency and proportionality aware VM scheduling approach. EASE first executes customized computing intensive, memory intensive, and hybrid benchmarks to calculate a server's energy efficiency and EP. Then it schedules VMs to servers to keep them working at their peak energy efficiency point (or optimal working range). This step improves the overall energy efficiency of the cluster and the data center. For performance guarantee, EASE migrates VMs from servers under highly contending conditions. The experimental results on real clusters show that power consumption can be saved 37.07% ~ 49.98% in the homogeneous cluster. The average completion time of the computing intensive VMs increases only 0.31 % ~ 8.49%. In the heterogeneous nodes, the power consumption of the computing intensive VMs can be reduced by 44.22 %. The job completion time can be saved by 53.80%. Congfeng Jiang, Yumei Wang, Dongyang Ou, Yeliang Qiu, Youhuizi Li, Jian Wan 0001, Weisong Shi, Christophe Cérin |
SBAC-PAD | 6 |
| 2018 | Characterizing the Effectiveness of Query Optimizer in SparkabstractIn the big data community, Spark has been widely used for processing interactive queries. Spark employs a query optimizer, called Catalyst, to provides a set of optimization rules and supports Cost-Based Optimization (CBO). In this paper, we investigated the effectiveness of the optimization rules and cost-based optimization in Catalyst. We conducted comprehensive validation experiments by varying the data volume and cluster scale, and found that the execution time of most TPC-H queries were reduced slightly even when query optimizations are applied. We derived some interesting observations on Catalyst, which can help the community better understand and improve the query optimizer of Spark in future. Zujie Ren, Na Yun, Weisong Shi, Youhuizi Li, Jian Wan 0001, Lihua Yu, Xinxin Fan |
SERVICES | 5 |
| 2017 | Quantifying the Isolation Characteristics in Container Environments
Yusen Wu 0001, Zujie Ren, Weisong Shi, Jian Wan 0001 |
NPC | 6 |
| 2017 | Realistic and Scalable Benchmarking Cloud File Systems: Practices and Lessons from AliCloudabstractThe past decade has witnessed the rapid boom of cloud computing. Many public cloud infrastructures have been implemented and serve millions of tenants. Cloud file systems, which take charge of petabyte-scale data storage, play a crucial role in the performance of cloud infrastructures. Typical cloud file systems, including GFS, HDFS and Ceph, have attracted notable research efforts for performance evaluation and optimization. However, due to the heterogeneity and complexity of I/O workload characteristics in cloud environments, it is still challenging to conduct an accurate and efficient performance evaluation. To address this problem, we collected a two-week I/O workload trace from a 2,500-node production cluster in AliCloud, which is one of the largest cloud providers in Asia. Using the AliCloud trace, we characterized the I/O workload and data distribution, and compared two cloud services in multiple perspectives, including the request arrival pattern, request size, data population and so on. A list of observations and implications were derived and applied to help design a cloud file system benchmarking suite, called Porcupine. Porcupine aims to deploy a scalable and efficient performance evaluation on cloud file systems using realistic I/O workloads. We conducted a group of validation experiments, which demonstrated that Porcupine can achieve high accuracy and scalability. This paper provides our experiences and lessons in generating I/O workloads and deploying performance tests on cloud file systems, which we believe will be insightful to the cloud computing community in general. Zujie Ren, Weisong Shi, Jian Wan 0001, Jiangbin Lin |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Virtual Page Behavior Based Page Management Policy for Hybrid Main Memory in Cloud ComputingabstractA new generation memory, Non-Volatile Memory (NVM), such as Phase-Change Memory (PCM), has been adopted together with DRAM in the main memory to form the hybrid main memory for low energy consumption and high capacity. The biggest challenge of hybrid memory is how to decrease the average memory access cost for the higher cost of NVM's read/write operation. Currently, most researches are based on migration. However, the page migration itself is a high cost operation. And the migration based policy produces many migration operations, which induces high cost in memory access. Therefore, in order to decrease the cost, we present a virtual page behavior based page management policy (VBPM) in this paper. According to the virtual pages' behavior, we allocate virtual pages into DRAM or PCM physical pages correspondingly. The whole process is migration independent. The experimental results show our VBPM decreases the average memory access time by 24%, moreover, VBPM improves real-time performance in critical path. Jie Huang 0014, Guangjie Han, Gangyong Jia, Huizi Liyou, Jian Wan 0001 |
MSN | 6 |
| 2016 | Efficient parallel implementation of incompressible pipe flow algorithm based on SIMPLEabstractSummary Parallel semi‐implicit method for pressure‐linked equations(SIMPLE) algorithm is used to solve the 3‐D incompressible pipe flow problem. In this paper, we proposed a novel parallel SIMPLE algorithm that uses the alternate tiling technique. Firstly, a parallel SIMPLE algorithm based on domain decomposition method was established, and the implementation of domain partition and data exchange was presented. Then, we presented serial finite difference stencil algorithm based on alternate tiling. Furthermore, an iteration space parallel two‐way finite difference stencil algorithm based on alternate tiling was proposed, introducing the sequence of iterative space tiles as the sequence of execution and using time skewing technique to partition the iteration space, thus to improve the data locality of algorithm. The cache misses and the cost of communication and synchronization are reduced by reordering the tiles of iteration space. Finally, the effectiveness of the two parallel SIMPLE algorithms were compared. The results showed that the parallel SIMPLE algorithm that uses the two‐way finite difference stencil algorithm based on alternate tiling has good data locality, performance, and scalability in the Deepcomp7000 cluster computing environment. Copyright © 2013 John Wiley & Sons, Ltd. Junfeng Yuan, Jian Wan 0001, Jie Mao, Li-Ting Zhu, Li Zhou 0008, Congfeng Jiang, Peng Di, Jue Wang 0013 |
Concurr. Comput. Pract. Exp. | 3 |
| 2016 | Efficient sparse matrix-vector multiplication using cache oblivious extension quadtree storage format
Jian Wan 0001, Jie Mao, Li Zhuang, Junfeng Yuan, Enyi Liu, Zhuoer Yu |
Future Gener. Comput. Syst. | 2 |
| 2016 | Boosting video popularity through keyword suggestion and recommendation systems
Samamon Khemmarat, Lixin Gao 0001, Jian Wan 0001, Yuyu Yin, Jun Yu 0002 |
Neurocomputing | 4 |
| 2016 | How YouTube videos are discovered and its impact on video views
Samamon Khemmarat, Lixin Gao 0001, Jian Wan 0001 |
Multim. Tools Appl. | 4 |
| 2015 | Human pose recovery by supervised spectral embedding
Jun Yu 0002, Yukun Guo, Dapeng Tao, Jian Wan 0001 |
Neurocomputing | 4 |
| 2015 | l2, 1 Norm regularized fisher criterion for optimal feature selection
Jian Zhang 0026, Jun Yu 0002, Jian Wan 0001 |
Neurocomputing | 3 |
| 2015 | Multimodal Deep Autoencoder for Human Pose RecoveryabstractVideo-based human pose recovery is usually conducted by retrieving relevant poses using image features. In the retrieving process, the mapping between 2D images and 3D poses is assumed to be linear in most of the traditional methods. However, their relationships are inherently non-linear, which limits recovery performance of these methods. In this paper, we propose a novel pose recovery method using non-linear mapping with multi-layered deep neural network. It is based on feature extraction with multimodal fusion and back-propagation deep learning. In multimodal fusion, we construct hypergraph Laplacian with low-rank representation. In this way, we obtain a unified feature description by standard eigen-decomposition of the hypergraph Laplacian matrix. In back-propagation deep learning, we learn a non-linear mapping from 2D images to 3D poses with parameter fine-tuning. The experimental results on three data sets show that the recovery error has been reduced by 20%-25%, which demonstrates the effectiveness of the proposed method. Jun Yu 0002, Jian Wan 0001, Dacheng Tao, Meng Wang 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Temperature-Aware Scheduling Based on Dynamic Time-Slice Scaling
Gangyong Jia, Youwei Yuan, Jian Wan 0001, Congfeng Jiang, Xi Li 0003, Dong Dai 0001 |
ICA3PP (1) | 3 |
| 2014 | Combine thread with memory scheduling for maximizing performance in multi-core systemsabstractThe growing gap between microprocessor speed and DRAM speed is a major problem that computer designers are facing. In order to narrow the gap, it is necessary to improve DRAM's speed and throughput. Moreover, on multi-core platforms, DRAM memory shared by all cores usually suffers from the memory contention and interference problem, which can cause serious performance degradation and unfairness among parallel running threads. To address these problems, this paper proposes techniques to take both advantages of partitioning cores, threads and memory banks into groups to reduce interference among different groups and grouping the memory accesses of the same row together to reduce cache miss rate. A memory optimization framework combined thread scheduling with memory scheduling (CTMS) is proposed in this paper, which simultaneously minimizes memory access schedule length, memory access time and reduce interference to maximize performance for multi-core systems. Experimental results show CTMS is 12.6% shorter in memory access time, while improving 11.8% throughput on average. Moreover, CTMS also saves 5.8% of the energy consumption. Gangyong Jia, Guangjie Han, Liang Shi 0001, Jian Wan 0001, Dong Dai 0001 |
ICPADS | 4 |
| 2014 | PUMA: Pseudo unified memory architecture for single-ISA heterogeneous multi-core systemsabstractSingle-ISA heterogeneous multi-core processors have advantages over cost-equivalent homogeneous ones, which integrate cores having the same instruction set architecture (ISA) but offer different performance and power characteristics. When these cores share the off-chip main memory, requests from different cores will interfere with each other, leading to low system performance and unfairness even starvation. Unfortunately, state-of-the-art memory scheduling and thread scheduling algorithms are ineffective at solving these problems. This paper proposes a fundamentally new memory architecture of pseudo unified memory (PUMA), which partitions the memory into regions according cores' different performance, each core mostly requests only one memory region seldom exceeding, reducing interfere among cores while retaining bank level parallelism for improving performance and fairness. We evaluate the design trade-offs involved in our PUMA and compare it against three state-of-the-art memory management methods. Our experimental results show that PUMA improves both system performance and fairness among cores while reducing memory power. Gangyong Jia, Liang Shi 0001, Jian Wan 0001, Youwei Yuan, Xi Li 0003, Dong Dai 0001 |
RTCSA | 3 |
| 2014 | Workload Analysis, Implications, and Optimization on a Production Hadoop Cluster: A Case Study on TaobaoabstractUnderstanding the characteristics of MapReduce workloads in a Hadoop cluster is the key to making optimal configuration decisions and improving the system efficiency and throughput. However, workload analysis on a Hadoop cluster, particularly in a large-scale e-commerce production environment, has not been well studied yet. In this paper, we performed a comprehensive workload analysis using the trace collected from a 2000-node Hadoop cluster at Taobao, which is the biggest online e-commerce enterprise in Asia, ranked 10th in the world as reported by Alexa. The results of the workload analysis are representative and generally consistent with the data warehouses for e-commerce web sites, which can help researchers and engineers understand the workload characteristics of Hadoop in their production environments. Based on the observations and implications derived from the trace, we designed a workload generator Ankus, to expedite the performance evaluation and debugging of new mechanisms. Ankus supports synthesizing an e-commerce style MapReduce workload at a low cost. Furthermore, we proposed and implemented a job scheduling algorithm, Fair4S , which is designed to be biased towards small jobs. Small jobs account for the majority of the workload, and most of them require instant and interactive responses, which is an important phenomenon at production Hadoop systems. The inefficiency of Hadoop fair scheduler for handling small jobs motivates us to design the Fair4S, which introduces pool weights and extends job priorities to guarantee the rapid responses for small jobs. Experimental evaluation verified that the Fair4S accelerates the average waiting times of small jobs by a factor of 7 compared with the fair scheduler. Zujie Ren, Jian Wan 0001, Weisong Shi, Xianghua Xu |
IEEE Trans. Serv. Comput. | 2 |
| 2013 | Coordinate Task and Memory Management for Improving Power Efficiency
Gangyong Jia, Xi Li 0003, Jian Wan 0001, Chao Wang 0003, Dong Dai 0001, Congfeng Jiang |
ICA3PP (1) | 3 |
| 2012 | An Efficient Parallel Implementation for Three-Dimensional Incompressible Pipe Flow Based on SIMPLEabstractSIMPLE (Semi-Implicit Method for Pressure-Linked Equations) algorithm is important in the simulation of steady flows. As the traditional 3-D SIMPLE algorithm is time-consuming, we propose a parallel SIMPLE algorithm based on a novel tiling strategy -- alternate tiling, through replacing the original linear system and reordering the iteration space tiles. The novelty of our parallel algorithm lies in the introduction of the sequence of iteration space tiles as the sequence of execution, the time skewing technique to partition the iteration space, update operations of the grids from two directions alternately, and the improvement of the data locality. The effectiveness of the parallel algorithm and serial model of finite difference stencil algorithm are validated. Numerical experiments on distributed clusters show that the cache misses and the cost of communication and synchronization are reduced by reordering the tiles of iteration space, and the parallel SIMPLE algorithm based on alternate tiling has a good data locality and parallel efficiency in the three-dimensional incompressible pipe flow project. Li-Ting Zhu, Jian Wan 0001, Jie Mao, Xianghua Xu, Congfeng Jiang, Peng Di |
CCGRID | 3 |
| 2012 | Dual-JT: Toward the high availability of JobTracker in HadoopabstractMapReduce is a state-of-the-art computation paradigm that is becoming widely used for processing large-scale datasets. Hadoop is an open-source implementation of MapReduce and follows a masterCslave architecture. This architecture makes Hadoop suffer from a single point of failure in the JobTracker. In this paper, we design a solution to resolve the single point of failure of the Job Tracker and then enhance its availability. In this solution, a standby Job Tracker is introduced to act as a hot backup node of the active Job Tracker. The standby Job Tracker synchronizes the job execution process with the active Job Tracker by collecting and parsing the job log. If the active Job Tracker fails, the standby Job Tracker can take over quickly. This solution is implemented in Hadoop 0.20.x. Extensive experiments illustrate that this solution effectively enhances the availability of Job Tracker. A big production cluster in a large e-Commerce company has adopted this solution, which avoids interrupting job submission and execution when the Job Tracker fails or restarts. Jian Wan 0001, Minggang Liu, Xixiang Hu, Zujie Ren, Weisong Shi |
CloudCom | 1 |
| 2012 | An Optimized Degree Strategy for Persistent Sensor Network Data DistributionabstractNodes of wireless sensor networks (WSN) may fail easily due to the lack of energy or disaster scenarios. Such failures can severely reduce the persistence and collection efficiency of sensed data. Network coding technology can be employed to enhance the data persistence of wireless sensor network, but it may cause serious "cliff effect" in decoding process. In this paper, the influence of prioritized coding degree distribution strategy on cliff effect is observed, and a distributed storage algorithm PLTD-Alpha is proposed. With PLTD-Alpha, the data in sensor network nodes present a trend that their degree distribution increase along with the degree level predefined, and the persistent data packets can be submitted to the sink node according to its degree in order. Experiment results show that PLTD-Alpha can greatly improve the data collection and decoding efficiency of sensor network while data persistence is not notably affected. Wei Zhang 0138, Qinchao Zhang, Xianghua Xu, Jian Wan 0001 |
PDP | 4 |
| 2011 | An Adaptive Management Mechanism for Resource Scheduling in Multiple Virtual Machine System
Jian Wan 0001, Laurence T. Yang, Yunfa Li 0001, Xianghua Xu, Naixue Xiong |
ATC | 1 |
| 2011 | Network Coding Data Collecting Mechanism Based on Prioritized Degree Distribution in Wireless Sensor NetworkabstractWireless sensor network (WSN) is a typical distributed storage system, and network coding technology is developed to enhance the data persistence of WSN. However, the traditional distributed coding strategy may cause serious "cliff effect" in the decoding process, that is to say, few source data can be recovered before sufficient encoded packets are received. Moreover, nodes may fail due to the lack of energy or the influence of the switching of external environment, such as a disaster scenario. Such failures may concentrate in a small region or distribute in the whole deployment area which can severely reduced the decoding efficiency of the persistent data in WSN. In this paper, we propose the PLTCDS (prioritized LT codes based distributed storage) algorithm to improve the data decoding efficiency when the data persistence is assured. The main idea of PLTCDS is that the predefined node broadcasts a beacon to stimulate the nodes to form the network with degree distribution priority. To ensure the effectiveness of storage nodes, PLTCDS introduces a class of cumulative counter scheme to avoid empty storage. Also we discuss about the idea of another type of PLTCDS. Experimental results show that PLTCDS algorithm can enhance the network data collection performance and reduce the influence of "cliff effect". Wei Zhang 0138, Xianghua Xu, Qinchao Zhang, Jian Wan 0001, Naixue Xiong |
EUC | 4 |
| 2010 | MSNAP: Fault Tolerant Event Localization in Wireless Sensor NetworksabstractThis paper investigates event localization in wireless sensor networks. We improve the SNAP (Subtract on Negative Add on Positive) localization algorithm and propose the MSNAP (Modified Subtract on Negative Add on Positive) localization algorithm with higher localization accuracy and better performance of fault tolerance. First, every sensor node obverses the event signal and compares its observed reading with a threshold. If the reading is above the threshold, the node will send it to the sink station. Otherwise, it remains silent. Based on the observed readings which the nodes report, the sink station constructs the likelihood matrix by simply adding ± 1 contributions in the area around the nodes, whose maximum value points to the event location. Compared with the SNAP algorithm, when constructing the likelihood matrix, MSNAP dynamically adjusts the size of estimated region depending on the observed readings the nodes reported. Experimental results show that the algorithm effectively improves the localization accuracy and fault tolerance. Xianghua Xu, Xueyong Gao, Jian Wan 0001 |
APSCC | 3 |
| 2010 | Regulative Growth Codes: Enhancing Data Persistence in Sparse Sensor NetworksabstractGrowth Codes (GC) enhances the data persistence in dense sensor networks. However, GC exchanges data with neighbors in a completely random way, which may lead to uneven sensor data distribution in sparse sensor network. This significantly reduces the efficiency of GC data acquisition and fault-tolerance in sparse sensor network with less connectivity. In this paper, we propose Regulative Growth Codes (RGC) which use a random sequencing policy in data exchange operation instead of the completely random policy used in GC, and introduce the self-detection mechanism to reduce the redundant exchange. Simulation results show that the performance of RGC is better than GC in sparse sensor networks. Xianghua Xu, Jian Wan 0001 |
APSCC | 3 |
| 2010 | Improve the Completeness of Passive Monitoring Trace in Wireless Sensor NetworkabstractThe performance evaluation of wireless sensor network based on passive monitoring restricted by two respects: first, passive monitoring trace is incomplete, because monitors can not capture every transmission in the network. Second, the information directly extracted from trace is insufficient. Performance evaluation always needs some implicit information, e.g., packet reception. To solving the problems above, we use two processing steps to construct an enhanced trace of network. First, an online merging procedure combines the incomplete traces of various monitors into a single more complete trace. Next, an inference procedure based on finite state machine reconstructs packets that were not captured by any monitor and determines whether a packet was received by its destination. Merging and inference procedure are realized in CTP Network which is incorporated with TinyOS. The evaluation is performed on simulation. The experiment results show that, if the trace contains more than 70% of the total packets, the engine can infer 20%-25% more loss packets and 90% packets' reception. Xianghua Xu, Jian Wan 0001 |
APSCC | 3 |
| 2010 | Power aware job scheduling with QoS guarantees based on feedback controlabstractWith the scale of computing system increases, power consumption has become the major challenge to system performance, reliability and IT management costs. Specifically, system performance and reliability, described by various Quality of Service(QoS) metrics, cannot be guaranteed if the objective is to minimize the total power consumption solely, despite of the violations of QoS. Various methods have been developed to control power consumption to avoid system failures and thermal emergencies through coarse-grained designs. However, the existing methods can be improved and more power can be saved if fine-grained job level adaptation is integrated into them. In this paper a feedback control based power aware job scheduling algorithm is proposed to minimize power consumption in computing system and to provide QoS guarantees. In the proposed algorithm, jobs are scheduled according to the realtime and historical power consumption as well as the QoS requirements. Simulations and experiments on real multi core computing system show that the power potential of the system can be deeply explored while still providing QoS guarantees and the performance degradation is acceptable. The experiment results also show that fine-grained job-level power aware scheduling can achieve better power/performance balancing between multiple processors or cores than coarse-grained methods. Congfeng Jiang, Xianghua Xu, Jian Wan 0001, Xindong You, Ritai Yu |
IWQoS | 3 |
| 2010 | A Control Mechanism about Quality of Service for Resource Scheduling in Multiple Virtual Machine SystemabstractWith the growth of hardware and software resources, It becomes a challenge that how to ensure the Quality of Service (QoS) of resource scheduling in multiple virtual machine system. In order to solve this problem, the authors propose a control mechanism about QoS for resource scheduling. In the control mechanism, we first propose a theoretical model. Then, we present an elicitation algorithm to resolve the optimal model. In order to justify the feasibility and availability of this control mechanism about QoS for resource scheduling in multiple virtual machine system, a series of experiments have been done. The results show that it is feasible to schedule the system resources and control the QoS of system resources for tasks in multiple virtual machine system. Yunfa Li 0001, Xianghua Xu, Jian Wan 0001, Wanqing Li 0003 |
PDCAT | 3 |
| 2009 | A Utility-Based Adaptive Resource Allocation Policy in Virtualized EnvironmentabstractHow to meet the QoS of the applications and improve resource utilization is an important problem in virtualized environment. In this paper, we propose a utility based resource allocation policy with QoS constrained in virtualized environment. Firstly, we build a model reflecting the mapping relation between performance metrics and resource allocation through Web server benchmarking experiments. Then, we devise a utility function as objective to optimizing the total utility and achieving a reasonable resource allocation. Finally, based on the model and the requirements of the applications QoS, we present an optimized policy for resource allocation in virtualized system. The simulation results show that the utility-based policy is effective to allocate resource from user perspective while improving the resource utilization. Jian Wan 0001, Peipei Shan, Xianghua Xu |
DASC | 1 |
| 2009 | Grey Prediction Control of Adaptive Resources Allocation in Virtualized Computing SystemabstractIn order to improve the resource utilization of virtual machine and control the resource allocation online effectively, in this paper, we present a grey prediction control model used for dynamic resource allocation in virtual machine as workloads changing. First, we forecast the allocation of virtualized resources by the grey control model. We also adjust the boundary conditions of grey prediction model to make the prediction more accurately. Then, the control theory is used to feedback control resource utilization to obtain desired resource utilization levels by regulating the value of allocation of virtualized resources automatically. Our experimental results show the grey control model is effective in the virtualized resource allocation. The control model and algorithm can be applied to other resource allocation. Xianghua Xu, Yanna Yan, Jian Wan 0001 |
DASC | 3 |
| 2008 | Aeolus: Reconcilable Key Management Mechanism for Secure Group Communication in GridabstractIn traditional grid, grid user can not validate whether grid service registered in grid platform can execute correctly or not because the system does not provide the measurement mechanism for grid service. In order to solve this problem, we propose a series of strategies and methods for trusted grid. These strategies and methods include: an access control policy for reference database, an integrity attestation for reference datasheet, a construction method of reference datasheet, a trusted storage algorithm of reference datasheet and a report algorithm of reference datasheet. All these constitute the measurement mechanism of service for trusted grid. Some experiments are done in order to validate the feasibility and the availability of this mechanism. The results show that it is feasible and efficient to validate whether grid service registered in trusted grid platform can execute correctly or not. Yunfa Li 0001, Xianghua Xu, Jian Wan 0001, Hai Jin 0001, Zongfen Han |
APSCC | 3 |
| 2008 | Grid computing based large scale Distributed Cooperative Virtual Environment SimulationabstractGrids provide infrastructures and solutions to solve large scale cooperative problems such as large scale distributed virtual environment simulation, multi-institutional scientific computing and data analysis, etc. In this paper, the key techniques for Grid computing based large scale Distributed Cooperative Virtual Environment Simulation (GDCVES) are discussed and a hierarchical architecture of GDCVES is proposed. The solutions of GDCVES, such as resource management, massive data management, security aware task scheduling, and fault-tolerance are also discussed. To evaluate the feasibility and scalability of GDCVES, a prototype was implemented and the simulation workflow framework is also analyzed. Congfeng Jiang, Xianghua Xu, Jian Wan 0001 |
CSCWD | 3 |
| 2008 | A Peer-to-Peer Assisting Scheme for Live Streaming Services
Jian Wan 0001, Liangjin Lu, Xianghua Xu, Xueping Ren |
GPC | 1 |
| 2006 | An Approach of Image Retrieval based on Bayesian and AAMabstractSemantic-based image retrieval using low-level visual features is a challenging and important issue in content-based image retrieval. In this paper, we cast the image retrieval issue in a Bayesian framework and AAM (the active appearance model). Specifically, we propose an approach for complex semantic-based image retrieval, for example selecting the grassland images including horses. That is, the approach is used for selecting images including specific scene and model. In the approach, we integrate low-level features and spatial distribution into Bayesian frame. The approach uses Bayesian framework to select the images including the scene (forest, grassland), and uses AAM to select the images including the specific model (horse). Experimental results indicate that our approach is effective in complex semantic-based image retrieval and provides a sound retrieval performance. Xueping Ren, Jian Wan 0001, Xianghua Xu |
SMC | 2 |