EDBT 2026 Demo / reviewers in the wild / expert
Xingjun Zhang
dblp:10/1263
· DBLP profile ↗
86ranked-venue papers
7as first author
49since 2021 · last 2027
0000-0003-1434-7016ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 42 · 1 first-author · 27 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Computer networks · 7 · 2 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | STAO: Adaptive spatio-temporal orchestration for efficient and QoS-aware DNN inference on GPU
Shaoxun Wang, Xingjun Zhang, Tongchao Miao |
Future Gener. Comput. Syst. | 2 |
| 2026 | Efficient disk read and recovery cost reduction approach in heterogeneous liberation-coded storage systems
Ningjing Liang, Genqing Bian, Songchen Huang, Xingjun Zhang |
Future Gener. Comput. Syst. | 6 |
| 2026 | MAFNet: Multi-scale active fusion network for long-term time series forecasting
Qianyang Li, Xingjun Zhang, Shaoxun Wang, Jia Wei 0002 |
Neurocomputing | 2 |
| 2026 | Dual-Pronged Deep Learning Preprocessing on Heterogeneous Platforms With CPU, Accelerator and CSDabstractFor image-related deep learning tasks, the first step often involves reading data from external storage and performing preprocessing on the CPU. As accelerator speed increases and the number of single compute node accelerators increases, the computing and data transfer capabilities gap between accelerators and CPUs gradually increases. Data reading and preprocessing become progressively the bottleneck of these tasks. Our work, DDLP, addresses the data computing and transfer bottleneck of deep learning preprocessing using Computable Storage Devices (CSDs). DDLP allows the CPU and CSD to efficiently parallelize preprocessing from both ends of the datasets, respectively. To this end, we propose two adaptive dynamic selection strategies to make DDLP control the accelerator to automatically read data from different sources. The two strategies trade-off between consistency and efficiency. DDLP achieves sufficient computational overlap between CSD data preprocessing and CPU preprocessing, accelerator computation, and accelerator data reading. In addition, DDLP leverages direct storage technology to enable efficient SSD-to-accelerator data transfer. In addition, DDLP reduces the use of expensive CPU and DRAM resources with more energy-efficient CSDs, alleviating preprocessing bottlenecks while significantly reducing power consumption. Extensive experimental results show that DDLP can improve learning speed by up to 23.5% on ImageNet Dataset while reducing energy consumption by 19.7% and CPU and DRAM usage by 37.6%. DDLP also improves the learning speed by up to 27.6% on the Cifar-10 dataset. Jia Wei 0002, Xingjun Zhang, Witold Pedrycz, Jie Zhao 0002 |
IEEE Trans. Computers | 2 |
| 2026 | Dpwmixer: dual-path wavelet mixer for long-term time series forecasting
Qianyang Li, Xingjun Zhang, Shaoxun Wang |
J. Supercomput. | 2 |
| 2025 | USFF: A Unified Sales Forecasting Framework for Vending Machines
Qianyang Li, Xingjun Zhang, Shaoxun Wang |
IEEE Big Data | 2 |
| 2025 | KANETAS: an elastic scheduler for heterogeneous many-core systems
Zhao Mao, Xingjun Zhang |
CCF Trans. High Perform. Comput. | 2 |
| 2025 | A trajectory privacy protection method based on the replacement of points of interest in hotspot regions
Ruowei Gui, Xiaolin Gui, Xingjun Zhang |
Comput. Secur. | 3 |
| 2025 | Unifying and revisiting Sharpness-Aware Minimization with noise-injected micro-batch scheduler for efficiency improvement
Xingjun Zhang, Zhendong Tan |
Neural Networks | 2 |
| 2025 | Dynamic Fuzzy Sampler for Graph Neural NetworksabstractGraph neural networks (GNNs) have been widely used in many fields. Inductive learning has replaced transductive learning as the current mainstream paradigm for GNN training due to its higher memory efficiency, computing speed, and stronger generalization. Neighbor node sampling as a key step in GNN inductive learning is crucial to the model performance. However, existing samplers only focus on how to sample nodes from the adjacency matrix, ignoring the fact that different neighbors have different impacts on the target node at different moments. They usually aggregate the neighbor information in a simple way such as averaging or summing, which limits the information representation, robustness, and generalization. In order to address the shortcomings of existing graph inductive learning samplers, this article proposes a dynamic fuzzy sampler (DFS) based on a Gaussian fuzzy system. The DFS fully takes into account the diversity of nodes in the graph-structured data, and efficiently models and handles the uncertainties and fuzziness of the mutual information of various nodes at different moments. Specifically, DFS first innovatively constructs a learnable Gaussian fuzzy set system for determining the membership degree of different neighbors to the target node at different moments. Subsequently, DFS aggregates the target node embeddings and membership-weighted neighbor embeddings to update the target node's features, which makes the target node utilize the sampling information more effectively. The aggregated target node effectively captures the graph structure information and neighbor node information, which can facilitate the subsequent GNN-based graph representation model with stronger representation and generalization capabilities. Our supervised and self-supervised experimental results on graph datasets of different sizes show that DFS has consistently excellent performance, significantly outperforming other state-of-the-art sampling schemes. DFS achieves up to 1.90% and 9.52% F1-score improvement compared to the state-of-the-art schemes on small- and large-scale graphs, respectively. Jia Wei 0002, Xingjun Zhang, Witold Pedrycz, Weiping Ding 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | WRP: Weight Recover Prune for Structured SparsityabstractAs the scale of Large Language Models (LLMs) increases, it is necessary to compress the models to reduce the substantial demand on computational resources.Network pruning significantly reduces the model size by converting the weight matrix from dense to sparse data format.Current methodologies advocate for one-shot pruning to avoid the expense of retraining, ensuring the maintenance of model performance under conditions of 50%-60% unstructured pruning.Nevertheless, matrices characterized by this level of sparsity could not be treated as sparse matrices, because the indices would incur significant costs.To mitigate this problem, NVIDIA introduced the 2:4 structured sparsity.However, we observe a notable decline in model performance when adopting 2:4 structured sparsity due to group constraints.In this paper, we introduce the Weight Recover Prune (WRP) approach.By recovering a minimal set of critical weights, WRP aims to enhance model performance while maintaining the efficiency of the compression.Our evaluation of the WRP method on the LLAMA2 and OPT models shows that it outperforms other 2:4 pattern one-shot pruning methods.Meanwhile, WRP can guarantee that the size of the pruned model is about 60% of the dense model.Our code is available at Zhendong Tan, Xingjun Zhang |
ACL (1) | 2 |
| 2024 | Improving the Relationship Between B+-Tree and Memory Allocator for Persistent MemoryabstractIn the traditional volatile DRAM, the B+-Tree inter-acts with the memory allocator by specifying space requirements for scaling. However, the emergence of persistent memory (PM) as a potential substitute for DRAM heap implementation presents new challenges, including crash-consistent issues and limited performance. Prevalent solutions in both the allocators and B+- Trees are confronted with similar dilemmas between runtime performance and recovery efficiency. In this paper, we propose a novel crash-consistent protocol called ST-protocol (share and talk protocol) to address these dilemmas. The core idea of ST-protocol is overlapping the post-crash garbage collection of the allocator with the recovery of the B+-Tree, according to the shared node status. This overlapped process can be performed by multiple threads, according to the shared space layout. Consequently, both the allocator and index can achieve high runtime performance from the volatile components in DRAM without much compromising the recovery efficiency. We further propose a B+-Tree model and a new PM allocator as two interacting characters in ST-protocol, and the experimental results demonstrate that our design outperforms existing solutions by a large margin in both runtime performance and recovery efficiency. Xingjun Zhang |
ICDE | 2 |
| 2024 | A local differential privacy extension scheme for sensitive locations of hotspot areasabstractWith the popularization of smartphones with GPS, user information attached to the locations is facing the risk of leakage. In the real world, even if the locations of a user are well protected, attackers can also mine user privacy by analyzing the user correlations in common hotspots. To solve the above problems, we propose a local differential privacy extension scheme in hotspot areas. In this scheme, we firstly obtain the user’s hotspots by mining the user trajectories based on the sliding time window, and then, we extract the correlation degrees among users from these hotspots using the Jaccard correlation coefficient, and finally we introduce the personalization correlation sensitivity to extend the local differential privacy so as to protect sensitive locations in hotspot areas. Experiments show that, compared with existing methods, our scheme can improve the usability of trajectories up to $\mathbf{6. 6 4 \%}$ at the same privacy level. Ruowei Gui, Xingjun Zhang, Xiaolin Gui |
ICPADS | 2 |
| 2024 | Roughness tooth surface loaded contact pressure forecasting model for face-hobbed hypoid gears
Han Ding 0003, Shifeng Rong, Xingjun Zhang, Keliang Li, Kaibin Rong |
Adv. Eng. Informatics | 3 |
| 2024 | A highly write-optimized concurrent B+-tree for persistent memory
Xingjun Zhang |
Future Gener. Comput. Syst. | 2 |
| 2024 | A Location Correlation Differential Privacy Extension Scheme Based on User Spatiotemporal CharacteristicsabstractWith the popularity of mobile terminals with GPS functions, location-based services are widely used, and all kinds of user information attached to the location are facing the risk of disclosure, and privacy protection is being challenged. In current researches, it is usually assumed that the locations of different users are independent of each other. However, in real world, the locations of different users have some certain internal correlation. Even if the locations of a single user are well protected, attackers can still mine user privacy through the correlation analysis of locations. To solve the above problems, this article proposes a multiuser location-correlated differential privacy extension scheme under strict privacy budget. In this scheme, we first extract the user spatiotemporal characteristics by mining the stay points and stay areas from trajectories based on locations with timestamps, and then, we calculate the correlation degree among users using the Jaccard correlation coefficient according to the spatiotemporal characteristics, and further, we realize the adaptive differential privacy protection of different users by introducing the concept of individual correlation sensitivity, and finally, we design the differential privacy extension method to protect sensitive locations in the stay areas. Experimental results show that, compared with the existing methods, our proposed scheme not only can improve the usability of the trajectories after privacy protection, but also can enhance the privacy protection of the sensitive locations in the stay areas. Ruowei Gui, Xingjun Zhang, Xiaolin Gui, Jinsong Han |
IEEE Internet Things J. | 2 |
| 2024 | A Concise Concurrent B+-Tree for Persistent MemoryabstractPersistent memory (PM) presents a unique opportunity for designing data management systems that offer improved performance, scalability, and instant restart capability. As a widely used data structure for managing data in such systems, B + -Tree must address the challenges presented by PM in both data consistency and device performance. However, existing studies suffer from significant performance degradation when maintaining data consistency on PM. To settle this problem, we propose a new concurrent B + -Tree, CC-Tree, optimized for PM. CC-Tree ensures data consistency while providing high concurrent performance, thanks to several technologies, including partitioned metadata, log-free split, and lock-free read. We conducted experiments using state-of-the-art indices, and the results demonstrate significant performance improvements, including approximately 1.2–1.6x search, 1.5–1.7x insertion, 1.5–2.8x update, 1.9–4x deletion, 0.9–10x range scan, and up to 1.55–1.82x in hybrid workloads. Xingjun Zhang |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | Revisit and Benchmarking of Automated Quantization Toward Fair ComparisonabstractAutomated quantization has emerged as an entirely new design paradigm to automate the optimal configuration of bitwidth for deep neural networks (DNNs), making the DNN more memory-efficient and faster to execute on hardware with limited resources. Reinforcement learning (RL) and differentiable neural architecture search (DNAS) are two main solution paths that have shown their superiority. Yet, there are countless methods with various implementations within each path. It has been hard to comprehend their differences and make a relatively fair comparison due to the lack of a benchmark framework and a clear analysis of which aspects are common, respectively distinct, between different implementations. To this end, we introduce BenQ to pave the way towards fair comparisons in two separate race tracks, i.e., intra-comparison of the RL-based and the DNAS-based methods, respectively. We provide a systematic approach, which helps to reveal relatively vital aspects of different implementations. Finally, we conduct comprehensive experi-ments on VGG, AlexNet, ResNet, GoogleNet, MobileNet-V2, and Vision Transformer (ViT), and the new observations shed light on potential future directions for automated quantization to move forward. Xingjun Zhang, Zeyu Ji, Jia Wei 0002 |
IEEE Trans. Computers | 2 |
| 2024 | Energy-efficient DAG scheduling with DVFS for cloud data centers
Wenbing Yang, Mingqiang Zhao, Xingjun Zhang |
J. Supercomput. | 4 |
| 2023 | Leader population learning rate schedule
Jia Wei 0002, Xingjun Zhang, Zhimin Zhuo, Zeyu Ji, Qianyang Li |
Inf. Sci. | 2 |
| 2023 | Fastensor: Optimise the Tensor I/O Path from SSD to GPU for Deep Learning TrainingabstractIn recent years, benefiting from the increase in model size and complexity, deep learning has achieved tremendous success in computer vision (CV) and (NLP). Training deep learning models using accelerators such as GPUs often requires much iterative data to be transferred from NVMe SSD to GPU memory. Much recent work has focused on data transfer during the pre-processing phase and has introduced techniques such as multiprocessing and GPU Direct Storage (GDS) to accelerate it. However, tensor data during training (such as Checkpoints, logs, and intermediate feature maps), which is also time-consuming, is often transferred using traditional serial, long-I/O-path transfer methods. In this article, based on GDS technology, we built Fastensor, an efficient tool for tensor data transfer between the NVMe SSDs and GPUs. To achieve higher tensor data I/O throughput, we optimized the traditional data I/O process. We also proposed a data and runtime context-aware tensor I/O algorithm. Fastensor can select the most suitable data transfer tool for the current tensor from a candidate set of tools during model training. The optimal tool is derived from a dictionary generated by our adaptive exploration algorithm in the first few training iterations. We used Fastensor’s unified interface to test the read/write bandwidth and energy consumption of different transfer tools for different sizes of tensor blocks. We found that the execution efficiency of different tensor transfer tools is related to both the tensor block size and the runtime context. We then deployed Fastensor in the widely applicable Pytorch deep learning framework. We showed that Fastensor could perform superior in typical scenarios of model parameter saving and intermediate feature map transfer with the same hardware configuration. Fastensor achieves a 5.37x read performance improvement compared to torch.save () when used for model parameter saving. When used for intermediate feature map transfer, Fastensor can increase the supported training batch size by 20x, while the total read and write speed is increased by 2.96x compared to the torch I/O API. Jia Wei 0002, Xingjun Zhang |
ACM Trans. Archit. Code Optim. | 2 |
| 2023 | Fine-Grained Conditional Convolution Network With Geographic Features for Temperature PredictionabstractShort-to-medium term temperature prediction in high resolution is a very challenging task, involving meteorology, physics, mathematics, geography, and many other subjects. Its purpose is to fit a complex function from historical meteorological data to predict the future 1–5 days temperature, which is a typical spatio-temporal prediction problem. Meteorological data show complex correlations in local space. Most of the existing machine learning methods are based on image pixel-level tasks or spatio-temporal prediction tasks, which model meteorological data without considering the characteristics of meteorological data and use rough global patterns to model local space which would lose many details. To address the above issues, our work fine-grained conditional convolution network (FCCN) proposes a novel grid-level conditional convolution module, including a local geographic adaptive weight (GAW) and a local data adaptive weight (DAW). These two components are integrated into a multiscale meteorological fusion gated recurrent unit (GRU) architecture for the end-to-end temperature prediction. Experiments in real-world datasets from ERA-5 show our FCCN model has a better performance than all other baseline methods. Guoshuai Zhao 0001, Junjiao Liu, Xingjun Zhang, Xueming Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Towards Handling Sudden Changes in Feature Maps During Depth EstimationabstractDepth estimation aims to predict depth map from RGB images without high cost equipments. Deep learning based depth estimation methods have shown their effectiveness. However in existing methods, depth information is represented by a per-pixel depth map. Such depth map representation is fragile facing different kinds of depth changes. This paper proposes a Compressive Sensing based Depth Representation (CSDR) scheme, which formulates the problem of depth estimation in pixel space into the task of fixed-length vector regression in representation space. In this way, deep model training errors will not directly interfere depth estimation, and distortions in estimated depth maps can be restrained in the greatest extent. In addition, we improve depth estimation from two other aspects: model structure and loss function. To capture the features in different scales, we propose a Multiscale Encoder \& Multiscale Decoder (MEMD) structure as the vector regression model. To further deal with depth change, we also modify the loss function, where the curvature difference between ground truth and estimation is directly incorporated. With the support of CSDR, MEMD and the curvature loss, the proposed approach achieves superior performance on a challenging depth estimation dataset: NYU-Depth-v2. A range of experiments support our claim that regression in CSDR space performs better than traditionally direct depth map estimation in pixel space. Yu Cao 0016, Xubin Feng, Meilin Xie, Ke Li 0032, Xingjun Zhang, Xueming Qian |
IEEE Trans. Multim. | 6 |
| 2022 | BenQ: Benchmarking Automated Quantization on Deep Neural Network AcceleratorsabstractHardware-aware automated quantization promises to unlock an entirely new algorithm-hardware co-design paradigm for efficiently accelerating deep neural network (DNN) inference by incorporating the hardware cost into the reinforcement learning (RL) -based quantization strategy search process. Existing works usually design an automated quantization algorithm targeting one hardware accelerator with a device-specific performance model or pre-collected data. However, determining the hardware cost is non-trivial for algorithm experts due to their lack of cross-disciplinary knowledge in computer architecture, compiler, and physical chip design. Such a barrier limits reproducibility and fair comparison. Moreover, it is notoriously challenging to interpret the results due to the lack of quantitative metrics. To this end, we first propose BenQ, which includes various RL-based automated quantization algorithms with aligned settings and encapsulates two off-the-shelf performance predictors with standard OpenAI Gym API. Then, we leverage cosine similarity and manhattan distance to interpret the similarity between the searched policies. The experiments show that different automated quantization algorithms can achieve near equivalent optimal trade-offs because of the high similarity between the searched policies, which provides insights for revisiting the innovations in automated quantization algorithms. Xingjun Zhang, Zeyu Ji, Jia Wei 0002 |
DATE | 2 |
| 2022 | How Much Storage Do We Need for High Performance ServerabstractProcessor, memory, and storage are the three most critical components of a High Performance Server(HPS). For a long time, given the processor and memory, how much storage capacity do we need? has been a problem that has deeply troubled the academic and enterprise communities. Especially in recent years, with the rapid development of Artificial Intelligence(AI), AI-based data-intensive tasks such as Deep Learning(DL), Re-inforcement Learning(RL), and High Performance Data Analy-sis(HPDA) have taken up the vast majority of processors, mem-ory, and storage overhead. The ability to support AI applications has become a key evaluation metric for the performance of HPS. Therefore, we propose an HPS storage design solution for typical AI applications. Furthermore, as AI models continue to grow more giant and the GPU Memory Wall problem becomes increasingly significant, using storage for offloading models and intermediate variables becomes the mainstream approach for training and inferring Extreme-Scale AI Models in the future. We need to consider the static overhead of models and datasets and the dynamic offloading requirements that may arise when AI tasks are run. We propose the Server Storage Computing Ratio (SSCR) model. The model uses DL training capabilities to characterize processor and memory performance. When config-ured for AI tasks, it can get the maximum server performance and the least amount of storage space. In other words, in the HPS mainly oriented to AI tasks, our model answers the question What is the minimum amount of Storage space that needs to be configured to maximize server performance for a given processor and memory? Xingjun Zhang |
ICDE | 2 |
| 2022 | Status, challenges and trends of data-intensive supercomputing
Jia Wei 0002, Pei Ren, Yujia Lei, Yuqi Qu, Qiyu Jiang, Xiaoshe Dong, Weiguo Wu, Qiang Wang 0062, Xingjun Zhang |
CCF Trans. High Perform. Comput. | 12 |
| 2022 | A high-applicability heterogeneous cloud data centers resource management algorithm based on trusted virtual machine migration
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008, Xingjun Zhang |
Expert Syst. Appl. | 4 |
| 2022 | GARLSched: Generative adversarial deep reinforcement learning task scheduling optimization for large-scale high performance computing systems
Xingjun Zhang, Jia Wei 0002, Zeyu Ji |
Future Gener. Comput. Syst. | 2 |
| 2022 | LogSC: Model-based one-sided communication performance estimation
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Xingjun Zhang |
Future Gener. Comput. Syst. | 5 |
| 2022 | DPLRS: Distributed Population Learning Rate Schedule
Jia Wei 0002, Xingjun Zhang, Zeyu Ji |
Future Gener. Comput. Syst. | 2 |
| 2022 | SA-RSR: a read-optimal data recovery strategy for XOR-coded distributed storage systemsabstractTo ensure the reliability and availability of data, redundancy strategies are always required for distributed storage systems. Erasure coding, one of the representative redundancy strategies, has the advantage of low storage overhead, which facilitates its employment in distributed storage systems. Among the various erasure coding schemes, XOR-based erasure codes are becoming popular due to their high computing speed. When a single-node failure occurs in such coding schemes, a process called data recovery takes place to retrieve the failed node’s lost data from surviving nodes. However, data transmission during the data recovery process usually requires a considerable amount of time. Current research has focused mainly on reducing the amount of data needed for data recovery to reduce the time required for data transmission, but it has encountered problems such as significant complexity and local optima. In this paper, we propose a random search recovery algorithm, named SA-RSR, to speed up single-node failure recovery of XOR-based erasure codes. SA-RSR uses a simulated annealing technique to search for an optimal recovery solution that reads and transmits a minimum amount of data. In addition, this search process can be done in polynomial time. We evaluate SA-RSR with a variety of XOR-based erasure codes in simulations and in a real storage system, Ceph. Experimental results in Ceph show that SA-RSR reduces the amount of data required for recovery by up to 30.0% and improves the performance of data recovery by up to 20.36% compared to the conventional recovery method. Xingjun Zhang, Ningjing Liang, Changjiang Zhang |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2022 | C-Lop: Accurate contention-based modeling of MPI concurrent communication
Ziheng Wang 0002, Heng Chen 0002, Weiling Cai, Xiaoshe Dong, Xingjun Zhang |
Parallel Comput. | 5 |
| 2022 | Optimizing Small-Sample Disk Fault Detection Based on LSTM-GAN ModelabstractIn recent years, researches on disk fault detection based on SMART data combined with different machine learning algorithms have been proven to be effective. However, these methods require a large amount of data. In the early stages of the establishment of a data center or the deployment of new storage devices, the amount of reliability data for disks is relatively limited, and the amount of failed disk data is even less, resulting in the unsatisfactory detection performances of machine learning algorithms. To solve the above problems, we propose a novel small sample disk fault detection (SSDFD) 1 optimizing method based on Generative Adversarial Networks (GANs). Combined with the characteristics of hard disk reliability data, the generator of the original GAN is improved based on Long Short-Term Memory (LSTM), making it suitable for the generation of failed disk data. To alleviate the problem of data imbalance and expand the failed disk dataset with reduced amounts of original data, the proposed model is trained through adversarial training, which focuses on the generation of failed disk data. Experimental results on real HDD datasets show that SSDFD can generate enough virtual failed disk data to enable the machine learning algorithm to detect disk faults with increased accuracy under the condition of a few original failed disk data. Furthermore, the model trained with 300 original failed disk data has a significant effect on improving the accuracy of HDD fault detection. The optimal amount of generated virtual data are, 20–30 times that of the original data. Yufei Wang 0008, Xiaoshe Dong, Weiduo Chen, Xingjun Zhang |
ACM Trans. Archit. Code Optim. | 5 |
| 2022 | Implementation and optimization of ChaCha20 stream cipher on sunway taihuLight supercomputer
Weilin Cai, Heng Chen 0002, Ziheng Wang 0002, Xingjun Zhang |
J. Supercomput. | 4 |
| 2022 | SunwayURANS: 3D full-annulus URANS simulations of transonic axial compressors on Sunway TaihuLight
Heng Chen 0002, Ziheng Wang 0002, Xiaoshe Dong, Xingjun Zhang |
J. Supercomput. | 6 |
| 2022 | EP4DDL: addressing straggler problem in heterogeneous distributed deep learning
Zeyu Ji, Xingjun Zhang, Jia Wei 0002 |
J. Supercomput. | 2 |
| 2022 | Thou code: a triple-erasure-correcting horizontal code with optimal update complexity
Ningjing Liang, Xingjun Zhang, Heng Chen 0002, Changjiang Zhang |
J. Supercomput. | 2 |
| 2022 | Extending τ-Lop to model MPI blocking primitives on shared memory
Ziheng Wang 0002, Heng Chen 0002, Xiaoshe Dong, Weilin Cai, Yan Kang 0005, Xingjun Zhang |
J. Supercomput. | 6 |
| 2022 | Domain Adaptive Box-Supervised Instance Segmentation Network for Mitosis DetectionabstractThe number of mitotic cells present in histopathological slides is an important predictor of tumor proliferation in the diagnosis of breast cancer. However, the current approaches can hardly perform precise pixel-level prediction for mitosis datasets with only weak labels (i.e., only provide the centroid location of mitotic cells), and take no account of the large domain gap across histopathological slides from different pathology laboratories. In this work, we propose a Domain adaptive Box-supervised Instance segmentation Network (DBIN) to address the above issues. In DBIN, we propose a high-performance Box-supervised Instance-Aware (BIA) head with the core idea of redesigning three box-supervised mask loss terms. Furthermore, we add a Pseudo-Mask-supervised Semantic (PMS) head for enriching characteristics extracted from underlying feature maps. Besides, we align the pixel-level feature distributions between source and target domains by a Cross-Domain Adaptive Module (CDAM), so as to adapt the detector learned from one lab can work well on unlabeled data from another lab. The proposed method achieves state-of-the-art performance across four mainstream datasets. A series of analysis and experiments show that our proposed BIA and PMS head can accomplish mitosis pixel-wise localization under weak supervision, and we can boost the generalization ability of our model by CDAM. Liangfu Li, Xingjun Zhang, Xueming Qian |
IEEE Trans. Medical Imaging | 4 |
| 2022 | SPGNet: Serial and Parallel Group NetworkabstractNeural-network Processing Units (NPU), which specializes in the acceleration of deep neural networks (DNN), is of great significance to latency-sensitive areas like robotics or edge computing. However, there are few works focusing on the network design for NPU in recent studies. Most of the popular lightweight structures (e.g. MobileNet) are designed with depthwise convolution, which has less computation in theory but is not friendly to existing hardwares, and the speed tested on NPU is not always satisfactory. Even under similar FLOPs (the number of multiply-accumulates), vanilla convolution operation is always faster than depthwise one. In this paper, we will propose a novel architecture named Serial and Parallel Group Network (SPGNet), which can capture discriminative multi-scale information and at the same time keep the structure compact. Extensive evaluations have been conducted on different computer vision tasks, e.g. image classification (CIFAR and ImageNet), object detection (PASCAL VOC and MS COCO) and person re-identification (Market-1501 and DukeMTMC-ReID). The experimental results show that our proposed SPGNet can achieve comparable performance with the state-of-the-art networks while the speed is 120% faster than MobileNetV2 under similar FLOPS and over 300% faster than GhostNet with similar accuracy on NPU. Xuan Wang 0018, Shenqi Lai, Zhenhua Chai, Xingjun Zhang, Xueming Qian |
IEEE Trans. Multim. | 4 |
| 2021 | HaDPA: A Data-Partition Algorithm for Data Parallel Applications on Heterogeneous HPC Platforms
Yuqi Qu, Xingjun Zhang |
ICA3PP (2) | 4 |
| 2021 | Double deep Q-learning network-based path planning in UAV-assisted wireless powered NOMA communication networksabstractThis paper studies an unmanned aerial vehicle (UAV)-enabled wireless power communication networks (WPCN-s), where the UAV provides energy for mobile user nodes (M-UNs) and receives information from M-UNs. The movement of M-UN complies with a Gauss-Markov random model. To ensure acceptable quality-of-service (QoS), we consider dynamically planning the flight path of the UAV according to the movements of M-UNs. Since the flight time of UAV is restricted by limited energy, nonorthogonal multiple access (NOMA) is adopted to access a large number of M-UNs for simultaneous information transmission. Based on the above considerations, we aim to maximize the throughput via path planning of the UAV, subject to the QoS requirements of M-UNs and the UAV's energy constraint. To handle the challenges brought by dynamically changing channels to solving the problem, we propose a QoS-based double deep Q-learning network (DDQN). Numerical simulation results show that, compared with the conventional algorithms, the proposed framework achieves higher throughput. Ming Lei 0003, Scott Fowler, Juzhen Wang, Xingjun Zhang, Bocheng Yu, Bin Yu 0008 |
VTC Fall | 4 |
| 2021 | Efficient Computation Offloading for Edge-cloud Collaborative NetworksabstractMobile edge computing is a novel paradigm that provides computing capabilities at the edge of the radio access network close to end devices to support latency-critical applications and services. However, the benefits will be canceled out by the limited computing capacity of edge servers. An edge-cloud paradigm has been studied to improve computing capabilities to solve the above problem. In this paper, a multi-cell edge-cloud architecture is considered to meet the demands of latency-sensitive applications and deal with large-scale data offloading. The optimized offloading scheme is studied to minimize the devices' overhead which is measured as a function of energy consumption and computational cost. We formulate the problem as a Mixed Integer Linear Programming Problem and adopt the Branch-and-Bound algorithm to solve it. Due to the high time overhead of the method, we first transform the problem into a more tractable form and then adopt a learning approach to imitate the branching strategy to improve the Branch-and-Bound algorithm. Experiments of results show that our approach can reduce the time-cost of the Branch-and-Bound and the result is close to the traditional scheme. Bocheng Yu, Xingjun Zhang, Juzhen Wang, Ming Lei 0003 |
VTC Fall | 2 |
| 2021 | Energy-aware task scheduling optimization with deep reinforcement learning for large-scale heterogeneous systems
Xingjun Zhang, Jia Wei 0002, Zeyu Ji |
CCF Trans. High Perform. Comput. | 2 |
| 2021 | A tile-fusion method for accelerating Winograd convolutions
Zeyu Ji, Xingjun Zhang, Jia Wei 0002 |
Neurocomputing | 2 |
| 2021 | Crowd Scene Analysis Encounters High Density and Scale VariationabstractCrowd scene analysis receives growing attention due to its wide applications. Grasping the accurate crowd location is important for identifying high-risk regions. In this article, we propose a Compressed Sensing based Output Encoding (CSOE) scheme, which casts detecting pixel coordinates of small objects into a task of signal regression in encoding signal space. To prevent gradient vanishing, we derive our own sparse reconstruction backpropagation rule that is adaptive to distinct implementations of sparse reconstruction and makes the whole model end-to-end trainable. With the support of CSOE and the backpropagation rule, the proposed method shows more robustness to deep model training error, which is especially harmful to crowd counting and localization. The proposed method achieves state-of-the-art performance across four mainstream datasets, especially achieves excellent results in highly crowded scenes. A series of analysis and experiments support our claim that regression in CSOE space is better than traditionally detecting coordinates of small objects in pixel space for highly crowded scenes. Xingjun Zhang, Xueming Qian |
IEEE Trans. Image Process. | 4 |
| 2021 | Performance evaluation of convolutional neural network on Tianhe-3 prototype
Weiduo Chen, Xiaoshe Dong, Heng Chen 0002, Qiang Wang 0062, Xingda Yu, Xingjun Zhang |
J. Supercomput. | 6 |
| 2021 | OKCM: improving parallel task scheduling in high-performance computing systems using online learning
Xingjun Zhang, Zeyu Ji, Xiaoshe Dong, Chenglong Hu |
J. Supercomput. | 2 |
| 2021 | Throughput maximization for UAV-assisted wireless powered D2D communication networks with a hybrid time division duplex/frequency division duplex scheme
Ming Lei 0003, Xingjun Zhang, Bocheng Yu, Scott Fowler, Bin Yu 0008 |
Wirel. Networks | 2 |
| 2020 | Memory-aware resource management algorithm for low-energy cloud data centers
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008, Xingjun Zhang |
Future Gener. Comput. Syst. | 4 |
| 2020 | H2Pregel : A partition-based hybrid hierarchical graph computation approach
Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
Future Gener. Comput. Syst. | 4 |
| 2020 | Sketch-Based Image Retrieval With Multi-Clustering Re-RankingabstractTo improve the performance of sketch-based image retrieval (SBIR) methods, most existing SBIR methods develop brand new SBIR methods. In fact, a re-ranking approach, which can refine the retrieval results of SBIR methods, is also beneficial. Inspired by this, in this paper, an SBIR re-ranking approach based on multi-clustering is proposed. In order to make the re-ranking approach invisible to users and adaptive to different types of image datasets, we made it an unsupervised method using blind feedback. Distinguished from the existing methods, this re-ranking approach uses the semantic information of three types of images: edge maps, object images (images with black background and natural images' foreground objects) and natural images themselves. With the initial retrieval results of an SBIR method, our approach first does the clustering operation for three types of images. Then, we utilize the clustering results to generate a cluster score for each initial retrieval result. Finally, the cluster score is used to calculate the final retrieval scores for the initial retrieval results. The experiments on different SBIR datasets are conducted. Experimental results demonstrate that, by implementing our re-ranking approach, the retrieval accuracy of a variety of SBIR methods is increased. Furthermore, the comparisons between our re-ranking method and the existing re-ranking methods are given. Luo Wang, Xueming Qian, Xingjun Zhang, Xingsong Hou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Structured mesh-oriented framework design and optimization for a coarse-grained parallel CFD solver based on hybrid MPI/OpenMP programming
Xiaoshe Dong, Nianjun Zou, Weiguo Wu, Xingjun Zhang |
J. Supercomput. | 5 |
| 2020 | A low-power task scheduling algorithm for heterogeneous cloud computing
Bin Liang 0005, Xiaoshe Dong, Yufei Wang 0008, Xingjun Zhang |
J. Supercomput. | 4 |
| 2020 | NADE: nodes performance awareness and accurate distance evaluation for degraded read in heterogeneous distributed erasure code-based storage
Xingjun Zhang, Xiaoshe Dong |
J. Supercomput. | 1 |
| 2020 | Fine-grained scheduling in multi-resource clusters
Mosong Zhou, Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
J. Supercomput. | 4 |
| 2019 | Optimal Resource Allocation Through Joint VM Selection and Placement in Private Clouds
Hongkun Chen, Feilong Tang 0001, Linghe Kong, Wenchao Xu 0002, Xingjun Zhang, Yanqin Yang |
NPC | 5 |
| 2019 | Optimal Resource Allocation for NOMA-Enabled Cache Replacement and Content DeliveryabstractIn a content-delivery network, files’ popularity and users’ requests change fast. Conventional caching schemes, e.g., caching (re)placement once per day during the off-peak hours, may not capture the up-to-date popularity. In this case, the contents in caches have to be regularly updated to prevent information becoming outdated, and at the same time users’ requested files must be delivered. These two tasks are challenging in practical heavy-traffic and multi-user scenarios when the network resources are limited. In this paper, we apply non-orthogonal multiple access (NOMA) to facilitate concurrent caching replacement and content delivery in downlink transmission. We formulate a resource allocation problem to investigate how to efficiently push proactive files to the cache at the small base station and deliver the requested files to users. The resource-allocation problem is formulated as a mixed-integer exponential conic optimization problem. To enable a computationally-efficient optimal solution with finite convergence, we develop an iterative algorithm based on polyhedral outer approximation, where a polyhedral relaxation subproblem and a convex subproblem are constructed and iteratively solved to tighten the lower and upper bounds for the optimum, respectively. The numerical results demonstrate significant performance gains of the NOMA-enabled data transmission scheme in power and resource savings compared to the baseline scheme. Lei Lei 0001, Thang X. Vu, Lin Xiang 0001, Xingjun Zhang, Symeon Chatzinotas, Björn Ottersten 0001 |
PIMRC | 4 |
| 2019 | NoT: a high-level no-threading parallel programming method for heterogeneous systems
Shusen Wu, Xiaoshe Dong, Xingjun Zhang, Zhengdong Zhu |
J. Supercomput. | 3 |
| 2019 | Close-Proximity Detection for Hand Approaching Using Backscatter CommunicationabstractSmart environments and security systems require automatic detection of human behaviors including approaching to or departing from an object. Existing human motion detection systems usually require human beings to carry special devices, which limits their applications. In this paper, we present a system called APID to detect hand approaching behaviors by analyzing backscatter communication signals from a passive RFID tag on the object. APID does not require human beings to carry any device. The idea is based on the influence of hand movements to the vibration of backscattered tag signals. APID is compatible with commodity off-the-shelf devices and the EPCglobal Class-1 Generation-2 protocol. In APID, a commercial RFID reader continuously queries tags through emitting RF signals and tags simply respond with their IDs. A USRP monitor passively analyzes the communication signals and reports the approach and departure behaviors. We have implemented the APID system for both single-object and multi-object scenarios. Extensive evaluations demonstrate that APID can achieve high detection accuracy in both scenarios. Han Ding 0002, Chen Qian 0001, Jinsong Han, Jian Xiao 0002, Xingjun Zhang, Ge Wang 0003, Wei Xi 0003, Jizhong Zhao |
IEEE Trans. Mob. Comput. | 5 |
| 2018 | Max-Min Fairness Scheme in Wireless Powered Communication Networks with Multi-user Cooperation
Xingjun Zhang, Bocheng Yu |
WASA | 2 |
| 2018 | A Runtime Available Resource Capacity Evaluation Model Based on the Concept of Similar TasksabstractA mismatch between resource supply and demand in cloud computing leads to inefficient utilization of resources or performance degradation. Therefore, this paper establishes a runtime model to evaluate the available capacity of computing resources on the basis of similar tasks. This model takes advantage of a characteristic of cloud workload; that is, similar tasks in cloud computing have a similar execution logic. The model evaluates the available resource capacity according to task similarity, thus avoiding any impact on the resource consumption of existing benchmarks. We apply the model to propose a resource capacity evaluation method called Caipan, which considers numerous factors according to resource type. This method obtains accurate results in a timely manner at little cost. We use the results of Caipan to develop some algorithms that aim to match resource supply and demand, and improve cloud platform performance. We test the Caipan method and the Caipan-based algorithms in both dedicated and real-world cloud environments. The test results show that the Caipan method obtains the available resource capacity both accurately and in a timely manner, and effectively supports the optimization of both algorithms and platforms. Moreover, algorithms based on Caipan reduce the mismatch between resource supply and demand, and significantly improve cloud platform performance. Mosong Zhou, Xiaoshe Dong, Heng Chen 0002, Xingjun Zhang |
Comput. J. | 4 |
| 2018 | Power and discrete rate adaptation in BER constrained wireless powered communication networksabstractOptimal system throughput is a crucial design issue in wireless powered communication networks. Unlike related literature, this study considers the problem of maximising throughput (MTP) for a practical scenario, in which each user node (UN) selects its own rate from a discrete rate set and controls its transmission power according to its own channel state. The MTP with a bit‐error‐rate constraint is investigated based on two general uplink access methods. The formulation size is exponentially large with respect to the problem input. A novel greedy algorithm based on the column generation method (GA‐CGM) is proposed to solve the problem efficiently. The GA‐CGM decomposes the problem into a master problem and a subproblem. The master problem is solved by the simplex method. The subproblem is solved by a greedy algorithm due to its non‐linearity programming model. The max–min throughput problem (MMTP) is also considered, because of the ‘doubly near‐far’ phenomenon which leads to the unfair throughput among different UNs. Experimental results demonstrate that the proposed solutions are close to the optimal solutions for MTP and MMTP problems. They also show increasing the transmission power of a base station or decreasing the path‐loss exponent improves the throughput performance. Ming Lei 0003, Xingjun Zhang, Bocheng Yu, Xiaoshe Dong |
IET Commun. | 2 |
| 2018 | Stochastic geometry modeling and energy efficiency analysis of millimeter wave cellular networks
Song Cen, Xingjun Zhang, Ming Lei 0003, Scott Fowler, Xiaoshe Dong |
Wirel. Networks | 2 |
| 2017 | Performance analysis of packet layer FEC codes and interleaving in FSO channelsabstractThe combination of forward error correction (FEC) and interleaving can be used to improve free‐space optical communication systems. Recent research has optimised the codeword length and interleaving depth under the assumption of a fixed buffering size; however, how the buffering size influences the system performance remains unsolved. This study models the system performance as a function of buffering size and FEC recovery threshold, which allows system designers to determine optimum parameters in consideration of the overhead. The modelling is based on statistics of temporal features of correct data reception and burst error length through the measurement of the channel good time and outage time. The experimental results show good coherence with the theoretical values. This method can also be applied in other channels if a continuous‐time‐Markov‐chain model of the channel can be derived. Xingjun Zhang, Keith J. Blow, Scott Fowler |
IET Commun. | 2 |
| 2016 | Exploiting latency variation for access conflict reduction of NAND flash memoryabstractNAND flash memory has been widely used in storage systems by offering greater read/write performance and lower power consumption than mechanical hard drives. Recently, the tradeoff between endurance, write speed, and read speed has been exploited from many ways for I/O performance improvement, which also induce the read/write latency variation. In this paper, the latency variation is exploited in I/O scheduling for access characteristic guided read and write latency minimization. First, with the understanding of the relationship among read latency, write latency and raw bit error rates (RBER), different ways to exploit the relationship for read and write latency reduction is discussed. Then, an I/O scheduling scheme is proposed by using hotness and retention age of accessed data to determine the speed of writes or reads, giving scheduling priority to fast writes and fast reads for conflict reduction. Experiments with various traces reveal that the proposed technique achieves significant read and write performance improvements. Jinhua Cui 0001, Weiguo Wu, Xingjun Zhang, Jianhang Huang, Yinfeng Wang |
MSST | 3 |
| 2016 | Successive Interference Cancellation for Throughput Maximization in Wireless Powered Communication NetworksabstractIn wireless powered communication networks (WPCNs), each user node, e.g., wireless powered sensor, is capable of either harvesting energy from a power station or transmitting data to a sink node. In the previous works, time division multiple access (TDMA) is typically used for transmission scheduling in WPCNs, that is, only one node can transmit data in one time slot. The spectrum efficiency is therefore limited by this orthogonality in time-domain scheduling. In this paper, to maximize the throughput in WPCNs, we present a new scheduling approach for energy harvesting and data transmission. Unlike TDMA, we consider that multiple nodes can simultaneously transmit their data in the same time slot, and the signals are separated at the sink node by performing successive interference cancellation (SIC). We formulate the throughput maximization problem as a linear programming problem. For solving the large scale instances, we design an algorithmic framework based on column generation. Numerical results demonstrate that compared to the TDMA based scheduling approach, substantial throughput improvement is achieved by the proposed algorithm. Xingjun Zhang, Lei Lei 0001, Qing He 0002, Di Yuan 0001 |
VTC Fall | 2 |
| 2016 | Improving the Reliability of the Operating System Inside a VMabstractVirtualization technology can provide reusability and strong isolation between different virtual machines (VMs). However, there is no effective isolation mechanism inside a VM to solve an operating system's reliability problems, including driver faults. This paper describes Chariot, an architecture that provides effective and transparent driver isolation inside the VM, achieves fine-grained driver isolation and retains the reusability advantage of virtualization technology. First, Chariot transparently monitors an isolated driver with monitoring wrappers, and establishes an access control table (ACT) in a timely manner that records the driver write permissions. Secondly, Chariot protects the shadow page table of the VM (where the driver resides) in due time to capture its write operations. Next, the ACT examines the correctness of the write operations. Finally, if an illegal write operation is detected, Chariot recovers the faulty driver and prevents the spread of driver faults in the VM. The experimental results show that Chariot effectively isolates more than 90% of injected faults (with performance losses of |$<$|20% in most benchmarks) and effectively improves the reliability of the VM. In addition, Chariot can be easily extended to isolate new drivers and ported to other versions of OSs in the virtualization environment. Hao Zheng 0004, Xiaoshe Dong, Zhengdong Zhu, Baoke Chen, Xiuxiu Bai, Xingjun Zhang, Endong Wang |
Comput. J. | 6 |
| 2016 | TextGen: a realistic text data content generation method for modern storage system benchmarksabstractModern storage systems incorporate data compressors to improve their performance and capacity. As a result, data content can significantly influence the result of a storage system benchmark. Because real-world proprietary datasets are too large to be copied onto a test storage system, and most data cannot be shared due to privacy issues, a benchmark needs to generate data synthetically. To ensure that the result is accurate, it is necessary to generate data content based on the characterization of real-world data properties that influence the storage system performance during the execution of a benchmark. The existing approach, called SDGen, cannot guarantee that the benchmark result is accurate in storage systems that have built-in word-based compressors. The reason is that SDGen characterizes the properties that influence compression performance only at the byte level, and no properties are characterized at the word level. To address this problem, we present TextGen, a realistic text data content generation method for modern storage system benchmarks. TextGen builds the word corpus by segmenting real-world text datasets, and creates a word-frequency distribution by counting each word in the corpus. To improve data generation performance, the word-frequency distribution is fitted to a lognormal distribution by maximum likelihood estimation. The Monte Carlo approach is used to generate synthetic data. The running time of TextGen generation depends only on the expected data size, which means that the time complexity of TextGen is O ( n ). To evaluate TextGen, four real-world datasets were used to perform an experiment. The experimental results show that, compared with SDGen, the compression performance and compression ratio of the datasets generated by TextGen deviate less from real-world datasets when end-tagged dense code, a representative of word-based compressors, is evaluated. Xiaoshe Dong, Xingjun Zhang, Yinfeng Wang, Tao Ju 0002, Guofu Feng |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2015 | Research on Algorithms to Capture Drivers' Write OperationsabstractFull virtualization technology is highly reusable. Using this property, various types and versions of existing operating systems and drivers can be reused in a virtual machine to customize users’ application environments. However, these environments are threatened by drivers’ write operation faults, which are caused by bugs in reused drivers. Chariot is a reliability architecture that has been developed to solve this problem. This architecture captures a driver's write operations by maintaining the write permissions of shadow pages as read-only to examine their correctness. Nevertheless, this capture method produces many page faults in the virtual machine monitor and has an adverse impact on the performance of isolated drivers. To reduce performance losses, this paper examines two algorithms that cache recently used shadow pages using different structures to avoid frequent page faults. The experimental results show that the performance of isolated drivers can be greatly improved using these shadow page caches without significantly impacting the isolation efficiency of Chariot. Hao Zheng 0004, Xiaoshe Dong, Zhengdong Zhu, Baoke Chen, Yizhi Zhang, Xingjun Zhang |
Comput. J. | 6 |
| 2015 | A scalability prediction approach for multi-threaded applications on manycore processors
Xiuxiu Bai, Endong Wang, Xiaoshe Dong, Xingjun Zhang |
J. Supercomput. | 4 |
| 2014 | An Availability Approached Task Scheduling Algorithm in Heterogeneous Fault-Tolerant SystemabstractIn heterogeneous fault-tolerant system, especially high performance computer system, the issue of providing system with high availability assurance for real-time applications which have availability requirements has been widespread concerned. While, few research concentrates on combining real-time application availability requirement with scheduling algorithm. In this paper, an availability approached task scheduling algorithm is proposed. On the basis of heterogeneous fault-tolerant system scheduler and scheduling algorithm designation, we can improve the system availability without increasing additional hardware costs, and shorten task average response time, in addition, schedule task with high efficiency and reliability. Experiment results show that, such availability approached task scheduling algorithm has a system performance advantage over the traditional system task scheduling algorithms, it achieves the goal of balancing availability and task response time in heterogeneous fault-tolerant system, thus improve the system availability. Xiaoshe Dong, Xingjun Zhang, Yinfeng Wang |
NAS | 3 |
| 2014 | A run-time optimization approach for reducing data movements using locality-aware searching
Endong Wang, Xingjun Zhang, Tao Ju 0002, Xiaoshe Dong |
J. Supercomput. | 3 |
| 2013 | A Novel Method for Computing Encoding Delay and Bandwidth on Network Coding NodeabstractComputing the encoding delay and bandwidth on the nodes is important for the deployment and implementation of network coding. However, the existing researches merely use the methods of experimental measurement or rough estimation to investigate it. In order to accurately compute the delay and bandwidth on the nodes with network coding implemented, this paper proposes a novel method to quantify encoding delay and encoding bandwidth. Firstly, network coding is applied to design a new network. And then, the encoding procedure is systematically investigated on the source node, the intermediate node, and also the decoding procedure is systematically investigated at the destination node. And the method, which can be used to compute the node encoding delay and encoding bandwidth, is presented, Thus the method of computing total delay of network with network coding is derived. Finally, the experiments validate the correctness of this method. Yuxing Wu, Xingjun Zhang, Song Cen, Xiaoshe Dong |
NAS | 2 |
| 2013 | Chariot: A High Compatible Architecture to Improve Virtual Machine ReliabilityabstractCurrently, the virtualization technologies can integrate multiple operating systems into a high-performance server to maximize the utilization of the server's resources. This server can serves more users. However, the driver faults in virtual machine still seriously affect the reliability of the virtual machine, and even affect the reliability of the entire server. This paper presents Chariot, a high compatible architecture to improve virtual machine reliability. If the driver is loaded by the Chariot's isolation loading mechanism, its memory usage will be timely monitored by Chariot, and its access control table will be established. Through setting the corresponding shadow page table of the whole kernel space of the virtual machine, Chariot captures the write operations of the isolated driver. Combing the access control table, Chariot can determine the correctness of these writing operations. Chariot has an effective errors isolation capability, and is easy to develop. Also Chariot has an excellent compatibility and needs not to modify the drivers and the operation system in the virtual machine. Experimental results show that Chariot can effectively isolate the driver faults, and improve the reliability of operation system in the virtual machine environments. Hao Zheng 0004, Xiaoshe Dong, Endong Wang, Baoke Chen, Weifeng Gong, Xingjun Zhang |
NAS | 6 |
| 2013 | An Undirected Graph Traversal Based Grouping Prediction Method for Data De-duplicationabstractThe data capacity of the de-duplication system, which is limited by the memory, is difficult to carry out a large-scale expansion. To solve this problem, the paper proposes a hash table grouping prediction method based on undirected graph traversal. This method exploits the indexing table replacement, which is similar to the virtual memory cache replacement, to expand the storage capacity of the data de-duplication system without increasing the system memory. The hit rate of the grouping prediction and system performance are improved by grouping index entries based on undirected graph traversal. Experimental results show that, based on the cache prefetching and the hash table grouping, the memory consuming takes up 10% of the index table size while the capacity equally rise to 10 times of original. The method can make the index table cache hit rate increased to 87.6%, comparing 47% without group in dataset 1 of our experiment, make the performance acceptable. Xingjun Zhang, Guofeng Zhu, Yueguang Zhu, Xiaoshe Dong |
SNPD | 2 |
| 2013 | Improving Virtual Machine Reliability with Driver Fault IsolationabstractWith the development of virtualization technologies, the server resources are maximized by integrating multiple operating systems into a high-performance server. So the server is possible to provide services to more users simultaneously. However, the driver fault still impacts the reliability of the operating system in the virtual machine, and impacts the continuity and stability of services. This paper proposes a architecture to improve the reliability of the virtual machine environments. By monitoring the driver's memory usage, the architecture creates the authorization table. Through setting the corresponding shadow page table in the virtual machine manager of the whole kernel space of the virtual machine, the architecture captures the write operation of the virtual machine. Combing with the authorization table, the correctness of the writing operations can be determined. Our architecture needn't to modify the drivers and is easy to develop. Experimental results show that the architecture can effectively isolate the driver faults, and improve the reliability of the virtual machine environments. Hao Zheng 0004, Xiaoshe Dong, Endong Wang, Baoke Chen, Xingjun Zhang |
SNPD | 6 |
| 2013 | A dual process redundancy approach to transient fault tolerance for ccNUMA architecture
Xingjun Zhang, Endong Wang, Feilong Tang 0001, Meishun Yang, Hengyi Wei, Xiaoshe Dong |
Neurocomputing | 1 |
| 2012 | Joint Source-Network Coding Optimization for Video Streaming over Wireless Multi-Hop NetworksabstractFor the reason of unreliable and shared media, supporting video streaming over wireless multi-hop networks faces greater technical challenges. In this paper, we investigate the optimization issue and propose a joint source-network coding scheme, which segments the streaming source into generations so as to maximize the video streaming quality. The factors influenced by the size of generation include the source rate, the efficiency of coding and the decoding delay. At the source node, the faster the source rate, the more packets generated. At the intermediate nodes, the number of packets transmitted into the network is decided by the network coding strategies. The experiment results indicate that with appropriate generation size, the joint source-network coding scheme can enhance the performance of video streaming over the wireless multi-hop networks. Huali Cui, Depei Qian 0001, Xingjun Zhang, Cuiping Jing |
VTC Spring | 3 |
| 2010 | mPlogP: A Parallel Computation Model for Heterogeneous Multi-core ComputerabstractDue to the heterogeneity and the multigrain parallelism of the heterogeneous multi-core computer, communication and memory access show hierarchical characteristics ignored by other models. In this paper, a new model named mPlogP, is presented on the basis of the PlogP model, in which communication and memory access is abstracted by considering these new characteristics of the heterogeneous multi-core computer. It uses memory access to model the behavior of computation, estimates the execution time of every part of applications and guides the optimization of effective parallel programs. Finally this proposed model is validated by experiments that it can precisely evaluate the execution of parallel applications under the heterogeneous multi-core computer. Xingjun Zhang, Jinghua Feng, Xiaoshe Dong |
CCGRID | 2 |
| 2010 | Video Streaming over Wireless Mesh Networks with Multi-Gateway SupportabstractSupporting video streaming over wireless mesh networks (WMNs) is a challenging task because of the limited network resource, severe signal interference and contention among neighbor traffic. Path and server diversities are proven feasible to provide solution for video streaming over lossy networks. In this paper, we propose MG-MDC, a multi-gateway technique with multiple description coding (MDC) scheme to enhance the quality of video streaming over wireless mesh networks. By taking advantage of multiple gateways in wireless mesh networks, the quality of video streaming can be improved. The simulation results demonstrate that the proposed scheme is more effective than video transport using single gateway with single path and multi-path. Huali Cui, Depei Qian 0001, Xingjun Zhang, Yi Liu 0013 |
EUC | 3 |
| 2009 | A Hierarchical Unequal Packet Loss Protection Scheme for Robust H.264/AVC TransmissionabstractIn this paper, we are concerned with robust H.264/AVC video transmission over lossy packet networks and present a hierarchical unequal packet loss protection (HULP) scheme in a transmission system that efficiently combines erasure coding, H.264/AVC error resilience techniques and importance measures in video coding schemes. The importance of the video stream packets is distinguished by three criteria: the frame sequence number in a group of pictures (GOP), the per-frame bitrate and the H.264/AVC data partition type. Using a fixed amount of redundancy, more important packets of a video stream are protected with a more powerful erasure code than the less important packets. We demonstrate the effectiveness of this approach by system implementation and performance evaluation. In the presence of packet loss, we show that the received video quality, as measured by PSNR, is significantly improved when the HULP scheme is used. More importantly, in our experiments, HULP can achieve higher PSNR values and better user perceived quality, but requiring less redundancy, than equal loss protection (ELP) schemes. Xingjun Zhang, Xiao-Hong Peng, Dajun Wu, Timothy Porter 0002, Richard Haywood |
CCNC | 1 |
| 2008 | Robust video transmission over lossy network by exploiting H.264/AVC data partitioningabstractIn this paper, a robust H.264/AVC video transmission system over lossy packet network is investigated and implemented. The H.264/AVC standard has defined a new data partition scheme which can be used to perform unequal loss protection (ULP) in video transmission systems. However, few actual implementation and experiments which integrate data partition and loss protection algorithms have been reported. Based on the H.264/AVC data partitioning mechanism, we implement three loss protection schemes using Reed-Solomon and XOR parity codes, namely equal loss protection (ELP), ULP and partition A protection only (PAP). The performance of each scheme is evaluated under different network environments. Experimental results and performance analysis show that ELP has higher peak signal-to-noise ratio (PSNR) values than ULP and PAP, but PAP requires the least redundancy and offers the best tradeoffs between PSNR values and redundancy among all the three schemes. Xingjun Zhang, Xiao-Hong Peng, Richard Haywood, Timothy Porter 0002 |
BROADNETS | 1 |
| 2008 | FORT: A decentralized automated trust negotiation framework for gridsabstractTrust has been recognized as an important factor for grid security. This paper proposes a decentralized automated trust negotiation framework, FORT, to establish trust relationship between service providers and service requesters in grids. FORT presents many innovative features. First, FORT is decentralized, so it scales well and is well-suited for large-scale grids. Second, FORT refines its policy language with attribute constraint, so it can provide the support for effective protection of sensitive information of the two negotiation parties and flexible limitation of delegation range. Last, we employ multithreaded technology to speed up negotiation. This paper depicts the implementation of FORT and designs experiments to evaluate its performance. Experimental results show that FORT can effectively protect sensitive services at the cost of little performance of systems and is scalable. Shangyuan Guan, Xiaoshe Dong, Yiduo Mei, Xingjun Zhang |
CSCWD | 5 |
| 2006 | Supplier Categorization with K-Means Type Subspace Clustering
Xingjun Zhang, Joshua Zhexue Huang, Depei Qian 0001, Liping Jing |
APWeb | 1 |
| 2003 | Site-Role Based GreedyDual-Size Replacement Algorithm
Xingjun Zhang, Depei Qian 0001, Dajun Wu, Yi Liu 0013, Tao Liu 0033 |
WAIM | 1 |