VLDB 2026 Research / reviewers in the wild / expert
Xiang Long
dblp:64/6328
· DBLP profile ↗
40ranked-venue papers
5as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 2 since 2021Systems, architecture and hardware · 9Applied, interdisciplinary, general and emerging computing · 4Computer networks · 3Security and privacy · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VerilogLAVD: LLM-Aided Pattern Generation for Verilog CWE DetectionabstractLLMs often fail in hardware vulnerability detection due to the intrinsic semantic concurrency of HDLs (Hardware Description Language), where vulnerabilities arise from the interaction of multiple concurrent execution statements rather than a single sequential execution path.Existing LLM-based methods struggle to capture the concurrency features.To address the problem, we propose VerilogLAVD, a LLM-Aided Vulnerability Detection framework by generating executable Traversal Detection Patterns (TDPs), i.e. the rules describing how to find the evidence of vulnerabilities in Verilog HDL.We first introduce a Unified Verilog Property Graph (VeriPG) that explicitly models parallel semantics by combining AST, CFG, and DDG.Furthermore, a semantic validation mechanism is designed to constrain and filter the LLM-generated TDPs.By executing these validated TDPs on VeriPG, our method produces stable and deterministic detection results.Experiments demonstrate that VerilogLAVD improves the F1 score by 133% compared to LLM-based methods.Furthermore, the framework successfully identifies real-world hardware vulnerabilities in open-source hardware design repositories.The code and datasets of this study are available at https://github.com/ Chip-Security-Lab/VerilogLAVD Xiang Long, Yingjie Xia, Li Kuang, Yao Wan 0001 |
ACL (1) | 1 |
| 2024 | MTS-DVGAN: Anomaly detection in cyber-physical systems using a dual variational generative adversarial network
Haili Sun, Yan Huang 0026, Lansheng Han, Cai Fu, Hongle Liu, Xiang Long |
Comput. Secur. | 6 |
| 2022 | IntTower: The Next Generation of Two-Tower Model for Pre-Ranking SystemabstractScoring a large number of candidates precisely in several milliseconds is vital for industrial pre-ranking systems. Existing pre-ranking systems primarily adopt the two-tower model since the "user-item decoupling architecture" paradigm is able to balance the efficiency and effectiveness. However, the cost of high efficiency is the neglect of the potential information interaction between user and item towers, hindering the prediction accuracy critically. In this paper, we show it is possible to design a two-tower model that emphasizes both information interactions and inference efficiency. The proposed model, IntTower (short for Interaction enhanced Two-Tower), consists of Light-SE, FE-Block and CIR modules. Specifically, lightweight Light-SE module is used to identify the importance of different features and obtain refined feature representations in each tower. FE-Block module performs fine-grained and early feature interactions to capture the interactive signals between user and item towers explicitly and CIR module leverages a contrastive interaction regularization to further enhance the interactions implicitly. Experimental results on three public datasets show that IntTower outperforms the SOTA pre-ranking models significantly and even achieves comparable performance in comparison with the ranking models. Moreover, we further verify the effectiveness of IntTower on a large-scale advertisement pre-ranking system. The code of IntTower is publicly available https://gitee.com/mindspore/models/tree/master/research/recommend/IntTower. Xiangyang Li 0004, Bo Chen 0023, Huifeng Guo, Chenxu Zhu, Xiang Long, Sujian Li, Yichao Wang 0002, Wei Guo 0006, Longxia Mao, Zhenhua Dong, Ruiming Tang |
CIKM | 6 |
| 2022 | Dressing in the Wild by Watching Dance VideosabstractWhile significant progress has been made in garment transfer, one of the most applicable directions of human-centric image generation, existing works overlook the in-the-wild imagery, presenting severe garment-person mis-alignment as well as noticeable degradation in fine texture details. This paper, therefore, attends to virtual try-on in real-world scenes and brings essential improvements in authenticity and naturalness especially for loose garment (e.g., skirts, formal dresses), challenging poses (e.g., cross arms, bent legs), and cluttered backgrounds. Specifically, we find that the pixel flow excels at handling loose gar-ments whereas the vertex flow is preferred for hard poses, and by combining their advantages we propose a novel generative network called wFlow that can effectively push up garment transfer to in-the-wild context. Moreover, former approaches require paired images for training. Instead, we cut down the laboriousness by working on a newly constructed large-scale video dataset named Dance50k with self-supervised cross-frame training and an online cycle op-timization. The proposed Dance50k can boost real-world virtual dressing by covering a wide variety of garments under dancing poses. Extensive experiments demonstrate the superiority of our w Flow in generating realistic garment transfer results for in-the-wild images without resorting to expensive paired datasets.11Xiaodan Liang is the corresponding author. The project page of wFlow is https://awesome-wflow.github.io. Fuwei Zhao, Zhenyu Xie, Xijin Zhang, Daniel K. Du, Xiang Long, Xiaodan Liang, Jianchao Yang |
CVPR | 7 |
| 2022 | Low Resource Style Transfer via Domain Adaptive Meta LearningabstractText style transfer (TST) without parallel data has achieved some practical success.However, most of the existing unsupervised text style transfer methods suffer from (i) requiring massive amounts of non-parallel data to guide transferring different text styles.(ii) colossal performance degradation when fine-tuning the model in new domains.In this work, we propose DAML-ATM (Domain Adaptive Meta-Learning with Adversarial Transfer Model), which consists of two parts: DAML and ATM.DAML is a domain adaptive meta-learning approach to learn general knowledge in multiple heterogeneous source domains, capable of adapting to new unseen domains with a small amount of data.Moreover, we propose a new unsupervised TST approach Adversarial Transfer Model (ATM), composed of a sequence-to-sequence pre-trained language model and uses adversarial style training for better content preservation and style transfer.Results on multi-domain datasets demonstrate that our approach generalizes well on unseen low-resource domains, achieving state-of-theart results against ten strong baselines. Xiang Long, Sujian Li |
NAACL-HLT | 2 |
| 2022 | Neural-FacTOR: Neural Representation Learning for Website Fingerprinting Attack over TOR AnonymityabstractTOR (The Onion Router) network is a widely used open source anonymous communication tool, the abuse of TOR makes it difficult to monitor the proliferation of online crimes such as to access criminal websites. Most existing approches for TOR network de-anonymization heavily rely on manually extracted features resulting in time consuming and poor performance. To tackle the shortcomings, this paper proposes a neural representation learning approach to recognize website fingerprint based on classification algorithm. We constructed a new website fingerprinting attack model based on convolutional neural network (CNN) with dilation and causal convolution, which can improve the perception field of CNN as well as capture the sequential characteristic of input data. Experiments on three mainstream public datasets show that the proposed model is robust and effective for the website fingerprint classification and improves the accuracy by 12.21% compared with the state-of-the-art methods. Haili Sun, Yan Huang 0026, Lansheng Han, Xiang Long, Hongle Liu, Chunjie Zhou |
TrustCom | 4 |
| 2022 | Purely Attention Based Local Feature Integration for Video ClassificationabstractRecently, substantial research effort has focused on how to apply CNNs or RNNs to better capture temporal patterns in videos, so as to improve the accuracy of video classification. In this paper, we investigate the potential of a purely attention based local feature integration. Accounting for the characteristics of such features in video classification, we first propose Basic Attention Clusters (BAC), which concatenates the output of multiple attention units applied in parallel, and introduce a shifting operation to capture more diverse signals. Experiments show that BAC can achieve excellent results on multiple datasets. However, BAC treats all feature channels as an indivisible whole, which is suboptimal for achieving a finer-grained local feature integration over the channel dimension. Additionally, it treats the entire local feature sequence as an unordered set, thus ignoring the sequential relationships. To improve over BAC, we further propose the channel pyramid attention schema by splitting features into sub-features at multiple scales for coarse-to-fine sub-feature interaction modeling, and propose the temporal pyramid attention schema by dividing the feature sequences into ordered sub-sequences of multiple lengths to account for the sequential order. Our final model pyramid×pyramid attention clusters (PPAC) combines both channel pyramid attention and temporal pyramid attention to focus on the most important sub-features, while also preserving the temporal information of the video. We demonstrate the effectiveness of PPAC on seven real-world video classification datasets. Our model achieves competitive results across all of these, showing that our proposed framework can consistently outperform the existing local feature integration methods across a range of different scenarios. Xiang Long, Gerard de Melo, Dongliang He, Fu Li 0003, Zhizhen Chi, Shilei Wen, Chuang Gan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | RSPNet: Relative Speed Perception for Unsupervised Video Representation LearningabstractWe study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such as action recognition. This task, however, is extremely challenging due to 1) the highly complex spatial-temporal information in videos and 2) the lack of labeled data for training. Unlike representation learning for static images, it is difficult to construct a suitable self-supervised task to effectively model both motion and appearance features. More recently, several attempts have been made to learn video representation through video playback speed prediction. However, it is non-trivial to obtain precise speed labels for the videos. More critically, the learned models may tend to focus on motion patterns and thus may not learn appearance features well. In this paper, we observe that the relative playback speed is more consistent with motion patterns and thus provides more effective and stable supervision for representation learning. Therefore, we propose a new way to perceive the playback speed and exploit the relative speed between two video clips as labels. In this way, we are able to effectively perceive speed and learn better motion features. Moreover, to ensure the learning of appearance features, we further propose an appearance-focused task, where we enforce the model to perceive the appearance difference between two video clips. We show that jointly optimizing the two tasks consistently improves the performance on two downstream tasks (namely, action recognition and video retrieval) w.r.t the increasing pre-training epochs. Remarkably, for action recognition on the UCF101 dataset, we achieve 93.7% accuracy without the use of labeled data for pre-training, which outperforms the ImageNet supervised pre-trained model. Our code, pre-trained models, and supplementary materials can be found at https://github.com/PeihaoChen/RSPNet. Peihao Chen, Deng Huang, Dongliang He, Xiang Long, Runhao Zeng, Shilei Wen, Mingkui Tan, Chuang Gan 0001 |
AAAI | 4 |
| 2021 | VSRNet: End-to-end video segment retrieval with text query
Xiang Long, Dongliang He, Shilei Wen, Zhouhui Lian |
Pattern Recognit. | 2 |
| 2020 | Multi-Label Classification with Label Graph SuperimposingabstractImages or videos always contain multiple objects or actions. Multi-label recognition has been witnessed to achieve pretty performance attribute to the rapid development of deep learning technologies. Recently, graph convolution network (GCN) is leveraged to boost the performance of multi-label recognition. However, what is the best way for label correlation modeling and how feature learning can be improved with label system awareness are still unclear. In this paper, we propose a label graph superimposing framework to improve the conventional GCN+CNN framework developed for multi-label recognition in the following two aspects. Firstly, we model the label correlations by superimposing label graph built from statistical co-occurrence information into the graph constructed from knowledge priors of labels, and then multi-layer graph convolutions are applied on the final superimposed graph for label embedding abstraction. Secondly, we propose to leverage embedding of the whole label system for better representation learning. In detail, lateral connections between GCN and CNN are added at shallow, middle and deep layers to inject information of label system into backbone CNN for label-awareness in the feature learning process. Extensive experiments are carried out on MS-COCO and Charades datasets, showing that our proposed solution can greatly improve the recognition performance and achieves new state-of-the-art recognition performance. Ya Wang 0002, Dongliang He, Fu Li 0003, Xiang Long, Jinwen Ma, Shilei Wen |
AAAI | 4 |
| 2020 | Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationabstractMulti-label image and video classification are fundamental yet challenging tasks in computer vision. The main challenges lie in capturing spatial or temporal dependencies between labels and discovering the locations of discriminative features for each class. In order to overcome these challenges, we propose to use cross-modality attention with semantic graph embedding for multi-label classification. Based on the constructed label graph, we propose an adjacency-based similarity graph embedding method to learn semantic label embeddings, which explicitly exploit label relationships. Then our novel cross-modality attention maps are generated with the guidance of learned label embeddings. Experiments on two multi-label image classification datasets (MS-COCO and NUS-WIDE) show our method outperforms other existing state-of-the-arts. In addition, we validate our method on a large multi-label video classification dataset (YouTube-8M Segments) and the evaluation results demonstrate the generalization capability of our method. Renchun You, Zhiyao Guo, Lei Cui 0009, Xiang Long, Sid Ying-Ze Bao, Shilei Wen |
AAAI | 4 |
| 2020 | Graph-PCNN: Two Stage Human Pose Estimation with Graph Pose Refinement
Jian Wang 0066, Xiang Long, Errui Ding, Shilei Wen |
ECCV (11) | 2 |
| 2020 | Deep Concept-wise Temporal Convolutional Networks for Action LocalizationabstractExisting action localization approaches adopt shallow temporal convolutional networks (i.e., TCN) on 1D feature map extracted from video frames. In this paper, we empirically find that stacking more conventional temporal convolution layers actually deteriorates action classification performance, possibly ascribing to that all channels of 1D feature map, which generally are highly abstract and can be regarded as latent concepts, are excessively recombined in temporal convolution. To address this issue, we introduce a novel concept-wise temporal convolutional network (C-TCN) as an alternative to TCN for training deeper action localization networks. To address this issue, we introduce a novel concept-wise temporal convolution (CTC) layer as an alternative to conventional temporal convolution layer for training deeper action localization networks. Instead of recombining latent concepts, CTC layer deploys a number of temporal filters to each concept separately with shared filter parameters across concepts. Thus can capture common temporal patterns of different concepts and significantly enrich representation ability. Via stacking CTC layers, we proposed a deep concept-wise temporal convolutional network (C-TCN), which boosts the state-of-the-art action localization performance on THUMOS'14 from 42.8 to 52.1 in terms of mAP(%), achieving a relative improvement of 21.7%. Favorable result is also obtained on ActivityNet. Xin Li 0106, Xiao Liu 0022, Wangmeng Zuo, Chao Li 0034, Xiang Long, Dongliang He, Fu Li 0003, Shilei Wen, Chuang Gan 0001 |
ACM Multimedia | 6 |
| 2020 | Real-time scheduling of parallel tasks with tight deadlines
Xu Jiang 0004, Nan Guan, Xiang Long, Yue Tang 0001, Qingqiang He |
J. Syst. Archit. | 3 |
| 2020 | Decomposition-Based Real-Time Scheduling of Parallel Tasks on Multicores PlatformsabstractMulticore processors have become mainstream computation platforms not only for general and high-performance computers but also for real-time embedded systems. To fully utilize the computation power of multicores, software must be parallelized. Recently, there has been a rapidly increasing interest in real-time scheduling of parallel real-time tasks, but the field is still much less mature than traditional real-time scheduling of sequential tasks. In this article, we study the real-time scheduling and techniques for parallel real-time tasks based on decomposition, where a task graph is transferred to a set of independent sporadic tasks. In particular, we propose new decomposition strategies that better explore the structure feature of each task to improve schedulability. We develop schedulability tests for the global earliest deadline first (EDF) scheduling algorithm based on decomposition and three types of its variants, with their own pros and cons in different aspects. We conduct experiments to evaluate the real-time performance of our proposed scheduling algorithms against the state-of-the-art scheduling and analysis methods of different types. Xu Jiang 0004, Nan Guan, Xiang Long, Han Wan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2018 | Multimodal Keyless Attention Fusion for Video ClassificationabstractThe problem of video classification is inherently sequential and multimodal, and deep neural models hence need to capture and aggregate the most pertinent signals for a given input video. We propose Keyless Attention as an elegant and efficient means to more effectively account for the sequential nature of the data. Moreover, comparing a variety of multimodal fusion methods, we find that Multimodal Keyless Attention Fusion is the most successful at discerning interactions between modalities. We experiment on four highly heterogeneous datasets, UCF101, ActivityNet, Kinetics, and YouTube-8M to validate our conclusion, and show that our approach achieves highly competitive results. Especially on large-scale data, our method has great advantages in efficiency and performance. Most remarkably, our best single model can achieve 77.0% in terms of the top-1 accuracy and 93.2% in terms of the top-5 accuracy on the Kinetics validation set, and achieve 82.2% in terms of GAP@20 on the official YouTube-8M test set. Xiang Long, Chuang Gan 0001, Gerard de Melo, Xiao Liu 0022, Yandong Li, Fu Li 0003, Shilei Wen |
AAAI | 1 |
| 2018 | Attention Clusters: Purely Attention Based Local Feature Integration for Video ClassificationabstractRecently, substantial research effort has focused on how to apply CNNs or RNNs to better capture temporal patterns in videos, so as to improve the accuracy of video classification. In this paper, however, we show that temporal information, especially longer-term patterns, may not be necessary to achieve competitive results on common trimmed video classification datasets. We investigate the potential of a purely attention based local feature integration. Accounting for the characteristics of such features in video classification, we propose a local feature integration framework based on attention clusters, and introduce a shifting operation to capture more diverse signals. We carefully analyze and compare the effect of different attention mechanisms, cluster sizes, and the use of the shifting operation, and also investigate the combination of attention clusters for multimodal integration. We demonstrate the effectiveness of our framework on three real-world video classification datasets. Our model achieves competitive results across all of these. In particular, on the large-scale Kinetics dataset, our framework obtains an excellent single model accuracy of 79.4% in terms of the top-1 and 94.0% in terms of the top-5 accuracy on the validation set. Xiang Long, Chuang Gan 0001, Gerard de Melo, Jiajun Wu 0001, Xiao Liu 0022, Shilei Wen |
CVPR | 1 |
| 2018 | Deep detection network for real-life traffic sign in vehicular networks
Xiang Long, Arun Kumar Sangaiah, Zhigao Zheng 0001, Chao Tong 0001 |
Comput. Networks | 2 |
| 2018 | A novel rating prediction method based on user relationship and natural noise
Chao Tong 0001, Yu Lian, Jianwei Niu 0002, Xiang Long |
Multim. Tools Appl. | 4 |
| 2018 | Video Captioning with Multi-Faceted AttentionabstractVideo captioning has attracted an increasing amount of interest, due in part to its potential for improved accessibility and information retrieval. While existing methods rely on different kinds of visual features and model architectures, they do not make full use of pertinent semantic cues. We present a unified and extensible framework to jointly leverage multiple sorts of visual features and semantic attributes. Our novel architecture builds on LSTMs with two multi-faceted attention layers. These first learn to automatically select the most salient visual features or semantic attributes, and then yield overall representations for the input and output of the sentence generation component via custom feature scaling operations. Experimental results on the challenging MSVD and MSR-VTT datasets show that our framework outperforms previous work and performs robustly even in the presence of added noise to the features and attributes. Xiang Long, Chuang Gan 0001, Gerard de Melo |
Trans. Assoc. Comput. Linguistics | 1 |
| 2017 | The Case for a Flexible Low-Level Backend for Software Data PlanesabstractRecent efforts to simplify network data plane programming focus on providing simple, high-level domain-specific languages (DSLs). In the case of software switches, data plane programs are written in these DSLs and then compiled to run on CPU-based architecture. However, the simplicity of these DSLs, along with the lack of low-level interfaces exposed by the software switch, restrict compilers from generating optimal data plane programs for CPU-based architecture. Sean Choi, Xiang Long, Muhammad Shahbaz 0001, Skip Booth, Andy Keep, John Marshall, Changhoon Kim |
APNet | 2 |
| 2017 | Introducing parallel computing concepts in computer system related coursesabstractAll semiconductor market domains are converging to concurrent platforms. This trend has certainly led real challenge to develop applications software that effectively uses these concurrent processors to achieve efficiency and performance goals. This paper argues that the Computer System related courses are natural places to introduce the parallelism, and the earlier to parallel computing concepts will have a wide reach. This paper showed how digital logic classes can motivate topics from parallel computing through common logic structures. We provided an alternative view of the digital logic topics, including: binary representation of integers using decision tree with recursive thinking; use a carry look-ahead adder to show how sequential operations can be parallelized. Another part to introduce parallel concepts focused on write high performance code with specific emphasis on graphic processing unit (GPU). In our teaching experience, parallel pattern teaching has been confirmed to be a useful pedagogical method for teaching parallel concepts. Finally, we report course experience in teaching parallel computing injected course, which resulted in positive student feedback. Han Wan, Xiaopeng Gao, Xiang Long, Bo Jiang 0001 |
FIE | 3 |
| 2017 | Semi-Federated Scheduling of Parallel Real-Time Tasks on MultiprocessorsabstractFederated scheduling is a promising approach to schedule parallel real-time tasks on multi-cores, where each heavy task exclusively executes on a number of dedicated processors, while light tasks are treated as sequential sporadic tasks and share the remaining processors. However, federated scheduling suffers resource waste since a heavy task with processing capacity requirement x+epsilon (where x is an integer and 0 epsilon 1) needs x+1 dedicated processors. In the extreme case, almost half of the processing capacity is wasted. In this paper we propose the semi-federate scheduling approach, which only grants x dedicated processors to a heavy task with processing capacity requirement x+epsilon, and schedules the remaining epsilon part together with light tasks on shared processors. Experiments with randomly generated task sets show the semi-federated scheduling approach significantly outperforms not only federated scheduling, but also all existing approaches for scheduling parallel real-time tasks on multi-cores. Xu Jiang 0004, Nan Guan, Xiang Long, Wang Yi 0001 |
RTSS | 3 |
| 2017 | Energy-aware scheduling on heterogeneous multi-core systems with guaranteed probability
Ying Li 0122, Jianwei Niu 0002, Mohammed Atiquzzaman, Xiang Long |
J. Parallel Distributed Comput. | 4 |
| 2016 | Real-Time Scheduling for Periodic Tasks in Homogeneous Multi-core System with Minimum Execution Time
Ying Li 0122, Jianwei Niu 0002, Mohammed Atiquzzaman, Xiang Long |
CollaborateCom | 5 |
| 2016 | An Optimized RM Algorithm by Task Affinity on Multi-Core ProcessorabstractScheduling of real-time tasks on a multi-core processor is challenging due to the execution time being a nondeterministic value. Through studying the relationship between task affinity and execution time, we propose an accelerated multi-core real-time scheduling algorithm (RM-λ) for periodic and dependent real-time tasks on a homogeneous multi-core processor based on acceleration between tasks to obtain a real-time scheduling scheme with less resource utilization. We adopt an acceleration factor matrix to represent the degree of affinity and develop a real-time scheduling model to find the best accelerated pair. The heterogeneous multi-core architectures can execute tasks by sharing their dependent data on L1 Cache. The results demonstrate our approach can loosens the schedulability constraints of RM (maximum improvement of 25%) so that an un-schedulable real-time tasks set on a single-core processor might be schedulable, and for those still hard to be scheduled, could be made schedulable on a multi-core processor. Ying Li 0122, Jianwei Niu 0002, Mohammed Atiquzzaman, Xiang Long |
ICPADS | 5 |
| 2016 | On the Decomposition-Based Global EDF Scheduling of Parallel Real-Time TasksabstractReal-time systems are shifting from single-core to multi-core processors, on which software must be parallelized to fully utilize the additional computation power. Recently different types of scheduling algorithms and analysis techniques have been proposed for parallel real-time tasks modeled as directed acyclic graphs (DAG). However, this field is still much less mature than traditional real-time scheduling of sequential tasks. In this paper, we study the decomposition-based scheduling for parallel real-time tasks, where a task graph is transferred to a set of independent sporadic tasks. In particular, we proposed a new decomposition strategy that better explores the feature of each task, represented by its structure characteristic value, to improve schedulability. The structure characteristic values do not only provide a clear guidance in task decomposition, but also can be directly used for schedulability tests, as well as to quantify the suboptimality of our scheduling algorithm in terms of capacity augmentation bounds. We conduct comprehensive experiments to evaluate the real-time performance of our proposed scheduling algorithm, against the state-of-the-art scheduling and analysis methods of different types. Experiment results show that our method consistently outperforms all of the previous methods under different parameter settings. Xu Jiang 0004, Xiang Long, Nan Guan, Han Wan |
RTSS | 2 |
| 2016 | FlexPoll: adaptive event polling for network-intensive applications
Xingbo Wu, Xiang Long, Lei Wang 0126 |
Frontiers Comput. Sci. | 2 |
| 2013 | An Efficient Grouped Virtual Mapreduce ClusterabstractVirtualization technology and MapReduce program model are sharp swords for the big data and cloud computing era. The combination of them exhibits powerful ability of easy-management, fast-deployment, feasible-scalability and high-efficiency. However, the downside is that the performance is limited by the I/O bottleneck of Virtual Machine(VM). A huge number of data should be handled in MapReduce cluster which is deployed in VMs. Luckily, data locality, a very crucial issue affecting performance in a shared clusters environment, is used to ease this conflict and improve the execution time of applications. We present a framework of Grouped Virtual MapReduce Cluster(GVMC) which takes fully advantage of VM data locality to exhibit high performance of Virtual MapReduce Cluster(VMC). The introduction of local-master nodes in GVMC not only offloads the pressure of the master node, but also lowers the communication cost. We compare the organization of three different VMC, describe the architecture of our cluster framework and do the performance analysis. Our experiments demonstrate that the framework of GVMC achieves higher locality and reduces the execution time in both CPU-intensive applications and I/O-intensive applications. Compared to Original Virtual MapReduce Cluster(OVMC), the performance of GVMC improvement is up to 16.5% and 36.2% for CPU-intensive applications and I/O-intensive applications respectively. Xiang Long, Bo Jiang 0001 |
AINA | 2 |
| 2013 | Research on SPH Parallel Acceleration Strategies for Multi-GPU Platform
Xukun Shen, Xiang Long |
APPT | 3 |
| 2013 | Towards RTOS: A Preemptive Kernel Basing on Barrelfish
Xiang Long, Xukun Shen, Lei Wang 0126, Shuaitao Feng, Siyao Zheng |
APPT | 2 |
| 2013 | Optimizing Event Polling for Network-Intensive Applications: A Case Study on RedisabstractIn today's data centers supporting Internet-scale computing and I/O services, increasingly more network-intensive applications are deployed on the network as a service. To this end, it is critical for the applications to quickly retrieve requests from the network and send their responses to the network. To facilitate this network function, operating system usually provides an event notification mechanism so that the applications (or the library) know if the network is ready to supply data for them to read or to receive data for them to write. As a widely used and representative notification mechanism, epoll in Linux provides a scalable and high-performance implementation by allowing applications to specifically indicate which connections and what events on them need to be watched. As epoll has been used in some major systems, including KV systems, such as Redis and Memcached, and web server systems such as NGINX, we have identified a substantial performance issue in its use. For the sake of efficiency, applications usually use epoll's system calls to inform the kernel exactly of what events they are interested in and always keep the information up-to-date. However, in a system with demanding network traffic, such a rigid maintenance of the information is not necessary and the excess number of system calls for this purpose can substantially degrade the system's performance. In this paper, we use Redis as an example to explore the issue. We propose a strategy of informing the kernel of the interest events in a manner adaptive to the current network load, so that the epoll system calls can be reduced and the events can be efficiently delivered. We have implemented the strategy, named as FlexPoll, in Redis without modifying any kernel code. Our evaluation on Redis shows that the query throughput can be improved by up to 46.9% on micro benchmarks, and even up to 67.8% on workloads emulating real-world access patterns. FlexPoll can be extended to other applications and event libraries built on the epoll mechanism in a straightforward manner. Xingbo Wu, Xiang Long, Lei Wang 0126 |
ICPADS | 2 |
| 2012 | Using Basic Block Based Instruction Prefetching to Optimize WCET Analysis for Real-Time ApplicationsabstractCache is an important component existing in modern computer system to bridge the performance gap between the fast CPU and the slow memory system. A variety of cache optimization technologies and mechanisms are proposed to improve the cache performance, such as instruction cache prefetching. Most instruction prefetching mechanisms existing are proposed to improve the average-case cache performance. However, real-time systems care more about the worst-case performance, and the worst-case execution time (WCET) analysis of real-time applications is critical for schedulability analysis of real-time systems. Due to its unpredictable behaviour, cache disastrously complicates the WCET analysis of real-time applications. In this paper, we proposed a basic block based instruction prefetching (BBIP) mechanism to improve both the average-case cache performance and the tightness of the WCET analysis of real-time applications. Measurements on typical real-time benchmarks show that BBIP can not only eliminate most of the instruction access misses, but also result in lower WCET estimations. To discuss the effectiveness of BBIP, we measured the WCET of the benchmarks for three processor configurations with and without BBIP: 1) processor with in-order pipeline and perfect branch prediction, 2) processor with out-of-order pipeline and perfect branch prediction, and 3) processor with out-of-order pipeline and 2-level branch prediction. The results show that BBIP can provide notable improvements in the tightness of WCET estimation, with the WCET values being 30.4% to 97.7% of the original ones. Our simulation results also reveal that 70% to 80% instruction access misses are eliminated with BBIP. Fan Ni, Xiang Long, Han Wan, Xiaopeng Gao |
PDCAT | 2 |
| 2012 | Robust wrinkle-aware non-rigid registration for triangle meshes of hand with rich and dynamic details
Ling Zhao 0006, Xukun Shen, Xiang Long |
Comput. Graph. | 3 |
| 2012 | Visible neighborhood graph of point clouds
Xiang Long, Lu Feng 0003, Pei Luo, Zhuangzhi Wu |
Graph. Model. | 2 |
| 2012 | Complex networks properties analysis for mobile ad hoc networksabstractRecently, research on complex network theory and applications draws a lot of attention in both academy and industry. In mobile ad hoc networks (MANETs) area of research, a critical issue is to design the most effective topology for given problems. It is natural and significant to consider complex networks topology when optimising the MANET topology. Current works usually transform MANET or sensor network topologies into either small-world or scale-free. However, some fundamental problems remain unsolved. Specifically, what are the average shortest path length, degree distribution and clustering characteristics of MANETs? Do MANETs have small-world effect and scale-free property? In this work, the authors introduce complex networks theory into the context of MANET topology and study complex network properties of the MANETs to answer the above questions. The authors have theoretically analysed the degree distribution and clustering coefficient of MANETs and proposed approach to computing them. The degree distribution and clustering coefficient of MANETs are theoretically deduced from node space probability distribution on different mobility models (including but not limited to random waypoint model). Simulation results on average shortest path length, clustering coefficient and degree distribution show that in most cases MANETs do not have the small-world effect and scale-free property. Chao Tong 0001, Jianwei Niu 0002, Guangzhi Qu, Xiang Long, Xiaopeng Gao |
IET Commun. | 4 |
| 2010 | I/O scheduling model of virtual machine based on multi-core dynamic partitioningabstractIn a virtual machine system, the scheduler within the virtual machine monitor (VMM) plays a key role in determining the overall fairness and performance characteristics of the whole system. However, traditional VMM schedulers focus on sharing the processor resources fairly among guest domains while leaving the scheduling of I/O missions as a secondary concern. This would cause serious degradation of I/O performance and make virtualization less desirable for I/O-intensive applications. In order to eliminate the I/O performance bottleneck caused by scheduling delay, this paper proposes a virtual machine I/O scheduling model based on multi-core dynamic partitioning, and implements a prototype based on Xen virtual machine. In this model, I/O operations of guest domains are monitored and the runtime information is analyzed. When the preset conditions are satisfied, the processor cores of the system are divided into three subsets to undertake different missions respectively. Each subset employs specific scheduling strategy to meet the requirement of different tasks. Experiment results demonstrate that our scheduling model can efficiently improve the I/O performance of virtual machine system: in comparison with the case using default Xen credit scheduler, the network and disk bandwidth increase by 35% and 12% respectively, and the average latency of ping operations drops by 37%. At the same time, our method only causes slight negative effect on the performance of compute-intensive applications, and the scheduling fairness can also be guaranteed. Yanyan Hu, Xiang Long |
HPDC | 2 |
| 2009 | GCSim: A GPU-Based Trace-Driven Simulator for Multi-level Cache
Han Wan, Xiaopeng Gao, Xiang Long |
APPT | 3 |
| 2009 | Toward Trustworthy Semantic Web Service Discovery and Selection
Jing Li 0075, Dianfu Ma, Xiang Long |
ATC | 4 |
| 2007 | SQS: A Secure and QoS Guaranteed Solution for Mobile ServiceabstractMobile IP protocol, a standard proposed by the Internet Engineering Task Force, was designed to support IP mobility. On the basis of analysis of Mobile IP protocol, it can be concluded that some problems (e.g., security flaws, limited agents deployments, triangular routing and just a little supports by operating systems, etc.) remain to be solved. In this paper, we present a novel mobility supporting scheme based on embedded technology which is called Secure and QoS Guaranteed Solution for Mobile Service (SQS). It implements mobility management for handling seamless handoffs with Embedded Mobile Agent (EmMA) and Mobility Management Server (MMS), which allows nodes to continue to receive datagrams when the nodes change their points of attachment to the Internet. It reduces the dependence on network infrastructure and provides nodes with transparent mobile service. Experimental results show that SQS effectively solves the problems mentioned above and improves Mobile IP protocol. Xiaopeng Gao, Xiang Long |
COMPSAC (1) | 4 |