Yuxing Peng 0001

dblp:42/2857-1 · also Yu-xing Peng 0001 · DBLP profile ↗
← Back
76ranked-venue papers
0as first author
15since 2021 · last 2023
0000-0002-8223-5588ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 9 since 2021Systems, architecture and hardware · 24 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 since 2021Computer networks · 8Applied, interdisciplinary, general and emerging computing · 7Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2023 Deep Graph-based Spatial Consistency for Robust Non-rigid Point Cloud Registration
abstract
We study the problem of outlier correspondence pruning for non-rigid point cloud registration. In rigid registration, spatial consistency has been a commonly used criterion to discriminate outliers from inliers. It measures the compatibility of two correspondences by the discrepancy between the respective distances in two point clouds. However, spatial consistency no longer holds in non-rigid cases and outlier rejection for non-rigid registration has not been well studied. In this work, we propose Graph-based Spatial Consistency Network (GraphSCNet) to filter outliers for non-rigid registration. Our method is based on the fact that non-rigid deformations are usually locally rigid, or local shape preserving. We first design a local spatial consistency measure over the deformation graph of the point cloud, which evaluates the spatial compatibility only between the correspondences in the vicinity of a graph node. An attention-based non-rigid correspondence embedding module is then devised to learn a robust representation of non-rigid correspondences from local spatial consistency. Despite its simplicity, GraphSCNet effectively improves the quality of the putative correspondences and attains state-of-the-art performance on three challenging benchmarks. Our code and models are available at https://github.com/qinzheng93/GraphSCNet.
Zheng Qin 0002, Hao Yu 0010, Yuxing Peng 0001, Kai Xu 0004
CVPR4
2023 CasViGE: Learning robust point cloud registration with cascaded visual-geometric encoding
Zheng Qin 0002, Yuxing Peng 0001, Kai Xu 0004
Comput. Aided Geom. Des.3
2023 GeoTransformer: Fast and Robust Point Cloud Registration With Geometric Transformer
abstract
We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods have shown great potential through bypassing the detection of repeatable keypoints which is difficult to do especially in low-overlap scenarios. They seek correspondences over downsampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer, or GeoTransformer for short, to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it invariant to rigid transformation and robust in low-overlap cases. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Extensive experiments on rich benchmarks encompassing indoor, outdoor, synthetic, multiway and non-rigid demonstrate the efficacy of GeoTransformer. Notably, our method improves the inlier ratio by 18 ∼ 31 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark.
Zheng Qin 0002, Hao Yu 0010, Yulan Guo, Yuxing Peng 0001, Slobodan Ilic, Dewen Hu, Kai Xu 0004
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Geometric Transformer for Fast and Robust Point Cloud Registration
abstract
We study the problem of extracting accurate correspondences for point cloud registration. Recent keypoint-free methods bypass the detection of repeatable keypoints which is difficult in low-overlap scenarios, showing great potential in registration. They seek correspondences over down-sampled superpoints, which are then propagated to dense points. Superpoints are matched based on whether their neighboring patches overlap. Such sparse and loose matching requires contextual features capturing the geometric structure of the point clouds. We propose Geometric Transformer to learn geometric feature for robust superpoint matching. It encodes pair-wise distances and triplet-wise angles, making it robust in low-overlap cases and invariant to rigid transformation. The simplistic design attains surprisingly high matching accuracy such that no RANSAC is required in the estimation of alignment transformation, leading to 100 times acceleration. Our method improves the inlier ratio by 17∼30 percentage points and the registration recall by over 7 points on the challenging 3DLoMatch benchmark. Our code and models are available at https://github.com/qinzheng93/GeoTransformer.
Zheng Qin 0002, Hao Yu 0010, Yulan Guo, Yuxing Peng 0001, Kai Xu 0004
CVPR5
2022 An Emotion Evolution Network for Emotion Recognition in Conversation
abstract
Emotion recognition in conversation (ERC) aims to detect the emotion in a conversation, which has drawn increasing interests due to its widely applications. Current methodologies mainly endeavor to capture a good representation of conversation context. However, we argue that the conversation context are not always consistent with the emotion evolution. This incongruity can greatly restrict the recognition performance. To address aforementioned challenges, in this paper, we propose an emotion evolution network for emotion recognition in conversation (E2Net). Specifically, a speaker-aware modeling methodology is firstly constructed to fuse the utterance from conversations. We employ the gated recurrent unit (GRU) encodes the utterance sequentially. For encoding the interaction between speakers, a listener state is introduced to aid in analyzing conversation context. Then, a Transformer-based method is proposed to capture the emotion evolution accompanying with the emotion transformation matrix. To demonstrate the superior performance of our proposed method, extensive experiments are conducted on four REC datasets and the experimental results suggest that our method is effective and outperforms the current state-of-the-art methods on multiple datasets.
Shimin Tang, Kele Xu, Zhen Huang 0006, Minpeng Xu, Yuxing Peng 0001
ICTAI6
2022 Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast
abstract
We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice and face. These methods neglect the semantic content of different videos, introducing false-negative pairs as training noise. Furthermore, the positive pairs are constructed based on the natural correlation between audio clips and visual frames. However, this correlation might be weak or inaccurate in a large amount of real-world data, which leads to deviating positives into the contrastive paradigm. To address these issues, we propose the cross-modal prototype contrastive learning (CMPC), which takes advantage of contrastive methods and resists adverse effects of false negatives and deviate positives. On one hand, CMPC could learn the intra-class invariance by constructing semantic-wise positives via unsupervised clustering in different modalities. On the other hand, by comparing the similarities of cross-modal instances from that of cross-modal prototypes, we dynamically recalibrate the unlearnable instances' contribution to overall loss. Experiments show that the proposed approach outperforms state-of-the-art unsupervised methods on various voice-face association evaluation protocols. Additionally, in the low-shot supervision setting, our method also has a significant improvement compared to previous instance-wise contrastive learning.
Boqing Zhu, Kele Xu, Zheng Qin 0002, Tao Sun 0005, Huaimin Wang 0001, Yuxing Peng 0001
IJCAI7
2022 Multi-representation knowledge distillation for audio classification
Liang Gao 0004, Kele Xu, Huaimin Wang 0001, Yuxing Peng 0001
Multim. Tools Appl.4
2022 ParaX : Bandwidth-Efficient Instance Assignment for DL on Multi-NUMA Many-Core CPUs
abstract
Commercial clouds now heavily use CPUs in DL (deep learning) because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN). This paper presents a novel approach called ParaX, which boosts the performance of DL on multi-NUMA (non-uniform memory access) many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the compute-intensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training/inference throughput for all tested models (for image recognition and natural language processing) by$1.73\times \sim 2.93{\times}$.
Yiming Zhang 0003, Lujia Yin, Dongsheng Li 0001, Yuxing Peng 0001, Kai Lu 0001
IEEE Trans. Computers4
2021 T-Mask: An Active and Accurate Dialogue State Tracking with Token Mask Prediction
abstract
Recent dialogue state tracking (DST) usually treats utterance, system action and ontology equally to estimate the slot types and values. In this way, the expression of slot in utterance is restricted. As the main way to directly express user semantics, utterance should receive further attention and its proportion in semantic expression should be dynamic according to the content. It’s common to recognize the different importance of information in all the DST models. However, most of them pay little attention to the position of slot in utterance. In fact, position and semantics are related due to human grammatical habits and expression habits. Therefore, we propose T-Mask1, a model to actively and accurately learn the token mask position of slot, and we further utilize the learned position information to influence the semantic expression of utterance. We verify the effectiveness of our model on DSTC2 and WoZ2.0. On WoZ2.0, we achieve 90.84 joint goal accuracy and 97.6 turn request accuracy, which is better than most existing models.
Shezheng Song, Dongsong Zhang, Zhen Huang 0006, Yuxing Peng 0001
ICTAI5
2021 SEED: A Cross-Layer Semantic Enhanced SLU Model With Role Context Differentiated Fusion
abstract
The mainstream SLU models, such as SDEN, take the joint training way of slot filling and intent detection because of their correlation and add contextual information to improve the model performance by the contextual vector. Although these models have proved effective, it also brings challenges for slot filling. The slot filling decoder is fed with the deep-layer semantic encoding without alignment information, which will affect the performance of slot filling. The alignment information of the history utterances is attenuated in the context vector because of the repeated fusion process, which is not conducive to the performance improvement of slot filling. In order to solve the above problems, we proposed a novel cross layer semantic enhanced SLU model with role context differentiated fusion, which contains two important improvements: 1) the word embedding information of the current utterance is introduced into the slot filling decoder to strengthen the alignment information based on the mutual attention mechanism; 2) the utterances of different roles are fused in different ways to reserve the alignment information of history utterances in the contextual vector. A large number of experiments were carried out on the standard dataset from SDEN, named KVRET*, and the results verify the effectiveness of our new model. Our model can increase the F1 score of slot filling by more than 7.5% than the existing models.
Dongsong Zhang, Shezheng Song, Zhen Huang 0006, Yuxing Peng 0001
ICTAI5
2021 Semi-supervised medical image classification based on CamMix
abstract
Collecting a large amount of labeled data is crutial for training deep neural network, which is a limitation for medical image classification because it necessarily involves expert knowledge. To mitigate this problem of insufficient labeled medical data, in this work, we propose a novel semi-supervised framework for medical image classification. For unlabeled data, we apply the consistency-based strategy to produce high-quality pseudo label, which encourages model to output the same predictions under different perturbations. In addition, we present a novel mixed sample data augmentation CamMix to effectively exploit the relation between samples, mixing pairs of input data and labels according to the class activation map mask. We have evaluated our proposed method on two public medical image datasets, interstitial lung disease dataset and ISIC 2018 skin lesion analysis dataset. The results demonstrate superior performance of our method over other existing methods on the two datasets. Meanwhile, our proposed CamMix performs better than the current mixed sample data augmentation methods.
Lingchao Guo, Dongsong Zhang, Kele Xu, Zhen Huang 0006, Yuxing Peng 0001
IJCNN7
2021 USET : A network based on Utterance hidden State transfEr for Task-oriented dialogue
abstract
Multi-turn dialogue is challenging because semantic information is not only contained in the current utterance, but also in the dialogue context. In fact, understanding multiturn dialogue is a dynamic process. With the increase of dialogue turn, users' understanding is also changing. In this case, we propose a network based on utterance hidden state transfer for task-oriented dialogue (USET). In our model, we first extract the hidden state of previous utterance as the previous comprehension. Then, this comprehension is passed to the next turn. We take the comprehension as a prior knowledge to understand the semantic information in the dialogue context. Finally, we put the previous comprehension and current utterance together to understand current utterance. In order to realize the transfer of comprehension in dialogue, we propose a continuous sample training method: CST, which takes a multi-turn dialogue as a whole to understand. All the sentences in a dialogue are put into a batch for training. Our method makes use of previous comprehension and achieves the information exchange among dialogue. Experimental results on Stanford Multi-Domain dataset demonstrate that our model is superior to existing models. Code is available at https://github.com/season1blue/USET
Shezheng Song, Dongsong Zhang, Yuxing Peng 0001, Yuan Yuan 0034
IJCNN4
2021 MapperX: Adaptive Metadata Maintenance for Fast Crash Recovery of DM-Cache Based Hybrid Storage Devices
Lujia Yin, Yiming Zhang 0003, Yuxing Peng 0001
USENIX ATC4
2021 Elastic scheduler: Heterogeneous and dynamic deep Learning in the cloud
abstract
Abstract GPUs and CPUs have been widely used for model training of deep learning (DL) in the cloud, where both DL workloads and resource usage might heavily change over time. Traditional training methods require beforehand specification on the type (either GPUs or CPUs) and amount of computing devices, and thus cannot elastically schedule the dynamic DL workloads onto available GPUs/CPUs. In this paper, we propose Elastic Scheduler (ES), a novel approach that efficiently supports both heterogeneous training (with different device types) and dynamic training (with varying device numbers). ES (i) accumulates local gradients and simulates multiple virtual workers on one GPU to alleviate the performance gap between GPUs and CPUs for achieving similar accuracy in heterogeneous GPU‐CPU‐hybrid training as in homogeneous training and (ii) uses local gradients stabilizes batch sizes for high accuracy without long compensation. Experiments show that ES achieves significantly higher performance than existing methods for heterogeneous and dynamic training as well as inference.
Lujia Yin, Yiming Zhang 0003, Yuxing Peng 0001, Dongsheng Li 0001
Concurr. Comput. Pract. Exp.3
2021 ParaX: Boosting Deep Learning for Big Data Analytics on Many-Core CPUs
abstract
Despite the fact that GPUs and accelerators are more efficient in deep learning (DL), commercial clouds like Facebook and Amazon now heavily use CPUs in DL computation because there are large numbers of CPUs which would otherwise sit idle during off-peak periods. Following the trend, CPU vendors have not only released high-performance many-core CPUs but also developed efficient math kernel libraries. However, current DL platforms cannot scale well to a large number of CPU cores, making many-core CPUs inefficient in DL computation. We analyze the memory access patterns of various layers and identify the root cause of the low scalability, i.e., the per-layer barriers that are implicitly imposed by current platforms which assign one single instance (i.e., one batch of input data) to a CPU. The barriers cause severe memory bandwidth contention and CPU starvation in the access-intensive layers (like activation and BN). This paper presents a novel approach called ParaX, which boosts the performance of DL on many-core CPUs by effectively alleviating bandwidth contention and CPU starvation. Our key idea is to assign one instance to each CPU core instead of to the entire CPU, so as to remove the per-layer barriers on the executions of the many cores. ParaX designs an ultralight scheduling policy which sufficiently overlaps the access-intensive layers with the compute-intensive ones to avoid contention, and proposes a NUMA-aware gradient server mechanism for training which leverages shared memory to substantially reduce the overhead of per-iteration parameter synchronization. We have implemented ParaX on MXNet. Extensive evaluation on a two-NUMA Intel 8280 CPU shows that ParaX significantly improves the training/inference throughput for all tested models (for image recognition and natural language processing) by 1.73X ~ 2.93X.
Lujia Yin, Yiming Zhang 0003, Zhaoning Zhang 0001, Yuxing Peng 0001
Proc. VLDB Endow.4
2020 RAD: Reinforced Attention Decoder Model On Question Generation
abstract
Question Generation (QG) aims to construct questions from given text automatically. Recently, QG has received widely concerned. The mainstream method is still based on the fixed sequence generation model of Seq2Seq model, and few people consider the influence of generation order in the result. In this paper, we present a novel Reinforced Attention Decoder Neural Network for the QA-SRL task. First, our model draws on the idea of the reinforcement learning algorithm with baseline, which is using the accuracy of each slot and sentence as the reward, updating the Policy Network to predict the optimal generation order. Second, we apply the Attention mechanism on the baseline to get more relevant information about the entire sentence. Addition experiments explore that distilling knowledge from RAD (the teacher) model to guide RAD-Reborn model (the student) training can achieve better performance. Extensive experiments on QA-SRL Bank 2.0 show that our model outperforms previous systems of all of the evaluation metrics. In particular, the metric EM (Exact Match) increased significantly by over 3%.
Zhen Huang 0006, Shiyi Xu, Yuxing Peng 0001
IJCNN7
2020 A Distributed Computing Framework Based on Variance Reduction Method to Accelerate Training Machine Learning Models
abstract
To support large-scale intelligent applications, distributed machine learning based on JointCloud is an intuitive solution scheme. However, the distributed machine learning is difficult to train due to that the corresponding optimization solver algorithms converge slowly, which highly demand on computing and memory resources. To overcome the challenges, we propose a computing framework for L-BFGS optimization algorithm based on variance reduction method, which can utilize a fixed big learning rate to linearly accelerate the convergence speed. To validate our claims, we have conducted several experiments on multiple classical datasets. Experimental results show that the computing framework accelerate the training process of solver and obtain accurate results for machine learning algorithms.
Zhen Huang 0006, Mingxing Tang, Jinyan Qiu, Hangjun Zhou, Yuan Yuan 0034, Dongsheng Li 0001, Yuxing Peng 0001
JCC9
2020 Accelerating SGD using flexible variance reduction on large-scale datasets
Mingxing Tang, Linbo Qiao, Zhen Huang 0006, Xinwang Liu 0002, Yuxing Peng 0001, Xueliang Liu
Neural Comput. Appl.5
2020 Audio Tagging by Cross Filtering Noisy Labels
abstract
High quality labeled datasets have allowed deep learning to achieve impressive results on many sound analysis tasks. Yet, it is labor-intensive to accurately annotate large amount of audio data, and the dataset may contain noisy labels in the practical settings. Meanwhile, the deep neural networks are susceptive to those incorrect labeled data because of their outstanding memorization ability. In this article, we present a novel framework, named CrossFilter, to combat the noisy labels problem for audio tagging. Multiple representations (such as, Logmel and MFCC) are used as the input of our framework for providing more complementary information of the audio. Then, though the cooperation and interaction of two neural networks, we divide the dataset into curated and noisy subsets by incrementally pick out the possibly correctly labeled data from the noisy data. Moreover, our approach leverages the multi-task learning on curated and noisy subsets with different loss function to fully utilize the entire dataset. The noisy-robust loss function is employed to alleviate the adverse effects of incorrect labels. On both the audio tagging datasets FSDKaggle2018 and FSDKaggle2019, empirical results demonstrate the performance improvement compared with other competing approaches. On FSDKaggle2018 dataset, our method achieves state-of-the-art performance and even surpasses the ensemble models.
Boqing Zhu, Kele Xu, Qiuqiang Kong, Huaimin Wang 0001, Yuxing Peng 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Read + Verify: Machine Reading Comprehension with Unanswerable Questions
abstract
Machine reading comprehension with unanswerable questions aims to abstain from answering when no answer can be inferred. In addition to extract answers, previous works usually predict an additional “no-answer” probability to detect unanswerable cases. However, they fail to validate the answerability of the question by verifying the legitimacy of the predicted answer. To address this problem, we propose a novel read-then-verify system, which not only utilizes a neural reader to extract candidate answers and produce no-answer probabilities, but also leverages an answer verifier to decide whether the predicted answer is entailed by the input snippets. Moreover, we introduce two auxiliary losses to help the reader better handle answer extraction as well as no-answer detection, and investigate three different architectures for the answer verifier. Our experiments on the SQuAD 2.0 dataset show that our system obtains a score of 74.2 F1 on test set, achieving state-of-the-art results at the time of submission (Aug. 28th, 2018).
Furu Wei, Yuxing Peng 0001, Zhen Huang 0006, Nan Yang 0002, Dongsheng Li 0001
AAAI3
2019 Retrieve, Read, Rerank: Towards End-to-End Multi-Document Reading Comprehension
abstract
This paper considers the reading comprehension task in which multiple documents are given as input.Prior work has shown that a pipeline of retriever, reader, and reranker can improve the overall performance.However, the pipeline system is inefficient since the input is re-encoded within each module, and is unable to leverage upstream components to help downstream training.In this work, we present RE 3 QA, a unified question answering model that combines context retrieving, reading comprehension, and answer reranking to predict the final answer.Unlike previous pipelined approaches, RE 3 QA shares contextualized text representation across different components, and is carefully designed to use high-quality upstream outputs (e.g., retrieved context or candidate answers) for directly supervising downstream modules (e.g., the reader or the reranker).As a result, the whole network can be trained end-to-end to avoid the context inconsistency problem.Experiments show that our model outperforms the pipelined baseline and achieves state-ofthe-art results on two versions of TriviaQA and two variants of SQuAD.
Yuxing Peng 0001, Zhen Huang 0006, Dongsheng Li 0001
ACL (1)2
2019 Open-Domain Targeted Sentiment Analysis via Span-Based Extraction and Classification
abstract
Open-domain targeted sentiment analysis aims to detect opinion targets along with their sentiment polarities from a sentence.Prior work typically formulates this task as a sequence tagging problem.However, such formulation suffers from problems such as huge search space and sentiment inconsistency.To address these problems, we propose a span-based extract-then-classify framework, where multiple opinion targets are directly extracted from the sentence under the supervision of target span boundaries, and corresponding polarities are then classified using their span representations.We further investigate three approaches under this framework, namely the pipeline, joint, and collapsed models.Experiments on three benchmark datasets show that our approach consistently outperforms the sequence tagging baseline.Moreover, we find that the pipeline model achieves the best performance compared with the other two models.
Yuxing Peng 0001, Zhen Huang 0006, Dongsheng Li 0001, Yiwei Lv
ACL (1)2
2019 A Multi-Type Multi-Span Network for Reading Comprehension that Requires Discrete Reasoning
abstract
Minghao Hu, Yuxing Peng, Zhen Huang, Dongsheng Li. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yuxing Peng 0001, Zhen Huang 0006, Dongsheng Li 0001
EMNLP/IJCNLP (1)2
2019 ThunderNet: Towards Real-Time Generic Object Detection on Mobile Devices
abstract
Real-time generic object detection on mobile platforms is a crucial but challenging computer vision task. Prior lightweight CNN-based detectors are inclined to use one-stage pipeline. In this paper, we investigate the effectiveness of two-stage detectors in real-time generic detection and propose a lightweight two-stage detector named ThunderNet. In the backbone part, we analyze the drawbacks in previous lightweight backbones and present a lightweight backbone designed for object detection. In the detection part, we exploit an extremely efficient RPN and detection head design. To generate more discriminative feature representation, we design two efficient architecture blocks, Context Enhancement Module and Spatial Attention Module. At last, we investigate the balance between the input resolution, the backbone, and the detection head. Benefit from the highly efficient backbone and detection part design, ThunderNet surpasses previous lightweight one-stage detectors with only 40% of the computational cost on PASCAL VOC and COCO benchmarks. Without bells and whistles, ThunderNet runs at 24.1 fps on an ARM-based device with 19.2 AP on COCO. To the best of our knowledge, this is the first real-time detector reported on ARM platforms. Code will be released for paper reproduction.
Zheng Qin 0002, Zhaoning Zhang 0001, Yiping Bao, Gang Yu 0002, Yuxing Peng 0001, Jian Sun 0001
ICCV6
2019 TFPN: Twin Feature Pyramid Networks for Object Detection
abstract
FPN (Feature Pyramid Networks) is one of the most popular object detection networks, which can improve small object detection by enhancing shallow features. However, limited attention has been paid to the improvement of large object detection via deeper feature enhancement. One existing approach merges the feature maps of different layers into a new feature map for object detection, but can lead to increased noise and loss of information. The other approach adds a bottom-up structure after the feature pyramid of FPN, which superimposes the information from shallow layers into the deep feature map but weakens the strength of FPN in detecting small objects. To address these challenges, this paper proposes TFPN (Twin Feature Pyramid Networks), which consists of (1) FPN+, a bottom-up structure that improves large object detection; (2) TPS, a Twin Pyramid Structure that improves medium object detection; and (3) innovative integration of these two with FPN, which can significantly improve the detection accuracy of large and medium objects while maintaining the advantage of FPN in small object detection. Extensive experiments using the MSCOCO object detection datasets and the BDD100K automatic driving dataset demonstrate that TFPN significantly improves over existing models, achieving up to 2.2 improvement in detection accuracy (e.g., 36.3 for FPN vs. 38.5 for TFPN on COCO Val-17). Our method can obtain the same accuracy as FPN with ResNet-101 based on ResNet-50 and needs fewer parameters.
Fangzhao Li, Yuxing Peng 0001, Qin Lv, Yuan Yuan 0034, Zhen Huang 0006
ICTAI4
2019 Towards a delivery scheme for speedup of data backup in distributed storage systems using erasure codes
Pengfei You, Zhen Huang 0006, Yuxing Peng 0001, Guofeng Yan
J. Supercomput.3
2018 Attention-Guided Answer Distillation for Machine Reading Comprehension
abstract
Despite that current reading comprehension systems have achieved significant advancements, their promising performances are often obtained at the cost of making an ensemble of numerous models. Besides, existing approaches are also vulnerable to adversarial attacks. This paper tackles these problems by leveraging knowledge distillation, which aims to transfer knowledge from an ensemble model to a single model. We first demonstrate that vanilla knowledge distillation applied to answer span prediction is effective for reading comprehension systems. We then propose two novel approaches that not only penalize the prediction on confusing answers but also guide the training with alignment information distilled from the ensemble. Experiments show that our best student model has only a slight drop of 0.4% F1 on the SQuAD test set compared to the ensemble teacher, while running 12x faster during inference. It even outperforms the teacher on adversarial SQuAD datasets and NarrativeQA benchmark.
Yuxing Peng 0001, Furu Wei, Zhen Huang 0006, Dongsheng Li 0001, Nan Yang 0002, Ming Zhou 0001
EMNLP2
2018 Fd-Mobilenet: Improved Mobilenet with a Fast Downsampling Strategy
abstract
We present Fast-Downsampling MobileNet (FD-MobileNet), an efficient and accurate network for very limited computational budgets (e.g., 10-140 MFLOPs). Our key idea is applying a fast downsampling strategy to Mobile Net framework. In FD-Mobile Net, we perform 32× downsampling within 12 layers, only half the layers in the original MobileNet. This design brings three advantages: (i) It remarkably reduces the computational cost. (ii) It increases the information capacity and achieves significant performance improvements. (iii) It is engineering-friendly and provides fast actual inference speed. Experiments on ILSVRC 2012 and PASCAL VOC datasets demonstrate that FD-Mobile Net consistently outperforms MobileNet and achieves comparable results with ShufflieNet under different computational budgets, for instance, surpassing Mobile-Net by 5.5% on the ILSVRC 2012 top-l accuracy and 8.3% on the VOC 2007 mAP under a complexity of 12 MFLOPs. On an ARM-based device, FD-Mobile Net achieves 1.11× inference speedup over Mobile Net and 1.82× over Shufflie Net under the same complexity.
Zheng Qin 0002, Zhaoning Zhang 0001, Xiaotao Chen, Yuxing Peng 0001
ICIP5
2018 FVR-SGD: A New Flexible Variance-Reduction Method for SGD on Large-Scale Datasets
Mingxing Tang, Zhen Huang 0006, Linbo Qiao, Shuyang Du, Yuxing Peng 0001
ICONIP (2)5
2018 A Quick Survey on Large Scale Distributed Deep Learning Systems
abstract
Deep learning have been widely used in various fields and has worked very well as a major role. While the gradual penetration into various fields, data quantity of each applications is increasing tremendously, and so as the computation complexity and model parameters. As an obvious result, the training and inference is time consuming. For example, a classic Resnet50 classification model will be trained in 14 days on a NVIDIA M40 GPU with ImageNet data set. Thus, distributed acceleration is a very useful way to dispatch the computation of training and even inference to scale of nodes in parallel and accelerate the whole process. Facebook's work and UC Berkeley's acceleration can training the Resnet-50 model within hour and minutes by distributed deep learning algorithm and system, representatively. As other distributed accelerations, it gives a possibility to accelerate large models on large data sets from weeks to minutes, which gives researchers and developers more space to explore and search. However, besides acceleration, what other issues will be confronted of the distributed deep learning system? Where is the upper limit of acceleration? What application will acceleration be used for? What is the price and cost of acceleration? In this paper, we will take a simple and quick survey on the distributed deep learning system from algorithm perspective, distributed system perspective and applications perspective. We will present several recent excellent works, and bring analysis on the restricts and prospects of the distributed methods.
Zhaoning Zhang 0001, Lujia Yin, Yuxing Peng 0001, Dongsheng Li 0001
ICPADS3
2018 Multi-LCNN: A Hybrid Neural Network Based on Integrated Time-Frequency Characteristics for Acoustic Scene Classification
abstract
Acoustic scene classification (ASC) is an important task in audio signal processing and can be useful in many real-world applications. Recently, several deep neural network models have been proposed for ASC, such as LSTMs based on temporal analysis and CNNs based on frequency spectrum, as well as hybrid models of LSTM and CNN to further improve classification performance. However, existing hybrid models fail to properly preserve the temporal information when transferring data between different models. In this work, we first analyze the cause of such temporal information loss. We then propose Multi-LCNN, a new hybrid model with two important mechanisms: (1) a LCNN architecture to effectively preserve temporal information; and (2) a multi-channel feature fusion mechanism (MCFF) that combines enhanced temporal information and frequency spectrogram information to learn highly integrated and discriminative features for ASC. Evaluations on the TUT ASC 2016 dataset show that our model can achieve an improvement of 10.23% over the baseline method, and is currently the best-performing end-to-end model on this dataset.
Jin Lei, Boqing Zhu, Qin Lv, Zhen Huang 0006, Yuxing Peng 0001
ICTAI6
2018 Reinforced Mnemonic Reader for Machine Reading Comprehension
abstract
In this paper, we introduce the Reinforced Mnemonic Reader for machine reading comprehension tasks, which enhances previous attentive readers in two aspects. First, a reattention mechanism is proposed to refine current attentions by directly accessing to past attentions that are temporally memorized in a multi-round alignment architecture, so as to avoid the problems of attention redundancy and attention deficiency. Second, a new optimization approach, called dynamic-critical reinforcement learning, is introduced to extend the standard supervised method. It always encourages to predict a more acceptable answer so as to address the convergence suppression problem occurred in traditional reinforcement learning algorithms. Extensive experiments on the Stanford Question Answering Dataset (SQuAD) show that our model achieves state-of-the-art results. Meanwhile, our model outperforms previous systems by over 6% in terms of both Exact Match and F1 metrics on two adversarial SQuAD datasets.
Yuxing Peng 0001, Zhen Huang 0006, Xipeng Qiu, Furu Wei, Ming Zhou 0001
IJCAI2
2018 Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions
abstract
Depthwise convolutions provide significant performance benefits owing to the reduction in both parameters and mult-adds. However, training depthwise convolution layers with GPUs is slow in current deep learning frameworks because their implementations cannot fully utilize the GPU capacity. To address this problem, in this paper we present an efficient method (called diagonalwise refactorization) for accelerating the training of depthwise convolution layers. Our key idea is to rearrange the weight vectors of a depthwise convolution into a large diagonal weight matrix so as to convert the depthwise convolution into one single standard convolution, which is well supported by the cuDNN library that is highly-optimized for GPU computations. We have implemented our training method in five popular deep learning frameworks. Evaluation results show that our proposed method gains 15.4× training speedup on Darknet, 8.4× on Caffe, 5.4× on PyTorch, 3.5× on MXNet, and 1.4× on TensorFlow, compared to their original implementations of depthwise convolutions.
Zheng Qin 0002, Zhaoning Zhang 0001, Dongsheng Li 0001, Yiming Zhang 0003, Yuxing Peng 0001
IJCNN5
2018 Merging and Evolution: Improving Convolutional Neural Networks for Mobile Applications
abstract
Compact neural networks are inclined to exploit “sparsely-connected” convolutions such as depthwise convolution and group convolution for employment in mobile applications. Compared with standard “fully-connected” convolutions, these convolutions are more computationally economical. However, “sparsely-connected” convolutions block the inter-group informa-tion exchange, which induces severe performance degradation. To address this issue, we present two novel operations named merging and evolution to leverage the inter-group information. Our key idea is encoding the inter-group information with a narrow feature map, then combining the generated features with the original network for better representation. Taking advantage of the proposed operations, we then introduce the Merging-and- Evolution (ME) module, an architectural unit specifically designed for compact networks. Finally, we propose a family of compact neural networks called MENet based on ME modules. Extensive experiments on ILSVRC 2012 dataset and PASCAL VOC 2007 dataset demonstrate that MENet consistently outperforms other state-of -the-art compact networks under different computational budgets. For instance, under the computational budget of 140 MFLOPs, MENet surpasses ShuffleNet by 1% and MobileNet by 1.95% on ILSVRC 2012 top-l accuracy, while by 2.3% and 4.1% on PASCAL VOC 2007 mAP, respectively.
Zheng Qin 0002, Zhaoning Zhang 0001, Shiqing Zhang, Hao Yu 0010, Jincai Li, Yuxing Peng 0001
IJCNN6
2018 Learning Environmental Sounds with Multi-scale Convolutional Neural Network
abstract
Deep learning has dramatically improved the performance of sounds recognition. However, learning acoustic models directly from the raw waveform is still challenging. Current waveform-based models generally use time-domain convolutional layers to extract features. The features extracted by single size filters are insufficient for building discriminative representation of audios. In this paper, we propose multi-scale convolution operation, which can get better audio representation by improving the frequency resolution and learning filters cross all frequency area. For leveraging the waveform-based features and spectrogram-based features in a single model, we introduce two-phase method to fuse the different features. Finally, we propose a novel end-to-end network called WaveMsNet based on the multi-scale convolution operation and two-phase method. On the environmental sounds classification datasets ESC-10 and ESC-50, the classification accuracies of our WaveMsNet achieve 93.75% and 79.10% respectively, which improve significantly from the previous methods.
Boqing Zhu, Jin Lei, Zhen Huang 0006, Yuxing Peng 0001, Fei Li 0004
IJCNN6
2018 Matrix factorization for recommendation with explicit and implicit feedback
Shulong Chen, Yuxing Peng 0001
Knowl. Based Syst.2
2017 Fast Compressive Spectral Clustering
abstract
Compressive spectral clustering (CSC) efficiently leverages graph filter and random sampling techniques to speed up clustering process. However, we find that CSC algorithm suffers from two main problems: i) The direct use of the dichotomy and eigencount techniques for estimating laplacian matrix’s k-th eigenvalue is expensive. ii) The computation of polynomial approximation repeats in each iteration for every cluster in the interpolation process, which occupies most of the computation time of CSC. To address these problems, we propose a new approach called FCSC for fast compressive spectral clustering. FCSC addresses the first problem by assuming that the eigenvalues approximately satisfy local uniform distribution, and addresses the second problem by recalculating the pairwise similarity between nodes with low-dimensional representation to reconstruct denoised laplacian matrix. The time complexity of reconstruction is linear with the number of non-zeros in laplacian matrix. As experimentally demonstrated on artificial and real-world datasets, our approach significantly reduces the computation time while preserving high clustering accuracy comparable to previous designs, verifying the effectiveness of FCSC.
Yiming Zhang 0003, Dongsheng Li 0001, Xinwang Liu 0002, Yuxing Peng 0001
ICDM5
2017 SDF-NN: A Deep Neural Network with Semantic Dropping and Fusion for Natural Language Inference
abstract
Natural language inference (NLI) is an important task in natural language processing (NLP), and recently, several deep neural network based models have been proposed for NLI. In this work, we first make two important observations regarding NLI: (1) the existence of extra/interfering semantics and its negative impact on the correctness of final inference; and (2) the unbalanced importance of local inference results and the need to combine all local results for aggregation. Motivated by these two observations, we have designed SDF-NN, a new NLI model with two novel components: (1) a Semantic Dropping Network (SDN) to automatically discard some of the interfering semantics; and (2) a Semantic Fusion Alignment (SFA) method to effectively fuse all local inference results. Our model has achieved 88.2% accuracy on the SNLI corpus, which is currently the best performing single model.
Ludan Tan, Qin Lv, Yuxing Peng 0001, Xiang Zhao 0002, Zhen Huang 0006
ICTAI4
2017 PARIX: Speculative Partial Writes in Erasure-Coded Systems
Huiba Li, Yiming Zhang 0003, Shengyun Liu, Dongsheng Li 0001, Yuxing Peng 0001
USENIX ATC7
2017 Meeting deadlines for approximation processing in MapReduce environments
abstract
To provide timely results for big data analytics, it is crucial to satisfy deadline requirements for MapReduce jobs in today’s production environments. Much effort has been devoted to the problem of meeting deadlines, and typically there exist two kinds of solutions. The first is to allocate appropriate resources to complete the entire job before the specified time limit, where missed deadlines result because of tight deadline constraints or lack of resources; the second is to run a pre-constructed sample based on deadline constraints, which can satisfy the time requirement but fail to maximize the volumes of processed data. In this paper, we propose a deadline-oriented task scheduling approach, named ‘Dart’, to address the above problem. Given a specified deadline and restricted resources, Dart uses an iterative estimation method, which is based on both historical data and job running status to precisely estimate the real-time job completion time. Based on the estimated time, Dart uses an approach–revise algorithm to make dynamic scheduling decisions for meeting deadlines while maximizing the amount of processed data and mitigating stragglers. Dart also efficiently handles task failures and data skew, protecting its performance from being harmed. We have validated our approach using workloads from OpenCloud and Facebook on a cluster of 64 virtual machines. The results show that Dart can not only effectively meet the deadline but also process near-maximum volumes of data even with tight deadlines and limited resources.
Minghao Hu 0003, Yuxing Peng 0001
Frontiers Inf. Technol. Electron. Eng.3
2016 DISVMs: Fast SVMs Training on Large-Scale Data Sets
abstract
Support Vector Machines (SVMs) are powerful classification tools. However, the model training is very time-consuming when meeting large scale data sets. Some efforts have been devoted to screening out non-support vectors (non-SVs) to accelerate the training. But their processes rely on prior knowledge of other classifiers with different parameters to screen out non-SVs. In this paper, we propose Directional Indicator Support Vector Machines (DISVMs) to efficiently identify non-SVs. DISVMs employs a directional indicator, which points to the approximately orthogonal direction of the separating hyperplane, to qualitatively define the location of different samples and thus identify non-SVs. Furthermore, DISVMs leverages a two-stage algorithm: the first stage is to compute the directional indicator. The second stage is to identify non-SVs using the indicator. To avoid misjudgement, we propose CnSV method for non-SVs based on the majority rule. DISVMs screens out non-SVs with light computation and little accuracy loss. Experiments show that our approach significantly reduces the total computation cost.
Lijuan Cui, Ziyang Li 0003, Yuxing Peng 0001
ICTAI4
2016 OPTAS: Decentralized flow monitoring and scheduling for tiny tasks
abstract
Task-aware flow schedulers collect task information across the data center to optimize task-level performance. However, the majority of the tasks, which generate short flows and are called tiny tasks, have been largely overlooked by current schedulers. The large number of tiny tasks brings significant overhead to the centralized schedulers, while the existing decentralized schedulers are too complex to fit in commodity switches. In this paper we present OPTAS, a lightweight, commodity-switch-compatible scheduling solution that efficiently monitors and schedules flows for tiny tasks with low overhead. OPTAS monitors system calls and buffer footprints to recognize the tiny tasks, and assigns them with higher priorities than larger ones. The tiny tasks are then transferred in a FIFO manner by adjusting two attributes, namely, the window size and round trip time, of TCP. We have implemented OPTAS as a Linux kernel module, and experiments on our 37-server testbed show that OPTAS is at least 2.2× faster than fair sharing, and 1.2× faster than only assigning tiny tasks with the highest priority.
Ziyang Li 0003, Yiming Zhang 0003, Dongsheng Li 0001, Kai Chen 0005, Yuxing Peng 0001
INFOCOM5
2016 DFRS: A Large-Scale Distributed Fingerprint Recognition System Based on Redis
Zhen Huang 0006, Jinbang Chen, Yuxing Peng 0001
MMM (1)5
2016 Best Effort Task Scheduling for Data Parallel Jobs
abstract
The tasks of data-parallel computation jobs come up with diverse and time-varying resource requirements. The dynamic nature of task requirements brings challenges on making good scheduling decisions, due to it is hard to keep work-conserving. In this paper, we present BETS to cope with the requirement dynamics that aims at utilizing cluster resources fully. BETS employs a task model that represents for runtime task requirements, a coarse-grained task pipeline to make use of resources in a time-division multiplexing fashion, and fine-grained resource management to guarantee performance.
Ziyang Li 0003, Yiming Zhang 0003, Yuxing Peng 0001, Dongsheng Li 0001
SIGCOMM4
2015 DLG-Hypertree: A Low-Diameter, Server-centric Datacenter Network Architecture
abstract
The architecture of data center networks (DCN) is critical to the performance of cloud computing. Hyper tree is a novel interconnection structure which assures low diameter as well as high fault tolerance. We utilize the "distributed line graph" (DLG) technique on hyper tree to design an efficient DCN architecture, called DLG-Hyper tree (DH). We implement a prototype of DH using the Owlet language. The effectiveness of DH is demonstrated through simulation and evaluation on the prototype of DH.
Lin Gui 0005, Rui Shen 0003, Yuxing Peng 0001
CLOUD5
2015 Pallas: An Application-Driven Task and Network Simulation Framework
abstract
With the help of simulation tools, users can evaluate new proposals in cluster environment efficiently. However, current cloud simulators cannot meet the needs of application-driven simulation scenarios. In this paper, we propose Pallas, a task and network simulation framework that supports various cloud applications. Task-aware network scheduling and network-perceived task placement algorithms can be easily implemented in Pallas. We present the architecture and main components of Pallas and evaluate its effectiveness by comparing algorithm improvements to the actual results.
Yuming Ye, Ziyang Li 0003, Dongsheng Li 0001, Yiming Zhang 0003, Yuxing Peng 0001
CLUSTER6
2015 Deadline-Oriented Task Scheduling for MapReduce Environments
Pengfei You, Zhen Huang 0006, Yuxing Peng 0001
ICA3PP (2)5
2015 Parallel Data Regeneration Based on Multiple Trees with Network Coding in Distributed Storage System
Pengfei You, Zhen Huang 0006, Yuxing Peng 0001
ICA3PP (2)5
2015 Minimizing data redundancy for high reliable cloud storage systems
Zhen Huang 0006, Jinbang Chen, Yisong Lin, Pengfei You, Yuxing Peng 0001
Comput. Networks5
2014 Repairing Multiple Data Losses by Parallel Max-min Trees Based on Regenerating Codes in Distributed Storage Systems
Pengfei You, Yuxing Peng 0001, Zhen Huang 0006
ICA3PP (2)2
2014 Optimal Task Scheduling in MapReduce
abstract
The scheduling approach in MapReduce may result in the "long tail" problem because of the unreasonable task assignment and high scheduling overhead because of an amount of task scheduling operations. To address these problems, a new task scheduling approach for MapReduce, named "Iterative Task Scheduling Algorithm", is proposed. The new approach tries to schedule the map tasks according to the solution of the equation for the optimal task assignment. Thus the "long tail" problem can be mitigated effectively and the task scheduling operations can be significantly reduced. To support our new scheduling approach, two approaches are proposed: The first one is adopted to estimate task execution times of nodes and the second one is adopted to produce the optimal task assignment based on the known task execution times of nodes. Comprehensive experiments have been performed with the real log data from the Ali Cloud and the results verify the effectiveness of the new task scheduling approach. The map runtime of the job is reduced 23% in our experiments.
Yuxing Peng 0001, Mingxing Tang, Jinghua Feng, Pengfei You
NAS2
2014 RAFlow: Read Ahead Accelerated I/O Flow through Multiple Virtual Layers
abstract
Virtualization is the foundation for cloud computing, and the virtualization can not be achieved without software defined, elastic, flexible and scalable virtual layers. Unfortunately, if multiple virtual storage devices are chained together, the system may be subject to severe performance degradation. While the read-ahead (RA) mechanism in storage devices plays a very important role to improve I/O performance, RA may not be effective as expected for multiple virtualization layers, since it is originally designed for one layer only. When I/O requests are passed through a long I/O path, they may trigger a chain reaction and lead to unnecessary data transmission and thus bandwidth waste. In this paper, we study the dynamic behavior of RA through multiple I/O layers and demonstrate that if controlled well, RA can greatly accelerate I/O speed. We present RAFlow, a RA control mechanism, to effectively improve I/O performance by strategically expanding RA window at each layer. Our real-world experiments show that it can achieve 20% to 50% performance improvement in I/O paths with up to 8 virtualized storage devices.
Zhaoning Zhang 0001, Kui Wu 0001, Huiba Li, Jinghua Feng, Yuxing Peng 0001, Xicheng Lu
NAS5
2014 Performance Analysis of End-to-End Services in Virtualized Computing Environments
Guofeng Yan, Yuxing Peng 0001
NPC2
2014 VMThunder: Fast Provisioning of Large-Scale Virtual Machine Clusters
abstract
Infrastructure as a service (IaaS) allows users to rent resources from the Cloud to meet their various computing requirements. The pay-as-you-use model, however, poses a nontrivial technical challenge to the IaaS cloud service providers: how to fast provision a large number of virtual machines (VMs) to meet users' dynamic computing requests? We address this challenge with VMThunder, a new VM provisioning tool, which downloads data blockson demandduring the VM booting process and speeds up VM image streaming by strategically integrating peer-to-peer (P2P) streaming techniques with enhanced optimization schemes such as transfer on demand, cache on read, snapshot on local, and relay on cache. In particular, VMThunder stores the original images in a share storage and in the meantime it adopts a tree-based P2P streaming scheme so that common image blocks are cached and reused across the nodes in the cluster. We implement VMThunder in CentOS Linux and thoroughly test its performance. Comprehensive experimental results show that VMThunder outperforms the state-of-the-art VM provisioning methods, with respect to scalability, latency, and VM runtime I/O performance.
Zhaoning Zhang 0001, Ziyang Li 0003, Kui Wu 0001, Dongsheng Li 0001, Huiba Li, Yuxing Peng 0001, Xicheng Lu
IEEE Trans. Parallel Distributed Syst.6
2013 OPTAS: Optimal Data Placement in MapReduce
abstract
The data placement strategy greatly affects the efficiency of MapReduce. The current strategy only takes the map phase into account to optimize the map time. But the ignored shuffle phase may increase the total running time significantly in many jobs. We propose a new data placement strategy, named OPTAS, which optimizes both the map and shuffle phases to reduce their total time. However, the huge search space makes it difficult to find out an optimal data placement instance (DPI) rapidly. To address this problem, an algorithm is proposed which can prune most of the search space and find out an optimal result quickly. The search space firstly is segmented in ascending order according to the potential map time. Within each segment, we propose an efficient method to construct a local optimal DPI with the minimal total time of both the map and shuffle phases. To find the global optimal DPI, we scan the local optimal DPIs in order. We have proven that the global optimal DPI can be found as the first local optimal DPI whose total time stops decreasing, thus further pruning the search space. In practice, we find that at most fourteen local optimal DPIs are scanned in tens of thousands of segments with the pruning strategy. Extensive experiments with real trace data verify not only the theoretic analysis of our pruning strategy and construction method but also the optimality of OPTAS. The best improvements obtained in our experiments can be over 40% compared with the existing strategy used by MapReduce.
Yongrui Qin, Zhen Huang 0006, Yuxing Peng 0001, Dongsheng Li 0001, Huiba Li
ICPADS4
2012 Providing Information Services for Wireless Sensor Networks through Cloud Computing
abstract
Wireless sensor networks (WSN) is a critical technology for information gathering covering many areas, including health-care, transportation, air traffic control and environment monitoring. Despite wide use, the fast increasing data emanating from WSN is not fully utilized due to the limitation for structure of WSN itself. Along with the further development of WSN, the data form which is not be efficiently managed and applied to supply information services for users. As the emerging IT technology, cloud computing supplies powerful utilization ability for IT resources, which makes many traditional applications migrate to cloud computing. In this paper, we propose a framework integrating cloud computing paradigm and WSN, which fully uses data process ability and service model for cloud computing. In the framework, data form WSN are efficiently utilized and managed, depending on which, information services for WSN are well provided to users.
Pengfei You, Yuxing Peng 0001
APSCC2
2012 A scalable code dissemination protocol in heterogeneous wireless sensor networks
Shaoliang Peng, Shanshan Li 0001, Xiangke Liao, Yuxing Peng 0001, Nong Xiao 0001
Sci. China Inf. Sci.4
2012 Robust Redundancy Scheme for the Repair Process: Hierarchical Codes in the Bandwidth-Limited Systems
Zhen Huang 0006, Yisong Lin, Yuxing Peng 0001
J. Grid Comput.3
2012 Fast Release/Capture Sampling in Large-Scale Sensor Networks
abstract
Efficient estimation of global information is a common requirement for many wireless sensor network applications. Examples include counting the number of nodes alive in the network and measuring the scale of physically correlated events. These tasks must be accomplished at extremely low overhead due to the severe resource limitation of sensor nodes, which poses a challenge for large-scale sensor networks. In this paper, we develop a novel protocol FLAKE to efficiently and accurately estimate the global information of large-scale sensor networks based on the sparse sampling theory. Specially, FLAKE disseminates a small number of messages called seeds to the network and issues a query about which nodes receive a seed. The number of nodes that have the information of interest can be estimated by counting the seeds disseminated, the nodes queried, and the nodes that receive a seed. FLAKE can be easily implemented in a distributed manner due to its simplicity. Moreover, desirable tradeoffs can be achieved between the accuracy of estimation and the system overhead. Our simulations show that FLAKE significantly outperforms several existing schemes on accuracy, delay, and message overhead.
Shaoliang Peng, Guoliang Xing, Shanshan Li 0001, Weijia Jia 0001, Yuxing Peng 0001
IEEE Trans. Mob. Comput.5
2011 Error Detection by Redundant Transaction in Transactional Memory System
abstract
This paper addresses the issue of error detection in transactional memory, and proposes a new method of error detection based on redundant transaction (EDRT). This method creates a transaction copy for every transaction, and executes both original transactions and transaction copies on adequate processor cores, and achieves error detection by comparing the execution results. EDRT utilizes the data-versioning mechanism of transactional memory to achieve the acquisition of an approximate minimum error detection comparing data set, and the acquisition is transparent and online. At last, this paper validates the EDRT through 5 test programs, including 4 SPLASH-2 benchmarks. The experimental results show that, the average error detecting cost is about 3.68% relative to the whole program, and it's only about 12.07% relative to the transaction parts of the program.
Yuxing Peng 0001
NAS3
2011 Reducing Repair Traffic in P2P Backup Systems: Exact Regenerating Codes on Hierarchical Codes
abstract
Peer to peer backup systems store data on “unreliable” peers that can leave the system at any moment. In this case, the only way to assure durability of the data is to add redundancy using either replication or erasure codes. Erasure codes are able to provide the same reliability as replication requiring much less storage space. Erasure coding breaks the data into blocks that are encoded and then stored on different nodes. However, when storage nodes permanently abandon the system, new redundant blocks must be created, which is referred to as repair. For “classical” erasure codes, generating a new block requires the transmission of k blocks over the network, resulting in a high repair traffic. Recently, two new classes of erasure codes, Regenerating Codes and Hierarchical Codes, have been proposed that significantly reduce the repair traffic. Regenerating Codes reduce the amount of data uploaded by each peer involved in the repair, while Hierarchical Codes reduce the number of nodes participating in the repair. In this article we propose to combine these two codes to devise a new class of erasure codes called ER-Hierarchical Codes that combine the advantages of both.
Zhen Huang 0006, Ernst W. Biersack, Yuxing Peng 0001
ACM Trans. Storage3
2010 Nexus: Speculative Execution for Event-Driven Networking Programs
abstract
The efficiency of communication is a key factor to the performance of networking applications, and concurrent communication is an important approach to the efficiency of communication. However, many concurrency opportunities are very difficult to exploit because they depend on some undeterministic conditions. If these conditions are highly predictable, speculative execution can be a very effective approach to cope with the uncertainties. Existing researches on speculation seldom target at networking systems, and none of them can handle the event-driven model that is very popular in such systems. In this paper, we propose Nexus, a novel speculation scheme that supports event-driven networking applications. Nexus analyzes the dependence relationship of events, and performs speculation according to the duality of events and threads. Evaluation on a prototype implementation of nexus shows that this approach can significantly reduces the time needed to complete an event-driven program.
Huiba Li, Xicheng Lu, Yuxing Peng 0001
ICPADS3
2010 Automatic Concurrency Management for distributed applications
abstract
Building distributed applications is difficult mostly because of concurrency management. Existing approaches primarily include events and threads. Researchers and developers have been debating for decades to prove which is superior. Although the conclusion is far from obvious, this long debate clearly shows that neither of them is perfect. One of the problems is that they are both complex and error-prone. Both events and threads need the programmers to explicitly manage concurrency, and we believe it is just the source of difficulties. In this paper, we propose a novel approach—automatic concurrency management by the runtime system. It dynamically analyzes the programs to discover potential concurrency opportunities; and it dynamically schedules the communication and the computation tasks, resulting in automatic concurrent execution. This approach is inspired by the instruction scheduling technologies used in modern microprocessors, which dynamically exploits instruction-level parallelism. However, hardware scheduling algorithms do not fit software in many aspects, thus we have to design a new scheme completely from scratch. automatic concurrency management is a runtime technique with no modification to the language, compiler or byte code, so it is good at backward compatibility. It is essentially a dynamic optimization for networking programs.
Huiba Li, Shengyun Liu, Yuxing Peng 0001, Dongsheng Li 0001
ISCC3
2010 Fish a lake: Fast release/capture sampling in large-scale sensor networks
abstract
Efficient estimation of global information is a common requirement for many wireless sensor network applications. Examples include counting the number of nodes alive in the network and measuring the scale of physically correlated events. These tasks must be accomplished at extremely low overhead due to the severe resource limitation of sensor nodes, which poses a challenge for large-scale sensor networks. In this paper, we develop a novel protocol called FLAKE that can efficiently and accurately estimate the global information of large-scale sensor networks based on the sparse sampling theory. Specially, FLAKE disseminates a small number of messages called seeds to the network and issues a query about which nodes receive a seed. The number of nodes that have the information of interest can be estimated by counting the seeds disseminated, the nodes queried, and the nodes that receive a seed. FLAKE can be easily implemented in a distributed manner due to its simplicity. Moreover, desirable trade-offs can be achieved between the accuracy of estimation and the system overhead. Our simulations show that FLAKE significantly outperforms several existing schemes on accuracy, delay and message overhead.
Shaoliang Peng, Guoliang Xing, Shanshan Li 0001, Weijia Jia 0001, Yuxing Peng 0001
IWQoS5
2010 PerturbationAnalyzer: a tool for investigating the effects of concentration perturbation on protein interaction networks
abstract
UNLABELLED: The propagation of perturbations in protein concentration through a protein interaction network (PIN) can shed light on network dynamics and function. In order to facilitate this type of study, PerturbationAnalyzer, which is an open source plugin for Cytoscape, has been developed. PerturbationAnalyzer can be used in manual mode for simulating user-defined perturbations, as well as in batch mode for evaluating network robustness and identifying significant proteins that cause large propagation effects in the PINs when their concentrations are perturbed. Results from PerturbationAnalyzer can be represented in an intuitive and customizable way and can also be exported for further exploration. PerturbationAnalyzer has great potential in mining the design principles of protein networks, and may be a useful tool for identifying drug targets. AVAILABILITY: PerturbationAnalyzer can be accessed from the Cytoscape web site http://www.cytoscape.org/plugins/index.php or http://biotech.bmi.ac.cn/PerturbationAnalyzer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fei Li 0004, Wenjian Xu, Yuxing Peng 0001, Xiaochen Bo, Shengqi Wang
Bioinform.4
2010 Superscalar communication: A runtime optimization for distributed applications
Huiba Li, Shengyun Liu, Yuxing Peng 0001, Dongsheng Li 0001, Hangjun Zhou, Xicheng Lu
Sci. China Inf. Sci.3
2009 A Peer-to-Peer Media Streaming System Based on the iVCE Platform
abstract
With the advancement of peer-to-peer technology, media streaming applications become more and more popular in the Internet. However, the traditional development methods for this kind of applications need developers not only to consider the application logic but also to manage the dynamics of Internet resources, thus increasing the difficulty of development and limiting the deployment of personal video distribution applications. In this paper, we design and implement a peer-to-peer streaming system in a much easier way. In this way we can concentrate on the application itself without distraction from the dynamics of Internet resources. Such simplification owes to the Internet-based Virtual Computing Environment (iVCE), which provides programming abstractions and runtime utilities that can encapsulate the complexity of managing transient resources into the platform, thus facilitating the construction of Internet applications. When we build our streaming application based on the iVCE, we only need to define the interaction protocols among distributed nodes with the Owlet programming language. Also, we implement a JavaBean, which can be used by the Owlet program, to assist the transferring and rendering of the content. Our implementation shows that peer-to-peer applications such as media streaming, can be elegantly built using the iVCE platform, and it can serve as a reference implementation for developing similar applications.
Jiqing Wu, Yuxing Peng 0001, Rui Shen 0003
ICPADS2
2009 Providing Responsiveness Requirement Based Consistency in DVE
abstract
Consistency and responsiveness are two important factors in providing the sense of reality in Distributed Virtual Environment (DVE). However, it is not easy to optimize both aspects because of the trade-off between these two factors. As a result, most existing consistency maintenance methods ignored the responsiveness requirements, or just assumed a simple responsiveness requirement model which cannot meet the real need of DVE systems. In this paper, we first present a new responsiveness requirement model. The model can describe requirement satisfaction situation of each node. Base on this model, we propose a responsiveness requirement based consistency method. The method can adjust the utilization of time resource according to the requirements of different nodes and improve the overall responsiveness performance by at least 20%. Therefore, it provides a good support to increase the applicability of DVE systems.
Wei Zhang 0027, Hangjun Zhou, Yuxing Peng 0001, Sikun Li
ICPADS3
2009 FOCUS: A Cost-Effective Approach for Large-Scale Crop Monitoring with Sensor Networks
abstract
Current investment in crop monitoring consumes a large amount of financial cost, and how to reduce this cost has been a long-standing problem in agriculture. Traditional crop monitoring approaches are not cost-effective, because they rely on either heavy human labor or intensive computation with expensive instruments. In this paper, we explore the possibility of deploying networked sensor nodes for low-cost crop monitoring. As an example, we compute an important agricultural metric called global leaf area index (LAI) to illustrate the benefit of using sensor networks. We propose an approach called FOCUS that incrementally deploys sensor nodes into farmland to improve the accuracy of global LAI measurements. We design and implement a novel algorithm that calculates the total size of crop leaves with light intensity readings captured by the sensors under the crop canopies. FOCUS not only lowers the deployment cost considerably but also reduces the number of sensors for the long-term monitoring. Through a small-scale field test and large-scale simulations, we validate our design and show its effectiveness in crop monitoring.
Yuan Yuan 0034, Shanshan Li 0001, Kui Wu 0001, Weijia Jia 0001, Yuxing Peng 0001
MASS5
2009 Transmission Scheduling in Data-Driven Peer-to-Peer Streaming towards Optimal Throughput
abstract
Peer-to-peer media streaming has been an important service on the internet in recent years. The Data-driven (or mesh-based) structure is adopted by most working systems,in which data scheduling is one of the important problems.However, those frequently used scheduling algorithms are often faced with such a case: A neighbor peer takes up its bandwidth to deliver the packets that other neighbors can also supply, but some packets only held by it are not delivered.These packets can not be delivered in the current scheduling cycle, even though that the other neighbors have surplus bandwidth. This is a kind of waste of bandwidth and decreases the throughput of transmission. In this paper we propose anew scheduling algorithm aiming at the optimal throughput:Bipartite-matching based Block Scheduling algorithm(BBS).We convert the original data scheduling problem to a problem of finding a maximum match on the correspond bipartite graph, then assign data packets to neighbors according to the maximum match. We evaluate the performance of BBS with extensive experiments and the results show that BBS throughput and provides better streaming quality than those frequently used scheduling algorithms.
Jiqing Wu, Yuxing Peng 0001
NAS2
2009 Estimation of a Population Size in Large-Scale Wireless Sensor Networks
Shaoliang Peng, Shanshan Li 0001, Xiangke Liao, Yuxing Peng 0001, Nong Xiao 0001
J. Comput. Sci. Technol.4
2008 SenCast: Scalable multicast in wireless sensor networks
abstract
Multicast is essential for wireless sensor network (WSN) applications. Existing multicast protocols in WSNs are often designed in a P2P pattern, assuming small number of destination nodes and frequent changes on network topologies. In order to truly adopt multicast in WSNs, we propose a base-station model- based multicast, SenCast, to meet the general requirements of applications. SenCast is scalable and energy-efficient for large group communications in WSNs. Theoretical analysis shows that SenCast is able to approximate the Minimum Nonleaf Nodes (MNN) problem to a ratio of ln\R\ (R is the set of all destinations), best known lowest bound. We evaluate our design through comprehensive simulations. Experimental results demonstrate that SenCast outperforms previous multicast protocols including the most recent work uCast.
Shaoliang Peng, Shanshan Li 0001, Lei Chen 0002, Nong Xiao 0001, Yuxing Peng 0001
IPDPS5
2008 Scalable Base-Station Model-Based Multicast in Wireless Sensor Networks
Shaoliang Peng, Shanshan Li 0001, Lei Chen 0002, Yuxing Peng 0001, Nong Xiao 0001
J. Comput. Sci. Technol.4
2007 A Framework for Congestion Control for Reliable Data Delivery in Wireless Sensor Networks
abstract
WSN congestion occurs when offered traffic load exceeds available capacity. It causes overall channel quality to degrade and drop rates to rise. Furthermore, redundant transmissions are always adopted to guarantee reliable data delivery, which may deteriorate congestion since they bring on more contention and in reverse hampers the reliability. In this paper, we propose a framework to avoid, detect and mitigate congestion effectively. In this framework, a congestion aware traffic allocation (COTA) is used in multipath routing to balance traffic around the whole network COTA uses some heuristic information to analyze the potential congestion region and avoid traversing these regions. A runtime traffic adjustment CODEM is presented to use accurate metrics to detect and mitigate congestion. Compared with previous works, our work can control congestion while achieving the desired reliability at the same time. Comprehensive simulations have validated the distinguished performance in several aspects of our framework.
Shanshan Li 0001, Shaoliang Peng, Xiangke Liao, Peidong Zhu, Yuxing Peng 0001
Integrated Network Management5
2007 Real-Time Data Delivery in Wireless Sensor Networks: A Data-Aggregated, Cluster-Based Adaptive Approach
Shaoliang Peng, Shanshan Li 0001, Yuxing Peng 0001, Wen-sheng Tang, Nong Xiao 0001
UIC3
2006 A Model of Video Coding Based on Multi-agent
Zhiming Liu 0003, Yuxing Peng 0001
PRIMA3