Yusong Tan

dblp:42/1274 · DBLP profile ↗
← Back
77ranked-venue papers
5as first author
47since 2021 · last 2026
0000-0003-1233-5679ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 27 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 5 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Attention to Threat-Relevant Objects: Reasoning Detection in Autonomous Driving via Multimodal Large Language Models
abstract
Perceiving threats is an innate human instinct. During driving, humans naturally focus their attention on objects that pose real potential risks. Motivated by this observation, we shift the focus from traditional class-based detection to a novel task termed threat-oriented reasoning detection in autonomous driving. This task aims to localize threat objects and reason about their threat levels from a driver-centric perspective. To support this task, we build a benchmark comprising diverse corner-case scenarios, annotated by multiple experienced drivers to reflect human-aligned threat cognition. Given the reasoning demands of this task, we then explore the capabilities of multi-modal large language models (MLLMs) and introduce two methods based on whether the MLLM supports object detection: 1) For MLLMs lacking detection capability, we introduce ThreatCoT, a plug-and-play training-free method that combines chain-of-thought (CoT) with a visual expert toolchain to support step-by-step reasoning. 2) For MLLMs with detection support, we introduce ThreatReasoner, an end-to-end reinforcement learning (RL)-based method built on the GRPO algorithm, which enables per-object reasoning through a fully unsupervised reward strategy. Both quantitative and qualitative experiments show that our methods can effectively unlock the new capabilities of MLLM in threat-oriented reasoning detection.
Yu-Lin He, Wei Chen 0009, Xinbiao Gan, Siqi Wang 0001, Haotian Wang 0001, Yusong Tan
AAAI6
2026 Similarity-Aware Function Pre-Loading for Serverless Inference
abstract
The ubiquity of cold starts in serverless architectures poses a critical barrier to low-latency inference. While existing prewarming methods leverage idle container memory to pre-load functions, they often neglect resource contention among functions within the shared container, frequently resulting in severe request blocking. To address these challenges, this paper proposes SFP, a Similarity-Aware Function Pre-Loading strategy which optimizes function distribution by selecting target containers for pre-loading functions. SFP deploys functions with low invocation similarity within the same container while distributing those with high invocation similarity across distinct containers. Function invocation similarity, a metric proposed in this study, is derived from the Jaccard similarity coefficient and quantifies the temporal overlap between disparate function invocations. Experimental results based on real-world workloads demonstrate that, compared to state-of-the-art methods, the proposed strategy improves inference request throughput by up to 230% and achieves memory savings ranging from 6.7% to 34.8%.
Jichang Dong, Bao Li 0002, Yusong Tan
CF3
2026 Angel or devil: Discriminating hard samples and anomaly contaminations for unsupervised time series anomaly detection
Ruyi Zhang 0002, Hongzuo Xu, Songlei Jian, Yusong Tan, Haifang Zhou, Rulin Xu
Neural Networks4
2025 Achieving Speed-Accuracy Balance in Vision-based 3D Occupancy Prediction via Geometric-Semantic Disentanglement
abstract
Occupancy prediction plays a pivotal role in autonomous driving (AD) due to its capabilities of fine-grained 3D perception and general object recognition. However, existing methods often incur high computational costs, which conflict with AD's real-time demand. To this end, we redirect the focus from accuracy only to both accuracy and efficiency. By conducting a head-to-head comparison of existing methods, we find it challenging to balance accuracy and efficiency. We identify a core issue for this challenge: the strong coupling between geometry and semantics. Specifically, the predicted geometric structure (e.g., depth) guides the projection of 2D image features into 3D voxel space, which significantly affects feature discriminability and subsequent semantic learning. To address this issue, we focus on two key aspects: model design and learning strategies. 1) For model design, we propose a dual-branch network that disentangles the representation of geometry and semantics. The voxel branch utilizes a novel re-parameterized large-kernel 3D convolution to refine geometric structure efficiently, while the BEV branch employs temporal fusion and BEV encoding for efficient semantic learning. 2) For learning strategies, we propose to separate geometric learning from semantic learning by the mixup of ground-truth and predicted depths. Our method achieves 39.4% mIoU at 20 FPS on Occ3D-nuScenes, showcasing a state-of-the-art balance between accuracy and efficiency.
Yu-Lin He, Wei Chen 0009, Siqi Wang 0001, Tianci Xun, Yusong Tan
AAAI5
2025 Highly Parallelized Reinforcement Learning Training with Relaxed Assignment Dependencies
abstract
As the demands for superior agents grow, the training complexity of Deep Reinforcement Learning (DRL) becomes higher. Thus, accelerating training of DRL has become a major research focus. Dividing the DRL training process into sub-tasks and using parallel computation can effectively reduce training costs. However, current DRL training systems lack sufficient parallelization due to data assignment between sub-task components. This assignment issue has been ignored, but addressing it can further boost training efficiency. Therefore, we propose a high-throughput distributed RL training system called TianJi. It relaxes assignment dependencies between sub-task components and enables event-driven asynchronous communication. Meanwhile, TianJi maintains clear boundaries between sub-task components. To address convergence uncertainty from relaxed assignment dependencies, TianJi proposes a distributed strategy based on the balance of sample production and consumption. The strategy controls the staleness of samples to correct their quality, ensuring convergence. We conducted extensive experiments. TianJi achieves a convergence time acceleration ratio of up to 4.37 compared to related comparison frameworks. When scaled to eight computational nodes, TianJi shows a convergence time speedup of 1.6 and a throughput speedup of 7.13 relative to XingTian, emonstrating its capability to accelerate training and scalability. In data transmission efficiency experiments, TianJi significantly outperforms other frameworks, approaching hardware limits. TianJi also shows effectiveness in on-policy algorithms, achieving convergence time acceleration ratios of 4.36 and 2.95 compared to RLlib and XingTian.
Zhouyu He, Peng Qiao, Rongchun Li, Yong Dou, Yusong Tan
AAAI5
2025 GRWO: Toward Efficient Model Protection of Edge Inference via Very Few Weights Obfuscation Based on Gradient Ranking
abstract
The edge inference of deep neural networks (DNNs) raises considerable concerns regarding the security of DNN models. Using trusted execution environments (TEEs) to isolate model inference and thus protect model privacy has become a leading technology trend. However, existed methods for isolating entire neural network layers are constrained by the memory limitation of TEE. Meanwhile, the obfuscation of partial weights encounters challenges such as the complexity of weight selection and the high recovery overhead due to an excessive number of obfuscated weights. To address these issues, this paper introduces a novel two-phase global weight selection approach based on gradient ranking, designed to achieve optimal model protection with minimal obfuscation. The adversarial attack is also used to guide the weight noise processing, thereby greatly protecting the stealthiness of obfuscated weights. We validated the method on ARM TrustZone and optimized the memory allocation of the model in TEE. Experimental results demonstrate the effectiveness of our method, e.g., by obfuscating 115 out of 3.52 million weights (0.003%), the accuracy of the obfuscated MobileNet-V2 model on ImageNet drops to 0.1%. While the end-to-end latency of the method in this work is comparable to the state-of-the-art solutions, the TEE memory overhead is reduced by 64%.
Yusong Tan, Chunyan Chen, Yuanming Gao
CSCWD3
2025 Gated Cross-Attention Network for Depth Completion
abstract
Depth completion is a popular research direction in the field of depth estimation. The fusion of color and depth features is the critical challenge in this task, mainly due to the asymmetry between the rich scene details in color images and the sparse pixels in depth maps. To tackle this issue, we design an efficient Gated Cross-Attention Network that propagates confidence via a gating mechanism, simultaneously extracting and refining key information in both color and depth branches to achieve local spatial feature fusion. Additionally, we incorporate a Transformer-based attention network in low-dimensional space to effectively fuse global features and increase the network’s receptive field. At the same time, we use the Ray Tune mechanism with the AsyncHyperBandScheduler and the HyperOptSearch algorithm to automatically search for the optimal number of module iterations, which also allows us to achieve performance comparable to state-of-the-art methods. We conduct experiments on both indoor and outdoor scene datasets. Our fast network ranked first among real-time methods (below 30ms and 100ms), and our accurate network ranked first among all methods on the KITTI official website at the time of submission.
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang
ICASSP3
2025 MOVie: GPU Memory Optimization for Large-Scale DNNs Training with Virtual Memory Management
abstract
As the complexity of models and the scale of parameters grows rapidly, the limited memory capacity of intelligent acceleration devices such as GPUs has become a constraint on the development of large-scale deep neural networks (DNNs). Deep learning frameworks employ caching allocators to manage memory pools through a "split-merge" mechanism, enhancing GPU memory utilization while inevitably introducing memory fragmentation. While common memory optimization techniques effectively alleviate memory pressure during the training phase of DNNs, they also exacerbate the fragmentation issue. Existing methods rely on pre-acquired runtime information to guide memory allocation, which fails to adapt to dynamic and complex large-scale DNNs training scenarios. This paper proposes MOVie, a GPU memory optimization method for large-scale DNNs training based on virtual memory management. MOVie utilizes an over-subscribed virtual memory merging mechanism and a dynamic programming algorithm to combine non-contiguous memory blocks. Experimental results demonstrate that for various large-scale DNNs, MOVie can improve GPU memory utilization by up to 21.76% and achieves an average performance improvement of 1.5 times compared to the state-of-the-art methods.
Zijun Ma, Jiabao Tang, Songlei Jian, Yusong Tan
IJCNN5
2025 Hierarchical Neural Architecture Search for Fast and Accurate Depth Completion
Xiaogang Jia, Songlei Jian, Yusong Tan, Yonggang Che, Wei Chen 0009, Zhengfa Liang, Yu-Lin He
ICMR3
2025 Collaborative Data Aggregation Algorithm Considering User Preferences for Opportunistic Edge Computing
abstract
ABSTRACT In dynamic and unpredictable environments, opportunistic edge computing has emerged as a novel paradigm. It aims to leverage temporarily available computational resources to enable efficient data processing. To tackle the challenges of mobile group collaborative data aggregation in such contexts, this paper proposes a Multi‐Preference Redundant Data Collaborative Aggregation algorithm (MPRDCA), which jointly optimizes channel state and data preference weighting, and pioneers three key innovations: (a) a novel binary decision factor integrating dynamic user preferences with channel states, (b) a time‐varying utility function incorporating preference weights, and (c) Lyapunov‐based energy constraint transformation. Performance was rigorously evaluated via parameterized stochastic simulations incorporating probabilistic channel states and hierarchical preference coefficients, modeling a dynamic edge group. Key results demonstrate that under stringent energy constraints, MPRDCA achieves higher overall transmission utility while enhancing effective transmission volume of the most urgently required data by up to 41.99% compared to benchmarks. Theoretical and empirical results validate MPRDCA as a highly adaptive solution for resource‐constrained edge scenarios, achieving optimal utility‐efficiency trade‐offs.
Maojia Wang, Wenhua Xiao, Bixin Liu, Yusong Tan
Concurr. Comput. Pract. Exp.4
2025 Promoting Resource Utilization in HPC via Scheduling
abstract
ABSTRACT The increasing complexity of supercomputing workloads poses challenges to efficient resource management, especially in balancing computational and I/O demands. Shared burst buffers, as high‐speed intermediate storage, offer a promising avenue to mitigate I/O bottlenecks. However, naive job scheduling strategies often neglect the potential of burst buffers, relying on heuristic methods with limited adaptability. To address this, we propose an innovative burst‐buffer‐aware scheduling framework that integrates burst buffer capacity into job scheduling as a key resource. Through multi‐objective optimization, the framework intelligently balances trade‐offs among job waiting time, slowdown, and completion time, surpassing the rigidity of conventional approaches. Leveraging real‐world workload traces, the framework dynamically adapts to varying windows and job characteristics, combining computation and burst buffer demands to optimize scheduling decisions. Experimental results reveal that the proposed framework enhances scheduling efficiency and system adaptability, establishing a smarter and more effective approach to supercomputing job scheduling. This work underscores the importance of burst‐buffer‐aware strategies in advancing high‐performance computing, offering novel insights into intelligent resource management.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang, Bao Li 0002
Concurr. Comput. Pract. Exp.2
2025 SNCD: A fast and scalable distributed near-miss code clone detector for big code based on partial index
Rulin Xie, Yi Ren 0008, Jianbo Guan, Bao Li 0002, Jun Ma 0015, Yusong Tan
Future Gener. Comput. Syst.8
2024 MACA: Memory-aware convolution accelerating for CNN inference on edge devices
abstract
Deep learning inference tasks develop towards the edge due to their latency requirements and privacy issues. However, edge devices are limited by their power consumption and size, and generally have limited resources. The convolutional neural networks (CNN) is commonly used in image processing tasks which contains a large number of convolutional layers, accounting for more than 95% of the calculation time in most general used CNN model. This paper proposes a general convolutional layer optimization method called MACA, we implement and optimize a variety of convolution operators and design a memory-aware convolution operator automatic selection strategy to select appropriate operator, without modifying user code. Finally, we integrate MACA into PyTorch and conduct extensive experiments. The results show that when memory resources are sufficient, MACA can effectively increase the inference by 36.10% on average, and can reduce memory usage by an average of 29.14% to complete inference when resources are tight. This paper provides an effective solution for deploying deep learning models on resource-constrained edge devices.
Chaoxiong Yi, Songlei Jian, Yusong Tan, Yusen Zhang 0007
CSCWD3
2024 Examining SSD-correlated Failures within Racks in Production Data Centers
abstract
With flash-based solid-state drives (SSDs) becoming the primary storage medium in data centers, SSD failures are emerging as a significant factor affecting the reliability of data center storage systems. However, our understanding of the temporal and spatial distribution of field SSD failures remains limited, constraining the effectiveness of redundancy protection schemes in data center storage. To achieve high-quality rack-level fault tolerance, we conduct an in-depth analysis of correlated failures among 100K SSDs in Alibaba’s data center. The SSD-hosted applications exert a constant influence on the operation of the data center. We categorize intra-rack failures into heterogeneous application (hete-app) and homogeneous application (homo-app) failures based on the applications. We conduct a thorough investigation into the spatiotemporal trends of these two types of in-rack failures, exploring their causes and assessing the feasibility of rearranging racks to mitigate the occurrence of intra-rack failure chains. Furthermore, we employ a trace-driven simulator to verify the impact of different redundancy schemes on the reliability in clusters of hete-app racks and homo-app racks under high-failure-percentage environments.
Gang Xian, Yusong Tan, Jie Yu 0006, Wenxiang Yang
HPCC2
2024 DGLP: Incorporating Orientation Information for Enhanced Link Prediction in Directed Graphs
abstract
Link prediction in directed graphs offers a solution for uncovering detailed and accurate relationships among distinct entities. Unlike conventional link prediction in undirected graphs, the task becomes more intricate in directed graphs as it involves predicting both associations and orientations. Existing methods simply apply classic graph embedding techniques to learn node representations, followed by mapping representations of corresponding node pairs into probabilities indicating potential links. However, the inadequate capture of orientation information and sole reliance on node representations for prediction hinder the effective differentiation of orientation, thereby impeding the link prediction accuracy. In response, we introduce DGLP, an orientation-aware link prediction method tailored for directed graphs. DGLP utilizes the incidence matrix to learn both node and edge representations, effectively capturing structural and orientation information. By leveraging edge representations, DGLP achieves accurate link prediction in directed graphs without relying solely on implicit node representations. Experiments across six datasets demonstrate the effectiveness of DGLP, achieving a 1.2x improvement in prediction results.
Yusen Zhang 0007, Yusong Tan, Songlei Jian, Qingbo Wu 0003, Kenli Li 0001
ICASSP2
2024 Don't Turn a Blind Eye to Localization Noise: Localization Pseudo-label Correction and Learning for Semi-Supervised Object Detection
abstract
Pseudo-labeling has proven to be a simple yet effective technique for semi-supervised object detection (SSOD). However, the inevitable noise problem in pseudo-labels seriously hinders SSOD methods. Existing methods primarily focus on classification noise, while the specific and non-negligible localization noise remains not well-addressed. This paper analyzes the localization noise arising from the alternating learning and generation phases. For the generation phase, we innovatively explore the self-correction ability of models, stepping beyond the simple pseudo-label selection. We propose a localization pseudo-label correction (LPC) strategy to self-correct pseudo boxes and enhance prediction stability. In the learning phase, we propose a noisy localization loss (NLL) to enlarge the penalty of inconsistent predictions, thereby improving localization accuracy. Applied to two classic SSOD methods (Soft Teacher and Unbiased Teacher) and a recent state-of-the-art method (PseCo), our approach consistently improves accuracy across all of them.
Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Ke Liang 0006, Yusong Tan, Yulan Guo
ICME5
2024 Sniffing Threatening Open-World Objects in Autonomous Driving by Open-Vocabulary Models
abstract
Autonomous driving (AD) is a typical application that requires effectively exploiting multimedia information. For AD, it is critical to ensure safety by detecting unknown objects in an open world, driving the demand for open world object detection (OWOD). However, existing OWOD methods treat generic objects beyond known classes in the train set as unknown objects and prioritize recall in evaluation. This encourages excessive false positives and endangers safety of AD. To address this issue, we restrict the definition of unknown objects to threatening objects in AD, and introduce a new evaluation protocol, which is built upon a new metric named U-ARecall, to alleviate biased evaluation caused by neglecting false positives. Under the new evaluation protocol, we re-evaluate existing OWOD methods and discover that they typically perform poorly in AD. Then, we propose a novel OWOD paradigm for AD based on fine-tuning foundational open-vocabulary models (OVMs), as they can exploit rich linguistic and visual prior knowledge for OWOD. Following this new paradigm, we propose a brand-new OWOD solution, which effectively addresses two core challenges of fine-tuning OVMs via two novel techniques: 1) the maintenance of open-world generic knowledge by a dual-branch architecture; 2) the acquisition of scenario-specific knowledge by the visual-oriented contrastive learning scheme. Besides, a dual-branch prediction fusion module is proposed to avoid post-processing and hand-crafted heuristics. Extensive experiments show that our proposed method not only surpasses classic OWOD methods in unknown object detection by a large margin (∼× U-ARecall), but also notably outperforms OVMs without fine-tuning in known object detection (∼ 20% K-mAP). Our codes are available at https://github.com/harrylin-hyl/AD-OWOD.
Yu-Lin He, Siqi Wang 0001, Wei Chen 0009, Tianci Xun, Yusong Tan
ACM Multimedia5
2024 DOME: Dynamic Optimization of GPU Memory for Training DNN Models with Limited Resources
abstract
To improve data processing and learning capabilities, DNN models are developing towards increasingly complex structures and a huge number of parameters. However, the limited memory capacity of specialized acceleration chips such as GPUs has become a critical constraint in exploring advanced DNN models. Memory swapping and recomputation are two effective GPU memory optimization techniques that enable the training of larger models in memory-constrained hardware environments. Existing methods mostly require prior knowledge of the complete computational graph, which is not suitable for the dynamic computational graph mode adopted by mainstream deep learning frameworks. This paper proposes DOME, a tensor-level GPU memory optimization method for dynamic computational graphs. DOME combines tensor swapping and recomputation, inspired by cache replacement algorithms, to release the least costly tensors when GPU memory is insufficient, thereby supporting larger batch sizes during model training. DOME employs a heuristic evaluation function with higher training efficiency, estimating the cost of tensor swapping or recomputation using information obtained online at runtime. Experiments demonstrate that, for different DNN models, DOME can increase the maximum training batch size by 1.56 to 2.33 times. Additionally, DOME can improve training throughput by 9.8% to 35.0% compared to SOTA methods.
Zijun Ma, Songlei Jian, Yusong Tan
MSN3
2024 HMO: Host Memory Optimization for Model Inference Acceleration on Edge Devices
abstract
Deep learning (DL) is characterized by its demanding computational and memory requirements, which creates a significant challenge when deploying on edge devices. These devices often have limited computational capabilities and constrained resources. Most existing methods primarily focus on model-level techniques, such as model pruning or parameter quantization, to reduce model size and computation for accelerating inference. Considering the prevalent programming paradigms in DL, we propose a host memory optimization method, namely HMO, which can be integrated into DL programming framework, e.g., PyTorch, to improve the inference efficiency of DL models without modifying any model code. We particularly focus on memory optimization for intermediate variables in inference, aiming to enhance inference speed while maintaining a lower memory footprint. HMO involves a single profiling of inference to gather memory statistics about intermediate variables. These statistics are then used to guide subsequent inference. Additionally, we incorporate huge pages in operating systems to improve the memory access performance of HMO. Our experimental results show that HMO can achieve an average inference latency optimization ratio of 20.13 % compared with native PyTorch on six typical DL image representation models while effectively managing memory usage. Importantly, this is achieved without compromising model accuracy.
Chaoxiong Yi, Songlei Jian, Yusong Tan, Yusen Zhang 0007
SMC3
2024 OnceNAS: Discovering efficient on-device inference neural networks for edge devices
Yusen Zhang 0007, Yunchuan Qin, Yufeng Zhang 0001, Xu Zhou 0001, Songlei Jian, Yusong Tan, Kenli Li 0001
Inf. Sci.6
2024 Mobilizing underutilized storage nodes via job path: A job-aware file striping approach
Gang Xian, Wenxiang Yang, Yusong Tan, Jinghua Feng, Jie Yu 0006
Parallel Comput.3
2024 Exploring nonintrusive measurements of spatio-temporal portrait of microservices
abstract
Abstract As cloud native technology advances, the scale and complexity of applications built on microservice architecture continue to expand, leading to increasingly intricate differences between software within the same application. Microservice applications, offering high flexibility, are deployed in data centers as black boxes from the users' perspective, leaving them with no insight into the orchestration of cloud service providers. Consequently, users face challenges in promptly recognizing performance imbalances within their deployed applications. Meanwhile, cloud service providers may cut costs by offering a mix of qualified and unqualified services, potentially deceiving users. To enhance the understanding of microservice application organization, we propose a non‐intrusive measurement framework, termed NMPI. NMPI facilitates rapid identification of microservice application defects, offering insights into cloud services and detecting fraudulent behavior in microservice‐based applications. We model microservice applications using a queue analysis‐based approach and filter the dominant frequency components of average response time signals by employing k‐means on the fast fourier transform (FFT). Our model constructs a library of performance portraits for various software, with these portraits resembling human fingerprints that carry and mark the software's internal information. Utilizing a two‐tier microservices‐based application incorporating a database as a case study allows us to demonstrate the effectiveness of NMPI. Our experimental results show that NMPI can produce differentiable profiles of data service performance portraits across a diverse and extensive range of workloads, enabling the identification of software types and the analysis of performance conditions.
Zichen Xu 0001, Dan Wu 0010, Xiaoling Li 0002, Biyong Liu, Haichuan Hu, Shuang Tan, Yusong Tan, Chenren Xu, Christopher Stewart, Qihe Zhou
Softw. Pract. Exp.8
2023 Adversarial Learning-Based Stance Classifier for COVID-19-Related Health Policies
Feng Xie 0003, Xuechen Zhao, Jiaying Zou, Bin Zhou 0004, Yusong Tan
DASFAA (4)8
2023 Domain Generalized Fundus Image Segmentation via Dual-Level Mixing
abstract
Single domain generalization plays a vital role in signal processing tasks, which is capable of extracting domain-invariant knowledge from a single source domain such that the learned model can be generalizable to unseen domains. However, many existing methods require the help of auxiliary tasks, bringing extra computation cost. Aiming at this pitfall, this study proposes Dual-Level Mixing (DLM) to boost the diversity of the single source domain and enhance the generalization performance. Specifically, at the input level, we apply different image augmentations to get different variants of the same image. Patches from augmented images are spatially mixed to get a perturbed image as the input, which enlarges the scale of training data and boosts the diversity. Meanwhile, at the feature level, we characterize feature statistics as Gaussian distributions. Then, we resample the feature statistics to renormalize the latent features, which simulates the potential style variance of different domains. In summary, the proposed DLM method synergies input-level and feature-level mixing strategies, leading to enhanced generalization performance. Experimental results of cross-domain fundus image segmentation demonstrate that either input-level mixing or feature-level mixing can effectively promote the performance of domain generalization. Moreover, the collaboration of dual-level mixing strategies leads to superior or comparable performance to domain adaptation counterparts that rely on target data.
Xin Luo 0009, Wei Chen 0009, Chen Li 0034, Bin Zhou 0004, Yusong Tan
ICASSP5
2023 Improving Knowledge Graph Entity Alignment with Graph Augmentation
Feng Xie 0003, Bin Zhou 0004, Yusong Tan
PAKDD (2)4
2023 Robust unsupervised network intrusion detection with self-supervised masked context reconstruction
Wei Wang 0130, Songlei Jian, Yusong Tan, Qingbo Wu 0003, Chenlin Huang
Comput. Secur.3
2023 Adversarial style discrepancy minimization for unsupervised domain adaptation
abstract
Mainstream unsupervised domain adaptation (UDA) methods align feature distributions across different domains via adversarial learning. However, most of them focus on global distribution alignment, ignoring the fine-grained domain discrepancy. Besides, they generally require auxiliary models, bringing extra computation costs. To tackle these issues, this study proposes an UDA method that differentiates individual samples without the help of extra models. To this end, we introduce a novel discrepancy metric, termed style discrepancy, to distinguish different target samples. We also propose a paradigm for adversarial style discrepancy minimization (ASDM). Specifically, we fix the parameters of the feature extractor and maximize style discrepancy to update the classifier, which helps detect more hard samples. Adversely, we fix the parameters of the classifier and minimize the style discrepancy to update the feature extractor, pushing those hard samples near the support of the source distribution. Such adversary helps to progressively detect and adapt more hard samples, leading to fine-grained domain adaptation. Experiments on different UDA tasks validate the effectiveness of ASDM. Overall, without any extra models, ASDM reaches a 46.9% mIoU in the GTA5 to Cityscapes benchmark and an 84.7% accuracy in the VisDA-2017 benchmark, outperforming many existing adversarial-learning-based methods.
Xin Luo 0009, Wei Chen 0009, Zhengfa Liang, Chen Li 0034, Yusong Tan
Neural Networks5
2022 Adaptive Pseudo Labeling for Source-Free Domain Adaptation in Medical Image Segmentation
abstract
Domain adaptation is common but challenging in signal processing tasks due to the intrinsic discrepancy, especially in difficult-to-label medical image segmentation application scenarios. Pseudo labeling methods are widely utilized to compensate for the scarcity of annotation. However, most existing methods set the fixed thresholds to select highly-confident predictions as pseudo labels, inevitably generating false labels with noise. In this paper, we combine the dual-classifiers consistency and predictive category-aware confidence to form a novel regularization for pseudo-label denoising. The dual-classifiers consistency helps promote the robustness of pseudo labels. Meanwhile, category-aware confidence is utilized as adaptive pixel-wise weights, avoiding the need for handcrafted thresholds. The adapted model is refined by the rectified pseudo labels without source domain samples. The proposed method is model-independent and thus can be plug-and-play to improve existing UDA methods. We validated it on the cross-modality medical image segmentation and obtained more competitive results.
Chen Li 0034, Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Yusong Tan
ICASSP5
2022 SCORE: A Resource-Efficient Microservice Orchestration Model Based on Spectral Clustering in Edge Computing
Yusong Tan, Bao Li 0002
ICSOC2
2022 Evaluation Ranking is More Important for NAS
abstract
Search space, searching method, and candidate evaluation scheme are critical to the success of Neural Architecture Search (NAS), especially the evaluation strategy. An effective and efficient neural architecture performance evaluator could successfully save computing costs and search time while guiding the NAS process to the optimal solution as fast as possible. Most existing NAS algorithms attempt to compute the absolute accuracy of the candidate architecture, which is almost impossible to achieve and meaningless to the final performance improvement. In this paper, we propose ERNAS, a novel neural architecture performance evaluation approach that optimizes the ranking of the candidate architecture performance, rather than the absolute accuracy itself. With the help of ERNAS, many existing NAS methods could achieve better performance without further evaluation. The experimental results demonstrate that ERNAS can be trained effectively enough with extremely limited training data (423 neural architectures randomly sampled form NAS-Bench-101, which is only 0.1% of the entire search space). The accuracy of the neural architecture search result produced by ERNAS is greater than that of the SOTA methods.
Yusen Zhang 0007, Bao Li 0002, Yusong Tan, Songlei Jian
IJCNN3
2022 ProxyDWRR: A Dynamic Load Balancing Approach for Heterogeneous-CPU Kubernetes Clusters
abstract
Edge computing is booming as a promising paradigm to push the service and computation resources from the cloud to the edge of network. As the de-facto standard for container orchestration, Kubernetes is more and more widely used not only in cloud computing but also in edge computing. However, Kubernetes is designed for homogenous cloud data centers, and it does not take into account heterogeneous scenarios, which is ubiquitous is the edge. This will lead to load imbalance among containers with its default rough load balancing mechanism. To deal with this problem, we firstly propose a Dynamically Weighted Random Routing (DWRR) algorithm based on the default random algorithm in Kubernetes. Besides, we design and implement ProxyDWRR, a load balancing plugin for the Kubernetes cluster with heterogeneous CPU. It is fully compatible with the existing load balancing mechanism in Kubernetes. We validated our solution based on a cloud-native microservices application. The experimental results show that ProxyDWRR can effectively balance the load between containers in clusters with heterogeneous CPU. In our experiments, DWRR can improve the CPU utilization of the containers by about 25% and the throughput of the application by about 22.6% compared to the default load balancing algorithms, which enables the cluster to evacuate bursty load more effectively.
Qingkun Wang, Yi Ren 0008, Saqing Yang, Jianbo Guan, Bao Li 0002, Yusong Tan
JCC7
2022 Trusted-Committee- Based Secure and Scalable BFT Consensus for Consortium Blockchain
abstract
Compared with public blockchain, consortium blockchain is more secure and controllable deployed in an enterprise scenario. Byzantine fault tolerance (BFT) consensus is widely applied in consortium blockchain. Although PBFT is the most classic practical BFT consensus with message complexity O(n2), it still faces some security threats and has low consensus efficiency. To address these issues, we propose a secure and trusted BFT (S2BFT) consensus based on trusted committees. S2BFT generates a trusted anonymous number using trust execution environment (TEE) for each server node and selects committees by pseudo-random algorithm. S2BFT can efficiently reach consensus by the committees with an O(m*n) message complexity. In addition, correctness analysis proves that S2BFT can resist more attacks than traditional BFT consensus and tolerate 1/2 byzantine server nodes. Results further demonstrate the efficiency of the simulated S2BFT implementation.
Liaoliao Feng, Yusong Tan, Xiang Fu 0002, Keming Wang, Junsheng Chang
MSN3
2022 EpiGNN: Exploring Spatial Transmission with Graph Neural Network for Regional Epidemic Forecasting
Feng Xie 0003, Bin Zhou 0004, Yusong Tan
ECML/PKDD (6)5
2022 Inter- and Intra-Series Embeddings Fusion Network for Epidemiological Forecasting
abstract
The accurate forecasting of infectious epidemic diseases is the key to effective control of the epidemic situation in a region.Most existing methods ignore potential dynamic dependencies between regions or the importance of temporal dependencies and inter-dependencies between regions for prediction.In this paper, we propose an Interand Intra-Series Embeddings Fusion Network (SEFNet) to improve epidemic prediction performance.SEFNet consists of two parallel modules, named Inter-Series Embedding Module and Intra-Series Embedding Module.In Inter-Series Embedding Module, a multiscale unified convolution component called Region-Aware Convolution is proposed, which cooperates with self-attention to capture dynamic dependencies between time series obtained from multiple regions.The Intra-Series Embedding Module uses Long Short-Term Memory to capture temporal relationships within each time series.Subsequently, we learn the influence degree of two embeddings and fuse them with the parametric-matrix fusion method.To further improve the robustness, SEFNet also integrates a traditional autoregressive component in parallel with nonlinear neural networks.Experiments on four real-world epidemic-related datasets show SEFNet is effective and outperforms state-of-the-art baselines.
Feng Xie 0003, Xuechen Zhao, Bin Zhou 0004, Yusong Tan
SEKE5
2022 Fine-tuning more stable neural text classifiers for defending word level adversarial attacks
Zibo Yi, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003
Appl. Intell.3
2022 Representation learning-based network intrusion detection system by capturing explicit and implicit feature interactions
Wei Wang 0130, Songlei Jian, Yusong Tan, Qingbo Wu 0003, Chenlin Huang
Comput. Secur.3
2022 Towards an Efficient and Robust Adversarial Attack Against Neural Text Classifier
abstract
Adversarial attack is a serious threat to neural network-based natural language processing applications. Adversarial attack uses tiny well-crafted perturbations to mislead neural networks. While existing adversarial text attacks can achieve good attack effects, they still do not guarantee efficiency and robustness. The adversarial text attacks are more efficient if they use less perturbation to achieve a higher attack success rate. The attacks are more robust if they can achieve a higher success rate when defense strategies are applied. To improve the efficiency and robustness of the adversarial attack, we propose SMAL: Saliency Map Attack with Levenshtein-similarity. The proposed attack consists of two parts: (1) The saliency map measures the perturbation priority of each word. It considers not only the influence of each word on the classification result but also how to maintain the misled classification result to improve the robustness of the attack. (2) Levenshtein-similarity network embeds words into edit distance space. When perturbing sentences, some words are replaced by substitutions with less edit distance. This can reduce the amount of modification, which improves the efficiency of the attack. Since the words are embedded in edit distance space rather than semantic space, the semantic-based defense is not effective for this attack, which improves the robustness. The experiments show that SMAL achieves a higher attack success rate with fewer perturbations. Also, the proposed attack is better when attacking a classifier defended by adversarial training.
Zibo Yi, Shasha Li 0001, Jun Ma 0015, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003
Int. J. Pattern Recognit. Artif. Intell.5
2021 Tri-Directional Tasks Complementary Learning for Unsupervised Domain Adaptation of Cross-modality Medical Image Semantic Segmentation
abstract
Cross-modality adaptation is challenging due to the internal domain discrepancy in appearance and representation. When the trained model of source domain is transferred to the target domain, the domain shift will reduce accuracy. Meanwhile, unsupervised domain adaptation has the potential to recover this degradation among medical images of different modalities, so it is of clinical significance and meaningful in bioinformatics. However, previous related works usually try to align domains in a single direction or two directions, failing to take advantage of the complementary relationship between different directions and alignment tasks. In this paper, we propose the Tri-directional learning framework to solve domain shift in the task of medical image semantic segmentation. The proposed framework is able to synergize image style transformation, mask segmentation and edge segmentation. The above three tasks are mutually boosted through complementary training in each iteration. In this way, our method performs cross-modality medical images semantic segmentation from labeled source domain (MRI) to unlabeled target domain (CT). The experimental results demonstrate the effectiveness of the proposed method. For the task of cardiac structure segmentation from cross-modality medical images, our proposed framework achieves state-of-the-art performance. The code is available at https://github.com/lichen14/TriDL.
Chen Li 0034, Wei Chen 0009, Mingfei Wu, Xin Luo 0009, Yu-Lin He, Yusong Tan
BIBM6
2021 AttENT: Domain-Adaptive Medical Image Segmentation via Attention-Aware Translation and Adversarial Entropy Minimization
abstract
Due to the intrinsic domain shift among different modalities, it is nontrivial to directly apply a well-trained model into other cross-modality medical images. Unsupervised domain adaptation (UDA) has the potential to reduce such domain shift. However, existing UDA methods try to align domains in either image level or in feature level, failing to consider the unified relationship between cross-modality images and their corresponding features. In this paper, we propose a novel UDA framework for domain adaptive medical image segmentation. The proposed framework synergizes both pixel space and entropy space for domain alignment. Specifically, in the pixel space, we introduce the attention mechanism into CycleGAN, and enhance the semantic and geometric consistency of the target organs during the image style transformation. In entropy space, we utilize entropy minimization principle to force consistent image segmentation between well-annotated source domain and non-annotated target domain. The aligned ensemble of two representation spaces enables a well-trained segmentation model to effectively transfer from source domain to target domain. The experimental results demonstrate the effectiveness of the proposed method. For the task of multi-organs segmentation from cross-modality medical images, our proposed framework achieves state-of-the-art performance, with some specific metric even superior to those of supervised methods. The code is available at https://github.com/lichen14/AttENT.
Chen Li 0034, Xin Luo 0009, Wei Chen 0009, Yu-Lin He, Mingfei Wu, Yusong Tan
BIBM6
2021 Multi-Scale Cascade Disparity Refinement Stereo Network
abstract
Stereo matching has attracted much attention in recent years. Traditional methods can quickly generate a disparity result, but the accuracy is low. On the contrary, methods based on neural networks can achieve a high accuracy level, but they are difficult to reach the real-time level. Therefore, this paper presents MCDRNet, which combines traditional methods with neural networks to achieve real-time and accurate stereo matching results. Concretely, our network first generates a rough disparity map based on the traditional ADCensus algorithm. Then we design a novel Multi-Scale Cascade Network to refine the disparity map from coarse to fine. We evaluate our best-trained model on the KITTI official website. The results show that our network is much faster than most current top-performing methods(31×than CSPN, 56×than GANet, etc.). Meanwhile, it is more accurate than traditional stereo methods(SGM, SPS-St) and other fast 2D convolution networks(Fast DS-CS, DispNetC, etc.), demonstrating the rationalities and feasibilities of our method.
Xiaogang Jia, Wei Chen 0009, Zhengfa Liang, Xin Luo 0009, Mingfei Wu, Yusong Tan, Libo Huang 0002
ICASSP6
2021 Many-To-Many Chinese ICD-9 Terminology Standardization Based on Neural Networks
Shasha Li 0001, Jie Yu 0008, Yusong Tan, Jun Ma 0015, Qingbo Wu 0003
ICIC (2)4
2021 Span Representation Generation Method in Entity-Relation Joint Extraction
Yongtao Tang, Jie Yu 0008, Shasha Li 0001, Bin Ji 0002, Yusong Tan, Qingbo Wu 0003
ICIC (2)5
2021 Multi-Scale Cost Volumes Cascade Network for Stereo Matching
abstract
Stereo matching is essential for robot navigation. However, the accuracy of current widely used traditional methods is low, while methods based on CNN need expensive computational cost and running time. This is because different cost volumes play a crucial role in balancing speed and accuracy. Thus we propose MSCVNet, which combines traditional methods and neural networks to improve the quality of cost volume. Concretely, our network first generates multiple 3D cost volumes with different resolutions and then uses 2D convolutions to construct a novel cascade hourglass network for cost aggregation. Meanwhile, we design an algorithm to distinguish and calculate the loss for discontinuous areas of the disparity result. According to the KITTI official website, our network is much faster than most top-performing methods (24than CSPN, 44than GANet, etc.). Meanwhile, compared to traditional methods (SPS-St, SGM) and other real-time stereo matching networks (Fast DS-CS, DispNetC, and RTSNet, etc.), our network achieves a big improvement in accuracy, demonstrating the feasibility and capability of the proposed method.
Xiaogang Jia, Wei Chen 0009, Chen Li 0034, Zhengfa Liang, Mingfei Wu, Yusong Tan, Libo Huang 0002
ICRA6
2021 Fast and Accurate Lane Detection via Frequency Domain Learning
abstract
It is desirable to maintain both high accuracy and runtime efficiency in lane detection. State-of-the-art methods mainly address the efficiency problem by direct compression of high-dimensional features. These methods usually suffer from information loss and cannot achieve satisfactory accuracy performance. To ensure the diversity of features and subsequently maintain information as much as possible, we introduce multi-frequency analysis into lane detection. Specifically, we propose a multi-spectral feature compressor (MSFC) based on two-dimensional (2D) discrete cosine transform (DCT) to compress features while preserving diversity information. We group features and associate each group with an individual frequency component, which incurs only 1/7 overhead of one-dimensional convolution operation but preserves more information. Moreover, to further enhance the discriminability of features, we design a multi-spectral lane feature aggregator (MSFA) based on one-dimensional (1D) DCT to aggregate features from each lane according to their corresponding frequency components. The proposed method outperforms the state-of-the-art methods (including LaneATT and UFLD) on TuSimple, CULane, and LLAMAS benchmarks. For example, our method achieves 76.32% F1 at 237 FPS and 76.98% F1 at 164 FPS on CULane, which is 1.23% and 0.30% higher than LaneATT. Our code and models are available at https://github.com/harrylin-hyl/MSLD.
Yu-Lin He, Wei Chen 0009, Zhengfa Liang, Dan Chen 0001, Yusong Tan, Xin Luo 0009, Chen Li 0034, Yulan Guo
ACM Multimedia5
2021 FastDCF: A Partial Index Based Distributed and Scalable Near-Miss Code Clone Detection Approach for Very Large Code Repositories
Yi Ren 0008, Jianbo Guan, Bao Li 0002, Jun Ma 0015, Yusong Tan
PDCAT7
2021 On-demand cut off the covert channel to mitigate meltdown
Yusong Tan, Baozi Chen, Liehuang Zhu, Qingbo Wu 0003, Yuanzhang Li 0001
Sci. China Inf. Sci.1
2021 Toward security as a service: A trusted cloud service architecture with policy customization
Chenlin Huang, Wei Chen 0009, Songlei Jian, Yusong Tan, Dan Chen 0001
J. Parallel Distributed Comput.6
2020 Span-based Joint Entity and Relation Extraction with Attention-based Span-specific and Contextual Semantic Representations
abstract
Span-based joint extraction models have shown their efficiency on entity recognition and relation extraction.These models regard text spans as candidate entities and span tuples as candidate relation tuples.Span semantic representations are shared in both entity recognition and relation extraction, while existing models cannot well capture semantics of these candidate entities and relations.To address these problems, we introduce a span-based joint extraction framework with attention-based semantic representations.Specially, attentions are utilized to calculate semantic representations, including span-specific and contextual ones.We further investigate effects of four attention variants in generating contextual semantic representations.Experiments show that our model outperforms previous systems and achieves state-of-the-art results on ACE2005, CoNLL2004 and ADE.
Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003
COLING6
2020 Dynamic Co-located VM Detection and Membership Update for Residency Aware Inter-VM Communication in Virtualized Clouds
Yi Ren 0008, Jianbo Guan, Ziqi You, Saqing Yang, Yusong Tan
ICA3PP (3)6
2020 Attention Unet++: A Nested Attention-Aware U-Net for Liver CT Image Segmentation
abstract
Liver cancer is one of the cancers with the highest mortality. In order to help doctors diagnose and treat liver lesion, an automatic liver segmentation model is urgently needed due to manually segmentation is time-consuming and error-prone. In this paper, we propose a nested attention-aware segmentation network, named Attention UNet++. Our proposed method has a deep supervised encoder-decoder architecture and a redesigned dense skip connection. Attention UNet++ introduces attention mechanism between nested convolutional blocks so that the features extracted at different levels can be merged with a task-related selection. Besides, due to the introduction of deep supervision, the prediction speed of the pruned network is accelerated at the cost of modest performance degradation. We evaluated proposed model on MICCAI 2017 Liver Tumor Segmentation (LiTS) Challenge Dataset. Attention UNet++ achieved very competitive performance for liver segmentation.
Chen Li 0034, Yusong Tan, Wei Chen 0009, Xin Luo 0009, Yuanming Gao, Xiaogang Jia, Zhiying Wang 0003
ICIP2
2020 ANU-Net: Attention-based nested U-Net to exploit full resolution features for medical image segmentation
Chen Li 0034, Yusong Tan, Wei Chen 0009, Xin Luo 0009, Yu-Lin He, Yuanming Gao
Comput. Graph.2
2020 Research on Chinese medical named entity recognition based on collaborative cooperation of multiple neural network models
Bin Ji 0002, Shasha Li 0001, Jie Yu 0008, Jun Ma 0015, Jintao Tang, Qingbo Wu 0003, Yusong Tan, Huijun Liu 0003, Yun Ji
J. Biomed. Informatics7
2020 SLR-SELinux: Enhancing the Security Footstone of SEAndroid with Security Label Randomization
abstract
The root privilege escalation attack is extremely destructive to the security of the Android system. SEAndroid implements mandatory access control to the system through the SELinux security policy at the kernel mode, making the general root privilege escalation attacks unenforceable. However, malicious attackers can exploit the Linux kernel vulnerability of privilege escalation to modify the SELinux security labels of the process arbitrarily to obtain the desired permissions and undermine system security. Therefore, investigating the protection method of the security labels in the SELinux kernel is urgent. And the impact on the existing security configuration of the system must also be reduced. This paper proposes an optimization scheme of the SELinux mechanism based on security label randomization to solve the aforementioned problem. At the system runtime, the system randomizes the mapping of the security labels inside and outside the kernel to protect the privileged security labels of the system from illegal obtainment and tampering by attackers. This method is transparent to users; therefore, users do not need to modify the existing system security configuration. A tamper-proof detection method of SELinux security label is also proposed to further improve the security of the method. It detects and corrects the malicious tampering behaviors of the security label in the critical process of the system timely. The above methods are implemented in the Linux system, and the effectiveness of security defense is proven through theoretical analysis and experimental verification. Numerous experiments show that the effect of this method on system performance is less than 1%, and the success probability of root privilege escalation attack is less than 10−9.
Pan Dong, Yusong Tan, Chenlin Huang, Lifeng Wei, Yudan Zuo
Wirel. Commun. Mob. Comput.4
2019 CLASC: A Changelog Based Automatic Code Source Classification Method for Operating System Packages
abstract
Open source represents an important way in which today's software is developed. The adoption of open source software continues to accelerate because of the great potential it offers, such as productivity improvement, cost savings and quicker innovation. While the complexity and the size of software composition grow, it becomes difficult to effectively scan and track the code source, especially for software with tremendous scale of code, such as operating systems. So far, existing work on open source components mainly focus on how to mitigate potential license incompliance, to reduce potential security risks introduced by open source vulnerabilities, and to detect and match open source components in the code. To ensure code traceability and manageability for large scale mixed-source operating system, we believe it is beneficial to automatically distinguish sources of the system code in the granularity of software packages and manage them separately. However, according to the literature, there is a lack of relevant work in this area. In this paper, we first classify the packages into three categories in terms of code source from the perspective of OS developers and maintainers. Then we propose CLASC, an efficient code source classification algorithm. With the capability of package info extraction and analysis, CLASC can classify software packages into the defined categories according to their changelog info. And we design and implement KyAnalyzer, a Web-based package management and code source analysis platform. It provides automatic code source analyzing services and is capable of managing OS packages differentially according to their different categories of code source with CLASC incorporated as a component of it. Experimental results show the correctness and efficiency of the Web-enabled package source classifier.
Yi Ren 0008, Jianbo Guan, Jun Ma 0015, Yusong Tan, Qingbo Wu 0003
APSEC4
2019 Incremental Learning of GAN for Detecting Multiple Adversarial Attacks
Zibo Yi, Jie Yu 0008, Shasha Li 0001, Yusong Tan, Qingbo Wu 0003
ICANN (3)4
2019 An Efficient and Transparent Approach for Adaptive Intra-and Inter-Node Virtual Machine Communication in Virtualized Clouds
abstract
Network I/O workloads are dominating as one of the leading costs for most of the virtualized clouds. One way to improve the inter virtual machine (VM) inefficiency is to build shared memory channels between VMs co-located on the same physical node to by-pass traditional TCP/IP network stack, so that the overhead is reduced by shorter communication path and fewer kernel interactions. However, it is a key challenge for existing work to achieve high performance inter-VM communication while keeping the capability of VM live migration, and most of existing work are neither seamlessly agile in the presence of VM live migration nor transparent to upper users as well as to operating system kernels, which limits the application of current co-location aware shared-memory based approaches. In this paper, we present the design and implementation of XenVMC, an adaptive and transparent inter-VM communication system for high performance network I/O in virtualized clouds. With proposed dynamic co-located VM membership update mechanism, XenVMC is applicable not only to intra-node VM communication, but also to cross-node communication. It also supports adaptive switching between shared-memory based channel and traditional network-based channel in case of VM live migration, with the aid of proposed VM migration perception and handling mechanisms. XenVMC enables efficient data transmission for both TCP and UDP workloads, with multilevel transparency guaranteed. Extensive experiments show that XenVMC achieves better performance for both TCP and UDP workloads with high transparency, compared with both native virtualized environment and representative existing work. Experimental results also show that it is capable of automatically handling VM migration correctly with acceptable latency.
Yi Ren 0008, Renshi Liu, Qi Zhang 0009, Jianbo Guan, Ziqi You, Yusong Tan, Qingbo Wu 0003
ICPADS6
2019 Multi-resource workload mapping with minimum cost in cloud environment
abstract
Summary Workload mapping in cloud environment refers to map multiple workloads provided by the cloud users/tenants to the substrate network provided by the cloud providers, which is a NP‐hard problem. The workload is a service demand made to the cloud, which is modeled as a logical network consists of virtual nodes and virtual links. Substrate network is a physical network consists of physical nodes that are inter‐connected via communication links. Devising heuristic methods has become the mainstream of workload mapping problem, which can obtain a feasible solution, but the quality of the solution is not guaranteed. Pointing to this issue, this paper takes the mapping cost of the workloads as the solving objective, and models the workload mapping as a constraint optimization problem. Based on the constraint optimization model, we devise two algorithms to solve the problem. These algorithms can not only find the feasible solution, but also ensure the solution is optimal. Lastly, we have demonstrated the optimality of the proposed algorithms through theoretical proof and evaluated the performance of them through simulation experiment.
Xiaoling Li 0002, Xiaoyong Li 0002, Yusong Tan, Shuang Tan
Concurr. Comput. Pract. Exp.3
2019 Resource stealing: a resource multiplexing method for mix workloads in cloud system
Yusong Tan, Fuhui Wu, Qingbo Wu 0003, Xiangke Liao
J. Supercomput.1
2018 A virtual cluster embedding approach by coordinating virtual network and software-defined network
Yusong Tan, Rongzhen Li, Qingbo Wu 0003
Soft Comput.1
2017 Drug-Drug Interaction Extraction via Recurrent Neural Network with Multiple Attention Layers
Zibo Yi, Shasha Li 0001, Jie Yu 0008, Yusong Tan, Qingbo Wu 0003, Ting Wang 0009
ADMA4
2017 MicRun: A framework for scale-free graph algorithms on SIMD architecture of the Xeon Phi
abstract
Graph algorithms currently play increasingly important roles, especially in social networks and language modeling scenarios. Recently, accelerating graph algorithms by heterogeneous high performance computers with the integrated cores and expanded SIMD lanes has been becoming the mainstream. However, the existing methods, restricted by the low-efficiency grouping strategy and the non-optimized selection mechanism of tile size of a graph, are far below our expectations in many ways. Moreover, there are few convenient integrated tools provided for deploying the graph algorithms on MIC architecture. In this paper, we propose a high-efficiency framework MicRun, which is flexible to be used for graph algorithms on SIMD architecture of the Xeon Phi. There are two key components in MicRun, the Bucket Grouping module and Auto-tuning module. In the Grouping module, an optimization algorithm is designed for splitting graph tiles into conflict-free groups, which can be directly processed on SIMD parallelism. In the Auto-tuning module, a novel strategy is proposed for optimizing the tile size to boost execution efficiency of the graph computation. MicRun currently supports Bellman-Ford and PageRank algorithms, we also conduct extensive validation experiments on MicRun. Experimental results show that MicRun outperforms existing mechanisms in terms of storage and time overhead. As a consequence, both graph algorithms achieve an average speedup of 1.1× by MicRun, compared with the state-of-the-art.
Qingbo Wu 0003, Yusong Tan, Jie Yu 0008, Qi Zhang 0028, Xiaoling Li 0002, Lei Luo 0002
ASAP3
2016 PCP-B2: Partial critical path budget balanced scheduling algorithms for scientific workflow applications
Fuhui Wu, Qingbo Wu 0003, Yusong Tan, Rongzhen Li, Wei Wang 0130
Future Gener. Comput. Syst.3
2016 micMR: An efficient MapReduce framework for CPU-MIC heterogeneous architecture
Wenzhu Wang, Yusong Tan, Qingbo Wu 0003, Yaoxue Zhang
J. Parallel Distributed Comput.2
2015 Optimizing the MapReduce Framework for CPU-MIC Heterogeneous Cluster
Wenzhu Wang, Qingbo Wu 0003, Yusong Tan, Yaoxue Zhang
APPT3
2015 Maximize Throughput Scheduling and Cost-Fairness Optimization for Multiple DAGs with Deadline Constraint
Wei Wang 0130, Qingbo Wu 0003, Yusong Tan, Fuhui Wu
ICA3PP (2)3
2015 Unified Multi-constraint and Multi-objective Workflow Scheduling for Cloud System
Fuhui Wu, Qingbo Wu 0003, Yusong Tan, Wei Wang 0130
ICA3PP (2)3
2015 An Evolutionary Algorithm with Double-Level Archives for Multiobjective Optimization
abstract
Existing multiobjective evolutionary algorithms (MOEAs) tackle a multiobjective problem either as a whole or as several decomposed single-objective sub-problems. Though the problem decomposition approach generally converges faster through optimizing all the sub-problems simultaneously, there are two issues not fully addressed, i.e., distribution of solutions often depends on a priori problem decomposition, and the lack of population diversity among sub-problems. In this paper, a MOEA with double-level archives is developed. The algorithm takes advantages of both the multiobjective-problem-level and the sub-problem-level approaches by introducing two types of archives, i.e., the global archive and the sub-archive. In each generation, self-reproduction with the global archive and cross-reproduction between the global archive and sub-archives both breed new individuals. The global archive and sub-archives communicate through cross-reproduction, and are updated using the reproduced individuals. Such a framework thus retains fast convergence, and at the same time handles solution distribution along Pareto front (PF) with scalability. To test the performance of the proposed algorithm, experiments are conducted on both the widely used benchmarks and a set of truly disconnected problems. The results verify that, compared with state-of-the-art MOEAs, the proposed algorithm offers competitive advantages in distance to the PF, solution coverage, and search speed.
Ni Chen, Weineng Chen, Yue-Jiao Gong, Zhi-hui Zhan, Jun Zhang 0003, Yun Li 0002, Yusong Tan
IEEE Trans. Cybern.7
2015 Workflow scheduling in cloud: a survey
Fuhui Wu, Qingbo Wu 0003, Yusong Tan
J. Supercomput.3
2013 A Vectorized K-Means Algorithm for Intel Many Integrated Core Architecture
Fuhui Wu, Qingbo Wu 0003, Yusong Tan, Lifeng Wei, Lisong Shao
APPT3
2013 Speeding Up SIFT Algorithm by Multi-core Processor Supporting SIMD Instruction Sets
abstract
Scale Invariant Feature Transform (SIFT) method plays a critical role in a wide variety of vision applications. But it is now facing the real-time computational challenge. Parallel computing is one of the most promising solutions to overcome the computational challenge. In this paper, we target at parallelizing SIFT by multi-core architecture with per-core SIMD support. We focus on the SIMDization of data parallel parts of SIFT to fully utilize per-core computing power. At Orientation Assignment and Key point Descriptor stages, we observe that load balance is an important factor. We also implement the optimized algorithm on multi-core system with SIMD support from Tianhe-2 Supercomputer and make comparison with the State-of-the-Art parallel SIFT algorithms.
Fuhui Wu, Qingbo Wu 0003, Yusong Tan
CAD/Graphics3
2013 Multi-resource Aware Congestion Control in Data Centers
abstract
Network has been widely reported as a bottleneck of data center applications. However, current researches of congestion control are unaware of multiple resources consuming and decrease all flows when congestion, ignoring some involved flows may not be the faults. In this paper, we propose a novel multi-resources aware congestion control framework MRTCP to provide a fine-grain control on flows when congestions appear. MRTCP exploits a multi-tuple vector model to measure multi-resources provision and consumption, and develops a novel metric RB (Resource Balance) to denote the heterogeneous amounts of resources employed by each flow. It analyzes which resources are being the bottlenecks that lead to congestions, calculates the responsibility of each flow to this congestion, and then adjusts their sending rates respectively. Our experiment results demonstrate that MRTCP is able to optimize network multi-resources utilization and improve network throughput without adding obvious packets delays.
Deke Guo, Qingbo Wu 0003, Shanshan Li 0001, Yusong Tan, Quanyuan Wu
ICPADS4
2011 FPGA accelerator for protein secondary structure prediction based on the GOR algorithm
abstract
BACKGROUND: Protein is an important molecule that performs a wide range of functions in biological systems. Recently, the protein folding attracts much more attention since the function of protein can be generally derived from its molecular structure. The GOR algorithm is one of the most successful computational methods and has been widely used as an efficient analysis tool to predict secondary structure from protein sequence. However, the execution time is still intolerable with the steep growth in protein database. Recently, FPGA chips have emerged as one promising application accelerator to accelerate bioinformatics algorithms by exploiting fine-grained custom design. RESULTS: In this paper, we propose a complete fine-grained parallel hardware implementation on FPGA to accelerate the GOR-IV package for 2D protein structure prediction. To improve computing efficiency, we partition the parameter table into small segments and access them in parallel. We aggressively exploit data reuse schemes to minimize the need for loading data from external memory. The whole computation structure is carefully pipelined to overlap the sequence loading, computing and back-writing operations as much as possible. We implemented a complete GOR desktop system based on an FPGA chip XC5VLX330. CONCLUSIONS: The experimental results show a speedup factor of more than 430x over the original GOR-IV version and 110x speedup over the optimized version with multi-thread SIMD implementation running on a PC platform with AMD Phenom 9650 Quad CPU for 2D protein structure prediction. However, the power consumption is only about 30% of that of current general-propose CPUs.
Fei Xia 0003, Yong Dou, Guo-Qing Lei, Yusong Tan
BMC Bioinform.4
2010 Vapor: Virtual Machine Based Parallel Program Profiling Framework
abstract
It is hard to execute parallel program efficiently on man-core platform because we could not divide program into appropriate granularity executed simultaneously. Based on virtual machine and binary translation technologies the article proposes the vapor profiling framework that uses SBIRP instruction in-place replacement method to collect program's run-time control flow and data flow information precisely. Moreover, it explains how to create control flow and data flow dependency graphs. Experiment results prove that vapor has better performance than traditional methods.
Yusong Tan, Wei Chen 0009, Qingbo Wu 0003
ICPADS1
2010 An Energy Efficient Clustering Scheme with Self-Organized ID Assignment for Wireless Sensor Networks
abstract
In wireless sensor networks, how to efficiently use the energy of the nodes while assigning global unique ID to each node is a challenging problem. By analyzing the communication cost of the clustering and topological features of a sensor network, we present a distributed scheme of Energy Efficient Clustering with Self-organized ID Assignment (EECSIA). In the context of EECSIA, a network first selects the nodes in the high-density areas as cluster heads, and then assigns an unique ID to each node based on local information. In addition, EECSIA periodically updates cluster heads according to the nodes' residual energy and density. The method is independent of time synchronization, and it does not rely on the nodes' geographic locations either. Simulation results show that the scheme performs well in terms of cluster scale, and number of nodes alive over rounds.
Qingchao Zheng, Zhixin Liu 0001, Yusong Tan, Dan Chen 0001, Xin-Ping Guan
ICPADS4
2009 Block-Based In-Place Replacement Strategy for x86 Sensitive Instructions in Virtual Machine
abstract
It is trendy that virtualization technology is adopted by server and desktop computers recently. Binary translation is an important method to implement full virtualization supporting any guest operating system without modification. Traditional methods use trap or interrupt to catch sensitive instruction's execution. Its performance is influenced by trap's context switch overhead. This article proposes a novel code scanning and replacing strategy, named as Block-based In-Place Replacement. BIPR tries to find a code block whose length is longer than 5 bytes and replaces the block with 5-bytes JMP instruction. The translated code block has same run-time mode as original code. As a result, BIPR's cost is lower than traditional trap methods. Moreover, it gives an optimize strategy, i.e. Super Block-based In-Place Replacement, to reduce unnecessary translation overhead of BIPR and get better performances. Experiment results prove that SBIPR performs pretty.
Yusong Tan, Qingbo Wu 0003
ISPA1
2009 System Monitoring and Controlling Mechanism Based on Hypervisor
abstract
Current commodity operating systems allow a privileged user to run some programs in kernel mode by installing a kernel module or a device driver, but there isnpsilat an available method to verify the reliability of these programs. As a result, malware leverages this way to corrupt system services, defeat anti-malware and even get control of the whole system. It makes operating-system-based security tools undergo all kinds of hardships in face of various attacks. Virtualization technology brings a new chance to protect operating systems out-of-box in hypervisors. This article proposes a system monitoring and controlling framework, named as BMCS, which monitors low-level states and hardware accessing events of operating system in a hypervisor. BMCS abstracts low-level states and events and profiles OS-level semantics with the help of a hosted daemon. Moreover, it can control the execution of operating system software through the hypervisor. Our out-of box method can collect real run-time operating system profile and control the execution of the OS according some defined rules. As a result, it makes the operating system stronger. We implement BMCS based on KVM and experiment results prove that BMCS performs well.
Qingbo Wu 0003, Yusong Tan
ISPA3
2005 Dynamic Thread Management in Kernel Pipeline Web Server
Shanshan Li 0001, Xiangke Liao, Yusong Tan, Jin-Yuan Liu
NPC3