Tong Shu

dblp:53/5387 · DBLP profile ↗
← Back
24ranked-venue papers
15as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 9 first-authorSystems, architecture and hardware · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 PredTOP: Latency Predictor Utilizing DAG Transformers for Distributed Deep Learning Training with Operator Parallelism
abstract
With the increasing sizes of deep learning (DL) models, distributed training with various parallelization techniques, such as pipeline, model, and tensor parallelism, has become an absolute necessity. Accurately and efficiently predicting iteration latency for distributed DL training is critical for high-quality application scheduling, resource allocation, and automatic performance optimization. Unfortunately, existing latency predictors for distributed DL training mainly consider data parallelism, but fail to be applied for DL training optimized by other parallelization techniques. Also, the latency prediction of large-scale distributed DL training with hybrid parallelism is very challenging due to the extremely high measurement cost for one-time execution and the prohibitively huge search space of combinations from multiple parallelization techniques. In this paper, we propose PredTOP, a framework to accurately and efficiently predict the iteration latency of distributed DL training with different types of parallelization techniques. First, we numerically formulate the end-to-end time of pipeline parallelism using stage latency to leverage the low cost of white-box modeling. Then, we adopt a black-box modeling technique to predict the latency of each stage in pipeline parallelism. Specifically, we utilize Transformers over directed acyclic graphs to improve prediction accuracy. The performance superiority of our PredTOP is validated by extensive experimental results with two real-world benchmarks in comparison with state-of-the-art approaches and illustrated in a realistic automatic parallelization plan generation system.
Dipak Acharya, Tong Shu
IPDPS2
2025 Slide-Based Graph Collaborative Training for Histopathology Whole Slide Image Analysis
abstract
The development of computational pathology lies in the consensus that pathological characteristics of tumors are significant guidance for cancer diagnostics. Most existing research focuses on the inner-contextual information within each WSI yet ignores the possible inter-correlations between slides. As the development of tumors is a continuous process involving a series of histological, morphological, and genetic changes that accumulate over time, the similarities and differences between WSIs across various stages, grades, locations and patients should potentially contribute to the representation of WSIs and deserve to be taken into account in WSI modeling. To verify the advancement of introducing the slide inter-correlations into the representation learning of WSIs, we proposed a generic WSI analysis pipeline SlideGCD that can be adapted to any existing Multiple Instance Learning (MIL) frameworks and improve their performance. With the new paradigm, the prior knowledge of cancer development can participate in the end-to-end workflow, which concurrently initializes and refines the slide representation, as a guide for message passing in the slide-based graph. Extensive comparisons and experiments are conducted to validate the effectiveness and robustness of the proposed pipeline across 4 different tasks, including cancer subtyping, cancer staging, survival prediction, and gene mutation prediction, with 8 representative SOTA WSI analysis frameworks as backbones. The code is available at https://github.com/HFUT-miaLab/SlideGCD.
Jun Shi 0006, Tong Shu, Zhiguo Jiang 0001, Wei Wang 0380, Yushan Zheng
IEEE Trans. Medical Imaging2
2024 Exploration of TPU Architectures for the Optimized Transformer in Drainage Crossing Detection
abstract
Understanding hydrologic connectivity within landscapes is crucial for managing environmental challenges. Despite advancements in high-resolution Digital Elevation Models (DEMs) derived from Light Detection and Ranging (LiDAR) technology, accurately delineating hydrologic connectivity remains challenging due to disruptions caused by virtual flow barriers, such as roads and bridges. This study addresses this issue by enhancing the detection performance and reducing the latency of Transformer models for image detection of drainage crossings. We retrained a Detection Transformer (DETR) with a specialized recipe to improve culvert detection performance. Owing to the high susceptibility of LiDAR-based DEMs to measurement noise and varying data modalities, we conducted extensive data preprocessing to ensure DETR compatibility with the culvert dataset. Ablation studies on input size indicate that the model performs optimally with 800×800 pixel inputs, demonstrating its adaptability to new data modalities. Additionally, we employed Tensor Processing Units (TPUs) to decrease the model’s latency. We developed a novel strategy to optimize TPU architecture, utilizing genetic algorithms to expedite the discovery of optimal TPU configurations for detection deployment. Our model surpasses the performance of previous models on the same task. This work not only addresses the computational complexities of deploying advanced object detection in environmental contexts but also significantly contributes to the precise and efficient monitoring of hydrologic connectivity.
Amirhossein Nazeri, Denys W. Godwin, Aikaterini Maria Panteleaki, Iraklis Anagnostopoulos, Michael Edidem, Ruopu Li, Tong Shu
IEEE Big Data7
2024 A Deep Learning Approach to Maximizing Electrostatic Sieve Efficiency in Regolith Beneficiation
abstract
This study investigates the optimization of an electrostatic sieve designed for lunar regolith beneficiation. Two parameters of the electrostatic sieve, 1) the voltage amplitude and 2) angle of inclination, were chosen as variables in the optimization process. Numerical simulations revealed that increasing voltage amplitude significantly enhances sieve performance over the sieve angle. However, optimal separation required careful voltage adjustment for specific sieve angles. A comprehensive dataset incorporating additional parameters was then created to train Machine Learning (ML) and Deep Learning (DL) models for further optimization. The ML/DL models were trained on a small subset of the original dataset to predict the yield. We showcase the benefits of leveraging DL techniques to improve the electrostatic sieve for regolith beneficiation via tailored evaluations. Our model, trained on lower-yield examples, accurately (92%) identifies parameter combinations that increase yields above 30%. It leads to a near-optimal yield with 10× reduction on runtime when compared with exhaustive simulations. This not only reduces the reliance on resource-intensive numerical simulations but also offers a rapid, validated approach to optimizing equipment for lunar mining operations.
Kalpit M. Vadnerkar, Emmanuela Amen Eze, Rinoj Gautam, Daoru Han, Xin Liang 0001, Tong Shu
IEEE Big Data6
2024 Modeling Lunar Surface Charging Using Physics-Informed Neural Networks
abstract
Modeling the electric potential profile above the lunar surface is critical for understanding surface charging and interactions with the space environment. Traditional methods like Particle-in-Cell (PIC) simulations are highly accurate but computationally expensive. To address this, we propose a hybrid approach using a Multi-Layer Perceptron (MLP) architecture in both data-driven neural networks and Physics-Informed Neural Networks (PINNs). The PINN component incorporates physical laws directly into the training process, ensuring physical consistency, while the data-driven component captures complex patterns. This combination offers a significant reduction in computational cost compared to PIC methods while maintaining high modeling accuracy. Our results show that the proposed method effectively represents the electric potential profile above the lunar surface, even with limited data.
Niloofar Zendehdel, Adib Mosharrof, Katherine Delgado, Daoru Han, Xin Liang 0001, Tong Shu
IEEE Big Data6
2024 SlideGCD: Slide-Based Graph Collaborative Training with Knowledge Distillation for Whole Slide Image Classification
Tong Shu, Jun Shi 0006, Dongdong Sun, Zhiguo Jiang 0001, Yushan Zheng
MICCAI (4)1
2023 HIOS: Hierarchical Inter-Operator Scheduler for Real-Time Inference of DAG-Structured Deep Learning Models on Multiple GPUs
abstract
Neural-network-enabled data analysis in real-time scientific applications imposes stringent requirements on inference latency. Meanwhile, recent deep learning (DL) model design trends to replace a single branch with multiple branches for high prediction accuracy and robustness, which makes inter-operator parallelization become an effective approach to improve inference latency. However, existing inter-operator parallelization techniques for inference acceleration are mainly focused on utilization optimization in a single GPU. With the data size of an input sample and the scale of a DL model ever-growing, the limited resource of a single GPU is insufficient to support the parallel execution of large operators. In order to break this limitation, we study hybrid inter-operator parallelism both among multiple GPUs and in each GPU. In this paper, we design and implement a hierarchical inter-operator scheduler (HIOS) to automatically distribute large operators onto different GPUs and group small operators in the same GPU for parallel execution. Particularly, we propose a novel scheduling algorithm, named HIOS-LP, which consists of inter-GPU operator parallelization through iterative longest-path (LP) mapping and intra-GPU operator parallelization based on a sliding window. In addition to extensive simulation results, experiments with modern convolutional neural network benchmarks demonstrate that our HIOS-LP outperforms the state-of-the-art inter-operator scheduling algorithm IOS by up to 17% in real systems.
Turja Kundu, Tong Shu
CLUSTER2
2021 In-situ workflow auto-tuning through combining component models
abstract
In-situ parallel workflows couple multiple component applications via streaming data transfer to avoid data exchange via shared file systems. Such workflows are challenging to configure for optimal performance due to the huge space of possible configurations. Here, we propose an in-situ workflow auto-tuning method, ALIC, which integrates machine learning techniques with knowledge of in-situ workflow structures to enable automated workflow configuration with a limited number of performance measurements. Experiments with real applications show that ALIC identify better configurations than existing methods given a computer time budget.
Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç
PPoPP1
2021 Bootstrapping in-situ workflow auto-tuning via combining performance models of component applications
abstract
In an in-situ workflow, multiple components such as simulation and analysis applications are coupled with streaming data transfers. The multiplicity of possible configurations necessitates an auto-tuner for workflow optimization. Existing auto-tuning approaches are computationally expensive because many configurations must be sampled by running the whole workflow repeatedly in order to train the auto-tuner surrogate model or otherwise explore the configuration space. To reduce these costs, we instead combine the performance models of component applications by exploiting the analytical workflow structure, selectively generating test configurations to measure and guide the training of a machine learning workflow surrogate model. Because the training can focus on well-performing configurations, the resulting surrogate model can achieve high prediction accuracy for good configurations despite training with fewer total configurations. Experiments with real applications demonstrate that our approach can identify significantly better configurations than other approaches for a fixed computer time budget.
Tong Shu, Yanfei Guo, Justin M. Wozniak, Xiaoning Ding, Ian T. Foster, Tahsin M. Kurç
SC1
2020 Energy-efficient mapping of large-scale workflows under deadline constraints in big data computing systems
Tong Shu, Chase Qishi Wu
Future Gener. Comput. Syst.1
2019 MPI jobs within MPI jobs: A practical way of enabling task-level fault-tolerance in HPC workflows
Justin M. Wozniak, Matthieu Dorier, Robert B. Ross, Tong Shu, Tahsin M. Kurç, Li Tang 0007, Norbert Podhorszki, Matthew Wolf
Future Gener. Comput. Syst.4
2017 Energy-Efficient Dynamic Scheduling of Deadline-Constrained MapReduce Workflows
abstract
Big data workflows comprised of moldable parallel MapReduce programs running on a large number of processors have become a main consumer of energy at data centers. The degree of parallelism of each moldable job in such workflows has a significant impact on the energy efficiency of parallel computing systems, which remains largely unexplored. In this paper, we validate with experimental results the moldable parallel computing model where the dynamic energy consumption of a moldable job increases with the number of parallel tasks. Based on our validation, we construct rigorous cost models and formulate a dynamic scheduling problem of deadline-constrained MapReduce workflows to minimize energy consumption in Hadoop systems. We propose a semi-dynamic online scheduling algorithm based on adaptive task partitioning to reduce dynamic energy consumption while meeting performance requirements from a global perspective, and also design the corresponding system modules for algorithm implementation in Hadoop architecture. The performance superiority of the proposed algorithm in terms of dynamic energy saving and deadline violation is illustrated by extensive simulation results in Hadoop/YARN in comparison with existing algorithms, and the core module of adaptive task partitioning is further validated through real-life workflow implementation and experimental results using the Oozie workflow engine in Hadoop/YARN systems.
Tong Shu, Chase Qishi Wu
eScience1
2017 Performance optimization of Hadoop workflows in public clouds through adaptive task partitioning
abstract
Cloud computing provides a cost-effective computing platform for big data workflows where moldable parallel computing models such as MapReduce are widely applied to meet stringent performance requirements. The granularity of task partitioning in each moldable job has a significant impact on workflow completion time and financial cost. We investigate the properties of moldable jobs and design a big-data workflow mapping model, based on which, we formulate a workflow mapping problem to minimize workflow makespan under a budget constraint in public clouds. We show this problem to be strongly NP-complete and design i) a fully polynomial-time approximation scheme (FPTAS) for a special case with a pipeline-structured workflow executed on virtual machines in a single class, and ii) a heuristic for a generalized problem with an arbitrary directed acyclic graph-structured workflow executed on virtual machines in multiple classes. The performance superiority of the proposed solution is illustrated by extensive simulation-based results in Hadoop/YARN in comparison with existing workflow mapping models and algorithms.
Tong Shu, Chase Qishi Wu
INFOCOM1
2017 Bandwidth Scheduling for Energy Efficiency in High-Performance Networks
abstract
The transfer of big data in various applications across high-performance networks (HPNs) in a national or international scope consumes a significant amount of energy on a daily basis. However, most existing bandwidth scheduling algorithms only consider traditional objectives, such as data transfer time minimization, and very limited efforts have been devoted to energy efficiency in HPNs. In this paper, we consider two widely adopted power models, i.e., power-down and speed-scaling, and formulate two instant bandwidth scheduling problems to minimize energy consumption under data transfer deadline and reliability constraints. We prove the NP-completeness of both problems, and design a fully polynomial time approximation scheme for the problem using the power-down model. We also design an approximation algorithm and a heuristic approach that considers the tradeoff between objective optimality and time cost in practice for the problem using the speed-scaling model. The performance superiority of the proposed solutions is illustrated by extensive results based on both simulated and real-life networks in comparison with existing methods.
Tong Shu, Chase Qishi Wu
IEEE Trans. Commun.1
2016 An improved grey neural network model for predicting transportation disruptions
Chunxia Liu, Tong Shu, Shou Chen, Shou-Yang Wang, Kin Keung Lai
Expert Syst. Appl.2
2014 GBOM-oriented management of production disruption risk and optimization of supply chain construction
Tong Shu, Shou Chen, Shou-Yang Wang, Kin Keung Lai
Expert Syst. Appl.1
2013 Advance bandwidth reservation for energy efficiency in high-performance networks
abstract
An increasing number of high-performance networks provision dedicated channels through circuit-switching or MPLS/GMPLS tunneling techniques to support large data transfer. The link bandwidths of these networks are typically shared by multiple users through advance scheduling and reservation. The sheer volume of data transfer across such networks in a national or international scope requires a significant amount of energy on a daily basis. However, most existing bandwidth scheduling algorithms only concern traditional objectives such as data transfer time minimization, and very limited efforts have been devoted to energy efficiency in high-performance networks. In this paper, we adopt a practical power model and formulate an advance instant bandwidth scheduling problem to minimize energy consumption under a data transfer deadline constraint. We design a polynomial-time optimal solution to this problem and provide a rigorous correctness proof. The performance superiority of the proposed solution in terms of energy saving is illustrated by extensive results based on both simulated and real-life networks in comparison with existing methods.
Tong Shu, Chase Qishi Wu, Daqing Yun
LCN1
2011 Spectrum Allocation for Distributed Throughput Maximization under Secondary Interference Constraints in Wireless Mesh Networks
abstract
Secondary interference constraints are important, because of representing the transmission constraints of the widespread and promising IEEE 802.11 wireless technology. Under secondary interference constraints, distributed link scheduling algorithms for multihop wireless networks can only achieve a fraction of the maximum possible throughput in general, but distributed Greedy Maximal Scheduling (GMS) algorithms can achieve optimal throughput in some network graph structures. It is possibly helpful for the improvement of distributed throughput to partition a network into subnetworks such that the subnetwork assigned to each frequency channel achieves distributed throughput maximization. In this paper, we investigate the structure characteristics of the subnetwork in which GMS achieves optimal throughput under secondary-interference constraints, and define a type of network subgraph structures meeting the requirement - special chordal subgraphs. Based on this, we propose a channel assignment algorithm, including a network partitioning algorithm and a topology balancing algorithm. By simulation, we evaluate the achievable throughput and fairness in a distributed matter using our algorithm, in comparison with the existing Max K-cut based channel assignment algorithm.
Tong Shu, Min Liu 0001, Zhongcheng Li
ICCCN1
2011 Exploiting the full potential of multi-AP diversity in centralized WLANs through back-pressure scheduling
abstract
Centralized WLANs widely deployed in enterprise environment or university campus often have high density of Access Point (AP). The high density leads to multi-AP diversity, which brings possibility to improve network performance. Previ ous studies have proposed different schemes to exploit multi-AP diversity, however, these schemes are all based on heuristic and cannot guarantee an optimal exploitation of multi-AP diversity. In this paper, we propose a Theory Based Centralized Scheduling (TBCS) to exploit the full potential of multi-AP diversity. TBCS is based on the well-known back-pressure scheduling. Although back-pressure scheduling is proved to be throughput-optimal, most of previous studies are purely theoretical. To make a practical use of the theoretical back-pressure scheduling, we design new mechanisms in TBCS to handle the problem caused by the wired/wireless mixed scenario of centralized WLANs and to synchronize the scheduling. We evaluate TBCS through NS 2 simulations and show that compared with previous methods, TBCS can support the largest capacity region and greatly improves the throughput of a network.
Anfu Zhou, Min Liu 0001, Tong Shu, Yilin Song, Zhongcheng Li
LCN3
2011 Interference pair-based distributed spectrum allocation in wireless mesh networks with frequency-agile radios
abstract
Spectrum allocation algorithms are able to improve the performance of wireless mesh networks by exploiting the frequency agility of modern radios, and several such algorithms have been proposed. However, their interference constraints are at a coarse-grained level, which results in a low spectrum efficiency. To achieve higher spectrum resource utilization, we use interference pairs as a finer granularity to model the interference constraints in wireless mesh networks, and derive a sufficient and necessary condition for interference-free spectrum allocation. Based on a set of rigorous models, we formulate spectrum allocation as an optimization problem and divide it into two subproblems, for which we propose a two-phase interference pair-based distributed spectrum allocation (IPDSA) algorithm. In IPDSA, a negotiation-based frequency hierarchy mechanism heuristically determines the relation between the center frequencies of links in each interference pair; and then a dual decomposition-based spectrum allocation algorithm converges to the optimal allocation of center frequencies and spectral widths of all links. Extensive simulation results show that IPDSA is able to significantly improve spectrum utilization and thus increase network utility and aggregate throughput, thanks to a high accuracy in modeling interference constraints.
Tong Shu, Min Liu 0001, Zhongcheng Li, Chase Qishi Wu
SECON1
2010 Joint Variable Width Spectrum Allocation and Link Scheduling for Wireless Mesh Networks
abstract
In wireless mesh networks with frequency-agile radios, an algorithm of dynamically combining consecutive channels has recently been proposed. However, the available channel widths are limited in the algorithm. In order to further improve the fairness or the throughput under given fairness, we propose a joint variable width spectrum allocation and link scheduling optimization algorithm. Our algorithm is composed of time division multiple access for no interface conflict and frequency division multiple access for no signal interference. In the first phase, we use as few time slots as possible to assign at least one time slots to each radio link with Max-Min fairness. In the second phase, our design jointly allocates the lengths of time slots as well as the spectral widths and center frequencies of radio links in each time slot. Numerical results indicate that compared to the existing algorithm, our algorithm significantly increases the fairness or the throughput under given fairness.
Tong Shu, Min Liu 0001, Zhongcheng Li, Anfu Zhou
ICC1
2010 A Diagnosis-Based Soft Vertical Handoff Mechanism for TCP Performance Improvement
abstract
Most existing soft handoff approaches lead to plenty of out-of-order packets during downward vertical handoffs (VHOs). We have presented a soft VHO scheme, called SHORDER, to avoid packet reordering caused by downward VHOs. In this paper, we analyze the effects of our SHORDER scheme and another typical existing soft VHO method on the handoff latency and the received data size during a downward VHO for TCP applications. Then, we approximately derive the applicable conditions of the two approaches, and further propose a diagnosis-based soft vertical handoff (DSVH) mechanism which can self-adaptively deal with reordering packets. The mechanism has practical advantages of no changes to correspondent nodes and compatibility with various enhanced TCP variants. With numerical analysis and test-bed experiments, we show that the DSVH mechanism has better performance than the SHORDER scheme and the typical existing method. Furthermore, experimental and analytical results are consistent with each other.
Tong Shu, Min Liu 0001, Zhongcheng Li, Anfu Zhou
ICCCN1
2009 A performance evaluation model for RSS-based vertical handoff algorithms
abstract
Many RSS-based vertical handoff algorithms have recently been proposed. However, there are only a few models to evaluate the performance of vertical handoff algorithms and none of the existing models reflect the effect of a doorway on received signal strength (RSS). Considering that RSS from heterogeneous networks cannot be directly compared with each other, we firstly present an effective method to compare RSS of different networks, based on the corresponding bandwidth in each network. Then, we take into account signal abrupt attenuation near a doorway and construct a novel performance evaluation model for RSS-based vertical handoff algorithms. This model also reflects the correlation between RSS at two adjacent locations in a WLAN. Following that, we propose an integrative performance evaluation function based on two metrics - the decision delay and the number of handoffs. Furthermore, we analyze hysteresis and dwell-timer algorithms with our model. The results show a good match between simulation and analysis.
Tong Shu, Min Liu 0001, Zhongcheng Li
ISCC1
2009 Network-layer soft vertical handoff schemes without packet reordering
abstract
Existing soft handoff techniques lead to plenty of out-of- sequence packets during downward vertical handoffs (DVHOs). In this paper, we present two new network-layer soft vertical handoff schemes, called SHORDER and E-SHORDER. The former can prevent mobile nodes from receiving reordered packets during DVHOs with a low overhead. The latter further hinders their correspondent nodes from receiving out-of-order packets caused by mobile nodes' DVHOs. Then, we analyze the performance of our proposed approaches. By experiments, we show that they have a good effect in practice.
Tong Shu, Min Liu 0001, Zhongcheng Li
LCN1