VLDB 2026 Research / reviewers in the wild / expert
Hui Tian 0001
dblp:57/1592-1
· DBLP profile ↗
81ranked-venue papers
18as first author
37since 2021 · last 2026
0000-0002-7952-571XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 15 · 2 first-author · 11 since 2021Systems, architecture and hardware · 14 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep reinforcement learning-based spectrum partition for elastic optical networks
Xin Wang 0142, Yue-Cai Huang, Hong Shen 0001, Hui Tian 0001 |
Comput. Commun. | 4 |
| 2025 | Cross-Modal Sequential Point-of-Interest Recommendation with Lightweight Hybrid Fusion Strategy
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001 |
DaWaK | 3 |
| 2025 | Local-Aware Convolutional Modulation for Short-Term Sequential Recommendation
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001 |
DaWaK | 3 |
| 2025 | NomaFdRaN: Performance Analysis of NOMA-Optimized Fully-Decoupled RAN for 6G Reliable Massive ConnectivityabstractIn order to meet unprecedented demands for reliable and massive connectivity (MC), sixth-generation (6 G) cellular Radio Access Networks (RANs) require architectural innovations. Conventional cellular RANs' scalability is limited by the tightly coupled control and user planes. Fully-Decoupled RANs (FD-RANs) are a promising architectural innovation that enables flexible plane separation. However, current architectures have limitations due to ineffective multiple access schemes. To this end, we introduce NomaFdRaN, an innovative Non-Orthogonal Multiple Access (NOMA)-optimized FD-RAN architecture in order to optimize reliability and MC. To achieve holistic system optimization, NomaFdRaN applies NOMA on all network planes, control plane and user plane, and transmission paths, uplink and downlink. To improve NOMA efficiency, we develop a user pairing optimization approach that minimizes total transmit power while maintaining linear computing complexity. Based on stochastic geometry, we develop analytical models to analyze NomaFdRaN's performance. Subsequently, we analytically derived closed-form expressions for key performance metrics. Our simulation results demonstrate the effectiveness of the NomaFdRaN architecture and provide insights into the deployment strategies for next-generation FD-RANs. Rawan A. Ameen, Haithm M. Al-Gunid, Xingfu Wang, Fuyou Miao 0001, Wei Zhao 0023, Ammar Hawbani, Hui Tian 0001, Nawaf Qasem Hamood Othman |
ICPADS | 7 |
| 2025 | A Comparative Performance of Wi-Fi Spatial Reuse Solutions in 802.11ax
Haonan Chi, Wee Lum Tan, Hui Tian 0001 |
PDCAT | 3 |
| 2025 | Efficient GNN-Based Client Selection for Optimizing Resource Allocation in Hierarchical Federated Learning
Chenghao Zhou, Huaiwen He, Hong Shen 0001, Hui Tian 0001 |
PDCAT | 4 |
| 2025 | Efficient Binary Task Offloading Optimization in Large-Scale IoT Networks via UAV-Enhanced Mobile Edge ComputingabstractUnmanned aerial vehicle (UAV)-enhanced mobile edge computing (U-MEC) integrates flexible deployment and wide coverage, effectively reducing computation latency for edge network devices. Whereas, optimization models that only consider a few or tens of nodes will hide the performance deficiencies of high-complexity algorithms in large-scale IoT networks. This paper aims to investigate the binary task offloading problem based on minimizing the shrinkage ratio in large-scale IoT networks. By optimizing the computation mode selection, bandwidth, and computing resource allocation of users, the goal is to maximize the sum of benefits brought to all users by the U-MEC network. To address the challenging mixed-integer nonlinear programming (MINLP) problem, an efficient solution based on the region coverage and a posterior method that eliminates unknown hovering constraints is adopted to decompose the problem into multiple independent small-scale subproblems. For each subproblem, the coordinate descent (CD) technique and a greedy strategy are employed, which is based on Karush-KuhnTucker (KKT) conditions to provide the closed-form solutions of bandwidth and computing resource allocation at the edge server. Experimental results demonstrate the effectiveness of our approach in terms of solution accuracy and execution time. Xiangdong Yang, Huaiwen He, Hong Shen 0001, Hui Tian 0001 |
WoWMoM | 5 |
| 2025 | Schedule multi-instance microservices to minimize response time under budget constraint in cloud HPC systemsabstractIn the emerging microservice-based architecture of cloud HPC systems, a challenging problem of critical importance for system service capability is how we can schedule microservices to minimize the end-to-end response time for user requests while keeping cost within the specified budget. We address this problem for multi-instance microservices requested by a single application to which no existing result is known to our knowledge. We propose an effective two-stage solution of first allocating budget (resources) to microservices within the budget constraint and then deploying microservice instances on servers to minimize system operational overhead. For budget allocation, we formulate it as the Discrete Time Cost Tradeoff (DTCT) problem which is NP-hard, present a linear program (LP) based algorithm, and provide a rigorous proof of its worst-case performance guarantee of 4 from the optimal solution. For microservice deployment, we show that it is harder than the NP-hard problem of 1-D binpacking through establishing its mathematical model, and propose a heuristic algorithm of Least First Mapping that greedily places microservice instances on fewest possible servers to minimize system operation cost. The experiment results of extensive simulations on DAG-based applications of different sizes demonstrate the superior performance of our algorithm in comparison with the existing approaches. • Formulate the problem of budget allocation to multi-instance microservices as the Discrete Time Cost Tradeoff (DTCT) problem, a well-known NP-hard problem. • Transform this problem to a linear program (LP) and present an approximation algorithm to minimize microservice completion time by determining the desired number of instances for each microservice within the given budget constraint. • Provide a rigorous proof of worst-case performance guarantee of 4 of our algorithm. • Present a heuristic algorithm of Least First Mapping for the problem of microservice deployment to place the microservice instances on fewest possible servers at minimum system operation cost. • Experimentally validate our algorithm and demonstrate its superiority to the existing approaches in terms of maximum completion time of microservices under budget constraint and the number of launched servers. Hong Shen 0001, Hui Tian 0001, Yuanhao Yang |
J. Parallel Distributed Comput. | 3 |
| 2025 | User-based clustering deep model for the sequential point-of-interest recommendation
Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Alan Wee-Chung Liew |
Knowl. Inf. Syst. | 3 |
| 2025 | SROdcn: Scalable and Reconfigurable Optical DCN Architecture for High-Performance ComputingabstractData Center Network (DCN) flexibility is critical for providing adaptive and dynamic bandwidth while optimizing network resources to manage variable traffic patterns generated by heterogeneous applications. To provide flexible bandwidth, this work proposes a machine learning approach with a new Scalable and Reconfigurable Optical DCN (SROdcn) architecture that maintains dynamic and non-uniform network traffic according to the scale of the high-performance optical interconnected DCN. Our main device is the Fiber Optical Switch (FOS), which offers competitive wavelength resolution. We propose a new top-of-rack (ToR) switch that utilizes Wavelength Selective Switches (WSS) to investigate Software-Defined Networking (SDN) with machine learning-enabled flow prediction for reconfigurable optical Data Center Networks (DCNs). Our architecture provides highly scalable and flexible bandwidth allocation. Results from Mininet experimental simulations demonstrate that under the management of an SDN controller, machine learning traffic flow prediction and graph connectivity allow each optical bandwidth to be automatically reconfigured according to variable traffic patterns. The average server-to-server packet delay performance of the reconfigurable SROdcn improves by 42.33% compared to inflexible interconnects. Furthermore, the network performance of flexible SROdcn servers shows up to a 49.67% latency improvement over the Passive Optical Data Center Architecture (PODCA), a 16.87% latency improvement over the optical OPSquare DCN, and up to a 71.13% latency improvement over the fat-tree network. Additionally, our optimized Unsupervised Machine Learning (ML-UnS) method for SROdcn outperforms Supervised Machine Learning (ML-S) and Deep Learning (DL). Kassahun Geresu, Huaxi Gu, Xiaoshan Yu 0001, Meaad Fadhel, Hui Tian 0001, Wenting Wei |
IEEE Trans. Cloud Comput. | 5 |
| 2025 | Optimal Partitioning of Traffic Demand for Coflow Scheduling in Hybrid SwitchesabstractIn contemporary data center networks (DCNs), scheduling groups of parallel flows (coflows) has emerged as a critical task for improving application-level communication efficiency. Recently, research interest has shifted to hybrid-switched DCN architectures that integrate optical circuit switches (OCSs) and electrical packet switches (EPSs) to respectively manage both high-volume and low-volume traffic efficiently. To minimize the overall communication latency, it is essential to effectively coordinate coflows over hybrid network links. The complexity of this task, however, is substantially greater than that of scheduling on monolithic network links of either OCS or EPS. The complexity arises from allocating flows between OCS and EPS in a coordinated way, while considering both the reconfiguration delay of circuit switching in OCS and the bandwidth limitation of packet switching in EPS, and completing the transmission in the shortest time. The current solutions for scheduling coflows in hybrid-switched DCNs are primarily based on heuristics and lack formal performance guarantees. In this paper, we first establish a coflow scheduling framework for hybrid-switched DCNs, which can transform any given circuit schedule (denoted as SC) designed for pure OCS into a corresponding hybrid scheduling scheme SH. On this basis, we further present two approximation algorithms w.r.t SC under two primary reconfiguration models (i.e., all-stop model and not-all-stop model) of OCS to minimize the coflow completion time (CCT) in a hybrid-switched DCN. Theoretical analysis demonstrates that our algorithms can achieve the optimal traffic partitioning for hybrid switch environments w.r.t SC. Furthermore, we theoretically prove that the proposed algorithms can transform any SC for pure OCS with an approximation rate of λ into a corresponding coflow scheduling scheme SH tailored for hybrid switches with the approximation rate of λ+1. Extensive simulations utilizing Facebook data traces show that our algorithm performs well in minimizing the CCT compared to state-of-the-art schemes. Xin Wang 0142, Hong Shen 0001, Hui Tian 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | Privacy-Preserving Deep Reinforcement Learning based on Differential PrivacyabstractDeep reinforcement learning, with its extensive applications and remarkable performance, is emerging as a pivotal technology garnering researchers’ attention. During the training process, there are frequent interaction and data exchange between agents and the environment, and the interaction information during training is closely tied to the training environment. Consequently, this process introduces a high risk of environmental privacy leakage. Malicious third parties may potentially steal state transition matrix or environmental information about the application domain of agent training, resulting in the compromise of user privacy. To address this issue, we propose novel differentially private value-based and policy-based deep reinforcement learning algorithms. Our methods have an advantage of being adaptable to various environmental privacy concerns. We also evaluate them in a customized experimental environment. Comparative experiments are conducted between the original and differentially private versions of the algorithms. The results indicate that our proposed approach can provide differential privacy protection to environmental information with minimal impact on algorithm performance, ultimately achieving a good balance between privacy and utility. Wenxu Zhao, Yingpeng Sang, Naixue Xiong, Hui Tian 0001 |
IJCNN | 4 |
| 2024 | Collaborative Traffic Offloading in Multi-UAV Cellular Networks via Hybridizing Optimization with Machine LearningabstractTraffic offloading via WiFi networks is an effective way to alleviate drone cellular network congestion. Existing studies are based on a limited model which assumes the presence of only a single drone and a single WiFi network in the region. We study the more general scenario of traffic offloading using multiple WiFi networks in the cellular networks composed of multiple drones which requires collaborative service provision, and address the problem of jointly minimizing the total distance traveled by the drones and the maximum latency. We propose an effective approach to solving this problem by combining optimization and machine learning. We first determine the user-drone allocation by applying meanshift clustering [1], solve the UAV navigation problem by Bellman-Floyd’s minimum-cost maximum-flow algorithm [2], and then deploy reinforcement learning to compute the optimal percentages of traffic offloading for different users. Simulation results demonstrate that our algorithm avoids violating the constraints and keeps the maximum delay in an acceptable range. Hong Shen 0001, Hui Tian 0001 |
NCA | 3 |
| 2024 | Multi-Coflow Scheduling in Not-All-Stop Optical Circuit Switches for Data Center NetworksabstractTo resolve the performance bottlenecks of power consumption and bandwidth shortage in data center networks (DCNs) using the traditional Electronic Packet Switching (EPS), Optical Circuit Switching (OCS) technology has gained extensive attention in recent years and showed great promises for the development of next-generation data centers capable of supporting escalating network traffic. In this context, coflow scheduling emerges as a pivotal technique for enhancing data transmission efficiency in DCNs. Nonetheless, the intrinsic constraints of OCS networks, such as port limitations and reconfiguration delays, introduce novel challenges to coflow scheduling. There are two optical path reconfiguration modes in OCS: All-Stop in which reconfiguration of a path will block the communications of all paths, and Not-All-Stop in which only the path under reconfiguration is blocked. Most of the existing studies focus on coflow scheduling in All-Stop mode, and little has been done for online scenarios in Not-All-Stop mode. This study delves into the online multi-coflow scheduling problem within DCNs supported by Not-All-Stop mode OCS, aiming to minimize the average Coflow Completion Time (CCT). We introduce an effective online algorithm comprising two main components: a coflow interscheduling priority strategy that takes full consideration of Not-All-Stop mode’s characteristics and network fairness, and a coflow intra-scheduling greedy algorithm (R-Greedy) that focuses on maximizing network resource utilization to determine the specific scheduling plans. The effectiveness of our algorithm is demonstrated through extensive simulation experiments. Hongkun Ren, Hong Shen 0001, Hui Tian 0001 |
NCA | 3 |
| 2024 | Privacy-Preserving in Medical Image Analysis: A Review of Methods and Applications
Yanming Zhu 0001, Xuefei Yin, Alan Wee-Chung Liew, Hui Tian 0001 |
PDCAT | 4 |
| 2024 | Effective graph-neural-network based models for discovering Structural Hole Spanners in large-scale and diverse networksabstractA Structural Hole Spanner (SHS) is a set of nodes in a network that act as a bridge among different otherwise disconnected communities. Numerous solutions have been proposed to discover SHSs that generally require high run time on large-scale networks. Another challenge is discovering SHSs across different types of networks for which the traditional one-model-fit-all approach fails to capture the inter-graph difference, particularly in the case of diverse networks. Therefore, there is an urgent need of developing effective solutions for discovering SHSs in large-scale and diverse networks. Inspired by the recent advancement of graph neural network approaches on various graph problems, we propose graph neural network-based models to discover SHS nodes in large scale networks and diverse networks. We transform the problem into a learning problem and propose an efficient model GraphSHS, that exploits both the network structure and node features to discover SHS nodes in large scale networks, endeavouring to lessen the computational cost while maintaining high accuracy. To effectively discover SHSs across diverse networks, we propose another model Meta-GraphSHS based on meta-learning that learns generalizable knowledge from diverse training graphs (instead of directly learning the model) and utilizes the learned knowledge to create a customized model to identify SHSs in each new graph. We theoretically show that the depth of the proposed graph neural network model should be at least Ω(n/logn) to accurately calculate the SHSs discovery problem. We evaluate the performance of the proposed models through extensive experiments on synthetic and real-world datasets. Our experimental results show that GraphSHS discovers SHSs with high accuracy and is at least 167.1 times faster than the comparative methods on large-scale real-world datasets. In addition, Meta-GraphSHS effectively discovers SHSs across diverse synthetic networks with an accuracy of 96.2%. Diksha Goel, Hong Shen 0001, Hui Tian 0001, Mingyu Guo 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Enhancing QoE in Large-Scale U-MEC Networks via Joint Optimization of Task Offloading and UAV TrajectoriesabstractUnmanned aerial vehicles (UAVs) have emerged as crucial components in advancing mobile edge computing (MEC), leveraging their proximity to edge nodes and scalable nature. This synergy holds significant promise within the Internet of Things (IoT) and Beyond 5G (B5G) domains. In this article, we concentrate on optimizing the shrinking ratio, a Quality of Experience (QoE) metric, within large-scale IoT networks empowered by UAV-enhanced MEC via joint optimizing task offloading, resource allocation, and UAV trajectories. This joint optimization problem presents significant challenges due to the intertwined nature of multiuser computing mode selection and strong coupling between user equipments (UEs) waiting time and UAV trajectory. To tackle these challenges, we formulate the problem as a mixed integer nonlinear programming (MINLP) problem and propose an iterative algorithm named BTOU by decomposing the original problem into two subproblems using the block coordinate descent (BCD) framework. For the task offloading and resource allocation subproblem, we present two algorithms: one employs a low-complexity greedy game-theoretic approach suitable for a large number of UEs, while the other leverages the penalty successive convex approximation (PSCA) technique along with first-order Taylor expansion approximation to achieve high-solution quality. For the UAV trajectory planning subproblem, we transform it into a Miller-Tucker–Zemlin (MTZ) model and devise a solution strategy. Extensive simulation results validate the effectiveness of our proposed algorithm, showcasing rapid convergence and a notable improvement in QoE of over 10% compared to benchmark methods. Huaiwen He, Xiangdong Yang, Hong Shen 0001, Hui Tian 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Energy-Efficiency Maximization for Relay-Aided Wireless-Powered Mobile Edge ComputingabstractMobile edge computing (MEC) integrated with wireless power transfer (WPT) has became a promising trend to shorten task delay and prolong battery life of wireless devices (WDs). Introducing the relay technique to WPT-MEC system can improve offloading capability and energy efficiency, particularly in the scenarios of poor wireless channel conditions between the server and WDs. In this paper, we focus on maximizing the energy efficiency (EE) of a multi-user relay-aided WPT-MEC system. The joint optimization of the configuration of relay, wireless charge time fraction, and decision of WDs’ offloading strategy presents significant challenges due to the combination of multi-user computing mode selection and strong coupling of transmission time allocation for each WD. To address these challenges, we formulate the problem as a mixed integer nonlinear programming (MINLP) problem and propose an efficient iterative algorithm called MSRA to solve it. Our approach leverages Dinkelbach’s method to transform the original problem into a tractable problem. Furthermore, we employ the alternating direction method of multipliers (ADMM) technique to decompose the problem into multiple subproblems, thereby avoiding the combination of computing mode selection at each WD and hence enabling parallel computation. Within each iteration step of the ADMM-based algorithm, we develop a DAI-Based algorithm to handle the strong coupling with offloading time allocation that incorporates a Bisection Search algorithm with constant time complexity and solves a standard convex problem. Extensive simulation results demonstrate the effectiveness of our proposed algorithm as evidenced by its rapid convergence and impressive energy efficiency imporvement of over 15% compared to benchmark methods. Huaiwen He, Hong Shen 0001, Hui Tian 0001 |
IEEE Internet Things J. | 4 |
| 2024 | NOMA-Enabled Integrated Space-Ground Cellular Networks Architecture Relying on Control- and User-Plane SeparationabstractWith the rapid expansion of Internet of Everything (IoE) devices and the increasing demand for high-speed data and reliable communication services, particularly within 6G cellular networks (CNs), the design of efficient and robust CNs has become a critical research area. Consequently, enabling massive connections, optimizing network resource utilization, and achieving cost-effective network operation pose significant challenges. To this end, integrated space-ground cellular networks based on control- and user-plane separation (ISGCN-CUPS) architecture has been proposed as a promising solution. Furthermore, it becomes an integral aspect of the broader paradigm of integrated space-air-ground CNs (ISAGCNs). However, scalability poses an issue when increasing the number of connected cellular users, especially when conventional orthogonal multiple access (OMA) is utilized. To address this challenge, this paper introduces the non-orthogonal multiple access (NOMA)-enabled ISGCN-CUPS architecture. Subsequently, we provide an analytical model to analyze the scenarios of proposed architecture. Utilizing stochastic geometry, we derive closed-forms for coverage probabilities over control and data channels, by considering the propagation channel models for control and data channels, both with and without interference. Furthermore, total area spectral and energy efficiencies are computed. The proposed architecture demonstrates significant enhancements in terms of the key evaluation metrics compared to conventional and OMA-enabled ISGCN-CUPS architectures. Haithm M. Al-Gunid, Xingfu Wang, Ammar Hawbani, Mohammed A. M. Sultan, Hui Tian 0001, Liang Zhao 0004 |
IEEE J. Sel. Areas Commun. | 6 |
| 2024 | Learning dynamic embeddings for temporal attributed networks
Luodi Xie, Hui Tian 0001, Hong Shen 0001 |
Knowl. Based Syst. | 2 |
| 2024 | LiDAR-Guided Cross-Attention Fusion for Hyperspectral Band Selection and Image ClassificationabstractThe fusion of hyperspectral and LiDAR data has been an active research topic. Existing fusion methods have ignored the high-dimensionality and redundancy challenges in hyperspectral images, despite that band selection methods have been intensively studied for hyperspectral image (HSI) processing. This paper addresses this significant gap by introducing a cross-attention mechanism from the transformer architecture for the selection of HSI bands guided by LiDAR data. LiDAR provides high-resolution vertical structural information, which can be useful in distinguishing different types of land cover that may have similar spectral signatures but different structural profiles. In our approach, the LiDAR data are used as the “query” to search and identify the “key” from the HSI to choose the most pertinent bands for LiDAR. This method ensures that the selected HSI bands drastically reduce redundancy and computational requirements while working optimally with the LiDAR data. Extensive experiments have been undertaken on three paired HSI and LiDAR data sets: Houston 2013, Trento and MUUFL. The results highlight the superiority of the cross-attention mechanism, underlining the enhanced classification accuracy of the identified HSI bands when fused with the LiDAR features. The results also show that the use of fewer bands combined with LiDAR surpasses the performance of state-of-the-art fusion models. Judy X. Yang, Jun Zhou 0001, Jing Wang 0062, Hui Tian 0001, Alan Wee-Chung Liew |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Scheduling Coflows in Hybrid Optical-Circuit and Electrical-Packet Switches With Performance GuaranteeabstractScheduling of coflows, each a collection of parallel flows sharing the same objective, is an important task of data transmission that arises in the networks supporting data-intensive applications such as data center networks (DCNs). The hybrid switch design combining the optical circuit switch (OCS) and electrical packet switch (EPS) for transmitting high-volume and low-volume traffic separately has received considerable research attention. To support this design, efficient scheduling of coflows on hybrid network links is crucial for reducing the overall communication time. However, because it needs to consider both reconfiguration delay of circuit switching in the OCS and bandwidth limitation of packet switching in the EPS, coflow scheduling on hybrid network links is more challenging than on monotonic network links of either OCS or EPS. The existing coflow scheduling algorithms in hybrid switches are all heuristic and provide no performance guarantees. In this work, we first propose an approximation algorithm with a worst-case performance guarantee of$2\tau$, where$\tau\le N$is the maximum number of non-zero elements of each row and column of coflow’s demand matrix, for single coflow scheduling in an$N\times N$hybrid switch to minimize the coflow completion time (CCT). We then extend the algorithm for scheduling multiple coflows to minimize the total weighted CCT with a provable performance guarantee of$\mu\tau_{\max}$, where$\mu=4M\cdot\frac{w_{\max}}{w_{\min}}$,$\tau_{\max}=\max_{1\le m\le M}\tau_{m}\leq N$,$w_{\max}$and$w_{\min}$are respectively the maximum and minimum weights of the$M$coflows. Extensive simulations using Facebook data traces show that our algorithms outperform the state-of-the-art coflow scheduling schemes. Specifically, our algorithms transmit a single coflow up to 1.08$\times$faster than Solstice (hybrid switch) and 1.42$\times$faster than Reco-Sin (pure OCS), and multiple coflows up to 1.11$\times$faster than Solstice and 1.17$\times$faster than Reco-Mul$+$(pure OCS). Xin Wang 0142, Hong Shen 0001, Hui Tian 0001 |
IEEE/ACM Trans. Netw. | 3 |
| 2024 | Scheduling Workflow Tasks With Unknown Task Execution Time by Combining Machine-Learning and Greedy-OptimizationabstractWorkflow tasks are time-sensitive and their task completion utility, i.e., value of task completion, is inversely proportional to their completion time. Existing solutions to the NP-hard problem of utility-maximization task scheduling were achieved under the assumptions of linear Time Utility Function (TUF), i.e., utility is inversely proportional to completion time following a linear function, and prior knowledge of task execution time, which is unrealistic for many applications and dynamic systems. This paper proposes a novel model of combining greedy optimization with machine learning for scheduling time-sensitive tasks with convex TUF and unknown task execution time on heterogeneous cloud servers offline nonpreemptively to maximize the total utility of input tasks. For a set of time-sensitive tasks with data dependencies, we first employ multi-layer perceptron neural networks to predict task execution time by utilizing historical data. Then, by solving a linear program after relaxing the disjunctive constraint introduced by the nonpreemption requirement to calculate maximum utility increment, we propose a novel greedy algorithm of marginal incremental utility maximization that jointly determines the task-to-processor allocation plan and tasks' execution sequence on each processor. We then show that our algorithm has an expected approximation ratio of$\frac{(e-1)(\tau -2)}{e\tau }$for convex TUF and$\frac{e-1}{3e}\approx 0.21$for linear TUF, where$\tau$is the ratio of total completion utility over total delay cost under optimal scheduling. Our result presents the first polynomial-time approximation solution for this problem that achieves a performance guarantee of bounded ratio for convex TUF and constant ratio for linear TUF respectively. Extensive experiment results through both simulation and real cloud implementation demonstrate significant performance improvement of our algorithm over the known results. Yuanhao Yang, Hong Shen 0001, Hui Tian 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Lightweight and Efficient Privacy-Preserving Multimodal Representation Inference via Fully Homomorphic Encryption
Zhaojue Li, Yingpeng Sang, Xinru Deng, Hui Tian 0001 |
ACIIDS (1) | 4 |
| 2023 | LHKV: A Key-Value Data Collection Mechanism Under Local Differential Privacy
Weihao Xue, Yingpeng Sang, Hui Tian 0001 |
DEXA (1) | 3 |
| 2023 | Performance Comparison of Distributed DNN Training on Optical Versus Electrical Interconnect Systems
Yawen Chen 0001, Zhiyi Huang 0001, Haibo Zhang 0001, Hui Tian 0001 |
ICA3PP (1) | 5 |
| 2023 | GeoMixer: The MLP-Based Sequential POI Recommender with Travel Routing ModellingabstractNowadays, with the rise of location-based services, the personalized sequential POI recommendation has become a pivotal element for enhancing customer experiences. Although many previous POI recommendation models have shown promising results and improvements in this area, several challenges still exist in this field. Firstly, the previous sequential recommenders do not well-utilize the geographical features that are highly affecting the user’s future choices of visits. Furthermore, the self-attention mechanism, which is a popular method used in sequential POI recommendation, has a limitation in treating the input user sequence as an unordered set. Using positional embedding is a typical way to overcome this limitation. However, the use of such embeddings may potentially restrict the model’s ability to learn meaningful patterns in user preferences among POIs. To address these challenges, we propose GeoMixer, a novel MLP-based sequential POI recommender that incorporates travel routing distance to capture geographical features and leverages Multi-layer Perceptron (MLP) architecture to model the spatial and sequential patterns in the sequential POI recommendations. By adopting MLP mixing layers, GeoMixer has the capability of memorizing the chronological order of the input POIs without the positional embedding and can emphasize the important latent features of each POI. The use of the travel routing information improves the model’s ability of capturing spatial patterns during the model learning process. Extensive experiments on real-world datasets show that GeoMixer outperforms state-of-theart methods in various metrics, highlighting the significance of incorporating travel routing distance and leveraging MLP architecture in sequential POI recommendation systems. Tianxing Wang 0004, Can Wang 0004, Hui Tian 0001, Hong Shen 0001 |
ICDM | 3 |
| 2023 | Resource Configuration for Cross-Server Deployment of Application-Oriented Microservices in Cloud-Edge Continuum with SLO ConstraintsabstractTo adapt to the emerging microservice-based architectures, user application requests are transformed from the traditional service-based monolithic configuration to that of multi-stage inner-dependent microservices. However, in a cloud-edge continuum, the distribution of microservices introduces extra communication overhead that could violate the Service Level Objectives (SLOs). Existing microservice-based resource allocation techniques have either considered only the case that instances belonging to the same microservice are deployed on the same server, or have not considered the case with SLO constraints. In this paper, we consider the case where instances belonging to the same microservice can be deployed across servers with SLO constraints. We propose a novel linear program (LP) based resource configurator to determine the desired number of instances of each microservice that edge servers should resource based on performance estimation for application requests such that the total cost of allocated resources of all needed microservices is minimized under the constraint that the maximum microservice completion time satisfies the given service level objective (SLO). We prove that our resource configurator achieves the worst-case performance guarantee ${(\sqrt 2 + 1)^2}$ by rigorous derivation. In comparison with the state-of-the-art microservice allocation methods, experimental results show that our proposed resource configurator reduces the computational resource usage by 7.6%, and the memory resources by 11.3% while guaranteeing the SLOs. Hong Shen 0001, Hui Tian 0001 |
ICPADS | 3 |
| 2023 | Differentially Private Functional Mechanism for Broad Learning SystemabstractTo avoid the complex structure of deep learning models and significant training costs associated with them, the Broad Learning System (BLS) based on Random Vector Functional Link Neural Network was developed. BLS only has one input layer and one output layer and uses ridge regression theory to approximate the pseudoinverse of input, greatly simplifying the network and reducing training expenses. Despite its benefits, an attacker may use membership inference attacks to determine if some sample belongs to the model's training data when model parameters are revealed, leading to information leakage. Till now there is no related work on protecting BLS against membership inference attacks. To overcome this privacy issue we present a Privacy-Preserving Broad Learning System (PPBLS), by perturbing the objective function based on Functional Mechanism (FM) in differential privacy. We theoretically prove that PPBLS can satisfy E-differential privacy, and also demonstrate its effectiveness on both regression and classification tasks. Yingpeng Sang, Hui Tian 0001 |
ISCC | 3 |
| 2023 | Adaptive multi-layer clustering strategies based on capacity weight for Internet of ThingsabstractSummary Due to the requirements brought by diversification of IoT applications, differentiation of nodes' capability, dynamic communication environment and demands, the real‐time information of nodes (actual energy consumption, living nodes' density, pairwise nodes' communication radius) should be considered comprehensively for the clustering strategies of wireless sensor networks to achieve efficient, stable, and flexible performance with limited energy and different quality of service (QoS). This article proposes an improved dynamic multi‐layer clustering strategy for various IoT applications with heterogeneous nodes' energy, unpredictable or fast‐changing distribution of alive nodes, and dynamic scenarios. In addition, an adaptive adjustment strategy based on capability weight for multi‐layer clustering network is proposed to reduce the impact of unreasonable head selection cycle of clustering. By analyzing the node energy, the change of node locations and historical data transmission of cluster head, different capability weights are assigned to each node to adaptively re‐cluster the clusters with heavy load and poor performance, further make the network topology better match current situation and specified QoS requirements. Experimental results have demonstrated that proposed strategy can achieve less energy consumption, longer network lifetime, and better load balancing, especially for the cases with heterogeneous initial energy, nonuniform distribution, and higher density of nodes. Xingchun Liu, Jingjing Yu 0003, Hongxv Wang, Hui Tian 0001 |
Concurr. Comput. Pract. Exp. | 5 |
| 2023 | Traffic flow privacy protection with performance guarantee for classification in large networks (minor revision of INS_D_21_805R3)abstractPrivacy-preserving traffic flow classification has attracted a significant amount of research interest because of its increasing importance to both network management and privacy protection. In this paper, we propose novel methods for effectively protecting network flow identifiers and attributes against privacy inference attacks in port-based and payload-based classifications respectively, and analyze their performance guarantee on data utility for flow classification and privacy (security). For protection of flow identifiers, we propose a partial identifier protection approach applying randomization and anonymization respectively on desired bit positions to conceal sensitive information, and show their expected-case performance guarantee. For protection of flow attributes, we propose a perturbation-based scheme that first selects the representative attributes by deploying an entropy-based attribute selection method to filter out redundant and insignificant attributes and reduce the problem space, then partitions the attribute domains into either equal-depth or equal-width intervals and perturbs attribute values in these intervals by swapping them with those in adjacent intervals and intervals with same value distribution respectively. We analyze the performance guarantee of the proposed methods and show the experiment results of classification accuracy obtained by implementing popular machine-learning based benchmark classifiers on our selected attributes against that on raw attributes, and on our perturbed attribute values against that on raw values, respectively. The experiment results show that our proposed methods for attribute selection and perturbation retain a high degree of data utility under the desired privacy guarantee for network traffic classification. Hui Tian 0001 |
Inf. Sci. | 1 |
| 2023 | Efficient and Fair: Information-Agnostic Online Coflow Scheduling by Combining Limited Multiplexing With DRLabstractIn shared data center networks, communications among users can be modeled as coflows, each comprising of a group of parallel data transmission flows. Efficient and fair scheduling of coflows is critical for improving both system performance and user satisfaction at the application level. Existing coflow scheduling methods maximizing efficiency (coflow completion time, CCT) and fairness (service isolation) simultaneously require prior knowledge of coflow (flow) size that is however not known before completion of coflow execution in reality, which limits their applicability. For information-agnostic scheduling, known results focus either solely on efficiency or fairness, but not both due to the hardness of achieving the desired compromise between them. In this paper, we first present an information-aware non-preemptive coflow scheduling algorithm, and show its provable long-term isolation guarantee under reasonable assumptions. We then adapt this algorithm to information-agnostic online coflow scheduling by combining limited multiplexing with Deep Reinforcement Learning (DRL) framework to achieve long-term isolation guarantee toward fair network sharing and lower average weighted CCT simultaneously. The simulation results show that our algorithm outperforms the state-of-the-art results of both fairness-optimal scheduling (NC-DRF) by 4.92in terms of average weighted CCT and performance-optimal scheduling (Aalo) in the metric of maximum normalized CCT. This fully demonstrates the superiority of our method in simultaneous optimization of efficiency and fairness for information-agnostic coflow scheduling. Xin Wang 0142, Hong Shen 0001, Hui Tian 0001 |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2022 | Advances in parallel and distributed computing and its applicationsabstractParallel and distributed computing has been the basis to many emerging areas, such as smart networks, cloud computing, big data analysis, and blockchain technology.Without the development of parallel and distributed computing technologies and various types of systems, it is not possible to meet the requirement on efficiency, accuracy, scalability, and reliability for various critical applications that support our modern economy and society.While parallel and distributed computing played a vital role in modern science, engineering, biology, medicine, pharmacy, astronomy, geology, and archaeology, its application has also been extended to business, finance, economics, management, government, and defense, covering all aspects of our modern society and life.Furthermore, parallel and distributed computing has emerged in recent advances of many hotspot research directions including artificial intelligence, machine learning, Internet of Things, bioinformatics, digital medicine, cybersecurity, and social computing, resulting in numerous ground-breading discoveries that are changing our society and life. Hui Tian 0001, Alan Wee-Chung Liew, Hong Shen 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | A multimodal differential privacy framework based on fusion representation learningabstractDifferential privacy mechanisms vary in modalities, and there have been many methods implementing differential privacy on unimodal data. Few studies focus on unifying them to protect multimodal data, though privacy protection of multimodal data is of great significance. In our work, we propose a multimodal differential privacy protection framework. Firstly, we use multimodal representation learning to fuse different modalities and map them to the same subspace. Then based on this representation, we use the Local Differential Privacy (LDP) mechanism to protect data. We propose two protection methods for low-dimensional and high-dimensional fusion tensors respectively. The former is based on Binary Encoding, and the latter is based on multi-dimensional Fourier Transform. To the best of our knowledge, we are the first to propose LDP-based methods for the representation learning of multimodal fusion. Experimental results demonstrate the flexibility of our framework where both approaches show efficient performance as well as high data utility. Chaoxin Cai, Yingpeng Sang, Hui Tian 0001 |
Connect. Sci. | 3 |
| 2022 | Online delay-guaranteed workload scheduling to minimize power cost in cloud data centers using renewable energy
Huaiwen He, Hong Shen 0001, Qing Hao, Hui Tian 0001 |
J. Parallel Distributed Comput. | 4 |
| 2022 | Improve the quality of charging services for rechargeable wireless sensor networks by deploying a mobile vehicle with multiple removable chargers
Zhansheng Chen, Hui Tian 0001, Hong Shen 0001 |
Wirel. Networks | 2 |
| 2021 | Maintenance of Structural Hole Spanners in Dynamic NetworksabstractStructural Hole (SH) spanners are the set of users who bridge different groups of users and are vital in numerous applications. Despite their importance, existing work for identifying SH spanners focuses only on static networks. However, real-world networks are highly dynamic where the underlying structure of the network evolves continuously. Consequently, we study SH spanner problem for dynamic networks. We propose an efficient solution for updating SH spanners in dynamic networks. Our solution reuses the information obtained during the initial runs of the static algorithm and avoids the recomputations for the nodes unaffected by the updates. Experimental results show that the proposed solution achieves a minimum speedup of 3.24 over recomputation. To the best of our knowledge, this is the first attempt to address the problem of maintaining SH spanners in dynamic networks. Diksha Goel, Hong Shen 0001, Hui Tian 0001, Mingyu Guo 0001 |
LCN | 3 |
| 2020 | Truth finding by reliability estimation on inconsistent entities for heterogeneous data sets
Hui Tian 0001, Wenwen Sheng, Hong Shen 0001, Can Wang 0004 |
Knowl. Based Syst. | 1 |
| 2020 | An efficient method for privacy-preserving trajectory data publishing based on data partitioning
Hong Shen 0001, Yingpeng Sang, Hui Tian 0001 |
J. Supercomput. | 4 |
| 2019 | TOM: A Threat Operating Model for Early Warning of Cyber Security Threats
Can Wang 0004, Yunwei Zhao, Kwok-Yan Lam, Chihung Chi, Hui Tian 0001 |
ADMA | 7 |
| 2019 | Variational Deep Collaborative Matrix Factorization for Social Recommendation
Teng Xiao, Hui Tian 0001, Hong Shen 0001 |
PAKDD (1) | 2 |
| 2019 | Privacy Preservation for Network Traffic ClassificationabstractWith the rapid development of Internet technology and massive demands of data sharing, data privacy issues have attracted more and more attention in recent years. The paper analyzes the network traffic classification methods and designs the features subset selection algorithm based on information entropy. The proposed privacy preserving algorithm is based on data perturbation. By applying the algorithm on the real network traffic data set, it is shown that the network traffic data protected by the algorithm can effectively ensure data security while maintaining data utility, which contributes to balance the contradiction between them in existing algorithms. It effectively solves the privacy leakage problem of network traffic in the process of data mining. Hui Tian 0001, Jingjing Yu 0003 |
PDCAT | 2 |
| 2019 | An Improved MCB Localization Algorithm Based on Received Signal Strength IndicatorabstractAn improved Monte Carlo Localization Boxed (MCB) localization Algorithm based on received signal strength indicator (RSSI) for mobile wireless sensor networks node localization is proposed. Aiming at the characteristics of instantaneity, mobility and complexity of mobile node location, this algorithm combines RSSI ranging model with MCL algorithm which has high positioning accuracy, good performance and wide application. It establishes anchor node sampling box by using the actual distance of signal propagation obtained, and effectively reduces the sampling range. The simulation results show that this algorithm improves the sampling efficiency, shortens the sampling time, and increases the positioning accuracy compared to the MCL algorithm. It is a low-cost, low-power and low-complexity localization algorithm without additional hardware and communication overhead. Qiao Yan, Chunyue Zhou, Baitong Zhong, Hui Tian 0001 |
PDCAT | 4 |
| 2019 | An Energy-Saving Routing Algorithm for Opportunity Networks Based on Sleeping ModeabstractOpportunistic network is a type of tolerant delay network (DTN). Establishing routing path from the source node to the destination node may not be possible. So the traditional routing algorithms are not suitable for this network except the Epidemic. But due to its replication characteristics, the energy of node will be used up quickly. Many researchers have achieved lots of results on how to forwarding data strategically in Epidemic, but pay little attention to the energy consumption of nodes. This paper proposes an energy-saving algorithm named ERSE which combines the asynchronous sleep mechanism. The ERSE is based on the node communication model, making the node enter the sleeping state when it's not active. It also controls the sleep duration dynamically to ensure that the node does not miss the communication opportunities. ERSE makes the low-energy state nodes enter a forced sleeping state at the same time to avoid them going to use up their energy quickly. The simulation results show that ERSE routing algorithm can greatly save the energy consumption of nodes. Chunyue Zhou, Hui Tian 0001, Yaocong Dong, Baitong Zhong |
PDCAT | 2 |
| 2019 | Foreword to the Special Section on Parallel and Distributed Computing, Applications and TechnologiesabstractThe increasing intensity and variety of computation in solving modern science, engineering, and social application problems have brought numerous challenges to parallel and distributed computing and enriched its research content in multiple folds from architecture design to computation paradigm and security with the focuses on computation efficiency for problem solving, data protection for large-scale applications, and energy efficiency for power savings and environment protection.The objective of this Special Section is to show some representative research results on these focuses.This Special Section includes three extended papers primarily selected from the papers presented at the 17th International Conference on Parallel and Distributed Computing, Applications and Technologies (PDCAT 2016).We received 10 papers submitted to this Special Section, among which three papers are finally selected after at least two rounds of strict review. Hong Shen 0001, Hui Tian 0001, Yingpeng Sang |
Concurr. Comput. Pract. Exp. | 2 |
| 2018 | Betweenness Centrality Based k-Anonymity for Privacy Preserving in Social NetworksabstractIn order to reduce information loss rate of the k-anonymous network, we propose an anonymous network reconstruction algorithm based on nodes' betweenness centrality. Betweenness centrality measures the nodes' significance in graph based on the shortest paths. In our algorithm, nodes are sorted according to betweenness centrality and the candidate nodes with great value are reconnected preferentially. Then the backbone structure of the network could be retained which guarantee the important features of the reconstruct network. We evaluate our method on real data sets by different metrics. The experiment results justify the efficiency and practical utility of our proposed method. Hui Tian 0001, Jingtian Liu, Jingjing Yu 0003 |
MoMM | 1 |
| 2018 | Privacy Preserving Classification Based on Perturbation for Network Traffic
Hui Tian 0001, Hong Shen 0001 |
PDCAT | 2 |
| 2018 | A Real-Time Routing Protocol in Wireless Sensor-Actuator Network
Hui Tian 0001, Jiajia Yin |
PDCAT | 2 |
| 2017 | An Improved (k, p, l)-Anonymity Method for Privacy Preserving Collaborative FilteringabstractCollaborative Filtering (CF) is a successful technique that has been implemented in recommender systems and Privacy Preserving Collaborative Filtering (PPCF) aroused increasing concerns of the society. Current solutions mainly focus on cryptographic methods, obfuscation methods, perturbation methods and differential privacy methods. But these methods have some shortcomings, such as unnecessary computational cost, lower data quality and hard to calibrate the magnitude of noise. This paper proposes a (k, p, I)-anonymity method that improves the existing k-anonymity method in PPCF. The method works as follows: First, it applies Latent Factor Model (LFM) to reduce matrix sparsity. Then it improves Maximum Distance to Average Vector (MDAV) microaggregation algorithm based on importance partitioning to increase homogeneity among records in each group which can retain better data quality and (p, I)-diversity model where p is attacker's prior knowledge about users' ratings and I is the diversity among users in each group to improve the level of privacy preserving. Theoretical and experimental analyses show that our approach ensures a higher level of privacy preserving based on lower information loss. Ruoxuan Wei, Hong Shen 0001, Hui Tian 0001 |
GLOBECOM | 3 |
| 2017 | Weighted Ensemble Classification of Multi-label Data Streams
Lulu Wang 0008, Hong Shen 0001, Hui Tian 0001 |
PAKDD (2) | 3 |
| 2017 | Efficient Algorithms for VM Placement in Cloud Data CentersabstractThe virtual machine (VM) placement problem is a major issue in optimizing resource ulitization of cloud data centers. With the rapid development of cloud computing, efficient algorithms are needed to reduce the power consumption and save energy in data centers. Many models and algorithms are designed with a objective to minimize the number of physical machines (PMs) used in a cloud data center. In this paper, we take into account the execution time of the PM, and formulat a new optimization problem of VM placement, which aims to minimize the total execution time of the PMs. We discuss the NP-hardness of the problem, and present heuristic algorithms to solve it in both offline and online scenarios. Furthermore, we conduct experiments to evaluate the performance of the proposed algorithms and the result show that our methods are able to perform better than other commonly used algorithms. Hui Tian 0001, Jiahuai Wu, Hong Shen 0001 |
PDCAT | 1 |
| 2017 | A Hybrid Network Traffic Prediction Model Based on Optimized Neural NetworkabstractWith growth of networks, it's demanding to predict the development of network traffic. In this paper, we analyze the network traffic based on the hybrid neural network model. The chaotic property of traffic data is verified by analyzing the chaos characteristics of the data. Based on the study of artificial neural network, wavelet transform theory and quantum genetic algorithm, we propose a neural network optimization method based on efficient global search capability of quantum genetic algorithm. The proposed quantum genetic artificial neural network model can predict the network traffic more accurately. The prediction results can be used to monitor the network anomaly in network security field, and improve the quality of service. The results will also benefit to search efficient network optimization solutions by predicting network behavior. Hui Tian 0001, Jingtian Liu |
PDCAT | 1 |
| 2016 | Study on Network Anomaly Localization TechniquesabstractWith the increased network scale and complexity of network topology, especially the constant variation of attacking modes, it becomes urgent to find efficient methods to localize anomalies quickly and accurately. Present studies on network performance are mostly limited to detecting anomalies, or balancing the missed rate and false alarm rate of anomaly detection. The existing small amount of localization researches are divided into two categories: localization methods based on end to end path probe and wavelet transform. This paper introduces some existing network anomaly localization methods and analyzes the performance of them. Finally, we discuss the existing technical difficulties in anomaly localization. Jingtian Liu, Hui Tian 0001 |
PDCAT | 2 |
| 2016 | Diffusion Wavelet-Based Anomaly Detection in NetworksabstractTraffic Matrix (TM) can contain information about irregular network topology structure and depict the traffic characteristics of global network. It is a critical parameter to network traffic engineering and attracts significant research interests. Diffusion Wavelet (DW) can perform an effective Multi-Resolution Analysis (MRA)on TM in both temporaland space domains because it intrinsically adapts to the underlying network structure. This paper shows how to apply DW to TM analysis and anomaly detection. By comparing with other anomaly detection methods, it is confirmed thatour method can detect anomaly effectively due to combining with the analysis results by DW. Hui Tian 0001, Meimei Ding |
PDCAT | 1 |
| 2016 | Comparison on Network Simulation TechniquesabstractNetwork simulation is an important technique for verifying new algorithms, analyzing network performance and deploying the practical networks. Different network simulation softwares are applied for different scenarios. In this paper, their performance in different applications are discussed in detail. Three kinds of main network simulation softwares are introduced in this paper: OPNET, Network Simulator (NS) and Objective Modular Network Testbed in C++ (OMNeT++). NS is widely used in network research. How to apply NS in network traffic analysis is discussed in this paper. Hui Tian 0001 |
PDCAT | 2 |
| 2016 | Corrigendum to 'On-demand data broadcast with deadlines for avoiding conflicts in wireless networks' [The Journal of Systems and Software 103 (2015) 118-127]
Hong Shen 0001, Hui Tian 0001 |
J. Syst. Softw. | 3 |
| 2016 | Achieving Probabilistic Anonymity in a Linear and Hybrid Randomization ModelabstractThe randomization methods that are applied for privacy-preserving data mining are commonly subject to reconstruction, linkage, and semantic-related attacks. Some existing works employed random noise addition to realize probabilistic anonymity, aiming only at linkage attacks. Random noise addition is vulnerable to reconstruction attacks, and is unable to achieve semantic closeness, particularly on high-dimensional data, to prevent semantic-related attacks. For linkage attacks, the main security vulnerability of their proposed probabilistic anonymity lies in the assumption that the attacker had a priori knowledge of the quasi-identifiers of all individuals. When only some individuals leak their quasi-identifiers, the proposed model will become incapable, because the attacker can deploy a different linkage attack that has not been studied before. This type of attack is much easier to deploy and is thus very harmful. In this paper, we propose new frameworks of probabilistic (1, k)and (k, k)-anonymity to defend against all these linkage attacks, and realize the frameworks on a hybrid randomization model. The model is also secure against reconstruction attacks. We further achieve statistical semantic closeness of high-dimensional data to prevent semantic-related attacks on the model. The frameworks also allow us to re-design the traditional K-nearest neighbor algorithm to leverage the introduced data uncertainty and improve the mining results. This paper demonstrates the promising applications in large-scale and high-dimensional data mining in clouds, by providing high efficiency and security to protect data privacy, guaranteeing high data utility for mining purposes, on-time processing, and non-interactive data publishing. Yingpeng Sang, Hong Shen 0001, Hui Tian 0001, Zonghua Zhang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2015 | On-demand data broadcast with deadlines for avoiding conflicts in wireless networks
Hong Shen 0001, Hui Tian 0001 |
J. Syst. Softw. | 3 |
| 2014 | Two-Phase Layered Learning Recommendation via Category Structure
Ke Ji, Hong Shen 0001, Hui Tian 0001, Yanbo Wu, Jun Wu 0007 |
PAKDD (2) | 3 |
| 2014 | A Selectively Re-train Approach Based on Clustering to Classify Concept-Drifting Data Streams with Skewed Distribution
Hong Shen 0001, Hui Tian 0001, Yidong Li, Jun Wu 0007, Yingpeng Sang |
PAKDD (2) | 3 |
| 2014 | Maximizing network lifetime in wireless sensor networks with regular topologies
Hui Tian 0001, Hong Shen 0001, Yingpeng Sang |
J. Supercomput. | 1 |
| 2013 | Efficient Approximation Algorithm for Data Retrieval with Conflicts in Wireless NetworksabstractGiven a set of data items broadcasting at multiple parallel channels, where each channel has the same broadcast pattern over a time period, and a set of client's requested data items, the data retrieval problem requires to find a sequence of channel access to retrieve the requested data items among the channels such that the total access latency is minimized, where both channel access (to retrieve a data item) and channel switch are assumed to take a single time slot. As an important problem of information retrieval in wireless networks, this problem arises in many applications such as e-commerce and ubiquitous data sharing, and is known two conflicts: requested data items are broadcast at same time slots or adjacent time slots in different channels. Although existing studies focus on this problem with one conflict, there is little work on this problem with two conflicts. So this paper proposes efficient algorithms from two views: single antenna and multiple antennae. Our algorithm adopts a novel approach that wireless data broadcast system is converted to DAG, and applies set cover to solve this problem. Through Experiments, this result presents currently the most efficient algorithm for this problem with two conflicts. Hong Shen 0001, Hui Tian 0001 |
MoMM | 3 |
| 2013 | Privacy-Preserving Ranked Fuzzy Keyword Search over Encrypted Cloud DataabstractAs Cloud Computing becomes popular, more and more data owners prefer to store their data into the cloud for great flexibility and economic savings. In order to protect the data privacy, sensitive data usually have to be encrypted before outsourcing, which makes effective data utilization a challenging task. Although traditional searchable symmetric encryption schemes allow users to securely search over encrypted data through keywords and selectively retrieve files of interest without capturing any relevance of data files or search keywords, and fuzzy keyword search on encrypted data allows minor typos and format inconsistencies, secure ranked keyword search captures the relevance of data files and returns the results that are wanted most by users. These techniques function unilaterally, which greatly reduces the system usability and efficiency. In this paper, for the first time, we define and solve the problem of privacy-preserving ranked fuzzy keyword search over encrypted cloud data. Ranked fuzzy keyword search greatly enhances system usability and efficiency when exact match fails. It returns the matching files in a ranked order with respect to certain relevance criteria (e.g., keyword frequency) based on keyword similarity semantics. In our solution, we exploit the edit distance to quantify keyword similarity and dictionary-based fuzzy set construction to construct fuzzy keyword sets, which greatly reduces the index size, storage and communication costs. We choose the efficient similarity measure of "coordinate matching", i.e., as many matches as possible, to obtain the relevance of data files to the search keywords. Qunqun Xu, Hong Shen 0001, Yingpeng Sang, Hui Tian 0001 |
PDCAT | 4 |
| 2013 | Multi-resolution Analysis on Traffic Matrix by Different Diffusion OperatorsabstractTraffic matrix describes the data flow between each pair of Origin-Destination (OD) over a measured period. However, it is very hard to be obtained in a large scale network. This paper compares two available diffusion operators. Based on the selection of good diffusion operator, we conduct multi-resolution analysis (MRA) on traffic matrices by diffusion wavelets. We also propose a method to detect the anomaly of the traffic matrix during a continuous period of time based on diffusion wavelet analysis. Binze Zhong, Hui Tian 0001 |
PDCAT | 2 |
| 2012 | On Robust Multicast in Multi-channel Multi-radio Wireless Mesh NetworksabstractThe multicast problem in multi-channel multi-radio wireless mesh networks has received much attention recently. Most recent studies on this problem focus on improving the network throughput. However, many real-world applications require routing algorithms to achieve low-delay and low-loss. In this paper, we tackle the problem of constructing a robust minimum-cost multicast tree that tolerates link interference. To save bandwidth resource and alleviate the interference in the communication, we propose a robust multicast algorithm for multi-channel multi-radio wireless mesh networks. Our experimental results show that our algorithm is very efficient to achieve better performances in network throughput and end-to-end delay than previous studies. Yidong Li, Yingpeng Sang, Hong Shen 0001, Hui Tian 0001, Yanbo Wu |
PDCAT | 5 |
| 2012 | A Clustering Algorithm Based on Density-Grid for Stream DataabstractMany real applications, such as network traffic monitoring, intrusion detection, satellite remote sensing, and electronic business, generate data in the form of a stream arriving continuously at high speed. Clustering is an important data analysis tool for knowledge discovery. Compared with traditional clustering algorithms, clustering stream data is an important and challenging problem which has attracted many researchers. Clustering stream data is facing two main challenges. First, as the data is continuously arriving with high rate and the computer storage capacity is limited, raw data can only be scaned in one pass. Second, stream data is always changing with time, so viewing a data stream as a set of static data can deteriorate the clustering quality. In fact, users are more concerned with the evolving behaviors of clusters which can help people making correct decisions. This paper proposes a density-grid based clustering algorithm, PKS-Stream-I, for stream data. It is an optimization of PKS-Stream in density detection period selection, sporadic grid detection and removal. Empirical results show the proposed method yields out better performance. Hui Tian 0001, Yingpeng Sang, Yidong Li, Yanbo Wu, Jun Wu 0007, Hong Shen 0001 |
PDCAT | 2 |
| 2012 | Effective Reconstruction of Data Perturbed by Random ProjectionsabstractRandom Projection (RP) has raised great concern among the research community of privacy-preserving data mining, due to its high efficiency and utility, e.g., keeping the euclidean distances among the data points. It was shown in [33] that, if the original data set composed of m attributes is multiplied by a mixing matrix of k\times m (m>;k) which is random and orthogonal on expectation, then the k series of perturbed data can be released for mining purposes. Given the data perturbed by RP and some necessary prior knowledge, to our knowledge, little work has been done in reconstructing the original data to recover some sensitive information. In this paper, we choose several typical scenarios in data mining with different assumptions on prior knowledge. For the cases that an attacker has full or zero knowledge of the mixing matrix R, respectively, we propose reconstruction methods based on Underdetermined Independent Component Analysis (UICA) if the attributes of the original data are mutually independent and sparse, and propose reconstruction methods based on Maximum A Posteriori (MAP) if the attributes of the original data are correlated and nonsparse. Simulation results show that our reconstructions achieve high recovery rates, and outperform the reconstructions based on Principal Component Analysis (PCA). Successful reconstructions essentially mean the leakage of privacy, so our work identify the possible risks of RP when it is used for data perturbations. Yingpeng Sang, Hong Shen 0001, Hui Tian 0001 |
IEEE Trans. Computers | 3 |
| 2011 | Diffusion Wavelets-Based Analysis on Traffic MatricesabstractTraffic matrix describes the traffic volumes traversing the network from the input nodes to the exit nodes over a measured period. Such a matrix is very hard, if not intractable, to be obtained for a large network. We apply a new technique to analyze the traffic matrix by use of diffusion wavelet in this paper. It is shown that diffusion wavelet can do an efficient multi-resolution analysis on TM. The original TM can be reconstructed by choosing the diffused traffic in a particular level. This paper also shows there are a lot of potential applications by use of diffusion wavelet-based analysis on traffic matrix. Hui Tian 0001, Matthew Roughan, Yingpeng Sang, Hong Shen 0001 |
PDCAT | 1 |
| 2009 | Dynamically Maintaining Duplicate-Insensitive and Time-Decayed Sum Using Time-Decaying Bloom Filter
Hong Shen 0001, Hui Tian 0001, Xianchao Zhang 0001 |
ICA3PP | 3 |
| 2009 | Reconstructing Data Perturbed by Random Projections When the Mixing Matrix Is Known
Yingpeng Sang, Hong Shen 0001, Hui Tian 0001 |
ECML/PKDD (2) | 3 |
| 2009 | Privacy-Preserving Tuple Matching in Distributed DatabasesabstractWe address the problems of privacy-preserving duplicate tuple matching (PPDTM) and privacy-preserving threshold attributes matching (PPTAM) in the scenario of a horizontally partitioned database among N parties, where each party holds a private share of the database's tuples and all tuples have the same set of attributes. In PPDTM, each party determines whether its tuples have any duplicate on other parties' private databases. In PPTAM, each party determines whether all attribute values of each tuple appear at least a threshold number of times in the attribute unions. We propose protocols for the two problems using additive homomorphic cryptosystem based on the subgroup membership assumption, e.g., Paillier's and ElGamal's schemes. By analysis on the total numbers of modular exponentiations, modular multiplications and communication bits, with a reduced computation cost which dominates the total cost, by trading off communication cost, our PPDTM protocol for the semihonest model is superior to the solution derivable from existing techniques in total cost. Our PPTAM protocol is superior in both computation and communication costs. The efficiency improvements are achieved mainly by using random numbers instead of random polynomials as existing techniques for perturbation, without causing successful attacks by polynomial interpolations. We also give detailed constructions on the required zero-knowledge proofs and extend our two protocols to the malicious model, which were previously unknown. Yingpeng Sang, Hong Shen 0001, Hui Tian 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2008 | Maximizing Networking Lifetime in Wireless Sensor Networks with Regular TopologiesabstractEnergy-constraint is a crucial problem in wireless sensor networks (WSNs). Many sensor node (SN) placement schemes and routing protocols are proposed to address this problem. In this paper, we first present how to place SNs by use of a minimal number to maximize the coverage area when the communication radius of the SN is not less than the sensing radius, which results in the application of regular topology to WSNs deployment. With nodes placed at an equal distance and equipped with an equal power supply, we discuss the energy imbalance problem and then give the mathematical formulation for maximizing network lifetime in grid-based WSNs. The formulation shows the problem of maximizing network lifetime is a non-linear programming problem and NP-hard even in the 1-D case. We discuss several heuristic solutions and show that the halving shift data collection scheme is the best solution among them. We also generalize the maximizing network lifetime problem to the randomly-deployed WSNs which shows the significance of our mathematical formulation for this crucial problem. Hui Tian 0001, Hong Shen 0001, Matthew Roughan |
PDCAT | 1 |
| 2006 | Reliable and Real-Time Data Gathering in Multi-hop Linear Wireless Sensor Networks
Haibo Zhang 0001, Hong Shen 0001, Hui Tian 0001 |
WASA | 3 |
| 2006 | Multicast-based inference for topology and network-internal loss performance from end-to-end measurements
Hui Tian 0001, Hong Shen 0001 |
Comput. Commun. | 1 |
| 2006 | Random Walk Routing in WSNs with Regular Topologies
Hui Tian 0001, Hong Shen 0001, Teruo Matsuzawa |
J. Comput. Sci. Technol. | 1 |
| 2005 | Hamming Distance and Hop Count Based Classification for Multicast Network Topology InferenceabstractTopology information of a multicast network benefits significantly to many applications such as resource management, loss and congestion recovery. In this paper we propose a new algorithm, namely binary hamming distance and hop count based classification algorithm (BHC), to infer multicast network topology from end-to-end measurements. The BHC algorithm identifies multicast network topology using hamming distance of the sequences on receipt/loss of probe packets maintained at each pair of nodes and incorporating the hop count available at each node. We analyze the inference accuracy of the algorithm and prove that the algorithm can obtain accurate inference at higher probability than previous algorithms for a finite number of probe packets. We implement the algorithm in a simulated network and validate the algorithm's performance in accuracy and efficiency. Hui Tian 0001, Hong Shen 0001 |
AINA | 1 |
| 2005 | Discover multicast network internal characteristics based on Hamming distanceabstractOne of the important techniques to monitor and control large-scale networks today is to implement only at the end. However end-based control needs to have the knowledge of network internal characteristics. The paper proposes a novel approach to discover network internal characteristics from end-to-end multicast traffic measurements, which requires no support from internal routers. Our approach is based on Hamming distance of sequences on receipt/loss of probe packets maintained at each pair of nodes. As we discuss in this paper, our approach mainly focuses on identification of network internal characteristics of routing topology and loss performance. The simulation shows that the Hamming distance-based approach can discover the routing topology which is more accurate and efficient with a finite number of probe packets than before. The Hamming distance matrix proposed in this paper can also effectively discover the loss performance of the network. Hui Tian 0001, Hong Shen 0001 |
ICC | 1 |
| 2005 | Developing Energy-Efficient Topologies and Routing for Wireless Sensor Networks
Hui Tian 0001, Hong Shen 0001, Teruo Matsuzawa |
NPC | 1 |
| 2005 | RandomWalk Routing for Wireless Sensor NetworksabstractTopology is important for any type of networks because it has great impact on the performance of the network. For wireless sensor networks (WSN), regular topologies, which can help to efficiently save energy and achieve long networking lifetime, have been well studied in [1, 4, 5, 7, 9]. However, little work is focused on routing in patterned WSNs except the shortest path routing with the knowledge of global location information. In this paper, we propose a routing protocol based on random walk. It doesn’t require global location information. Moreover, the random walk routing achieves load balancing property inherently for WSNs which is difficult to achieve for other routing protocols. We also prove that the random walk routing consumes the same amount of energy as the shortest path routing in the scenarios where the message required to be sent to the base station is in comparatively small size with the inquiry message among neighboring nodes. Since in many applications of WSNs, sensor nodes often send only beeplike small messages to the base station to report their status, our proposed random walk routing is a viable scheme. Though the random walk routing provides load balancing in the WSN, the nodes near to the base station (BS) are inevitably under heavier burden than the nodes far from the base station. Therefore we further propose a density-aware deployment scheme to guarantee that the heavy-load nodes do not affect the network lifetime even if they are exhausted. Hui Tian 0001, Hong Shen 0001, Teruo Matsuzawa |
PDCAT | 1 |
| 2004 | Analysis on binary loss tree classification with hop count for multicast topology discoveryabstractThe use of multicast inference on end-to-end measurement has recently been proposed as a means of obtaining the underlying multicast topology. We analyze the algorithm of binary loss tree classification with hop count (HBLT). We compare it with the binary loss tree classification algorithm (BLT) and show that the probability of misclassification of HBLT decreases more quickly than that of BLT as the number of probing packets increases. The inference accuracy of HBLT is always 1 (the inferred tree is identical to the physical tree) in the case of correct classification, whereas that of BLT is dependent on the shape of the physical tree and inversely proportional to the number of internal nodes with a single child. Our analytical result shows that HBLT is superior to BLT, not only on time complexity, but also on misclassification probability and inference accuracy. Hui Tian 0001, Hong Shen 0001 |
CCNC | 1 |
| 2004 | Lossy Link Identification for Multicast Network
Hui Tian 0001, Hong Shen 0001 |
PDCAT | 1 |