EDBT 2026 Demo / reviewers in the wild / expert
Yitu Wang
dblp:194/6915
· DBLP profile ↗
46ranked-venue papers
20as first author
29since 2021 · last 2026
0000-0003-4453-5966ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 12 first-author · 13 since 2021Systems, architecture and hardware · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Empowering Deterministic-Delay MEC via Intelligent Network Slicing
Xinglin Yang, Wei Wang 0021, Yitu Wang, Bing Hu 0002, Zhaoyang Zhang 0001 |
ICC | 3 |
| 2025 | AutoRAC: Automated Processing-in-Memory Accelerator Design for Recommender Systems
Tunhou Zhang, Junyao Zhang 0003, Jonathan Hao-Cheng Ku, Yitu Wang, Xiaoxuan Yang 0001, Hai Li 0001, Yiran Chen 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | A Hybrid Network Performance Measurement Framework for Deterministic Smart GridsabstractPerformance measurement is one of the key enablers towards the realization of future Smart Grids (SGs). However, balancing the trade-off among measurement completeness, accuracy, and timeliness under constrained resources, while providing targeted support for communication-intensive areas consisting of critical infrastructure remains a challenge. In this paper, we propose a hybrid network Performance Measurement (PM) framework for deterministic SGs in two steps. 1) To obtain PM information for the whole network without incurring overhead, passive measurement with missing data imputation is tailored for meeting the divergence of packet loss types of large-scale SGs based on the proposed Temporal-Spatial based Distinguishable Generative Adversarial Imputation Network (STD-GAIN), which explores and exploits both the spatial-temporal correlation and higher-order local correlation for refined imputation performance. 2) Regarding the circumstance that the reliability of the imputed data does not meet the requirement, active measurement is conducted at a finer granularity, i.e., along a specified path, where All-Pair Shortest Path (APSP) model is utilized to find the best paths. Finally, the simulation results show that the proposed framework could achieve complete, accurate and real-time network performance measurement for SGs. Junjie An, Wei Wang 0021, Yitu Wang, Xinglin Yang, Zhaoyang Zhang 0001 |
VTC2025-Fall | 3 |
| 2025 | Adaptive VR Video Transmission via Multimodal Viewpoint Prediction with Event InformationabstractDelivering high-quality 360-degree VR video challenges the communication system on low latency and high quality transmission, which can be resolved by tile-based transmission technology. However, existing works do not take into account unexpected events and consider field-of-view prediction and bandwidth allocation separately, resulting in the inability to match dynamic transmission conditions and further resulting in significant performance loss. This paper proposes a multimodal fusion-based framework to enhance viewpoint prediction which is jointly optimized with adaptive streaming. The proposed approach presents a neural network model to predict the viewpoint of the user based on multimodal fusion technique, which takes unexpected events into consideration for refined accuracy. After the prediction, we assign weights to the tiles based on the results and further allocate bandwidth and rate to optimize the user experience. Experimental results on real-world datasets demonstrate that the proposed method outperforms existing approaches, significantly enhancing user experience by 16% on average in adaptive VR streaming. Hsuanyi Lin, Wei Wang 0021, Yitu Wang, Zhaoyang Zhang 0001 |
VTC2025-Fall | 3 |
| 2025 | On Maximizing the Utility of Channel Forecast for Computation OffloadingabstractMobile Edge Computing (MEC) has been successful in proving solid support for delay-sensitive and computation-intensive applications, while invoking channel forecast enlightens a new dimension to further improve the performance. However, separately considering channel forecast and resource management fails in fully exploiting the merit of channel forecast. In this paper, we proceed in two steps. 1) By incorporating channel forecast, we extend the conventional Lyapunov optimization into multi-step-ahead Lyapunov optimization to minimize the queueing delay for non-causal scenario. 2) Based on the obtained insights, we tailor the conventional Long Short Term Memory (LSTM) into Differentiated Randomly Connected LSTM (DR-CLSTM) to obtain a desired trade-off between model complexity and forecast accuracy for the sake of delay minimization. Our simulation results highlight the performance gain of the proposed framework in terms of the system delay. Yitu Wang, Aixing Wang, Xuying Zhou, Wei Wang 0021, Takayuki Nakachi, Juin J. Liou |
VTC2025-Fall | 1 |
| 2025 | Fingerprint Adaptation for mmWave Vehicular Communications Based on Trajectory PredictionabstractMillimeter-wave (mmWave) vehicular communication brings new technical challenges on wireless resource management due to the sensitivity to blockages and the directionality property, as conventional beam alignment techniques suffer from large communication overhead. To enable fast base station (BS) association and beam alignment, we propose a lightweight online learning framework by embracing sparse representation (SR) and Gaussian process (GP). To obtain preliminary information of the transmission environment, fingerprint-based method is advocated for static scenarios, while its performance degrades in dynamic scenarios. To incorporate the influence of vehicle motion, we innovatively propose the idea of trajectory-aware fingerprint, which further triggers the following two designs: 1) Trajectory Prediction: We utilize GP to predict the trajectory of moving vehicles. Noticing the utility of the forecast information drops fast with the computational complexity, we propose a differentiated prediction framework to balance accuracy and model complexity to maximize such utility and 2) Fingerprint Adaptation: As the existence of infinite number of trajectories, we approximate a trajectory using grayscale image, and prove the influence of such approximation on throughput is limited. Then, given a predicted trajectory, SR is invoked to perform robust fingerprint adaptation that facilitating resource management. Finally, the simulation results demonstrate the superiority of the proposed framework. Guangchen Zhang, Xuying Zhou, Yitu Wang, Takayuki Nakachi, Wei Wang 0021, Juin J. Liou |
IEEE Internet Things J. | 3 |
| 2025 | When Average Delay Optimization Meets Deterministic Delay Constraint: A Renewal Framework for Resource AllocationabstractWhile the average delay is traditionally an importance metric for system performance, the emerging technologies have given birth to a variety of critical applications, for which the deterministic delay guarantee is highly desired. When optimizing the average delay objective is embraced with satisfying the deterministic delay constraint, it leads to complicated coupling and brings new challenges to resource allocation. In this paper, we propose a renewal framework for multi-user power control and subband allocation to improve the comprehensive delay performance. The average delay objective is optimized under the Markov decision process (MDP) problem, while the deterministic delay constraint is satisfied through Lyapunov optimization with virtual queues. Due to the conflict of the inter-slot influence in MDP and the i.i.d. state requirement in Lyapunov approach, we exploit the recurrent property of queue states and construct a renewal system. By solving the equivalent infinite-horizon MDP in the renewal framework, we propose a resource allocation algorithm, which is proved to be asymptotically optimal. Finally, the simulation results demonstrate that the proposed scheme meets the deterministic delay constraint and achieves better average delay performance than existing baselines. Yuze Jin, Wei Wang 0021, Ziwei Zheng, Yitu Wang, Rui Yin 0001, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 4 |
| 2025 | Deep-Unfolding Network Slicing for Deterministic Delay Services in Multi-Access Edge ComputingabstractDeterministic demand of mission-critical applications is essential in edge computing systems for realizing Industry 4.0. However, the conventional average-based network slicing schemes incur unexpected long-tail delay, resulting in the failure to meet strict deterministic delay guarantee. To resolve this issue, in this paper, we construct a two-scale TNS-Net architecture for the URLLC slice under the network slicing paradigm, aiming to meet deterministic end-to-end (E2E) delay requirements of multiple users with minimal resource usage. We consider multi-access edge computing (MEC) and model it as a many-to-one cascade queue, which includes the offloading queues at the user equipments (UEs) and a computation queue at the server. To analyze the delay performance, we decompose the offloading process into transmission and vacation periods, and employ the weighted approximation to address the multi-UE coupling in the computation process to derive the closed-form approximate E2E delay distribution. Based on the derived delay distribution, we propose an iterative two-scale network slicing (TNS) algorithm to guarantee deterministic delay, and construct a TNS-based deep-unfolding neural network, called TNS-Net, to improve the solution in presence of inaccurate channel statistics. Moreover, for the training of TNS-Net with deterministic delay as the network input, we apply extreme value theory (EVT) to analyze the distribution characteristic of delay bound violation. Finally, simulation results demonstrate that our theoretical analysis provides a relatively accurate estimate and the proposed TNS-Net ensures better delay guarantee with lower resource consumption. Xinglin Yang, Wei Wang 0021, Yitu Wang, Bing Hu 0002, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 3 |
| 2024 | ICGMM: CXL-enabled Memory Expansion with Intelligent Caching Using Gaussian Mixture ModelabstractCompute Express Link (CXL) emerges as a solution for wide gap between computational speed and data communication rates among host and multiple devices. It fosters a unified and coherent memory space between host and CXL storage devices such as such as Solid-state drive (SSD) for memory expansion, with a corresponding DRAM implemented as the device cache. However, this introduces challenges such as substantial cache miss penalties, sub-optimal caching due to data access granularity mismatch between the DRAM "cache" and SSD "memory", and inefficient hardware cache management. To address these issues, we propose a novel solution, named ICGMM, which optimizes caching and eviction directly on hardware, employing a Gaussian Mixture Model (GMM)-based approach. We prototype our solution on an FPGA board, which demonstrates a noteworthy improvement compared to the classic Least Recently Used (LRU) cache strategy. We observe a decrease in the cache miss rate ranging from 0.32% to 6.14%, leading to a substantial 16.23% to 39.14% reduction in the average SSD access latency. Furthermore, when compared to the state-of-the-art Long Short-Term Memory (LSTM)-based cache policies, our GMM algorithm on FPGA showcases an impressive latency reduction of over 10,000 times. Remarkably, this is achieved while demanding much fewer hardware resources. Hanqiu Chen, Yitu Wang, Vitorio Cargnini, Mohammadreza Soltaniyeh, Gongjin Sun, Pradeep Subedi, Yiran Chen 0001, Cong Hao |
DAC | 2 |
| 2024 | Improving the Efficiency of In-Memory-Computing Macro with a Hybrid Analog-Digital Computing Mode for Lossless Neural Network InferenceabstractAnalog in-memory-computing (IMC) is an attractive technique with a higher energy efficiency to process machine learning workloads. However, the analog computing scheme suffers from large interface circuit overhead. In this work, we propose a macro with a hybrid analog-digital mode computation to reduce the precision requirement of the interface circuit. Considering the distribution of the multiplication and accumulation (MAC) value, we propose a nonlinear transfer function of the computing circuits by only accurately computing low MAC value in the analog domain with a digital mode to deal with the high MAC value with smaller possibility. Silicon measurement results show that the proposed macro could achieve 160 GOPS/mm2 area efficiency and 25.5 TOPS/W for 8b/8b matrix computation. The architectural-level evaluation for real workloads shows that the proposed macro can achieve up to 2.92× higher energy efficiency than conventional analog IMC designs. Qilin Zheng, Ziru Li, Jonathan Hao-Cheng Ku, Yitu Wang, Brady Taylor, Deliang Fan, Yiran Chen 0001 |
DAC | 4 |
| 2024 | NDSEARCH: Accelerating Graph-Traversal-Based Approximate Nearest Neighbor Search through Near Data ProcessingabstractApproximate nearest neighbor search (ANNS) is a key retrieval technique for vector database and many data center applications, such as person re-identification and recommendation systems. It is also fundamental to retrieval augmented generation (RAG) for large language models (LLM) now. Among all the ANNS algorithms, graph-traversal-based ANNS achieves the highest recall rate. However, as the size of dataset increases, the graph may require hundreds of gigabytes of memory, exceeding the main memory capacity of a single workstation node. Although we can do partitioning and use solid-state drive (SSD) as the backing storage, the limited SSD I/O bandwidth severely degrades the performance of the system. To address this challenge, we present NDSEARCh, a hardware-software co-designed near-data processing (NDP) solution for ANNS processing. NDSeARCH consists of a novel in-storage computing architecture, namely, SEARSSD, that supports the ANNS kernels and leverages logic unit (LUN)-level parallelism inside the NAND flash chips. NDSEARCH also includes a processing model that is customized for NDP and cooperates with SearSSD. The processing model enables us to apply a two-level scheduling to improve the data locality and exploit the internal bandwidth in NDSearch, and a speculative searching mechanism to further accelerate the ANNS workload. Our results show that NDSEARCH improves the throughput by up to $31.7 \times, 14.6 \times, 7.4 \times 2.9 \times$ over CPU, GPU, a state-of-the-art SmartSSD-only design, and DeepStore, respectively. NDSEARCH also achieves two orders-of-magnitude higher energy efficiency than CPU and GPU. Yitu Wang, Shiyu Li 0001, Qilin Zheng, Linghao Song, Zongwang Li, Hai Li 0001, Yiran Chen 0001 |
ISCA | 1 |
| 2024 | Hybrid Digital/Analog Memristor-based Computing Architecture for Sparse Deep Learning AccelerationabstractFine-grained sparsity in recent bio-inspired models such as attention-based model could reduce the computation complexity dramatically. However, the unique sparsity pattern challenges the mapping efficiency of the conventional pure analog memristor-based computing architecture, as the conventional one uses a vector-matrix-multiplication primitives. To fill the gap between the memristor-based architecture and the sparse processing, in this paper, we would like to present our recent progress by using a hybrid digital/analog memristor-based computing architecture to improve the mapping efficiency. Our evaluation result shows that, over previous pure analog memristor-based architecture, our design could deliver up to 8.32× performance improvement and 3.4× energy efficiency improvement on a range of vision and language tasks for the recent attention-based bio-inspired model. Qilin Zheng, Shiyu Li 0001, Yitu Wang, Ziru Li, Yiran Chen 0001, Hai Li 0001 |
ISCAS | 3 |
| 2024 | Adaptive Modulation and Coding for URLLC RetransmissionabstractFor ultra-reliable low-latency communication (URLLC), retransmission data should be scheduled with different priorities due to the urgency, which brings new challenges in delay-oriented optimization, especially with deterministic delay constraint. In this paper, we propose an adaptive modulation and coding (AMC) scheme for both initial transmissions and retransmissions in separate data queues. Different from most of the existing works focusing on average delay, we jointly consider delay performance and deterministic delay requirement by combining the Markov decision process (MDP) with Lyapunov optimization technique. To overcome the coupling in the objective and the deterministic constraint, we transform this problem into an infinite horizon MDP by constructing a renewal system with sampling. Based on this, we propose a delay-optimal modulation and coding scheme (MCS) selection policy using reinforcement learning. Simulation results show that the proposed scheme achieves better delay performance than the conventional AMC schemes. Yuze Jin, Wei Wang 0021, Ziwei Zheng, Yitu Wang, Zhaoyang Zhang 0001 |
WCNC | 4 |
| 2024 | Privacy-Preserving Resource Management for Distributed Collaborative Edge Caching SystemsabstractCaching sheds a light on reducing long-distance data transmissions over networks, while raising significant privacy concerns. Moving one step ahead, collaborative edge caching is proposed to facilitate preserving user privacy via reducing the external data exposure. However, it still fails to avert the risk of privacy leakage from nearby edge devices. To tackle this issue, we develop an analytical framework for privacy preserving joint communication and content allocation algorithm for distributed collaborative edge caching systems, in which edge devices collaboratively cache and share the content items based on the dummy-based privacy preservation mechanism. Specifically, we define the system request uncertainty criterion from the perspective of information entropy to measure the privacy preservation performance. Consequently, the closed-form relationship between the system request uncertainty and the resource allocation decisions on both communication resources and content items can be derived. Then, we decompose the NP-hard resource management problem into two parts, and propose 1) an optimal dummy request allocation strategy through investigating special properties of the maximal allocation reward gain and 2) an asymptotically optimal content item allocation strategy with low complexity based on the extract penalty method (EPM), which are iterated to obtain a viable solution, followed by the proof of convergence and asymptotic monotone property. Finally, the performance improvements are verified by simulations. Qi Chen 0017, Yitu Wang, Wei Wang 0021, Takayuki Nakachi, Zhaoyang Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2024 | Content-Caching-Oriented Popularity Forecast and User ClusteringabstractContent popularity forecast is a key enabler toward the realization of proactive content caching, contributing to significant reduction of content fetching delay. Different from most of the existing literature that concentrating on enhancing the forecast accuracy, we tailor the popularity forecast and user clustering algorithms for improving the caching performance. Specifically, through analyzing the caching performance drop incurred by inaccurate popularity forecast from the Bayesian perspective, we obtain two critical insights, which trigger the following designs: 1) as the utility of forecast varies according to the content rank, we propose a content-caching-oriented popularity forecast algorithm based on Gaussian process (GP), where more computational resource is allocated to forecast the popularity of prioritized contents and 2) to alleviate the influence of forecast error on the rank of prioritized contents, we propose a content-caching-oriented user clustering algorithm based on the K-means algorithm. Since the involved optimization problem is NP-hard, we propose an iterative algorithm, whose convergence property in terms of region stability is proved, as the objective function may vary before a local minima is reached. Finally, the simulation results demonstrate the superiority of the proposed framework. Yitu Wang, Qi Chen 0017, Wei Wang 0021, Takayuki Nakachi, Guangchen Zhang, Juin J. Liou |
IEEE Internet Things J. | 1 |
| 2024 | NDRec: A Near-Data Processing System for Training Large-Scale Recommendation ModelsabstractRecent advances in deep neural networks (DNNs) have enabled highly effective recommendation models for diverse web services. In such DNN-based recommendation models, the embedding layer comprises the majority of model parameters. As these models scale rapidly, the embedding layer’s memory capacity and bandwidth requirements threaten to exceed the limits of current computing architectures. We observe the embedding layer’s computational demands increase much more slowly than its storage needs, suggesting an opportunity to offload embeddings to storage hardware. In this work, we present NDRec, a near-data processing system to train large-scale recommendation models. NDRec offloads both the parameters and the computation of the embedding layer to computational storage devices (CSDs), using coherence interconnects (CXLs) for communication between GPUs and CSDs. By leveraging the statistical properties of embedding access patterns, we develop an optimized CSD memory hierarchy and caching strategy. A lookahead embedding scheme enables concurrent execution of embeddings and other operations, hiding latency and reducing memory bandwidth requirements.We evaluate NDRec using real-world and synthetic benchmarks. Results demonstrate NDRec achieves up to 4.33× and 3.97× speedups over heterogeneous CPU-GPU platforms and GPU caching, respectively. NDRec also reduces per-iteration energy consumption by up to 54.9%. Shiyu Li 0001, Yitu Wang, Edward Hanson, Yang-Seok Ki, Hai Li 0001, Yiran Chen 0001 |
IEEE Trans. Computers | 2 |
| 2024 | Retransmission Aware Adaptive Modulation and Coding Toward Deterministic Delay PerformanceabstractUltra-reliable low-latency communication (URLLC) is an indispensable element towards supporting various latency-sensitive and reliability-critical applications. To optimize the average delay while satisfying the deterministic delay constraint, initial transmission and retransmission should be handled with different priorities due to the differentiated urgency, which creates complex interdependency and brings new technical challenges to delay-oriented optimization. In this paper, we propose a retransmission-aware adaptive modulation and coding (RAMC) scheme to improve the delay performance in URLLC scenarios. Specifically, we first establish a cascaded queue system, including an initial transmission queue and a retransmission queue. The deterministic delay constraint is satisfied through Lyapunov optimization, where we transform the Lyapunov drift-plus-penalty problem into an infinite horizon Markov decision process (MDP) by constructing a renewal system with sampling to overcome the challenge brought by queue coupling. Next, we propose the delay-optimal RAMC scheme by solving the associated Bellman equation by improved reinforcement learning, which is proved to be asymptotically optimal. Finally, the superiority of the proposed RAMC scheme is verified through simulations. Yuze Jin, Wei Wang 0021, Yitu Wang, Rui Yin 0001, Ziwei Zheng, Zhaoyang Zhang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2023 | Accelerating Sparse Attention with a Reconfigurable Non-volatile Processing-In-Memory ArchitectureabstractAttention-based neural networks have shown superior performance in a wide range of tasks. Non-volatile processing-in-memory (NVPIM) architecture shows its great potential to accelerate the dense attention model. However, the unique unstructured and dynamic sparsity pattern in the sparse attention model challenges the mapping efficiency of the NVPIM architecture, as the conventional NVPIM architecture uses a vector-matrix-multiplication primitives. In this paper, we propose a NVPIM architecture to accelerate a dynamic and unstructured sparse computation in the sparse attention. We aim to improve the mapping efficiency for both SDDMM and SpMM by introducing two vector-based primitives with a reconfigurable NVPIM bank. Further, based on our reconfigurable NVPIM bank, we further propose a hybrid stationary data flow to hide the latency. Our evaluation result shows that, over previous NVPIM accelerators, our design could deliver up to 12.36× performance improvement and 3.4× energy efficiency improvement on a range of vision and language tasks. Qilin Zheng, Shiyu Li 0001, Yitu Wang, Ziru Li, Yiran Chen 0001, Hai Li 0001 |
DAC | 3 |
| 2023 | Si-Kintsugi: Towards Recovering Golden-Like Performance of Defective Many-Core Spatial Architectures for AIabstractThe growing demand for higher compute and memory capacity driven by artificial intelligence (AI) applications pushes higher core counts in modern systems. Many-core architectures exhibiting spatial interconnects with high on-chip bandwidth are ideal for these workloads due to their data movement flexibility and sheer parallelism. However, the size of such platforms makes them particularly susceptible to manufacturing defects, prompting a need for designs and mechanisms that improve yield. Despite these techniques, nonfunctional cores and links are unavoidable. Although prior works address defective cores by disabling them and only scheduling workload to functional ones, communication latency through spatial interconnects is tightly associated with the locations of defective cores and cores with assigned work. Based on this observation, we present Si-Kintsugi, a defect-aware workload scheduling framework for spatial architectures with mesh topology. First, we design a novel and generalizable workload mapping representation and cost function that integrates defect pattern information. The mapping representation is formed into a 1D vector with simple constraints, making it an ideal candidate for open source heuristic-based optimization algorithms. After a communication latency optimized workload mapping is found, dataflow between the mapped cores is automatically generated to balance communication and computation cost. Si-Kintsugi is extensively evaluated on various workloads (i.e., BERT, ResNet, GEMM) across a wide range of defect patterns and rates. Experiment results show that Si-Kintsugi generates a workload schedule that is on average 1.34 × faster than the industry standard layer-pipelined schedule on defective platforms. Edward Hanson, Shiyu Li 0001, Guanglei Zhou, Yitu Wang, Rohan Bose, Hai Li 0001, Yiran Chen 0001 |
MICRO | 5 |
| 2023 | A Light-weight Online Learning Framework for Network Traffic Abnormality DetectionabstractNetwork traffic monitoring plays a crucial role in maintaining the security and reliability of the communication networks. Although Machine Learning (ML) assisted abnormal traffic detection has been emerged as a promising paradigm, the existing data-driven learning-based approaches are faced with challenges on inefficient traffic feature extraction and high computational complexity, especially when taking the evolving property of traffic process into consideration. To this end, we establish an online learning framework for abnormality traffic detection by embracing Gaussian Process (GP) and Sparse Representation (SR). The contributions of this paper are two-fold: 1). We utilize a special kernel, i.e., mixture of Gaussian, to better explore and exploit the evolving traffic characteristics, so as to more accurately model network traffic. 2). To combat noise and modeling error, we formulate a feature vector based on Kullback-Leibler (KL) divergence to measure the difference between normal and abnormal traffic, based on which SR is adopted to perform robust binary classification. Finally, we demonstrate the superiority of the proposed framework in terms of detection accuracy through simulation. Yitu Wang, Runqi Dong, Takayuki Nakachi, Wei Wang 0021 |
WCNC | 1 |
| 2023 | Stochastic Resource Allocation and Delay Analysis for Mobile Edge Computing SystemsabstractTo alleviate the local computation demands from the ever-increasing computation-intensive mobile applications, Mobile Edge Computing (MEC) has proved promising. Especially, by opportunistically offloading these computation tasks to the MEC server, the delay of computing could be significantly improved through communication. In this paper, we develop an analytical framework for joint communication and computation resources allocation for multi-user MEC systems. Specifically, to retrieve the combined effect of communication and computation capabilities, we establish a dual queue system, including a data queue sub-system and a computation queue sub-system. To address the associated stochastic resource optimization problem, we propose a low-complexity resource allocation algorithm by Lyapunov optimization to stabilize all the sub-queue systems. As the practical buffers are finite, the conventional delay analysis of Lyapunov optimization becomes inaccurate. Alternatively, we model the stochastic queue lengthes as discrete time controlled random walk processes, which are transformed to continuous time Stochastic Differential Equations (SDEs) with reflections by strong approximation. According to the steady state analysis on the SDEs, we derive closed-form steady state distributions of the queue lengths, and then obtain the average delay performance with finite buffers. Finally, the accuracy of the proposed delay analysis is verified through simulation. Yitu Wang, Wei Wang 0021, Vincent K. N. Lau, Takayuki Nakachi, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 1 |
| 2023 | EMS-i: An Efficient Memory System Design with Specialized Caching Mechanism for Recommendation InferenceabstractRecommendation systems have been widely embedded into many Internet services. For example, Meta’s deep learning recommendation model (DLRM) shows high prefictive accuracy of click-through rate in processing large-scale embedding tables. The SparseLengthSum (SLS) kernel of the DLRM dominates the inference time of the DLRM due to intensive irregular memory accesses to the embedding vectors. Some prior works directly adopt near data processing (NDP) solutions to obtain higher memory bandwidth to accelerate SLS. However, their inferior memory hierarchy induces low performance-cost ratio and fails to fully exploit the data locality. Although some software-managed cache policies were proposed to improve the cache hit rate, the incurred cache miss penalty is unacceptable considering the high overheads of executing the corresponding programs and the communication between the host and the accelerator. To address the issues aforementioned, we propose EMS-i , an efficient memory system design that integrates Solide State Drive (SSD) into the memory hierarchy using Compute Express Link (CXL) for recommendation system inference. We specialize the caching mechanism according to the characteristics of various DLRM workloads and propose a novel prefetching mechanism to further improve the performance. In addition, we delicately design the inference kernel and develop a customized mapping scheme for SLS operation, considering the multi-level parallelism in SLS and the data locality within a batch of queries. Compared to the state-of-the-art NDP solutions, EMS-i achieves up to 10.9× speedup over RecSSD and the performance comparable to RecNMP with 72% energy savings. EMS-i also saves up to 8.7× and 6.6 × memory cost w.r.t. RecSSD and RecNMP, respectively. Yitu Wang, Shiyu Li 0001, Qilin Zheng, Hai Li 0001, Yiran Chen 0001 |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2023 | Pattern Discovery and Multi-Slot-Ahead Forecast of Network Traffic: A Revisiting to Gaussian ProcessabstractThe forecast of network traffic with arbitrary predicting horizon is a key enabler of smart management in next-generation networks, as sufficient amount of time can be provided for the proactive manipulation of network resources to maintain high quality transmission. Nevertheless, the evolving characteristic of network traffic challenges the current learning-based and data-driven algorithms on both prediction accuracy and computational complexity. In this work, we explore special properties of network traffic, which are further encoded into the Gaussian Process (GP)-based online learning framework, so as to better comprehend and predict future network traffic from a Bayesian perspective. Specifically, we proceed by three steps, 1). Observing network traffic is evolving, to explore and exploit the dynamic traffic patterns at different times and time-scales, we try to approximate the optimal kernel function of GP by utilizing a mixture of Gaussian to encode the dominant and several nondominant patterns. 2). As network traffic at different time-scales share several common patterns, we adopt Process Convolution (PConv) to fully exploit correlations among multiple subsequent time-slots, so as to facilitate network traffic forecast with large predicting horizon. 3). To promote the tracking capability of the proposed GP-PConv framework without significantly increasing the number of hyper-parameters to train, we slightly modify the GP-based prediction through Lyapunov optimization, which brings performance improvements both in terms of accuracy and computational complexity. Finally, we demonstrate the superiority of the proposed algorithm through simulation. Yitu Wang, Takayuki Nakachi, Wei Wang 0021 |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2022 | FedCor: Correlation-Based Active Client Selection Strategy for Heterogeneous Federated LearningabstractClient-wise data heterogeneity is one of the major issues that hinder effective training in federated learning (FL). Since the data distribution on each client may vary dramatically, the client selection strategy can significantly influence the convergence rate of the FL process. Active client selection strategies are popularly proposed in recent studies. However, they neglect the loss correlations between the clients and achieve only marginal improvement compared to the uniform selection strategy. In this work, we propose FedCoran FLframework built on a correlation-based client selection strategy, to boost the convergence rate of FL. Specifically, we first model the loss correlations between the clients with a Gaussian Process (GP). Based on the GP model, we derive a client selection strategy with a significant reduction of expected global loss in each round. Besides, we develop an efficient GP training method with a low communication overhead in the FL scenario by utilizing the covariance stationarity. Our experimental results show that compared to the state-of-the-art method, FedCorr can improve the convergence rates by 34% ~ 99% and 26% ~ 51% on FMNIST and CIFAR-10, respectively. Minxue Tang, Xuefei Ning, Yitu Wang, Jingwei Sun 0002, Yu Wang 0002, Hai Li 0001, Yiran Chen 0001 |
CVPR | 3 |
| 2022 | Access Control for Privacy-Preserving Gaussian Process RegressionabstractIn this paper, we propose access control for privacy-preserving Gaussian process regression (GPR), in which the encrypted data are generated through a random unitary transform (RUT). The proposed secure GPR enables computation in both encrypted input and output domains, and the access to inputs and prediction results can be controlled. We prove that our GPR for encrypted data has the same prediction accuracy as GPR for non-encrypted data. Furthermore, we demonstrate the effectiveness of our method by experimenting with diabetes data from the medical analysis field. Takayuki Nakachi, Yitu Wang |
ICASSP | 2 |
| 2022 | QoE-driven Link Quality Prediction for Video Streaming in Mobile NetworksabstractThe link quality prediction facilitates high quality video streaming over mobile networks. However, the existing link quality prediction algorithms focus on minimizing the gap between the ground truth and the prediction result, while it remains a challenge to exploit such information to achieve high quality video streaming with minimum Quality of Experience (QoE) degradation. The accurate link quality prediction is one of keys to enable beyond 5G/6G world. In this paper, we produce artificial intelligence (AI) based link quality prediction which consists two steps: 1. We explore and exploit the temporal correlation in time series to adaptively learn and predict its short-term behavior based on Gaussian Process (GP). 2. The GP-based prediction is tailored to maximize QoE by finding a proper piece-wise convex envelope of the predicted link quality in an online manner. By using the measured uplink throughputs, the video streaming QoE of the proposed framework were evaluated. Yitu Wang, Riichi Kudo, Yuya Aoki, Yoshifumi Morihiro, Kahoko Takahashi, Hisashi Nagata |
VTC Spring | 1 |
| 2021 | Adaptive Multi-slot-ahead Prediction of Network Traffic with Gaussian ProcessabstractMulti-slot-ahead forecasting on network traffic provides an extra degree of freedom to proactively manipulate the network resources when immediate reconfiguration of networks is expensive or infeasible. In return, it challenges the existing data-driven learning-based approaches on accuracy, especially when considering the evolving property of the traffic process. To this end, we establish an adaptive learning framework for multi-slot-ahead network traffic prediction based on Gaussian Process (GP). GP facilitates learning and comprehending the traffic process from a Bayesian perspective, where the main characteristics can be encoded into the kernel function for performance enhancement. The contributions of this paper are two-fold: 1). To track the evolving traffic characteristics, we approximate the optimal kernel adapting to the current traffic. 2). To predict in a large time horizon without significantly hurt the performance, Linear Model of Co-regionalization (LMC) is utilized to better make use of the correlation among subsequent multiple time-slots. Finally, we demonstrate the high tracking capability as well as the superiority of the proposed framework in terms of prediction accuracy through simulation. Yitu Wang, Takayuki Nakachi, Takeru Inoue, Toru Mano |
GLOBECOM | 1 |
| 2021 | Correlation Discovery and Channel Prediction in Mobile Networks: A Revisiting to Gaussian ProcessabstractWith accurate knowledge of future Channel State Information (CSI), it becomes possible to better comprehend the radio propagating environment and manipulate the wireless resources in a proactive manner, so as to provide solid support to smart and high quality wireless transmission. However, in mobile environment, the evolving correlation patterns in CSI series challenge the existing data-driven algorithms to adaptively learn and predict its behavior. In this article, an adaptive learning algorithm is proposed based on Gaussian Process (GP), to discover and utilize the spatial correlation within a channel and across channels, and produce accurate CSI prediction. Specifically, 1). To track the evolving correlation of a channel, we tailor Spectrum Mixture (SM) kernel to not only approximate the optimal kernel adapting to the current CSI, but also capture the combined effect of path loss and User Equipment (UE) motion. 2). The correlation across channels is encoded into the GP-based learning framework through Linear Model of Co-regionalization (LMC). Finally, we verify the performance improvements through simulation. Yitu Wang, Takayuki Nakachi, Takeru Inoue, Toru Mano, Riichi Kudo |
GLOBECOM | 1 |
| 2021 | Rerec: In-ReRAM Acceleration with Access-Aware Mapping for Personalized RecommendationabstractPersonalized recommendation systems are widely used in many Internet services. The sparse embedding lookup in recommendation models dominates the computational cost of inference due to its intensive irregular memory accesses. Applying resistive random access memory (ReRAM) based process-in-memory (PIM) architecture to accelerate recommendation processing can avoid data movements caused by off-chip memory accesses. However, naïve adoption of ReRAM-based DNN accelerators leads to low computation parallelism and severe under-utilization of computing resources, which is caused by the fine-grained inner-product in feature interaction. In this paper, we propose Rerec, an architecture-algorithm co-designed accelerator, which specializes in fine-grained ReRAM-based inner-product engines with access-aware mapping algorithm for recommendation inference. At the architecture level, we reduce the size and increase the amount of crossbars. The crossbars are fully-connected by Analog-to-Digital Converters (ADCs) in one inner-product engine, which can adapt to the fine-grained and irregular computational patterns and improve the processing parallelism. We further explore trade-offs of (i) crossbar size vs. hardware utilization, and (ii) ADC implementation vs. area/energy efficiency to optimize the design. At the algorithm level, we propose a novel access-aware mapping (AAM) algorithm to optimize resource allocations. Our AAM algorithm tackles the problems of (i) the workload imbalance and (ii) the long recommendation inference latency induced by the great variance of access frequency of embedding vectors. Experimental results show that Rerecachieves 7.69x speedup compared with a ReRAM-based baseline design. Compared to CPU and the state-of-the-art recommendation accelerator, Rerecdemonstrates 29.26x and 3.48x performance improvement, respectively. Yitu Wang, Zhenhua Zhu 0002, Fan Chen 0001, Mingyuan Ma, Guohao Dai 0001, Yu Wang 0002, Hai Li 0001, Yiran Chen 0001 |
ICCAD | 1 |
| 2020 | ReBoc: Accelerating Block-Circulant Neural Networks in ReRAMabstractDeep neural networks (DNNs) emerge as a key component in various applications. However, the ever-growing DNN size hinders efficient processing on hardware. To tackle this problem, on the algorithmic side, compressed DNN models are explored, of which block-circulant DNN models are memory efficient and hardware-friendly; on the hardware side, resistive random-access memory (ReRAM) based accelerators are promising for in-situ processing of DNNs. In this work, we design an accelerator named ReBoc for accelerating block-circulant DNNs in ReRAM to reap the benefits of light-weight models and efficient in-situ processing simultaneously. We propose a novel mapping scheme which utilizes Horizontal Weight Slicing and Intra-Crossbar Weight Duplication to map block-circulant DNN models onto ReRAM crossbars with significant improved crossbar utilization. Moreover, two specific techniques, namely Input Slice Reusing and Input Tile Sharing are introduced to take advantage of the circulant calculation feature in block- circulant DNNs to reduce data access and buffer size. In REBOC, a DNN model is executed within an intra-layer processing pipeline and achieves respectively 96× and 8.86× power efficiency improvement compared to the state-of-the-art FPGA and ASIC accelerators for block-circulant neural networks. Compared to ReRAM-based DNN accelerators, REBOC achieves averagely 4.1× speedup and 2.6× energy reduction. Yitu Wang, Fan Chen 0001, Linghao Song, Chuanjin Richard Shi, Hai Li 0001, Yiran Chen 0001 |
DATE | 1 |
| 2020 | Privacy-Preserving Pattern Recognition Using Encrypted Sparse Representations in L0 Norm MinimizationabstractIn this paper, we propose a privacy-preserving pattern recognition method that uses encrypted sparse representations in L0 norm minimization. We prove, theoretically, that the proposal has exactly the same dictionary and sparse coefficient estimation performance as the Label Consistent K-Singular Value Decomposition (LC-KSVD) algorithm for non-encrypted signals. It can be directly implemented by the LC-KSVD algorithm without any modification. Finally, we demonstrate its excellent recognition performance and security strength for the face recognition task using the Extended YaleB database. Takayuki Nakachi, Yitu Wang, Hitoshi Kiya |
ICASSP | 2 |
| 2020 | Secure Face Recognition in Edge and Cloud Networks: From the Ensemble Learning PerspectiveabstractOffloading the computationally intensive workloads to the edge and cloud not only improves the quality of computation, but also creates an extra degree of diversity by collecting information from devices in service, which, in turn, has raised significant concerns on privacy as the aggregated information could be misused without the permission by the third party. Sparse coding, which has been successful in computer vision, is finding application in this new domain. In this paper, we develop a secure face recognition framework to orchestrate sparse coding in edge and cloud networks. Specifically, 1). To protect the privacy, we develop a low-complexity encrypting algorithm based on random unitary transform, where its influence on dictionary learning and sparse representation is analysed. We further prove that such influence will not affect the accuracy of face recognition. 2). To fully utilize the multi-device diversity, we extract deeper features in an intermediate space, expanded according to the dictionaries from each device, and perform classification in this new feature space to combat the noise and modeling error. Yitu Wang, Takayuki Nakachi |
ICASSP | 1 |
| 2020 | The Learning and Prediction of Network Traffic: A Revisiting to Sparse RepresentationabstractWith accurate network traffic prediction, future communication networks can realize self-management and enjoy intelligent and efficient automation. Benefiting from discovering the sparse property of network traffic in temporal domain, it becomes possible to develop compact algorithms with high accuracy and low computational complexity. For this purpose, we establish an analytical framework for network traffic prediction by extending traditional sparse representation to predictive sparse representation, and try to take the full advantage of such sparsity. Specifically, 1). To equip sparse representation with predictive capability, we divide the historical traffic records into two sets, and jointly train the representative/predictive dictionaries, such that the query point is embedded in terms of a sparse combination of dictionary atoms, and jointly coded with its T+1 time slot behind counterpart. 2). To estimate the sparse code of the query point, we only have to decompose its counterpart into a sparse combination of the representative dictionary atoms by adopting iterative projection method, which provides extra flexibility and adaptability in determining the dependence range. After this, the prediction is performed based on the predictive dictionary. 3). To promote the capability of capturing the rapidly changing traffic, we slightly modify the sparse representation-based prediction by adopting Lyapunov optimization, and minimize the time averaged prediction error. Finally, our proposed algorithm is evaluated by simulation to show its superiority over the conventional schemes. Yitu Wang, Takayuki Nakachi |
ICC | 1 |
| 2020 | Light-weight Machine Learning for mmWave Vehicular CommunicationsabstractExploiting the vacant spectrum resource at mmWave bands provides the potential for fulfilling the requirements of broadband services. Nevertheless, the sensitivity of mmWave to blockages together with its directionality bring new technical challenges in the vehicular context. In particular, traditional beam training is inadequate in satisfying low communication overhead and small latency, and the influence of the same blockage on a fixed position varies according to vehicle motion. To facilitate fast beam alignment, fingerprint-based method stands out as an efficient solution, where the utilities of selecting different beam pairs are recorded in the fingerprint database for reference at a given position. In order to better orchestrate fingerprint-based method with vehicular scenarios, we propose the idea of trajectory-aware fingerprint, which extracts the combined effect of mobility, propagation environment, and blockages, so as to faithfully reflect the transmitting condition in mobile scenarios. Then, a light-weight machine learning framework is established for intelligent adaptation among multiple fingerprints to find a near-optimal BS association and beam alignment solution. Finally, the simulation result verifies the performance improvements. Yitu Wang, Takayuki Nakachi |
VTC Fall | 1 |
| 2019 | Load scheduling for distributed edge computing: A communication-computation tradeoff
Wei Wang 0021, Yitu Wang, Zhaoyang Zhang 0001 |
Peer-to-Peer Netw. Appl. | 3 |
| 2019 | Closed-Form Delay-Optimal Computation Offloading in Mobile Edge Computing SystemsabstractMobile edge computing (MEC) has recently emerged as a promising technology to release the tension between computation-intensive applications and resource-limited mobile terminals (MTs). In this paper, we study the delay-optimal computation offloading in computation-constrained MEC systems. We consider the computation task queue at the MEC server due to its constrained computation capability. In this case, the task queue at the MT and that at the MEC server are strongly coupled in a cascade manner, which creates complex interdependences and brings new technical challenges. We model the computation offloading problem as an infinite horizon average cost Markov decision process (MDP) and approximate it to a virtual continuous time system (VCTS) with reflections. Different from most of the existing works, we develop the dynamic instantaneous rate estimation for deriving the closed-form approximate priority functions in different scenarios. Based on the approximate priority functions, we propose a closed-form multi-level water-filling computation offloading solution to characterize the influence of not only the local queue state information (LQSI) but also the remote queue state information (RQSI). Furthermore, we discuss the extension of our proposed scheme to multi-MT multi-server scenarios. Finally, the simulation results show that the proposed scheme outperforms the conventional schemes. Xianling Meng, Wei Wang 0021, Yitu Wang, Vincent K. N. Lau, Zhaoyang Zhang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2018 | Delay-Optimal Computation Offloading for Computation-Constrained Mobile Edge NetworksabstractMobile edge computing (MEC) has recently emerged as a promising technology to release the tension between computation-intensive applications and resource-limited mobile terminals (MTs). In this paper, we construct a Markov decision process (MDP) framework to optimize the delay performance in computation-constrained mobile edge networks. Different to most of the existing works, we consider the computation task queue at the MEC server due to its constrained computation capability. In this case, the task queue at the MT and that at the MEC server are mutually strongly coupled in a cascade manner, which creates complex interdependence and brings new technical challenges. To address the challenge, the infinite horizon average cost MDP is reformulated to a virtual continuous time system (VCTS) with reflections. We derive the closed-form approximate priority functions for the MDP in different system scenarios with dynamic rate estimation. Based on the approximate priority function, we propose a multi-level water filling solution to characterize the influence of not only the local queue state information (LQSI) but also the remote queue state information (RQSI) on the computation offloading policy in a closed form. Finally, the simulation results show that the proposed computation offloading scheme outperforms the conventional schemes. Xianling Meng, Wei Wang 0021, Yitu Wang, Vincent K. N. Lau, Zhaoyang Zhang 0001 |
GLOBECOM | 3 |
| 2018 | On the Cooperation for Content Caching from a Coalitional Game PerspectiveabstractCooperative content caching has been demonstrated to achieve significant performance gain over the conventional content caching paradigm by exploiting content diversity through the participation of multiple cooperative nodes. Although cooperative content caching has the potential to increase the efficiency, an improper coalition formation may result in severe performance degradation. Therefore, the cooperative nodes should be carefully selected according to their interests in different content objects. In this paper, we develop an analytical framework for cooperative content caching from a coalitional game perspective. The cooperation issue for content caching among nodes is studied by the coalitional game theory, and the associated problems are analyzed in different cases that the utility transfer among nodes is allowed or not. If the utility transfer is allowed, by exploiting the properties of the coalitional costs, we derive the non-empty property of the core of a transferable utility coalitional game, and prove that the grand coalition is stable in spite of the presence of coalition costs. If the utility transfer is not allowed, we adopt a non-transferable utility coalitional game model. The grand coalition is not always stable in the presence of coalition costs. A merge and split algorithm is proposed to form the coalitional structure for iteratively improving the caching performance. Finally, the simulation results demonstrate the cooperation gains on both the sum and individual utilities in different scenarios. Xuying Zhou, Wei Wang 0021, Yitu Wang, Zhaoyang Zhang 0001 |
GLOBECOM | 3 |
| 2018 | Relay Selection for Multi-Channel Cooperative Multicast: Lexicographic Max-Min OptimizationabstractCooperative multicast has been demonstrated to achieve significant performance gain over the classic source-destination transmission paradigm by exploiting spatial diversity through the participation of multiple relay nodes. As a major technical challenge, the selection of relays for a multicast session has significant impact on the multicast performance. The challenge is even more pronounced when the number of channels is limited as the relay selection is in this context coupled with channel allocation. The goal of this paper is to design a fair multicast relay selection scheme with limited channel resources. Specifically, we establish an analytical framework for this joint relay selection and channel allocation problem and develop a lexicographic max-min multicast relay selection scheme. Our design consists of two technical steps. First, we consider the maximization of the minimal data rate. By decoupling relay selection and channel allocation, the problem is transformed to a max-min-max problem, which is difficult to solve. To make this problem tractable, we reformulate it as a convex optimization problem via relaxation and smoothing, and prove the asymptotic equivalence from a geometrical perspective. Second, we propose an adjustment algorithm based on the initial max-min solution, and prove that the proposed scheme achieves lexicographic optimality. Finally, our proposed algorithm is evaluated by simulation to show its superiority over the conventional schemes. Yitu Wang, Wei Wang 0021, Lin Chen 0002, Pan Zhou 0001, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 1 |
| 2018 | Distributed Packet Forwarding and Caching Based on Stochastic Network Utility Maximization
Yitu Wang, Wei Wang 0021, Ying Cui 0001, Kang G. Shin, Zhaoyang Zhang 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2018 | Heterogeneous Spectrum Aggregation: Coexistence From a Queue Stability PerspectiveabstractSpectrum aggregation (SA) across heterogeneous channels, including both dedicated and shared channels, provides the potential for improving spectrum utilization and fulfilling the requirement of broadband services. Heterogeneous SA brings new technical challenges on multisystem coexistence on shared channels and the resource allocation over heterogeneous channels. In this paper, we develop an analytical framework for heterogeneous SA from a queue stability perspective. To make all systems on the shared channels stable, we design a resource allocation algorithm for the coexistence of multiple systems. Specifically, we derive the closed-form modified water-filling power control for the single-pair case by Lyapunov optimization and prove that it achieves the queue stability for all systems. Based on the results, we propose a low-complexity suboptimal resource allocation algorithm for multipair SA, which is a NP-hard problem. We partition user pairs into groups by using graph coloring and allocate the shared channels to pair groups according to the maximal weight bipartite matching model. The simulation results verify the queue stability and show that the proposed schemes outperform the conventional schemes. Yitu Wang, Wei Wang 0021, Vincent K. N. Lau, Lin Chen 0002, Zhaoyang Zhang 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2017 | Content Caching Clustering Based on Piecewise Interest SimilarityabstractCooperative caching is a promising technology for enhancing user experience and reducing redundant transmissions through the participation of multiple caching nodes. In this paper, we design a clustering algorithm for the sectionalized caching, in which each user divides its caching space into two parts and the contents cached in these two parts are determined according to the individual interest and the joint interest of all users in the same cluster respectively. Different to most of the existing works forming the clusters based on the interest similarity of all files, we adopt the piecewise interest similarity as the criterion of clustering, which takes advantage of the content diversity and contributes to the reduction of the transmission delay. We measure the gain of the cooperation between two users and obtain the piecewise interest similarity for two users accordingly. Since the gain of clustered caching highly depends on the formed cluster structure, we estimate the gain of the clustered caching based on the piecewise interest similarities by online learning and propose an affinity propagation (AP) based clustering algorithm. Finally, our proposed clustering algorithm is evaluated by simulation to show its superiority over the conventional clustering algorithms. Qi Chen 0017, Wei Wang 0021, Yitu Wang, Pan Zhou 0001, Zhaoyang Zhang 0001 |
GLOBECOM | 3 |
| 2017 | Computational resource constrained multi-cell joint processing in cloud radio access networksabstractCentralized joint processing has been demonstrated to achieve significant performance gain over the conventional per-cell processing by avoiding the inter-cell interference in cloud radio access networks (C-RANs), but full scale cooperation over the entire network is infeasible due to its huge computational complexity. The goal of this paper is to design remote radio head (RRH) clustering and the associated power allocation algorithm under computational resource constraint for C-RANs, which is challenging due to the combinatorial clustering problem with non-convex constraints. To overcome the challenge, our design consists of two technical steps. 1) Considering the maximization of the sum data rate, we reformulate both the objective function and the computational resource constraint via relaxation and successive convex approximation (SCA) technique. 2) By exploiting the structure of this problem, we propose a low-complexity algorithm to decouple RRH clustering and power allocation and solve them iteratively. Furthermore, we prove the convergence property of the proposed iterative algorithm. Finally, our proposed algorithm is evaluated by simulation to show its superiority on throughput performance over conventional algorithms. Wei Wang 0021, Yitu Wang, Zhaoyang Zhang 0001 |
ICC | 3 |
| 2017 | Lexicographic Relay Selection and Channel Allocation for Multichannel Cooperative MulticastabstractCooperative multicast has been demonstrated to achieve significant performance gain over the classic source-destination transmission paradigm by exploiting spatial diversity through the participation of multiple relay nodes. As a major technical challenge, the selection of relays for a multicast session has significant impact on the multicast performance. The challenge is even more pronounced when the number of channels are limited as the relay selection is in this context coupled with channel allocation. We establish an analytical framework for joint relay selection and channel allocation problem and develop a lexicographic max-min multicast relay selection scheme. Our design consists of two technical steps. 1) We consider the maximization of the minimal data rate. By decoupling relay selection and channel allocation, the problem is transformed to a max-min-max problem, which is difficult to solve. To make this problem tractable, we reformulate it as a convex optimization problem via relaxation and smoothing, and prove the asymptotic equivalence from a geometrical perspective. 2) We propose an adjustment algorithm based on the initial max-min solution, and prove that the proposed scheme achieves lexicographic optimality. Yitu Wang, Wei Wang 0021, Lin Chen 0002, Zhaoyang Zhang 0001 |
WCNC | 1 |
| 2017 | Moderate Incentive Design for Delay-Constrained Device-to-Device Relaying
Xuying Zhou, Wei Wang 0021, Yitu Wang, Lin Chen 0002, Zhaoyang Zhang 0001 |
Mob. Networks Appl. | 3 |
| 2016 | Energy Efficient Scheduling for Delay-Constrained Spectrum AggregationabstractIn this paper, we construct an analytical design framework for energy efficient scheduling for delay-constrained spectrum aggregation (ESSA), where the practical hardware limitations on SA capability bring various technical challenges. Specifically, the conventional water-filling power control cannot be adopted over all the channels, and the delay-aware scheduling solution should interact with the channel allocation. To overcome these challenges, we design the ESSA scheduling scheme in two steps. First, with given rate vector and channel allocation, we minimize the total power consumption for SA, including both the transmit power as well as the circuit power. Due to the properties of delay-constrained SA, we divide the scheduled users into conforming and nonconforming user sets, and design their water-filling power allocation strategies differentially. Second, based on the differentiated water-filling power control, we optimize the channel allocation and rate control iteratively via Lyapunov optimization to minimize the power consumption with the average delay constraint. The proposed ESSA scheme is finally evaluated by simulation results. Yitu Wang, Wei Wang 0021, Lin Chen 0002, Zhaoyang Zhang 0001 |
GLOBECOM | 1 |