EDBT 2026 Demo / reviewers in the wild / expert
Jun Wu 0006
dblp:20/3894-6
· DBLP profile ↗
88ranked-venue papers
2as first author
38since 2021 · last 2026
0000-0001-7090-8653ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 46 · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Security and privacy · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HFDFM: A Heterogeneous Credit Card Fraud Detection Model Based on Federated Learning With Membership PrivacyabstractCredit card fraud brings serious losses to both cardholders and card issuers. To reduce losses caused by fraudulent behaviors, banking institutions establish credit card fraud detection (CCFD) models to identify potential fraudulent behaviors. To develop more effective fraud detection models, banking institutions need to collaborate on model training. Federated learning (FL) enables collaboration on fraud detection model training without exchanging data between banking institutions. Nevertheless, the distribution of transaction data in the real-world is heterogeneous among banking institutions, which may lead to convergence issues in global fraud detection models. Furthermore, the behavior and weights of the model may implicitly contain the cardholders’ personal information, which makes existing federated models prone to transaction data leakage. In this article, we propose aheterogeneous credit cardfrauddetection model based onfederated learning withmembership privacy, called HFDFM. To protect the sensitive information of cardholders, we design a novel mechanism that ensures training-data confidentiality by minimizing the accuracy of the best black-box membership inference attack (MIA) against the model. Unlike previous FL frameworks that either optimize for client drift or for membership privacy, HFDFM simultaneously mitigates drift and provides certified membership privacy through a single min–max game. Additionally, we use the control variable to rectify the client drift in its local update in the scenario of heterogeneous data. Extensive experimental results on three mainstream real-world transaction datasets demonstrate that the proposed HFDFM has advantages in utility compared to 11 SOTA baselines, and the proposed HFDFM can mitigate the risks of MIAs (near random guess). Jun Wu 0006, Kun Zhu 0024, Rongkun Cui, Zhe Liu 0001, Changjun Jiang 0002 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2026 | Defrag: Reducing Resource Fragmentation in Large-Scale Heterogeneous GPU ClustersabstractTechnology companies have built large-scale heterogeneous GPU clusters to support various workloads. However, their cluster machines are found underutilized with severe resource fragmentation. The main causes are myopic online scheduling of incoming tasks and complex placement constraints specified by users or systems. In this paper, we propose to use task migration as a measure to alleviate resource fragmentation. Our trace-driven analysis on a production cluster with 12kmachines reveals that almost at any random snapshot, most of the concurrent tasks have long run-times and small migration times, thus justifying the feasibility of task migration. By making the complex constraints mathematically tractable, we formulate an integer linear programming problem. An efficient heuristic algorithm called Iterative Partitioned Defragmentation (IPD) is presented to perform task migration in multiple iterations of computation. We design and implementDefrag, an operational resource defragmentation system on Kubernetes, and deploy it on the production clusters. Trace-driven experiments show thatDefragcan reduce up to 80% idle CPUs and 29% idle GPUs on average. Furthermore,Defragcan refine the performance of online scheduling strategies by reducing up to 57% idle CPUs and 35% idle GPUs at the time of execution. Our real-world experiment also demonstrates the effectiveness ofDefrag. Yuedong Xu 0001, Jun Wu 0006, Yinghao Yu |
IEEE Trans. Netw. | 3 |
| 2025 | Distributed Multiagent Resource Allocation Based on Transformer and DRL for Cloud XR Hybrid Content TransmissionabstractExtended reality (XR) technologies and applications have grown rapidly in recent years. In addition to providing immersive ultrahigh-definition (UHD) video, XR also allows for a haptic experience where devices can be remotely manipulated to accomplish tasks. However, varying numbers of XR users accessing the communication system can strain limited spectrum resources, posing challenges in resource allocation. Therefore, this article studies resource blocks (RBs) allocation problem in a downlink transmission scenario where real-time cloud XR video and haptic contents need to be transmitted simultaneously. We also consider the random variation in the number of XR users and propose an adaptive distributed multiagent deep reinforcement learning (DRL) combined with Transformer (ADMA-DcT) for dynamic RB allocation method. This method addresses the dynamic change in network input dimensions due to user number variability using a state division module and a self-attention mechanism in the encoder module. To our knowledge, this is the first work to study the RB allocation problem in Cloud XR transmission considering simultaneous transmitting of video and haptic services with a dynamically changing user base. Our extensive simulations show that the ADMA-DcT model, end-to-end trained, outperforms other benchmarks in successfully serving a larger number of XR users under varying user number conditions, demonstrating excellent adaptivity and robustness. Zhaocheng Wang 0005, Jun Wu 0006, Rui Wang 0001, Ying Li 0020 |
IEEE Internet Things J. | 2 |
| 2025 | Enhancing Distributed Source Coding With Encoder-Centric Frequency Adaptation and Spatial TransformationabstractCurrent methodologies in distributed source coding have predominantly investigated decoder-focused strategies, emphasizing the alignment and exploitation of side information. This study introduces a paradigm shift by presenting an encoder-centric algorithm that conducts proactive optimization in the frequency domain. This shift is motivated by the current deep learning models' tendency to passively extract high-frequency elements, such as contours and content in the spatial domain at the encoder side, without considering the frequency characteristics of these spatial components. Unlike current trends, the proposed scheme actively selects the essential frequency components directly in the frequency domain by introducing an adaptive self-learning filter, enabling the encoder to discern and retain critical frequency components effectively and precisely. Furthermore, we align the side information in the spatial domain before feature extraction and implement an affine transformation-based alignment strategy to utilize the side information better. By leveraging the shared frequency domain components of the image pairs, the proposed algorithm adeptly learns affine coefficients to accomplish precise spatial alignment. This dual strategy of proactive encoder optimization and decoder alignment via affine transformations is highly efficient, outperforming existing state-of-the-art methods in distributed source coding when tested across two diverse datasets by an average of 0.5 dB in PSNR. Hao Xu 0027, Bin Tan 0001, Die Hu 0002, Jun Wu 0006 |
IEEE Trans. Multim. | 5 |
| 2025 | RuMono: Fuzz Driver Synthesis for Rust Generic APIsabstractFuzzing is a popular technique for detecting bugs, which can be extended to libraries by constructing executables that call library APIs, known as fuzz drivers. Automated fuzz driver synthesis has been an important research topic in recent years since it can facilitate the library fuzzing process. Nevertheless, existing approaches generally ignore generic APIs or simply treat them as non-generic APIs. As a result, they cannot generate effective fuzz drivers for generic APIs. This article explores the challenge of automating fuzz driver synthesis for Rust libraries with generic APIs. The problem is essential because Rust prioritizes security and generic APIs are widely employed in Rust libraries. We propose a novel approach and develop a prototype, RuMono, to tackle the problem. Our approach initially infers the API reachability from the generic API dependency graph, discovering the reachable and valid monomorphic APIs within the library. Further, we apply a similarity-based filter to eliminate redundant monomorphic APIs. Experimental results from 29 popular open source libraries demonstrate that RuMono can achieve promising generic API coverage with a low rate of invalid fuzz drivers. Besides, we have identified 23 previously unknown bugs in these libraries, with 18 related to generic APIs. Yehong Zhang, Jun Wu 0006, Hui Xu 0009 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2024 | A Congestion Control Algorithm for Live Video Streaming in Dynamic NetworkabstractTraditional TCP has been the dominant protocol for internet traffic after years of development in both academia and industry. However, the emergence of live video streaming applications and the increasing demand for low-latency video transmission, particularly in the context of sports and game live streaming, has posed a challenge to traditional TCP congestion control algorithms. As a representative TCP algorithm, BBR has been widely used in industry. However, BBR tends to inject more packets than the actual bottleneck in the dynamic network resulting in high latency. Because the maximum bandwidth in the past period is used to calculate the congestion window (CWND). To address this issue, we develop a recursive least squares (RLS) model to predict future bandwidth based on past bandwidth samples and update BBR's CWND periodically. To overcome the difficulty of modifying the kernel congestion control algorithm, we use extended Berkeley Packet Filter (eBPF) technology to rewrite TCP BBR and use eBPF MAP to exchange data between the kernel space and the data space. Experiments show that our algorithm can effectively reduce latency for live streaming in the dynamic network. Compared with CUBIC, our algorithm can achieve 76.9% average latency reduction with 2.3% average throughput loss and bring 42.4% average latency reduction with 1.9% average throughput loss compared with BBR. Wenqi Pan, Bin Tan 0001, Die Hu 0002, Jun Wu 0006 |
CSCloud | 4 |
| 2024 | PJSCC: A Puncturing-Based Joint Source Channel Coding Scheme with Hierarchical Down-Sampling LayerabstractIn this paper, we propose a puncturing-based joint source channel coding scheme with a hierarchical down-sampling layer (PJSCC). The proposed hierarchical down-sampling layer fully exploits both frequency and spatial priors. Moreover, to achieve adaptive compression ratio control, PJSCC utilizes a shared puncturing table as a global prior shared between the sender and receiver. This puncturing table plays a vital role in selectively pruning or padding symbols, and the adaptation of the compression ratio is achieved through the manual configuration of hyperparameters to adjust the puncturing rate. Experimental results show that the proposed scheme achieves superior reconstruction performance across several classic datasets. Bin Tan 0001, Jun Wu 0006, Die Hu 0002 |
ICASSP | 3 |
| 2024 | Surface-Constrained Progressive Feature Preserving Point Cloud CompressionabstractCurrent point cloud compression methods based on deep learning cannot guarantee that the reconstructed points are constrained to the surface, resulting in low reconstruction quality at low bitrates. Hence, this paper proposes an efficient deep learning-based point cloud geometry compression algorithm. Specifically, by introducing a two-dimensional plane at the decoder, the reconstructed local patch is constrained within a manifold, preserving sufficient surface features. This strategy ensures the decoder can reconstruct high-quality point clouds even at low bitrates. Moreover, we use the anchor features obtained by the neural network to compress the local features at the encoder. The experimental results show that, under the condition of the same restoration quality, the proposed method improves the point-to-plane PSNR by more than 2dB compared to the state-of-the-art methods. The code is available at https://github.com/zbaoye/SurfPCC. Baoye Zhang, Wenxiang Shen, Bin Tan 0001, Die Hu 0002, Jun Wu 0006 |
ICASSP | 5 |
| 2024 | Demo: CTSim: A Scalable and Flexible Cybertwin Network Simulator for Internet of Things ScenariosabstractThe future of networking must address the connectivity demands of billions of people and trillions of networked devices. The Intelligent Internet of Everything (IoE) is widely considered the next evolution of the Internet, yet it poses significant challenges to the current TCP/IP architecture. To overcome these challenges, the concept of the Cybertwin network has been proposed as a promising future Internet architecture. To facilitate the advancement of Cybertwin-related research, we have developed CTSim, a Cybertwin network simulator. Our demo demonstrates that the Cybertwin network could effectively enhances mobility, availability, and security in Internet of Things scenarios, making it a powerful tool for researchers in these fields. Yuke Ma, Shihan Lin, Yang Chen 0001, Jun Wu 0006 |
SenSys | 4 |
| 2024 | An eBPF-empowered Congestion Control System with Delay RequirementsabstractThe rapid development of new communication applications such as virtual reality and video conferencing has brought new challenges to congestion control algorithms. In particular, these applications have specific requirements in terms of delay. Meeting specific delay requirements without high throughput loss is difficult, especially in dynamic networks. In addition, it is important that the proposed congestion control algorithms can be easily deployed. In this paper, we propose a congestion control algorithm, namely TD-BBR, to meet the delay requirements of different applications. TD-BBR is built on BBR and can adapt to various network environments without high throughput loss. we employ an online algorithm based on the recursive least squares method to predict future bandwidth. We design a simple and effective algorithm to adjust the congestion window (CWND) to meet the specific delay requirements according to the value of bandwidth prediction and the distance between the current delay and the target delay. We implement a real congestion control system through extended Berkeley Packet Filter (eBPF) technology and have deployed it in the Linux kernel without recompiling the kernel. Extensive experiments show that TD-BBR can effectively meet different delay requirements in most cases, decrease the 95th percentile delay, and avoid high throughput loss compared to other congestion control algorithms, Wenqi Pan, Yuedong Xu 0001, Jun Wu 0006 |
SMC | 4 |
| 2024 | Attentive multi-granularity perception network for person search
Qixian Zhang, Jun Wu 0006, Duoqian Miao 0001, Cairong Zhao, Qi Zhang 0020 |
Inf. Sci. | 2 |
| 2024 | Multi-Space Point Geometry Compression With Progressive Relation-Aware TransformerabstractDeep network-based point cloud geometry compression is becoming more crucial and attractive due to constantly expanding 3D applications. The current strategy employing holistic point clouds as input imposes limitations on the compressed point cloud size and results in the loss of local geometry information, simplifying the repetition of local point cloud patterns. Nevertheless, employing point cloud patches as input results in a fresh set of erroneous points within the spliced point cloud. This occurrence can be attributed to the loss of global information and the relationship among the point cloud patches. Hence, this paper introduces a novel framework for progressive downsampling point cloud compression that follows the principles of two distinct methodologies. Our strategy still uses the autoencoder structure. Specifically, the encoder reduces the point cloud size and learns features incorporating local patch structure information and global semantic information. At the same time, the decoder upsamples the quantized and entropy-encoded features to reconstruct the original point cloud. More specifically, the geometrical details inside the patch and the relationship between patches are encoded to obtain the point cloud's global semantic and local geometry information. Concurrently, the computational burden is reduced by partitioning the complete point cloud input into segments and conducting continuous downsampling. Furthermore, we introduced an attention-based point cloud deconvolution module to address the localized repeating concentrations in reconstructed point clouds resulting from linear interpolation. This module samples the parent node and its neighboring relationships in the multi-space domain to improve the characteristics of the parent nodes to be sampled. It then uses the deconvolution to create sub-node characteristics that are more varied than the linear interpolation. Empirical evidence demonstrates that the proposed methodology achieves a effective equilibrium between the compression quality and ratio. Wenxiang Shen, Baoye Zhang, Hao Xu 0027, XiaoHan Li, Jun Wu 0006 |
IEEE Trans. Multim. | 5 |
| 2024 | Reconfigurable Intelligent Surface Empowered Federated Edge Learning With Statistical CSIabstractAs an emerging distributed learning framework, federated edge learning (FEEL) can efficaciously resolve the delay requirements and privacy concerns by enabling collaborative modeling among the edge devices under the premise of data localization. However, the communication bottlenecks, e.g., model damage and signal deviation, will critically diminish the convergence performance due to the restricted resources and the undesirable wireless fading. To overcome this challenge, one feasible way is to integrate the reconfigurable intelligent surface (RIS) into the FEEL system, to reinforce the communication quality by adaptively reconfiguring the signal propagation environment. However, the significant premise to effectively exploit the RIS in most of the prior works is the estimation of exact instantaneous channel state information (CSI), which is extremely thorny and potentially incurs additional communication overhead. To tackle this issue, we investigate in this paper the RIS-aided FEEL system under the realistic supposition where only the statistical CSI is available among devices. Specifically, considering the wireless outage caused by the uncertainty of non-line-of-sight components, we rigorously derive an explicit convergence upper bound of the RIS enabled FEEL framework with outage. Accordingly, a resource configuration problem with the goal of minimizing the sum of outage-probability is further formulated by jointly configuring the RIS configuration matrix and the bandwidth allocation. To seek the solutions, we carefully design a general Bernstein-Type inequality in this paper, and thus the probabilistic outage objective function can be effectively handled in an equivalent manner. Simulation experiments verify that our design can accomplish a significant promotion compared against state-of-the-art baselines. Heju Li, Rui Wang 0001, Jun Wu 0006, Wei Zhang 0001, Ismael Soto |
IEEE Trans. Wirel. Commun. | 3 |
| 2024 | Hierarchical Codebook Design and Analytical Beamforming Solution for IRS-Assisted CommunicationabstractIn intelligent reflecting surface (IRS) assisted communication, beam search is usually time-consuming as the multiple-input multiple-output (MIMO) of IRS is usually very large. The hierarchical codebook is a widely accepted method for reducing the complexity of searching time. The performance of this method strongly depends on the design scheme of beamforming of different beamwidths. In this paper, a non-constant phase difference (NCPD) beamforming algorithm is proposed. To implement the NCPD algorithm, we first model the phase shift of IRS as a continuous function and then determine the parameters of the continuous function through the analysis of its array factor. Then, we propose a hierarchical codebook and two beam training schemes, namely the joint searching (JS) scheme and direction-wise searching (DWS) scheme by using the NCPD algorithm which can flexibly change the width, direction, and shape of the beam formed by the IRS array. Numerical results show that the NCPD algorithm is more accurate with smaller side lobes, and also more stable on IRS of different sizes compared to other wide beam algorithms. The misalignment rate of the beam formed by the NCPD method is significantly reduced. The time complexity of the NCPD algorithm is constant, thus making it more suitable for solving the beamforming design problem with practically large IRS. Qingqing Wu 0001, Die Hu 0002, Rui Wang 0001, Jun Wu 0006 |
IEEE Trans. Wirel. Commun. | 5 |
| 2023 | SocialCache: A Pervasive Social-Aware Caching Strategy for Self-Operated Content Delivery Networks of Online Social NetworksabstractOnline Social Networks (OSNs) play a significant role in people's daily life. Increasing OSN traffic promotes the requirement for building self-operated Content Delivery Networks (CDNs) to deliver OSN media data efficiently and reduce traffic costs. OSN data in CDNs is heavily influenced by social connectivity, such as friendships. To reduce CDN network traffic by better using social connectivity information, we propose SocialCache, a pervasive social-aware caching strategy in self-operated CDNs. SocialCache supports several social connectivity metric options that reflect the importance of the OSN users and the popularity of the media files. Through the improvement of the cache replacement algorithm in the CDN node and the communication design between nodes, SocialCache realizes the optimization of network traffic. Worth mentioning, SocialCache can easily integrate into mainstream CDN architectures while protecting user privacy. We implement SocialCache on Mininet, using real-world network measurements for CDN hierarchy and a hill-climbing algorithm for parameter selection. SocialCache outperforms a range of state-of-the-art baselines on three real-world OSN datasets. On the Twitter dataset, SocialCache reduces the network traffic volume by 6.40% and improves the cache hit ratio by 14.22%. Tiancheng Guo, Yuke Ma, Xin Wang 0002, Jun Wu 0006, Yang Chen 0001 |
ICC | 5 |
| 2023 | Dynamic Offloading Strategy for Delay-Sensitive Task in Mobile-Edge Computing NetworksabstractMobile-edge computing (MEC) technology offers computing resources for mobile devices to conduct computationally heavy activities by putting servers at the wireless mobile network’s edge. This mitigates the scarcity of computing resources in mobile devices and enhances the intelligence of the Internet of Things (IoT), which is a crucial technology for achieving industrial digitalization. Considering the time-varying channel as well as the time-varying available computing resources of MEC servers, this article formulates a hybrid optimization problem that combines task offload and resource allocation. The goal is to minimize MEC servers’ overall power consumption. Since the channel state information (CSI) stored in the MEC system is not real time, we propose a reinforcement learning (RL) algorithm for predicting current CSI from historical CSI and obtain the optimal strategy for task offloading. On the other hand, convex optimization methods are used to accomplish the dynamic resource allocation strategy. In addition, an approach based on deep RL (DRL) is put forward to overcome the dimensionality curse in RL algorithms. The simulation experiments illustrate that the proposed algorithms outperform the nonpredictive schemes by a large margin, and their performance is close to that of the optimum scheme, which utilizes simultaneous CSI. Lihua Ai, Bin Tan 0001, Jiadi Zhang, Rui Wang 0001, Jun Wu 0006 |
IEEE Internet Things J. | 5 |
| 2023 | Cache-Aided MEC With the Assistance of Intelligent Reflecting SurfaceabstractTo address the large concerns of energy consumption caused by the rapid growth of online data traffic and network services such as the applications of the Internet of Things, many technologies have been developed to help wireless communication. Mobile edge computing (MEC) can develop the efficiency and reduce the energy consumption of the network through edge-cloud benefits. Intelligent reflecting surface (IRS) can improve spectrum efficiency and decrease power costs by changing the transmission environment. Considering that IRS also has some good physical characteristics, it can be integrated into MEC as auxiliary equipment and yield marked performance improvement. In this study, we design an IRS-assisted cache-aided MEC system by optimizing the beamformer of base station (BS) and the phase-shift vector of IRS jointly. We develop two algorithms based on the block coordinate descent (BCD) method to achieve this goal. First, we propose a branch-and-bound (BB)-based algorithm. By the algorithm, an approximately optimal solution of the IRS element optimization problem can be obtained under constant modulus constraints. Then, we develop a Lagrange multiplier method-based algorithm that has less complexity. The performance of IRS-assisted cache-aided MEC with the proposed algorithms is demonstrated by simulation results. Jiadi Zhang, Rui Wang 0001, Jun Wu 0006, Lihua Ai |
IEEE Internet Things J. | 3 |
| 2023 | Joint Localization and Communication Study for Intelligent Reflecting Surface Aided Wireless Communication SystemabstractThe intelligent reflecting surface (IRS) is promising in assisting user localization and wireless communication in the future wireless networks. In this paper, a novel IRS-aided joint localization and communication (L&C) scheme is designed in a millimeter-wave transmission system. For the proposed scheme, the user position/orientation estimation error bound (POEB) and the effective achievable data rate (EADR) are derived in closed-form as L&C performance metrics, which reveal the inherent trade-off between L&C capabilities. To achieve the joint optimal point of the POEB and EADR in consideration of the localization errors, a worst-case robust beamforming and time allocation optimization problem is formulated. To solve the original non-convex problem, a novel joint optimization approach is developed. Specifically, from an equivalent minimax problem, the local optimal solutions of the transceiver beamformers, the IRS phase-shift matrix, and the time allocation ratio between user localization stage (ULS) and effective data transmission stage (EDTS), are obtained in closed-form with respect to the localization errors. Then, the worst-case localization error is iteratively found by a dedicated majorize-minimization (MM) based algorithm. Subsequently, potential extensions to general wireless channels and discrete phase-shift models are discussed in detail. Finally, simulations are carried out to show the optimization results and the L&C performance trade-off. In comparison with the conventional non-robust method, the proposed approach is validated to be robust against the user localization uncertainty. Rui Wang 0001, Zhe Xing, Erwu Liu, Jun Wu 0006 |
IEEE Trans. Commun. | 4 |
| 2023 | Personalized Federated Learning With Differential Privacy and Convergence GuaranteeabstractPersonalized federated learning (PFL), as a novel federated learning (FL) paradigm, is capable of generating personalized models for heterogenous clients. Combined with with a meta-learning mechanism, PFL can further improve the convergence performance with few-shot training. However, meta-learning based PFL has two stages of gradient descent in each local training round, therefore posing a more serious challenge in information leakage. In this paper, we propose a differential privacy (DP) based PFL (DP-PFL) framework and analyze its convergence performance. Specifically, we first design a privacy budget allocation scheme for inner and outer update stages based on the Rényi DP composition theory. Then, we develop two convergence bounds for the proposed DP-PFL framework under convex and non-convex loss function assumptions, respectively. Our developed convergence bounds reveal that 1) there is an optimal size of the DP-PFL model that can achieve the best convergence performance for a given privacy level, and 2) there is an optimal tradeoff among the number of communication rounds, convergence performance and privacy budget. Evaluations on various real-life datasets demonstrate that our theoretical results are consistent with experimental results. The derived theoretical results can guide the design of various DP-PFL algorithms with configurable tradeoff requirements on the convergence performance and privacy levels. Kang Wei 0004, Jun Li 0004, Chuan Ma 0001, Ming Ding 0001, Wen Chen 0001, Jun Wu 0006, Meixia Tao, H. Vincent Poor |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | Accelerating Distributed DNN Training via Transport Layer SchedulingabstractCommunication scheduling is crucial to accelerate the training of large deep learning models, in which the transmission order of layer-wise deep neural network (DNN) tensors is determined for a better computation-communication overlap. Prior approaches adopt user-level tensor partitioning to enhance the priority scheduling with finer granularity. However, a startup time slot inserted before every tensor partition will neutralize this scheduling gain. Tuning hyper-parameters for tensor partitioning is difficult, especially when the network bandwidth is shared or time-varying in multi-tenant clusters. In this article, we propose Mercury, a simple transport layer scheduler that moves the priority scheduling to the transport layer at the packet granularity. The packets with the highest priority in the Mercury buffer will be transmitted first. Mercury achieves the near-optimal overlapping between communication and computation. It also leverages the immediate aggregation at the transport layer to enable the full overlapping of gradient push and pull. We implement Mercury in MXNet and conduct comprehensive experiments on five popular DNN models in various environments. Mercury can well adapt to dynamic communication and computation resources. Experiments show that Mercury accelerates the training by up to 130% compared to the classical PS architecture, and 104% compared to state-of-the-art tensor partitioning methods. Qingyang Duan, Zeqin Wang, Yuedong Xu 0001, Shaoteng Liu, Jun Wu 0006, John C. S. Lui |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2023 | One Bit Aggregation for Federated Edge Learning With Reconfigurable Intelligent Surface: Analysis and OptimizationabstractAs one of the most popular and attractive frameworks for model training, federated edge learning (FEEL) presents a new paradigm, which avoids direct data transmission by collaboratively training a global learning model across multiple distributed edge devices, thus overcoming the disadvantage of centralized machine learning in resource limitations, delay constraints, and privacy issues. However, due to the heavy cost of communicating gradient among edge devices, sharing the parameters of a large-scale neural network can still be time-intensive. To alleviate this bottleneck, an efficient scheme, called SignSGD has been recently proposed, where the one-bit gradient quantization with majority vote is featured at edge devices. Nevertheless, the performance of one-bit aggregation will inevitably deteriorate due to the undesirable propagation error introduced by wireless channels. To address this issue, we propose in this work a novel reconfigurable intelligent surface (RIS)-aided one-bit communication optimization scheme under orthogonal frequency division multiple access (OFDMA) to relieve the negative influence of communication error on the SignSGD-based FEEL. Specifically, a learning convergence analysis is firstly presented to quantitatively characterize the impact of wireless communication error measured by the union bound on pairwise bit error rate (BER) on the performance of SignSGD-based FEEL. Immediately, a unified communication-learning optimization problem is further formulated to jointly optimize the sub-band assignment strategy, the power allocation vector, and the RIS configuration matrix. Numerical experiments show that the proposed design achieves substantial performance improvement compared with the state-of-the-art approaches. Heju Li, Rui Wang 0001, Wei Zhang 0001, Jun Wu 0006 |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Location Information Assisted Beamforming Design for Reconfigurable Intelligent Surface Aided Communication SystemsabstractThe large overhead arising from conventional channel estimations in reconfigurable intelligent surface (RIS) aided millimeter-wave communication systems, may offset the performance gain brought by the RIS. To tackle this issue, we propose a location information assisted beamforming design without the requirement of the channel training process. First, we establish the geometrical relationship between the channel model and the user location, and mathematically derive an approximate channel state information (CSI) error bound based on the user location error region. Then, for combating the negative impact of the location error on the communication performance, we formulate a worst-case robust beamforming optimization problem to optimize the beamformer at the base station (BS) and the phase-shift matrix at the RIS. To solve this non-convex problem, we develop a novel relaxed alternating optimization process (RAOP) by utilizing various optimization tools, such as the Lagrange multiplier, the matrix inversion lemma, the semidefinite relaxation (SDR), as well as the branch-and-bound (BnB). Additionally, we prove sufficient conditions for the SDR to yield rank-one solutions, and modify the BnB to acquire the phase-shift solution under an arbitrary constraint of possible phase-shift values. Finally, we analyse the convergence and complexity of the proposed RAOP, and carry out simulations for performance evaluations. Compared to the conventional non-robust beamforming, our method performs better and shows strong robustness against the location-error-related CSI uncertainty. Compared to the robust beamforming based on the S-procedure and penalty convex-concave procedure (CCP), our method with BnB shows the advantages of being able to converge faster and handle arbitrary phase-shift argument sets. Zhe Xing, Rui Wang 0001, Xiaojun Yuan 0002, Jun Wu 0006 |
IEEE Trans. Wirel. Commun. | 4 |
| 2022 | Federated Edge Learning via Reconfigurable Intelligent Surface with One-Bit QuantizationabstractIn this paper, the problem of model aggregation for the federated edge learning (FEEL) over a realistic wireless network is investigated, where multiple distributed edge devices collaboratively train a global learning model using local data. In the considered model, the one-bit gradient quantization with majority vote is adopted at edge devices, which sends the local gradient sign to the edge server instead of transmitting high-dimensional stochastic gradients directly. After aggregating the quantified signs, the edge server sends back only the majority decision to significantly minimize the transmission overhead since all communications are compressed to one bit. Nevertheless, it turns out that the quality of training will be inevitably deteriorated by the undesirable propagation error introduced by wireless channels. To address this issue, we propose in this work a reconfigurable intelligent surface (RIS) assisted one-bit communication scheme under orthogonal frequency division multiplexing (OFDM) system to reduce the signal distortion during iterative model exchange of FEEL. Specifically, a learning convergence analysis with respect to wireless communication error is first established. After that, we further formulate a unified communication-learning design problem to jointly optimize the power allocation vector and the RIS configuration matrix. Numerical results demonstrate that the proposed design achieves substantial performance improvement compared with the state-of-the-art solutions. Heju Li, Rui Wang 0001, Jun Wu 0006, Wei Zhang 0001 |
GLOBECOM | 3 |
| 2022 | Mercury: A Simple Transport Layer Scheduler to Accelerate Distributed DNN TrainingabstractCommunication scheduling is crucial to improve the efficiency of training large deep learning models with data parallelism, in which the transmission order of layer-wise deep neural network (DNN) tensors is determined for a better computation-communication overlap. Prior approaches adopt tensor partitioning to enhance the priority scheduling with finer granularity. However, a startup time slot inserted before each tensor partition will neutralize this scheduling gain. Tuning the optimal partition size is difficult and the application-layer solutions cannot eliminate the partitioning overhead. In this paper, we propose Mercury, a simple transport layer scheduler that does not partition the tensors, but moves the priority scheduling to the transport layer at the packet granularity. The packets with the highest priority in the Mercury buffer will be transmitted first. Mercury achieves the near-optimal overlapping between communication and computation. It leverages immediate aggregation at the transport layer to enable the coincident gradient push and parameter pull. We implement Mercury in MXNet and conduct comprehensive experiments on five DNN models in an 8-node cluster with 10Gbps Ethernet. Experimental results show that Mercury can achieve about 1.18 ~ 2.18 × speedup over vanilla MXNet, and 1.08 ~ 2.04× speedup over the state-of-the-art tensor partitioning solution. Qingyang Duan, Zeqin Wang, Yuedong Xu 0001, Shaoteng Liu, Jun Wu 0006 |
INFOCOM | 5 |
| 2022 | Intelli-AR Preloading: A Learning Approach to Proactive Hologram Transmissions in Mobile ARabstractMobile augmented reality (AR), which integrates virtual objects (i.e., holographic contents) with 3-D real environments in real time, has been rapidly gaining popularity in the last five years. The delivery mechanisms of these holographic contents to mobile AR devices, however, are rarely investigated. To combat bandwidth limitations that preclude providing holographic contents to user devices on-demand, in this article, we propose the intelligent AR (Intelli-AR) preloading algorithm to improve transmission efficiency in the edge-assisted network, in which edge servers proactively transmit holographic contents to the devices. Without user devices’ future motion trajectories, the Intelli-AR preloading algorithm models the user devices’ motion trajectories as Markov decision process (MDP) and adaptively learns the optimal preloading policy. The Intelli-AR preloading is decomposed into two parts and separately deployed on the edge server and the user devices to reduce the computation complexity. The Intelli-AR solution improves the ratio of successful preloading by 11.52% compared to the best baseline in the practical data set when the users’ motion trajectories tend to be more random, and by 21.97% compared to the best baseline in the data set which is synthesized from a real-life mobile AR environment. Yuqi Han, Rui Wang 0001, Jun Wu 0006, Maria Gorlatova |
IEEE Internet Things J. | 4 |
| 2022 | Improved Instance Discrimination and Feature Compactness for End-to-End Person SearchabstractPerson search aims to locate and retrieve specific pedestrians in scene images, including two subtasks, pedestrian detection and person re-identification. Recently, triplet loss has been widely used in person re-identification, which effectively improves the pedestrian features embedding and achieves superior performance. However, forming triplet in the person search is not an easy task. Most of the existing end-to-end person search methods are based on Faster R-CNN. The training process of person re-identification part is affected by the detector. It is difficult to form pedestrian triplets within a limited batch size. Also, there are many pedestrian identities in the person search dataset, but each pedestrian identity only has a few samples. It is difficult to learn a robust pedestrian feature representation for person search. To resolve the problem discussed above, a novel Feature Compactness (FC) Loss for the person search is designed, which efficiently improves the inter-class discrimination and intra-class compactness of pedestrian features embedding without the need for positive or negative pairs. Besides, we propose a pedestrian attention module (PAM) to help the network focuses more on pedestrian information and suppresses irrelevant background information. Our method achieves comparable performance on two benchmarks, CUHK-SYSU and PRW, and achieves 91.96% of mAP and 93.34% of rank1 accuracy on CUHK-SYSU. Shaowei Hou, Cairong Zhao, Jun Wu 0006, Zhihua Wei 0001, Duoqian Miao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Context-Aware Feature Learning for Noise Robust Person SearchabstractPerson search aims to localize and identify specific pedestrians from numerous surveillance scene images. In this work, we focus on the noise in person search. We categorize the noise into scene-inherent noise and human-introduced noise. Scene-inherent noise comes from congestion, occlusion, and illumination changes. Human-introduced noise originates from the labeling process. For scene-inherent noise, we propose a novel context contrastive loss to take advantage of the latent contextual information from scene images. Features from context regions are utilized to construct contrastive pairs to constrain the feature discrimination among pedestrians in scene images while maintaining the feature consistency of the same identity. The network can thus learn to distinguish congested and overlapped pedestrians and more robust features can be obtained. For human-introduced noise, we propose a noise-discovery and noise-suppression training process for mislabeling robust person search. After the first training pass, the relation between feature prototypes of different identities is analyzed and the mislabeled pedestrians are discovered. During the second training pass, the label noise is suppressed to reduce the negative influence of mislabeled data. Experiments show that the proposed context-aware noise-robust (CANR) person search can achieve competitive performance. Further ablation studies confirm the effectiveness of CANR. Cairong Zhao, Shuguang Dou, Zefan Qu, Jiawei Yao, Jun Wu 0006, Duoqian Miao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Fitting the Search Space of Weight-sharing NAS with Graph Convolutional NetworksabstractNeural architecture search has attracted wide attentions in both academia and industry. To accelerate it, researchers proposed weight-sharing methods which first train a super-network to reuse computation among different operators, from which exponentially many sub-networks can be sampled and efficiently evaluated. These methods enjoy great advantages in terms of computational costs, but the sampled sub-networks are not guaranteed to be estimated precisely unless an individual training process is taken. This paper owes such inaccuracy to the inevitable mismatch between assembled network layers, so that there is a random error term added to each estimation. We alleviate this issue by training a graph convolutional network to fit the performance of sampled sub-networks so that the impact of random errors becomes minimal. With this strategy, we achieve a higher rank correlation coefficient in the selected set of candidates, which consequently leads to better performance of the final architecture. In addition, our approach also enjoys the flexibility of being used under different hardware constraints, since the graph convolutional network has provided an efficient lookup table of the performance of architectures in the entire search space. Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Longhui Wei, Yuhui Xu 0002, Qi Tian 0001 |
AAAI | 3 |
| 2021 | Scene Text Image Super-Resolution via Parallelly Contextual Attention NetworkabstractOptical degradation blurs text shapes and edges, so existing scene text recognition methods have difficulties in achieving desirable results on low-resolution (LR) scene text images acquired in real-world environments. The above problem can be solved by efficiently extracting sequential information to reconstruct super-resolution (SR) text images, which remains a challenging task. In this paper, we propose a Parallelly Contextual Attention Network (PCAN), which effectively learns sequence-dependent features and focuses more on high-frequency information of the reconstruction in text images. Firstly, we explore the importance of sequence-dependent features in horizontal and vertical directions parallelly for text SR, and then design a parallelly contextual attention block to adaptively select the key information in the text sequence that contributes to image super-resolution. Secondly, we propose a hierarchically orthogonal texture-aware attention module and an edge guidance loss function, which can help to reconstruct high-frequency information in text images. Finally, we conduct extensive experiments on TextZoom dataset, and the results can be easily incorporated into mainstream text recognition algorithms to further improve their performance in LR image recognition. Besides, our approach exhibits great robustness in defending against adversarial attacks on seven mainstream scene text recognition datasets, which means it can also improve the security of the text recognition pipeline. Compared with directly recognizing LR images, our method can respectively improve the recognition accuracy of ASTER, MORAN, and CRNN by 14.9%, 14.0%, and 20.1%. Our method outperforms eleven state-of-the-art (SOTA) SR methods in terms of boosting text recognition performance. Most importantly, it outperforms the current optimal text-orient SR method TSRN by 3.2%, 3.7%, and 6.0% on the recognition accuracy of ASTER, MORAN, and CRNN respectively. Cairong Zhao, Shuyang Feng, Brian Nlong Zhao, Zhijun Ding, Jun Wu 0006, Fumin Shen, Heng Tao Shen |
ACM Multimedia | 5 |
| 2021 | TyrLoc: a low-cost multi-technology MIMO localization system with a single RF chainabstractThis work presents the design and implementation of TyrLoc, an accurate multi-technology switching MIMO localization system that can be deployed on low-cost SDRs. TyrLoc only uses a single RF Chain to switch on each antenna in an antenna array within the coherence time asynchronously, thus mimicking a MIMO platform to pinpoint the positions of WIFI, Bluetooth Low Energy (BLE) and LoRa devices. TyrLoc makes three key technical contributions. First, TyrLoc modifies the firmware of inexpensive PlutoSDR that controls the antenna switching pattern and tags the signal associated with each antenna. Second, it develops a two-stage fine-grained carrier frequency offset (CFO) calibration algorithm that harnesses the agile antenna switching pattern and is 10× more accurate than the baseline method. Third, TyrLoc employs an interpolated transform approach to facilitate angle-of-arrival (AoA) estimation in the presence of missing antennas. The AoA-based localization experiments in a multipath-rich indoor environment show that TyrLoc with eight antennas achieves the median errors of 63cm for WIFI, 39cm for BLE and 32cm for LoRa, respectively. Taiwei He, Junwei Yin, Yuedong Xu 0001, Jun Wu 0006 |
MobiSys | 5 |
| 2021 | Progressive DARTS: Bridging the Optimization Gap for NAS in the Wild
Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Qi Tian 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | Cyclic CNN: Image Classification With Multiscale and Multilocation ContextsabstractImproving the capability of models at limited computational cost is an urgent demand in many vision-based Internet-of-Things applications. Recent progress on deep convolutional neural network (CNN) has largely accelerated the development of image classification. Although the hierarchical structure of CNN naturally helps to extract image features in different scales and locations progressively, conventional convolution can only handle contexts of one scale and on a limited area of a single location in a specific layer, limiting the utilization of multiscale and multilocation information. In this work, we present a cyclic CNN framework, which enables sufficient utilization of multiscale and multilocation contexts in a single layer of convolution. The cyclic CNN is an extremely simple but effective improvement upon conventional convolution, which occupies no additional parameter and negligible computation (even less than 0.1%). Moreover, cyclic CNN can be easily plugged into many existing CNN pipelines, e.g., the ResNet family, obtaining extremely low-cost performance gain upon them. Extensive experiments on both small-scale (CIFAR10 and CIFAR100) and large-scale (ILSVRC2012) image classification benchmarks demonstrate that a consistent performance promotion is obtained with the help of cyclic CNN. Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Qi Tian 0001 |
IEEE Internet Things J. | 3 |
| 2021 | Joint Design of Beamforming and Edge Caching in Fog Radio Access NetworksabstractIn this paper, we study a novel transmission framework based on statistical channel state information (SCSI) by incorporating edge caching and beamforming in a fog radio access network (F-RAN) architecture. By optimizing the statistical beamforming and edge caching, we formulate a comprehensive nonconvex optimization problem to minimize the backhaul cost subject to the BS transmission power, limited caching capacity, and quality-of-service (QoS) constraints. By approximating the problem using the l 0 -norm, Taylor series expansion, and other processing techniques, we provide a tailored second-order cone programming (SOCP) algorithm for the unicast transmission scenario and a successive linear approximation (SLA) algorithm for the joint unicast and multicast transmission scenario. This is the first attempt at the joint design of statistical beamforming and edge caching based on SCSI under the F-RAN architecture. Wenjing Lv, Rui Wang 0001, Jun Wu 0006, Zhijun Fang 0001, Songlin Cheng |
Secur. Commun. Networks | 3 |
| 2021 | Incremental Generative Occlusion Adversarial Suppression Network for Person ReIDabstractPerson re-identification (re-id) suffers from the significant challenge of occlusion, where an image contains occlusions and less discriminative pedestrian information. However, certain work consistently attempts to design complex modules to capture implicit information (including human pose landmarks, mask maps, and spatial information). The network, consequently, focuses on discriminative features learning on human non-occluded body regions and realizes effective matching under spatial misalignment. Few studies have focused on data augmentation, given that existing single-based data augmentation methods bring limited performance improvement. To address the occlusion problem, we propose a novel Incremental Generative Occlusion Adversarial Suppression (IGOAS) network. It consists of 1) an incremental generative occlusion block, generating easy-to-hard occlusion data, that makes the network more robust to occlusion by gradually learning harder occlusion instead of hardest occlusion directly. And 2) a global-adversarial suppression (G&A) framework with a global branch and an adversarial suppression branch. The global branch extracts steady global features of the images. The adversarial suppression branch, embedded with two occlusion suppression module, minimizes the generated occlusion's response and strengthens attentive feature representation on human non-occluded body regions. Finally, we get a more discriminative pedestrian feature descriptor by concatenating two branches' features, which is robust to the occlusion problem. The experiments on the occluded dataset show the competitive performance of IGOAS. On Occluded-DukeMTMC, it achieves 60.1% Rank-1 accuracy and 49.4% mAP. Cairong Zhao, Xinbi Lv, Shuguang Dou, Shanshan Zhang 0001, Jun Wu 0006, Liang Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | A Deep Image Coding Scheme With Generative Network to Learn From Correlated ImagesabstractThis paper provides a method to build a deep learning image coding system based on inverse problem, choosing a suitable measurement operator to reduce the amount of information transmitted at the sender, and reconstructing the original image by tackling the inverse problem at the receiver. Unlike most compressed sensing (CS) methods, the proposed coding scheme does not rely on sparsity but uses the structural priors of the generative adversarial networks (GAN) to solve the inverse problem. The proposed model trains the GAN to learn a mapping from the latent space to the sample space formed by correlated images on the cloud. Then the measurements are used to localize the optimal latent variable in the representation space which corresponding to the original image in the sample space. The proposed method encodes and transmits the measurements instead of the original image, which greatly reduces the cost of transmission while ensuring the quality of the reconstructed the image at high compression ratios. To the best of our knowledge, this is the first time to introduce the GAN-based inverse problem in the field of the deep image coding area. The experimental results show that the visual quality of the images generated by the proposed scheme is better than the traditional encoding scheme JPEG2000. Especially in the case of extremely high compression ratios, the proposed scheme can still maintain good performance. Bin Tan 0001, Jun Wu 0006, Zhifeng Zhang 0001, Haoqi Ren |
IEEE Trans. Multim. | 3 |
| 2021 | QoS-Aware User Grouping Strategy for Downlink Multi-Cell NOMA SystemsabstractIn multi-cell non-orthogonal multiple access (NOMA) systems, designing an appropriate user grouping strategy is an open problem due to diverse quality of service (QoS) requirements and inter-cell interference. In this paper, we exploit both game theory and graph theory to study QoS-aware user grouping strategies, aiming at minimizing power consumption in downlink multi-cell NOMA systems. Under different QoS requirements, we derive the optimal successive interference cancellation (SIC) decoding order with inter-cell interference, which is different from existing SIC decoding order of increasing channel gains, and obtain the corresponding power allocation strategy. Based on this, the exact potential game model of the user grouping strategies adopted by multiple cells is formulated. We prove that, in this game, the problem for each player to find a grouping strategy can be converted into the problem of searching for specific negative loops in the graph composed of users. Bellman-Ford algorithm is expanded to find these negative loops. Furthermore, we design a greedy based suboptimal strategy to approach the optimal solution with polynomial time. Extensive simulations confirm the effectiveness of grouping users with consideration of QoS and inter-cell interference, and show that the proposed strategies can considerably reduce total power consumption comparing with reference strategies. Fengqian Guo, Hancheng Lu, Xiaoda Jiang, Ming Zhang 0029, Jun Wu 0006, Chang Wen Chen |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Cache Placement Optimization in Mobile Edge Computing Networks With Unaware Environment - An Extended Multi-Armed Bandit ApproachabstractCaching high-frequency reuse contents at the edge servers in the mobile edge computing (MEC) network omits the part of backhaul transmission and further releases the pressure of data traffic. However, how to efficiently decide the caching contents for edge servers is still an open problem, which refers to the cache capacity of edge servers, the popularity of each content, and the wireless channel quality during transmission. In this paper, we discuss the influence of unknown user density and popularity of content on the cache placement solution at the edge server. Specifically, towards the implementation of the cache placement solution in the practical network, there are two problems needing to be solved. First, the estimation of unknown users’ preference needs a huge amount of records of users’ previous requests. Second, the overlapping serving regions among edge servers cause the wrong estimation of users’ preference, which hinders the individual decision of caching placement. To address the first issue, we propose a learning-based solution to adaptively optimize the cache placement policy without any previous knowledge of the user density and the popularity of the contents. We develop the extended multi-armed bandit (Extended MAB), which combines the generalized global bandit (GGB) and Standard Multi-armed bandit (MAB), to iteratively estimate both a global parameter, i.e., the user density, and individual parameters, i.e., the popularity of each content. For the second problem, a multi-agent Extended MAB based solution is presented to avoid the mis-estimation of parameters and achieve the decentralized cache placement policy. The proposed solution determines the primary time slot and secondary time slot for each edge server. The edge servers estimate expected satisfied user number of caching a content with the overlap information and determine the cache placement solution. The proposed strategies are proven to achieve the bounded regret according to the mathematical analysis. Extensive simulations verify the optimality of the proposed strategies when comparing with baselines. Yuqi Han, Lihua Ai, Rui Wang 0001, Jun Wu 0006, Dian Liu, Haoqi Ren |
IEEE Trans. Wirel. Commun. | 4 |
| 2021 | Achievable Rate Analysis and Phase Shift Optimization on Intelligent Reflecting Surface With Hardware ImpairmentsabstractIntelligent reflecting surface (IRS) is envisioned as a promising hardware solution to hardware cost and energy consumption in the fifth-generation (5G) mobile communication network. It exhibits great advantages in enhancing data transmission, but may suffer from performance degradation caused by inherent hardware impairment (HWI). For analysing the achievable rate (ACR) and optimizing the phase shifts in the IRS-aided wireless communication system with HWI, we consider that the HWI appears at both the IRS and the signal transceivers. On this foundation, first, we derive the closed-form expression of the average ACR and the IRS utility. Then, we formulate optimization problems to optimize the IRS phase shifts by maximizing the signal-to-noise ratio (SNR) at the receiver side, and obtain the solution by transforming non-convex problems into semidefinite programming (SDP) problems. Subsequently, we compare the IRS with the conventional decode-and-forward (DF) relay in terms of the ACR and the utility. Finally, we carry out simulations to verify the theoretical analysis, and evaluate the impact of the channel estimation errors and residual phase noises on the optimization performance. Our results reveal that the HWI reduces the ACR and the IRS utility, and begets more serious performance degradation with more reflecting elements. Although the HWI has an impact on the IRS, it still leaves opportunities for the IRS to surpass the conventional DF relay, when the number of reflecting elements is large enough or the transmitting power is sufficiently high. Zhe Xing, Rui Wang 0001, Jun Wu 0006, Erwu Liu |
IEEE Trans. Wirel. Commun. | 3 |
| 2020 | A Real-time Virtual Reality Adaptive Streaming SystemabstractCloud VR (Virtual Reality) is a VR scheme based on cloud computing, which can reduce the computing burden on terminal equipment. It uses edge computing for adaptive streaming to minimize the response latency and bandwidth consumption. However, traditional adaptive streaming approaches based on preprocessing have some limitations. In this paper, we proposed a novel adaptive Cloud VR system with real-time processing. We design and implement a GPU acceleration scheme to perform efficient projection and coding, which makes the computing latency acceptable. The scheme is further extended to pipeline execution with simple orientation prediction to support higher frame rate. The real-time processing not only reduces the storage size, but also eliminates the drawbacks of pre-generating limited video versions. According to the experimental results, our scheme can provide a VR stream matching user viewport more precisely. Through optimized GPU algorithm, we shorten the processing time to 1/10 of the original. Compared to classic pyramid projection scheme, our system effectively reduces the average orientation deviation by 90.23%, thus provides more robust service. Songyuan Zhao, Bin Tan 0001, Jun Wu 0006, Haoqi Ren, Zhifeng Zhang 0001 |
VTC Fall | 3 |
| 2020 | Reinforcement Learning-Based Optimal Computing and Caching in Mobile Edge NetworkabstractJoint pushing and caching are commonly considered an effective way to adapt to tidal effects in networks. However, the problem of how to precisely predict users' future requests and push or cache the proper content remains to be solved. In this paper, we investigate a joint pushing and caching policy in a general mobile edge computing (MEC) network with multiuser and multicast data. We formulate the joint pushing and caching problem as an infinite-horizon average-cost Markov decision process (MDP). Our aim is not only to maximize bandwidth utilization but also to decrease the total quantity of data transmitted. Then, a joint pushing and caching policy based on hierarchical reinforcement learning (HRL) is proposed, which considers both long-term file popularity and short-term temporal correlations of user requests to fully utilize bandwidth. To address the curse of dimensionality, we apply a divide-and-conquer strategy to decompose the joint base station and user cache optimization problem into two subproblems: the user cache optimization subproblem and the base station cache optimization subproblem. We apply value function approximation Q-learning and a deep Q-network (DQN) to solve these two subproblems. Furthermore, we provide some insights into the design of deep reinforcement learning in network caching. The simulation results show that the proposed policy can learn content popularity very well and predict users' future demands precisely. Our approach outperforms existing schemes on various parameters including the base station cache size, the number of users and the total number of files in multiple scenarios. Yichen Qian, Rui Wang 0001, Jun Wu 0006, Bin Tan 0001, Haoqi Ren |
IEEE J. Sel. Areas Commun. | 3 |
| 2020 | Deep Fusion Feature Representation Learning With Hard Mining Center-Triplet Loss for Person Re-IdentificationabstractPerson re-identification (Re-ID) is a challenging task in the field of computer vision and focuses on matching people across images from different cameras. The extraction of robust feature representations from pedestrian images through CNNs with a single deterministic pooling operation is problematic as the features in real pedestrian images are complex and diverse. To address this problem, we propose a novel center-triplet (CT) model that combines the learning of robust feature representation and the optimization of metric loss function. Firstly, we design a fusion feature learning network (FFLN) with a novel fusion strategy consisting of max pooling and average pooling. Instead of adopting a single deterministic pooling operation, the FFLN combines two pooling operations that can learn high response values, bright features, and low response values, discriminative features simultaneously. Our model obtains more discriminative fusion features by adaptively learning the weights of the features learned by the corresponding pooling operations. In addition, we design a hard mining center-triplet loss (HCTL), a novel improved triplet loss, which effectively optimizes the intra/inter-class distance and reduces the cost of computing and mining hard training samples simultaneously, thereby enhancing the learning of robust feature representation. Finally, we proved our method can learn robust and discriminative feature representations for complex pedestrian images in real scenes. The experimental results also illustrate that our method achieves an 81.8% mAP and a 93.8% rank-1 accuracy on Market1501, a 68.2% mAP and an 83.3% rank-1 accuracy on DukeMTMC-ReID, and a 43.6% mAP and a 74.3% rank-1 accuracy on MSMT17, outperforming most state-of-the-art methods and achieving better performance for person re-identification. Cairong Zhao, Xinbi Lv, Zhang Zhang 0001, Wangmeng Zuo, Jun Wu 0006, Duoqian Miao 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | Progressive Differentiable Architecture Search: Bridging the Depth Gap Between Search and EvaluationabstractRecently, differentiable search methods have made major progress in reducing the computational costs of neural architecture search. However, these approaches often report lower accuracy in evaluating the searched architecture or transferring it to another dataset. This is arguably due to the large gap between the architecture depths in search and evaluation scenarios. In this paper, we present an efficient algorithm which allows the depth of searched architectures to grow gradually during the training procedure. This brings two issues, namely, heavier computational overheads and weaker search stability, which we solve using search space approximation and regularization, respectively. With a significantly reduced search time (~7 hours on a single GPU), our approach achieves state-of-the-art performance on both the proxy dataset (CIFAR10 or CIFAR100) and the target dataset (ImageNet). Code is available at https://github.com/chenxin061/pdarts. Xin Chen 0033, Lingxi Xie, Jun Wu 0006, Qi Tian 0001 |
ICCV | 3 |
| 2019 | Fair Scheduling in Resonant Beam Charging for IoT DevicesabstractResonant beam charging (RBC) is the wireless power transfer (WPT) technology, which can provide high-power, long-distance, mobile, and safe wireless charging for Internet of Things (IoT) devices. Supporting multiple IoT devices charging simultaneously is a significant feature of the RBC system. To optimize the multiuser charging performance, the transmitting power should be scheduled for charging all IoT devices simultaneously. In order to keep all IoT devices working as long as possible for fairness, we propose the first access first charge (FAFC) scheduling algorithm. Then, we formulate the scheduling parameters quantitatively for algorithm implementation. Finally, we analyze the performance of FAFC scheduling algorithm considering the impacts of the receiver number, the transmitting power, and the charging time. Based on the analysis, we summarize the methods of improving the WPT performance for multiple IoT devices, which include limiting the receiver number, increasing the transmitting power, prolonging the charging time, and improving the single-user's charging efficiency. The FAFC scheduling algorithm design and analysis provide a fair WPT solution for the multiuser RBC system. Wen Fang 0001, Qingwen Liu 0001, Jun Wu 0006 |
IEEE Internet Things J. | 4 |
| 2019 | An Optimal Resource Allocation for Hybrid Digital-Analog With Combined MultiplexingabstractA generalized hybrid digital-analog (HDA) framework with the combination of orthogonal and nonorthogonal multiplexing is proposed, which can strike a balance between interference and resource for Internet of Things application. The optimal resource allocation for the proposed scheme is formalized as a 3-D mixed integer programming problem, which is a function of digital bandwidth, orthogonal power, and nonorthogonal power of analog signal. With divide and conquer strategy, we first search the space of digital bandwidth, which is constructed by the possible number of subcarriers in orthogonal frequency division multiplex system, then the optimization problem is reduced to a 2-D continuous optimization problem. We further decompose it into two 1-D continuous optimization problems, and prove they are convex 1-D functions unconditionally or conditionally, respectively. With their convexity, the 2-D optimization problem can be solved with iterative gradient descent algorithm. We design a resource allocation algorithm to solve the optimization problem in practical system. Our experimental results show that the proposed algorithm outperforms nonorthogonal multiplexing HDA by 1-3 dB in terms of peak signal to noise ratio. Bin Tan 0001, Jun Wu 0006, Rui Wang 0001, Wenlang Luo |
IEEE Internet Things J. | 2 |
| 2019 | Efficient Soft Video MIMO Design to Combine Diversity and Spatial Multiplexing GainabstractHow to strike a balance between diversity gain and spatial multiplexing gain in a soft video delivery system is an open problem. Due to the power limit, it is especially important to achieve the optimal balance during video transmission in the Internet of Things. In this paper, taking the multisimilarity feature of a soft video system into account, we design an adaptive multiple-input, multiple-output (MIMO) receiver that can utilize nearly optimally either diversity gain or multiplexing gain. At the transmitter, we arrange multisimilar video data according to the space-time coding style and transmit them through multiple antennas. At the receiver, we propose using two decoders, i.e., the multisimilar space-time block coding (Ms-STBC) decoder and the soft MIMO decoder. The decoder to be chosen is determined by the predicted performance gain. We show that the proposed Ms-STBC transmission scheme can be considered a joint source-channel design, which is proposed for a soft video multiantenna delivery system. The relationship between intracodewords similarity and the channel signal-to-noise ratio (SNR) gain is derived. The experimental results demonstrate that the proposed designs can achieve a significant improvement over either individual soft MIMO multiplexing decoding or individual space-time block coding decoding in terms of peak SNR (PSNR) under the condition of a time-varying channel and a wide range of SNR. We can obtain at most 4-dB PSNR gain compared with the soft MIMO system. Jian Wu 0021, Bin Tan 0001, Jun Wu 0006, Rui Wang 0001 |
IEEE Internet Things J. | 3 |
| 2019 | TDMA in Adaptive Resonant Beam Charging for IoT DevicesabstractResonant beam charging (RBC) can realize wireless power transfer (WPT) from a transmitter to multiple Internet of Things devices via resonant beams. The adaptive RBC (ARBC) can effectively improve its energy utilization. In order to support multiuser WPT in the ARBC system, we propose the time-division multiple access (TDMA) method and design the TDMA-based WPT scheduling algorithm. Our TDMA WPT method has the features of concurrently charging, continuous charging current, individual user power control, constant driving power, and flexible driving power control. The simulation shows that the TDMA scheduling algorithm has high efficiency, as the total charging time is roughly half (46.9% when charging 50 receivers) of that of the alternative scheduling algorithm. Furthermore, the TDMA for WPT inspires the ideas of enhancing the ARBC system, such as flow control and quality of service. Mingliang Xiong, Mingqing Liu 0002, Qingwen Liu 0001, Jun Wu 0006 |
IEEE Internet Things J. | 5 |
| 2019 | Adaptive Resonant Beam Charging for Intelligent Wireless Power TransferabstractAs a long-range high-power wireless power transfer (WPT) technology, resonant beam charging (RBC) can transmit watt-level power over long distance for the devices in the Internet of Things (IoT). Due to its open-loop architecture, RBC faces the challenge of providing dynamic current and voltage to optimize battery charging performance. In RBC, battery overcharge may cause energy waste, thermal effects, and even safety issues. On the other hand, battery undercharge may lead to charging time extension and significant battery capacity reduction. In this paper, we present an adaptive RBC (ARBC) system for battery charging optimization. Based on RBC, ARBC uses a feedback system to control the supplied power dynamically according to the battery preferred charging values. Moreover, in order to transform the received current and voltage to match the battery preferred charging values, ARBC adopts a dc-dc conversion circuit. Relying on the analytical models for RBC power transmission, we obtain the end-to-end power transfer relationship in the approximate linear closed-form of ARBC. Thus, the battery preferred charging power at the receiver can be mapped to the supplied power at the transmitter for feedback control. Numerical evaluation demonstrates that ARBC can save 61% battery charging energy and 53%-60% supplied energy compared with RBC. Furthermore, ARBC has high energy-saving gain over RBC when the WPT is unefficient. ARBC in WPT is similar to link adaption in wireless communications. Both of them play the important roles in their respective areas. Wen Fang 0001, Mingliang Xiong, Qingwen Liu 0001, Jun Wu 0006 |
IEEE Internet Things J. | 5 |
| 2019 | Optimal Resonant Beam Charging for Electronic Vehicles in Internet of Intelligent VehiclesabstractTo enable electric vehicles (EVs) to access to the Internet of Intelligent Vehicles (IoIV), charging EVs wirelessly anytime and anywhere becomes an urgent need. The resonant beam charging (RBC) technology can provide high-power and long-range wireless energy for EVs. However, the RBC system is unefficient. To improve the RBC power transmission efficiency, the adaptive RBC (ARBC) technology was introduced. In this paper, after analyzing the modular model of the ARBC system, we obtain the closed-form formula of the end-to-end power transmission efficiency. Then, we prove that the optimal power transmission efficiency uniquely exists. Moreover, we analyze the relationships among the optimal power transmission efficiency, the source power, the output power, and the beam transmission efficiency, which provide the guidelines for the optimal ARBC system design and implementation. Hence, perpetual energy can be supplied to EVs in IoIV virtually. Mingqing Liu 0002, Qingwen Liu 0001, Jun Wu 0006 |
IEEE Internet Things J. | 5 |
| 2019 | Contract-Based Small-Cell Caching for Data Disseminations in Ultra-Dense Cellular NetworksabstractEvidence indicates that demands from mobile users (MU) on popular cloud content, e.g., video clips, account for a dramatic increase in data traffic over cellular networks. The repetitive downloading of hot content from cloud servers will inevitably bring a vast quantity of redundant data transmissions to networks. A strategy of distributively pre-storing popular cloud content in the memories of small-cell base stations (SBS), namely, small-cell caching, is an efficient technology for reducing the communication latency whilst mitigating the redundant data streaming substantially. In this paper, we establish a commercialized small-cell caching system consisting of a network service provider (NSP), several video providers (VP), and randomly distributed MUs. We conceive this system in the context of 5G cellular networks, where the SBSs are ultra-densely deployed with the intensity much higher than that of the MUs. In such a system, the NSP, in charge of the SBSs, wishes to lease these SBSs to the VPs for the purpose of making profits, whilst the VPs, after pushing popular videos into the rented SBSs, can provide faster local video transmissions to the MUs, thereby gaining more profits. Specifically, we first model the MUs and SBSs as two independent Poisson point processes, and develop, via stochastic geometry theory, the probability of the specific event that an MU obtains the video of its choice directly from the memory of an SBS. Then, with the help of the probability derived, we formulate the profits of both the NSP and the VPs. Next, we solve the profit maximization problem based on the framework of contract theory, where the NSP acts as a monopolist setting up the optimal contract according to the statistical information of the VPs. Incentive mechanisms are also designed to motivate each VP to choose a proper resource-price item offered by the NSP. Numerical results validate the effectiveness of our proposed contract framework for the commercial caching system. Jun Li 0004, Shunfeng Chu, Feng Shu 0002, Jun Wu 0006, Dushantha N. K. Jayakody |
IEEE Trans. Mob. Comput. | 4 |
| 2019 | A Multi-Grained Parallel Solution for HEVC Encoding on Heterogeneous PlatformsabstractTo improve the parallel processing capability of video coding, the emerging high efficiency video coding (HEVC) standard introduces two parallel techniques, i.e., Wavefront Parallel Processing (WPP) andTiles, to make it much more parallel-friendly than its predecessors. However, these two techniques are designed to explore coarse-grained parallelism in HEVC encoding on multicore Central Processing Unit (CPU) platforms. As the computing architecture undergoes a trend toward heterogeneity in the last decade, multi-grained parallel computing methods can be designed to accelerate HEVC encoding on heterogeneous systems. In this paper, a multi-grained parallel solution (MPS) is proposed to optimize HEVC encoding on a typical heterogeneous platform. A massively parallel motion estimation algorithm is employed by MPS to parallelize part of HEVC encoding on Graphic Processing Unit (GPU). Meanwhile, several other HEVC encoding modules are accelerated on CPU through the cooperation of WPP and an adaptive parallel mode decision algorithm. The parallelism between CPU and GPU is well designed and implemented to guarantee an efficient concurrent execution of HEVC encoding on multi-grained parallel levels. The effectiveness of the proposed MPS for HEVC encoding is verified on a number of experiments. Bo Xiao 0005, Hanli Wang, Jun Wu 0006, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Multim. | 3 |
| 2019 | HARQ-Chaotic: Analog Chaotic Code Applied in HARQ Scheme of Wireless Communication SystemabstractThis paper proposes a novel symbol-level combining hybrid automatic repeat request (HARQ) scheme based on analog chaotic code and named HARQ-Chaotic. The transmitter of HARQ-Chaotic adopts analog chaotic code to encode Quadrature Amplitude Modulation (QAM) symbols of retransmission packets to combat fading and noise. As the analog chaotic code can only handle sources with amplitudes in the range of [-0.5, 0.5], QAM symbols must be scaled into this range. We derived the optimal scaling factor through theoretical analysis. A joint algorithm combining with novel soft chaotic decoder and novel soft QAM demapper is proposed for the receiver to enhance the performance of the whole communication system. We implemented HARQ-Chaotic with LDPC codes and 16-QAM/64-QAM to carry out simulations in both AWGN channels and multipath fading channels. Massive simulation results demonstrate that the proposed HARQ-Chaotic has 1dB-4dB gain over traditional HARQ-Chase combining (HARQ-CC) scheme in block error rate (BLER) performance. Fusheng Zhu, Jun Wu 0006, Rui Wang 0001, Haoqi Ren, Zhifeng Zhang 0001 |
Wirel. Commun. Mob. Comput. | 3 |
| 2018 | Optimal DoF Region of MIMO Y Channel with Hybrid Data ExchangesabstractWe study the optimal degrees of freedom (DoF) region of three-user asymmetric multiple-input multiple- output (MIMO) Y channel by considering a hybrid data exchange model. In the hybrid data exchange, we consider both the pairwise data exchange and full data exchange. To derive the optimal DoF region, we analyze the DoF region from both the converse and the achievability aspects. For the converse part, the tight DoF region outer bound is derived using cut-set theorem and genie aided approach. Further, a novel and systematic approach is proposed to analyze the achievability of DoF region. In proving the optimality of achievability, we propose to design distinct patterns to pack the overall transmit data over the channel. We find that the obtained achievable DoF coincides with derived DoF region outer bound. Rui Wang 0001, Xiaojun Yuan 0002, Jun Wu 0006, Wei Zhang 0001 |
GLOBECOM | 3 |
| 2018 | Improving Data Utility through Game Theory in Personalized Differential PrivacyabstractDue to dramatically increasing information published in social networks, privacy issues have given rise to public concerns. Although the presence of differential privacy provides privacy protection with theoretical foundations, the trade-off between privacy and data utility still demands further improvement. However, most existing works do not consider the impact of the adversary in the measurement of data utility. In this paper, we firstly propose a personalized differential privacy based on social distance. Then, we analyze the maximum data utility when users and adversaries are blind to the strategy sets of each other. We formulize all the payoff functions in the differential privacy sense, which is followed by the establishment of a Static Bayesian Game. The trade-off is calculated by deriving the Bayesian Nash Equilibrium. In addition, the in-place trade-off can maximize the user' data utility if the action sets of the user and the adversary are public while the strategy sets are unrevealed. Our extensive experiments on the real-world dataset prove the proposed model is effective and feasible. Youyang Qu, Lei Cui 0006, Shui Yu 0001, Wanlei Zhou 0001, Jun Wu 0006 |
ICC | 5 |
| 2018 | A Novel Forward-Link Multiplexed Scheme in Satellite-Based Internet of ThingsabstractSatellite communication has the potential to play a key role in many applications of Internet of Things (IoT). In this paper, we consider a satellite-based IoT and investigate the technology that can improve the spectral efficiency. In general, one beam in satellite systems serves one user. To serve multiple users, time division multiplexing or frequency division multiplexing is usually used. In this paper, we propose a novel forward-link multiplexed scheme, by which the signals of different users can be transmitted simultaneously using the same frequency band. Specifically, at the transmitter, we first map each combination of the users' constellation points to a higher-order constellation point, which is referred to as constellation coding, and then transmit such higher-order modulation signals. At the user side, after receiving and detecting the transmitted signal, each user obtain its own signal by the corresponding demapping, which is referred to as constellation decoding. The total system capacity over an additive white Gaussian noise channel is analyzed in this paper. Simulation results demonstrate that the proposed scheme can greatly improve the spectral efficiency. Die Hu 0002, Lianghua He, Jun Wu 0006 |
IEEE Internet Things J. | 3 |
| 2018 | Degrees of Freedom of the Circular Multirelay MIMO Interference Channel in IoT NetworksabstractIn this paper, we study the degrees of freedom (DoF) of a new network information flow model named the circular multirelay multiple-input multiple-output interference channel (CMMI). In this model, there are two clusters and each of them contains three users. Each user equipped with M antennas in one cluster intends to deliver data streams to another user in the same cluster in a circular one-way transmission via the common distributed K N-antenna relay nodes. The CMMI network model can be considered as a basic component to construct the complicated Internet of Things networks. By assuming linear processing at the users and the relays, we show that the original analysis of DoF comes down in finding solutions of some nonlinear matrix equations with rank constraints. Toward this end, by using linear precoding and post-processing techniques, we propose two different approaches to solve the nonlinear matrix equations based on different antenna configurations. We show that a √ DoF of max{min{M, (√6K/12)}, min{(M/3), (KN/2)}} is achievable for ∀(M/N) ∈ (0, +∞). In addition, to assess the optimal DoF, the cut-set approach is used for deriving the DoF upper bound by innovatively separating certain users to form two-pair two-way relay channels. We show that the DoF of CMMI is upper bounded by max{min{M, (KN/3)}, min{(2M/3), (KN/2)}}. By combining the achievable DoF and the upper bound, we finally show that the optimal DoF of CMMI can be achieved √ for (M/N)∈[0, (√6K/12)]∪[(3K/2), +∞), ∀K ≥ 1. Wenjing Lv, Rui Wang 0001, Jun Wu 0006, Jianwu Dou |
IEEE Internet Things J. | 3 |
| 2018 | Distributed Laser Charging: A Wireless Power Transfer ApproachabstractWireless power transfer (WPT) is a promising solution to provide convenient and perpetual energy supplies to electronics. Traditional WPT technologies face the challenge of providing Watt-level power over meter-level distance for Internet of Things (IoT) and mobile devices, such as sensors, controllers, smart-phones, laptops, etc. Distributed laser charging (DLC), a new WPT alternative, has the potential to solve these problems and enable WPT with the similar experience as WiFi communications. In this paper, we present a multimodule DLC system model, in order to illustrate its physical fundamentals and mathematical formula. This analytical modeling enables the evaluation of power conversion or transmission for each individual module, considering the impacts of laser wavelength, transmission attenuation, and photovoltaic-cell (PV-cell) temperature. Based on the linear approximation of electricity-to-laser and laser-to-electricity power conversion validated by measurement and simulation, we derive the maximum power transmission efficiency in closed-form. Thus, we demonstrate the variation of the maximum power transmission efficiency depending on the supply power at the transmitter, laser wavelength, transmission distance, and PVcell temperature. Similar to the maximization of information transmission capacity in wireless information transfer (WIT), the maximization of the power transmission efficiency is equally important in WPT. Therefore, this paper not only provides the insight of DLC in theory, but also offers the guideline of DLC system design in practice. Wen Fang 0001, Qingwen Liu 0001, Jun Wu 0006, Liuqing Yang 0001 |
IEEE Internet Things J. | 4 |
| 2018 | Improved KMV-Cast with BM3D Denoising
Xin-Lin Huang, Xiaowei Tang 0001, Xiaoning Huan, Ping Wang 0004, Jun Wu 0006 |
Mob. Networks Appl. | 5 |
| 2018 | Degrees of Freedom of a MIMO Multipair Two-Way Relay Channel With Delayed Channel State InformationabstractWe study the degrees of freedom (DoFs) of a multiple-input multiple-output K-pair two-way relay channel with delayed channel state information (CSI) with J distributed relays. In the considered model, we assume that the users are equipped with M antennas and the relay with N antennas. We propose two schemes, where the signaling design can be carried out either in each individual time slot or across multiple time slots with and without knowledge of CSI. The scheme involves a joint design of user beamforming matrices, relay beamforming matrices, and user postprocessing matrices, so as to meet interference neutralization and rank conditions. We show that the optimal DoF of N2K per user can be reached for N ≥ 1 for an arbitrary number of user pairs when J = 1. This implies that delayed CSI does not compromise the DoF performance of the considered model when N ≥ 1. Rui Wang 0001, Xiaojun Yuan 0002, Jun Wu 0006 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Cost-Distortion Optimization and Resource Control in Pseudo-Analog Visual CommunicationsabstractThe rate-distortion in conventional digital systems is replaced by cost distortion in pseudo-analog systems where the cost consists of power and bandwidth. In this paper, we formulate the cost-distortion optimization problem in terms of a power-bandwidth pair versus distortion to bring an insight to pseudo-analog transmission. Using a divide-and-conquer strategy, the 3-D optimization problem of a power-bandwidth pair versus distortion is decomposed into two subproblems: power distortion and bandwidth distortion optimization. To solve the integer nonlinear optimization problem, we propose two prediction models that transform the partial summation of variances and the square roots of variances into continuous functions. The proposed models are used to derive the closed-form solutions for both optimization subproblems, and a tradeoff between power and bandwidth is discussed. Accordingly, the resource control algorithm is designed to allocate the fewest resources required to obtain specific video quality; power can be traded for bandwidth and vice versa. Our experimental results show that the proposed optimization models achieve stable perceptual video quality comparable to that of Groups of Pictures and use resources more efficiently than do SoftCast. Dian Liu, Jun Wu 0006, Hao Cui 0001, Chong Luo 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | A Collaborative Scheduling-Based Parallel Solution for HEVC Encoding on Multicore PlatformsabstractIn order to meet the high computational demand to achieve superior coding efficiency and to explore the parallelism of parallel processing architectures, the emerging high efficiency video coding (HEVC) standard has been designed to be more parallelizable than previous video coding standards. However, it is still desirable to design an efficient parallel HEVC encoder to fully exploit the parallelism of the increasingly powerful multicore platforms, especially when considering the amount of parallelism, the scalability of parallelization, and the coding efficiency. In this work, a performance model of HEVC encoding is first introduced to investigate the speedup and the limitations of the technique of wavefront parallel processing (WPP) under various conditions. Then, a collaborative scheduling-based parallel solution (CSPS) for HEVC encoding is proposed, which includes adaptive parallel mode decision, asynchronous frame-level pixel interpolation, and multigrained task scheduling. The goal of the proposed CSPS is to defeat the disadvantages of WPP and further improve the parallelization of HEVC encoding on multicore platforms. Extensive experimental results demonstrate the efficiency of the proposed CSPS for parallelizing HEVC encoding as the computing resources of multicore architectures can be fully utilized. Hanli Wang, Bo Xiao 0005, Jun Wu 0006, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Multim. | 3 |
| 2017 | Multi-index fusion via similarity matrix pooling for image retrievalabstractDifferent kinds of features hold some distinct merits, making them complementary to each other. Inspired by this idea an index level multiple feature fusion scheme via similarity matrix pooling is proposed in this paper. We first compute the similarity matrix of each index, and then a novel scheme is used to pool on these similarity matrices for updating the original indices. Compared with the existing fusion schemes, the proposed scheme performs feature fusion at index level to save memory and reduce computational complexity. On the other hand, the proposed scheme treats different kinds of features adaptively based on its importance, thus improves retrieval accuracy. The performance of the proposed approach is evaluated using two public datasets, which significantly outperforms the baseline methods in retrieval accuracy with low memory consumption and computational complexity. Xin Chen 0033, Jun Wu 0006, Shaoyan Sun, Qi Tian 0001 |
ICC | 2 |
| 2017 | Adaptive Distributed Laser Charging for Efficient Wireless Power TransferabstractDistributed laser charging (DLC) is a wireless power transfer technology for mobile electronics. Similar to traditional wireless charging systems, the DLC system can only provide constant power to charge a battery. However, Li-ion battery needs dynamic input current and voltage, thus power, in order to optimize battery charging performance. Therefore, neither power transmission efficiency nor battery charging performance can be optimized by the DLC system. We at first propose an adaptive DLC (ADLC) system to optimize wireless power transfer efficiency and battery charging performance. Then, we analyze ADLC's power conversion to depict the adaptation mechanism. Finally, we evaluate the ADLC's power conversion performance by simulation, which illustrates its efficiency improvement by saving at least 60.4% of energy, comparing with the fixed-power charging system. Qingwen Liu 0001, Jun Wu 0006 |
VTC Fall | 4 |
| 2017 | An Optimal Resource Allocation for Superposition Coding-Based Hybrid Digital-Analog SystemabstractHybrid digital-analog (HDA) video transmission is a new cross-layer design, which can be widely used in Internet of Things. The key problem of HDA video transmission is to find the optimal resource allocation between the digital and analog part. This paper presents a new general resource allocation algorithm for superposition coding-based HDA system. On one hand, in order to achieve successful decoding in digital part, the bitrate is controlled by the quantization parameter (QP), and the channel coding rate and modulation order are determined by the signal to interference noise power ratio, in which analog part is considered as the interference. On the other hand, the overall video quality is directly determined by the mean square error of analog part, which depends jointly on the data variance of the analog part, the power allocated to the analog part and the channel noise power. We propose a prediction model to describe how the data variance of the analog part changes with the QP in the digital part. Based on the proposed model, the power allocation of two parts can be quantitatively connected to form an optimization problem. We prove the convexity of the resource allocation problem and the gradient descent method is utilized in system implementation. With extensive simulations, the proposed algorithm is validated, achieving 1.4-5.3 dB gain over the conventional digital system, and 6.2-7.4 dB gain over pseudo-analog system in peak signal-to-noise ratio. Bin Tan 0001, Hao Cui 0001, Jun Wu 0006, Chang Wen Chen |
IEEE Internet Things J. | 3 |
| 2017 | Cross-Platform Resource Scheduling for Spark and MapReduce on YARNabstractWhile MapReduce is inherently designed for batch and high throughput processing workloads, there is an increasing demand for non-batch processes on big data, e.g., interactive jobs, real-time queries, and stream computations. Emerging Apache Spark fills in this gap, which can run on an established Hadoop cluster and take advantages of existing HDFS. As a result, the deployment model of Spark-on-YARN is widely applied by many industry leaders. However, we identify three key challenges to deploy Spark on YARN, inflexible reservation-based resource management, inter-task dependency blind scheduling, and the locality interference between Spark and MapReduce applications. The three challenges cause inefficient resource utilization and significant performance deterioration. We propose and develop a cross-platform resource scheduling middleware, iKayak, which aims to improve the resource utilization and application performance in multi-tenant Spark-on-YARN clusters. iKayak relies on three key mechanisms: reservation-aware executor placement to avoid long waiting for resource reservation, dependency-aware resource adjustment to exploit under-utilized resource occupied by reduce tasks, and cross-platform locality-aware task assignment to coordinate locality competition between Spark and MapReduce applications. We implement iKayak in YARN. Experimental results on a testbed show that iKayak can achieve 50 percent performance improvement for Spark applications and 19 percent performance improvement for MapReduce applications, compared to two popular Spark-on-YARN deployment models, i.e., YARN-client model and YARN-cluster model. Dazhao Cheng, Xiaobo Zhou 0002, Palden Lama, Jun Wu 0006, Changjun Jiang 0002 |
IEEE Trans. Computers | 4 |
| 2017 | Knowledge-Enhanced Mobile Video Broadcasting Framework With Cloud SupportabstractThe convergence of mobile communications and cloud computing facilitates the cross-layer network design and content-assisted communication. Mobile video broadcasting can benefit from this trend by utilizing joint source-channel coding and strong information correlation in clouds. In this paper, a knowledge-enhanced mobile video broadcasting (KMV-Cast) is proposed. The KMV-Cast is built on a linear video transmission instead of a traditional digital video system, and exploits the hierarchical Bayesian model to integrate the correlated information into the video reconstruction at the receiver. The correlated information is distilled to obtain its intrinsic features, and the Bayesian estimation algorithm is used to maximize the video quality. The KMV-Cast system consists of both likelihood broadcasting and prior knowledge broadcasting. The simulation results show that the proposed KMV-Cast scheme outperforms the typical linear video transmission scheme called Softcast, and achieves 8 dB more of the peak signal-to-noise ratio (PSNR) gain at low-SNR channels (i.e., -10 dB), and 5 dB more of PSNR gain at high-SNR channels (i.e., 25 dB). Compared with the traditional digital video system, the proposed scheme has 7 dB more of PSNR gain than the JPEG2000 + 802.11a scheme at a 10-dB channel SNR. Xin-Lin Huang, Jun Wu 0006, Fei Hu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Analog Coded SoftCast: A Network Slice Design for Multimedia Broadcast/MulticastabstractThis paper presents a network slice design for ultra high definition (UHD) video broadcast/multicast to achieve higher network efficiency and improved quality of experience (QoE). The proposed network slice design consists of a rateless source compression scheme and an analog-coded SoftCast scheme. The rateless Spinal code is adopted to compress the video source at content server and the compressed source is transmitted from content server across wireless core network to the base station. Ana prioriinformation-assisted Spinal decoder is designed to utilize the sparsity of bit planes for compression. In the analog-coded SoftCast scheme, we design a new chaotic function-based analog code with negligible power penalty for the generalized Gaussian-distributed source in SoftCast because the existing chaotic functions designed for uniformly distributed sources suffer from serious power penalty in SoftCast. We also design a maximuma posterioriprobability decoding algorithm for the proposed analog code in order to exploit the statistics of video source asa prioriinformation to improve the performance. The experimental results show that the proposed rateless code-based compression scheme achieves efficient compression and approaches the bound of binary erasure channel. In particular, the 1/2 analog-coded SoftCast has almost 2 dB gain over conventional SoftCast with two repetitions, and the 1/3 analog-coded SoftCast has almost 3 dB gain over conventional SoftCast with three repetitions. The system simulations for the broadcast system show higher network capacity and improved QoE in the proposed UHD slice, because the reconstructed video quality of each user is commensurate with its channel condition. Bin Tan 0001, Jun Wu 0006, Ying Li 0020, Hao Cui 0001, Chang Wen Chen |
IEEE Trans. Multim. | 2 |
| 2017 | Joint Compression of Near-Duplicate VideosabstractThe expanding social network and multimedia technologies encourage more and more people to store and transmit information in visual format, such as image and video. However, the cost of this convenience brings about a shock to traditional video severs and exposes them under the risk of overloading. In the huge volume of online videos, there are a large amount of near-duplicate videos (NDVs). Although quite a number of research work have been proposed to detect NDVs, little research effort is made to compress these NDVs in a more effective manner than independent video compression. In this study, we make an in-depth exploration of the data redundancy of NDVs and propose a video analysis and coding framework to jointly compress NDVs. In order to employ the proposed NDV analysis and coding framework, a graph-based similar video grouping method and a number of preprocessing functions are designed to explore the correlation of visual information among NDVs and thus suit the requirement of joint video coding. Experimental results verify that the proposed NDV analysis and coding framework is able to effectively compress NDVs and thus save video data storage. Hanli Wang, Tao Tian, Jun Wu 0006 |
IEEE Trans. Multim. | 4 |
| 2016 | Robust Uncoded Video Transmission under Practical Channel EstimationabstractThis research solves the performance degradation problem of uncoded video system when the multiplicative noise caused by channel estimation error is not negligible. Through extensive analysis, we find that transmitting a specific part of source data by multiple channel use can decrease the mean end-to-end distortion and stabilize the video quality. However, under limited power and bandwidth constraints, allocating multiple channel use to one part of data means dropping another part of data. How to allocate the channel use to highly diverse video coefficients is an NP-hard problem. For the sake of practical implementation, a greedy iterative algorithm is proposed to achieve the optimal trade off between distortion decrease by multiple channel use and distortion increase from data dropping. Extensive simulations validate our analysis result. The proposed algorithm achieve 1.99~8.63dB gain compared to existing schemes without considering the multiplicative noise. Hao Cui 0001, Dian Liu, Yuqi Han, Jun Wu 0006 |
GLOBECOM | 4 |
| 2016 | Performance Analysis of KMV-Cast with Imperfect Prior KnowledgeabstractIt is predicted that the global mobile data traffic will increase nearly eightfold in the next five years, and the cloud applications will consist of 90 percent of total mobile data traffic by the end of 2019. There is no doubt that the increasing high quality video requirements and huge similar information stored in cloud servers enforce to setup a new communication paradigm. In this paper, we will take a brief review over the knowledge-enhanced mobile video broadcasting (KMV-Cast) framework, and analyze its performance with imperfect prior knowledge, which means that the prior knowledge will be transmitted through Gaussian channel. The simulation results have shown that KMV-Cast with imperfect prior knowledge performs the worst compared with the other two schemes in terms of PSNR. It indicates that the prior knowledge is very important for the video reconstruction in KMV-Cast. Meanwhile, with the SNR increasing for prior knowledge transmission, the performance still performs poor and has an optimum point. Xin-Lin Huang, Xiaoning Huan, Jun Wu 0006, Qingquan Sun, Yingchun Yuan |
GLOBECOM | 3 |
| 2016 | The Stable Channel State Analysis for Multimedia Packets Allocation over Cognitive Radio NetworksabstractIn cognitive radio networks (CRNs), the spectrum utilization can be dramatically improved with the secondary users (SUs) accessing the unoccupied licensed channels opportunistically. However, how to fully utilize the spectrum holes to meet the quality of service requirements of SUs is still an open issue. In this paper, we focus on multi-user, multi-channel case, and analyze the stable channel state after allocating packets over a licensed channel. We first assume that the SUs' packet arrival rate obeys Poisson distribution, and the lost packets will be retransmitted with exponential backoff delay. Then, we analyze the stable channel state when a SU selects and allocates certain percentage of its packets over one channel. The theoretic analysis shows that such stable channel state can be solved by the steady-state equations. Based on such stable channel state analysis, we propose a novel greedy packet allocation scheme in multi-user and multi-channel S-ALOHA system. The proposed packet allocation scheme will obtain a maximal spectrum utilization, which will support more SUs to share the spectrum holes or cause less access conflict with PUs. In the simulation part, we assume the PUs' channel access pattern follows Markov model, and study the spectrum efficiency after packet loading over licensed channels and make comparisons with different packet allocation schemes. The simulation results show that the proposed packet allocation scheme outperforms other related works in terms of successful packet delivery ratio, the number of collision packets, and spectrum efficiency. Xin-Lin Huang, Xiaowei Tang 0001, Jun Wu 0006 |
GLOBECOM | 4 |
| 2016 | Resource Allocation for Uncoded Multi-user Video Transmission over Wireless Networks
Dian Liu, Hao Cui 0001, Jun Wu 0006, Chong Luo 0001 |
Mob. Networks Appl. | 3 |
| 2016 | SWIFT: A Computationally-Intensive DSP Architecture for Communication Applications
Haoqi Ren, Zhifeng Zhang 0001, Jun Wu 0006 |
Mob. Networks Appl. | 3 |
| 2016 | Adaptive Hybrid Digital-Analog Video Transmission in Wireless Fading ChannelabstractWe propose an adaptive hybrid digital-analog video transmission (A-HDAVT) scheme for robust video streaming in mobile networks with realistic fading channels. This scheme is fundamentally different from recent research in hybrid digital-analog video transmission in which wireless channels are unrealistically assumed to be Gaussian. For fading channels, it is critical to take full advantage of diversity in both video contents and multiuser channels. Like all hybrid approaches, A-HDAVT is designed to exploit the benefits from both digital and analog systems. To achieve this goal, each group of pictures is first transformed into one low-pass frame and several high-pass frames with motion-compensated temporal filtering. The critical low-pass frame is reliably transmitted as base layer in a digital mode, while high-pass frames are transmitted as enhancement layers in an analog mode to achieve desired graceful degradation performance. In analog transmission, we introduce a channel prediction-based adaptive power-distortion optimization (P-APDO) scheme to combat channel fading in mobile networks. The basic idea behind P-APDO is to perform power allocation based on the video content as well as the predicted channel status. Furthermore, we also investigate the multiuser scenarios in which the content diversity and channel diversity among users are appropriately exploited. Extensive simulations have been carried out to evaluate the performance of A-HDAVT under various degrees of channel fading. The results show that A-HDAVT achieves significant performance gains over competing schemes Robust Uncoded Video Transmission, Parcast, and Softcast in both single-user and multiuser scenarios. Hancheng Lu, Chang Wen Chen, Jun Wu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Historical Spectrum Sensing Data Mining for Cognitive Radio Enabled Vehicular Ad-Hoc NetworksabstractIn vehicular ad-hoc network (VANET), the reliability of communication is associated with driving safety. However, research shows that the safety-message transmission in VANET may be congested under some urgent communication cases. More spectrum resource is an effective way to solve transmission congestion. Hence, we introduce cognitive radio (CR) enabled VANET (CR-VANET), where CR device can detect possible idle spectrum for VANET communications and assist to timely broadcast safety-message. Given high-speed mobility of vehicles and dynamically-changing availability of channels, a novel prediction algorithm is proposed to pick out the channel with the greatest probability of availability, which can meet the quality of service (QoS) requirement of urgent communications and effectively avoid conflict with licensed users. Specifically, the spatiotemporal correlations among historical spectrum sensing data are exploited to form prior knowledge of channel availability probability, and Bayesian inference is used to derive posterior probability of channel availability. Comparing with other spectrum detection methods, the proposed algorithm has more than 8 percent detection performance improvement at false alarm probability 0.2, and thus can avoid access conflict with licensed users dramatically. Furthermore, the proposed algorithm always has larger packet reception probability (PRP) and lower transmission delay compared with conventional VANET broadcasting. Hence, the proposed algorithm can improve reliability of safety-message transmission and enhance driving safety significantly. Xin-Lin Huang, Jun Wu 0006, Zhifeng Zhang 0001, Fusheng Zhu, Minghao Wu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2016 | DAC-Mobi: Data-Assisted Communications of Mobile Images with Cloud Computing SupportabstractThis research proposes a novel data assisted image transmission scheme, which utilizes a large amount of correlated images stored in the cloud to improve the spectrum efficiency and visual quality. First, a two-layer Coset coding is proposed for the DCT coefficients transmission. The most significant bits (MSB) of the coefficients are generated by the first layer Coset and together with a few low frequency coefficients are transmitted through the most reliable channel coding and digital modulation. The middle bits generated by the second layer Coset are discarded by the sender and the residual bits are transmitted through amplitude modulation. Based on the MSB and the residual bits, an approximation of the original image is reconstructed. With this approximation, a lot of correlated images can be retrieved from the cloud, which are used to recover the discarded middle bits. The two layer Coset coding can significantly decrease the data energy so as to improve the transmission power efficiency. Hence, the end to end distortion of amplitude modulation can be reduced. Second, the image quality can be further improved by joint internal and external denoising with the retrieved images. Simulations show that the proposed scheme outperforms conventional digital schemes about 4 dB in peak signal to noise power ratio (PSNR) and achieves 2 dB gain over the state-of-the-art uncoded transmission. At low signal to noise power ratio (SNR), an additional 2-3 dB gain is achieved. The visual quality comparison also validates the objective image assessment result. Jun Wu 0006, Jian Wu 0022, Hao Cui 0001, Chong Luo 0001, Xiaoyan Sun 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2016 | Rate-Adaptive Feedback With Bayesian Compressive Sensing in Multiuser MIMO Beamforming SystemsabstractMultiple-input multiple-output (MIMO) is a promising way to increase link capacity and energy efficiency in the next generation communication systems. However, the benefits of such an approach depend on proper channel state information (CSI) availability at the transmitter. The CSI is usually estimated at the receiver and fed back to the transmitter through a band-limited channel. Thus, an efficient feedback scheme is needed. In this paper, a comprehensive Bayesian compressive sensing (BCS) based feedback mechanism is proposed for time-varying spatially and temporally correlated vector autoregression (VAR) wireless channel, and the feedback rate distortion function is derived in closed form in statistics. The proposed BCS feedback scheme utilizes the sparse CSI features and prior knowledge to significantly compress the dimensionality of the feedback CSI. Furthermore, the relationship between the feedback rate and downlink capacity is derived in closed form in statistics to guide rate-adaptive feedback in MIMO system. We find out that the ergodic downlink capacity of a user is determined only by its own feedback rate in the proposed feedback scheme. Theoretical and simulation results all show that the proposed feedback scheme can realize efficient, rate-adaptive feedback based on downlink capacity requirement, and the proposed feedback performance is superior to other related works. Xin-Lin Huang, Jun Wu 0006, Yonggang Wen 0001, Fei Hu 0001, Yi Wang 0018, Tao Jiang 0002 |
IEEE Trans. Wirel. Commun. | 2 |
| 2015 | Accelerating Large-scale Image Retrieval on Heterogeneous Architectures with SparkabstractApache Spark is a general-purpose cluster computing system for big data processing and has drawn much attention recently from several fields, such as pattern recognition, machine learning and so on. Unlike MapReduce, Spark is especially suitable for iterative and interactive computations. With the computing power of Spark, a utility library, referred to as IRlib, is proposed in this work to accelerate large-scale image retrieval applications by jointly harnessing the power of GPU. Similar to the built-in machine learning library of Spark, namely MLlib, IRlib fits into the Spark APIs and benefits from the powerful functionalities of Spark. The main contributions of IRlib lie in two-folds. First, IRlib provides a uniform set of APIs for the programming of image retrieval applications. Second, the computational performance of Spark equipped with multiple GPUs is dramatically boosted by developing high performance modules for common image retrieval related algorithms. Comparative experiments concerning large-scale image retrieval are carried out to demonstrate the significant performance improvement achieved by IRlib as compared with single CPU thread implementation as well as Spark without GPUs employed. Hanli Wang, Bo Xiao 0005, Lei Wang 0063, Jun Wu 0006 |
ACM Multimedia | 4 |
| 2015 | Human Action Recognition With Trajectory Based Covariance Descriptor In Unconstrained VideosabstractHuman action recognition from realistic videos plays a key role in multimedia event detection and understanding. In this paper, a novel Trajectory Based Covariance (TBC) descriptor is proposed, which is formulated along the dense trajectories. To map the descriptor matrix to vector space and trim out the redundancy of data, the TBC descriptor matrix is projected to Euclidean space by the Logarithm Principal Components Analysis (LogPCA). Our method is tested on the challenging Hollywood2 and TV Human Interaction datasets. Experimental results show that the proposed TBC descriptor outperforms three baseline descriptors (i.e., histogram of oriented gradient, histogram of optical flow and motion boundary histogram), and our method achieves better recognition performances than a number of state-of-the-art approaches. Hanli Wang, Yun Yi, Jun Wu 0006 |
ACM Multimedia | 3 |
| 2015 | Correlative Filters for Convolutional Neural NetworksabstractThis paper introduces a regularization method called Correlative Filter (CF) for Convolutional Neural Network (CNN), which takes advantage of the relevance between the convolutional kernels belonging to the same convolutional layer. During the process of training with the proposed CF method, several pairs of filters are designed in a manner of randomness to contain opposite weights in low-level layers. Regarding higher level layers where synthetical features are processed, the relation between correlative filters is explored as translation of various directions. The proposed CF method attempts to optimize the inner structure of convolutional layers and it can work jointly with other regularization techniques, such as stochastic pooling, Dropout, etc. The experimental results on the competitive image classification benchmark dataset CIFAR-10 demonstrates the performance of the proposed CF method, additionally, it is also verified that the proposed CF method is wonderful to be employed to enhance several state-of-the-art regularization models. Peiqiu Chen, Hanli Wang, Jun Wu 0006 |
SMC | 3 |
| 2015 | Accelerating Support Vector Machine Learning with GPU-Based MapReduceabstractWith the exploding growth of data, the computational complexity required by learning Support Vector Machine (SVM) lays a heavy burden on real-world applications. To address this issue, parallel computational techniques can be employed such as the Graphics Processing Units (GPUs) and MapReduce model. As it is well known, GPUs are microprocessors on a multi-core architecture which reveal high performance in mass data parallel computing, and MapReduce allows computational tasks to be divided into a plurality of parts, distributed to various computing nodes and combined on a single node. In this paper, we propose a GPU-based MapReduce framework to accelerate SVM learning by jointly utilizing the parallel computing power of GPU and MapReduce. Extensive experimental results have verified the effectiveness and efficiency of the proposed approach. Tianyao Sun, Hanli Wang, Jun Wu 0006 |
SMC | 4 |
| 2015 | Intelligent Cooperative Spectrum Sensing via Hierarchical Dirichlet Process in Cognitive Radio NetworksabstractCognitive radio (CR) is a critical technology for improving spectrum utilization and solving the radio spectrum scarcity problem. In CR devices, spectrum sensing is important to implement opportunistic spectrum access. Many spectrum sensing schemes have been proposed, including uncooperative, cooperative, centralized, and distributed algorithms. However, they aimed to obtain a global consensus sensing result, which may not always be possible in large-scale cognitive radio networks (CRNs) due to heterogeneous spectrum availability in different areas. Hence, some new spectrum sensing schemes should be designed to discover idle heterogeneous spectrum in CRNs. In this paper, we propose an intelligent cooperative spectrum sensing algorithm based on a non-parametric Bayesian learning model, namely the hierarchical Dirichlet process, which groups spectrum sensing data without the need to know the number of hidden spectrum states, and discovers a common sparse spectrum within each group. Furthermore, a concisely distributed information exchange scheme is designed, where intra-cluster and inter-cluster spectrum information is shared for global spectrum cognition. Experimental results show that the proposed algorithm can exploit the spatial relationship among sensed data to achieve a better spectrum sensing performance in terms of detection probability and false alarm probability. Xin-Lin Huang, Fei Hu 0001, Jun Wu 0006, Hsiao-Hwa Chen, Gang Wang 0021, Tao Jiang 0002 |
IEEE J. Sel. Areas Commun. | 3 |
| 2015 | CHCF: A Cloud-Based Heterogeneous Computing Framework for Large-Scale Image RetrievalabstractThe last decade has witnessed a dramatic growth of multimedia content and applications, which in turn requires an increasing demand of computational resources. Meanwhile, the high-performance computing world undergoes a trend toward heterogeneity. However, it is never easy to develop domain-specific applications on heterogeneous systems while maximizing the system efficiency. In this paper, a novel framework, namely, cloud-based heterogeneous computing framework (CHCF), is proposed with a set of tools and techniques for compilation, optimization, and execution of multimedia mining applications on heterogeneous systems. With the aid of the compiler and the utility library provided by CHCF, users are able to develop multimedia mining applications rapidly and efficiently. The proposed framework employs a number of techniques, including adaptive data partitioning, knowledge-based hierarchical scheduling, and performance estimation, to achieve high computing performance. As one of the most important multimedia mining applications, large-scale image retrieval is investigated based on the proposed CHCF. The scalability, computing performance, and programmability of CHCF are studied for large-scale image retrieval by case studies and experimental evaluations. The experimental results demonstrate that CHCF can achieve good scalability and significant computing performance improvements for image retrieval. Hanli Wang, Bo Xiao 0005, Lei Wang 0063, Fengkuangtian Zhu, Yu-Gang Jiang 0001, Jun Wu 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2014 | Predicting zero coefficients for High Efficiency Video CodingabstractSimilar to previous video coding standards, transform and quantization are used in the most recent video coding standard High Efficiency Video Coding (HEVC) and a large number of transform coefficients are quantized to zeros. In order to reduce the computations involved in transform and quantization in HEVC, a prediction approach based on Gaussian model is proposed to predict zero quantized transform coefficients within the 4 × 4, 8 × 8, 16 × 16 and 32 × 32 blocks. Extensive experiments demonstrate that the proposed algorithm is able to effectively predict zero coefficients and thus reduce redundant computations while keeping the video quality and compression efficiency almost intact. Hanli Wang, Jun Wu 0006 |
ICME | 3 |
| 2014 | On reliability requirement for BSM broadcast for safety applications in DSRC systemabstractIn this paper, we derive application-level reliability metrics of various safety applications with worst-case settings and vehicular environments based on an accurate analytical model in one dimension vehicular communication networks. The reliability related QoS metrics in both the MAC-level and the Application-level including packet reception ratio, T-window reliability, and awareness probability are derived. Inspired by the correlation between road traffic and vehicle speed revealed by Greenshields model, we observe that the current reliability related QoS requirements for safety applications are set without considering change of road traffic. Based on the observation, a new reliability requirement setting is proposed to tune tolerance time for various safety applications with different road traffic. The QoS requirements for the three typical safety applications that are believed to have the most stringent QoS requirements are specified and discussed. Taking velocity, density and time into account, application-level reliability metrics are improved for high vehicle density. From the results obtained under worst case with different vehicle density, we can observe the feasibility of our proposed reliability requirement setting. Jun Wu 0006, Xiaomin Ma, Zhifeng Zhang 0001 |
Intelligent Vehicles Symposium | 2 |
| 2014 | Multimedia over cognitive radio networks: Towards a cross-layer scheduling under Bayesian traffic learning
Xin-Lin Huang, Gang Wang 0021, Fei Hu 0001, Sunil Kumar 0001, Jun Wu 0006 |
Comput. Commun. | 5 |
| 2014 | Early detection of all-zero 4×4 blocks in High Efficiency Video Coding
Hanli Wang, Weiyao Lin, Sam Kwong, Oscar C. Au, Jun Wu 0006, Zhihua Wei 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2013 | Compressive Coded Modulation for Seamless Rate AdaptationabstractThis paper presents a novel compressive coded modulation (CCM) which simultaneously achieves joint source-channel coding and seamless rate adaptation. The embedding of source compression into modulation brings significant throughput gain when the physical layer data contain non-negligible redundancy. The kernel of CCM is a new random projection (RP) code inspired by the compressive sensing (CS) theory. The RP code generates multilevel symbols from source binaries through weighted sum operations. Then, the generated RP symbols are mapped into a dense constellation for transmission. The receiver performs joint decoding based on received symbols. As the number of RP symbols can be adjusted in fine granularity, the rate adaptation becomes seamless. Two key design issues in the proposed CCM are addressed in this paper. First, we consider the RP code design for sources with different redundancies. Three principles are established and a concrete implementation is given. Second, we devise a linear-time decoding algorithm for the proposed RP code. In this belief propagation (BP) algorithm, we find that computing convolution in time domain is more efficient than that in frequency domain for binary variable nodes. Moreover, we invent a ZigZag deconvolution to further reduce the complexity. Analysis show that the proposed decoding algorithm is nearly 20 times faster than the state-of-the-art BP algorithm for CS called CS-BP. Emulations on traced data show that CCM achieves significant throughput gain, up to 33% and 70%, respectively, over the Hybrid ARQ with compression and BICM with compression, under practical time-varying wireless channels. Hao Cui 0001, Chong Luo 0001, Jun Wu 0006, Chang Wen Chen, Feng Wu 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2000 | Performance of MMSE Multiuser Detection for Downlink CDMAabstractPerformance analysis of minimum mean-square error (MMSE) multiuser detection over synchronous multipath channels is considered in this paper. Assuming the channel parameters are known, analysis shows that both the signal-to-noise ratio (SIR) at the output of the MMSE detector and near-far resistance are related to the cross-correlation matrix of combined spreading waveforms. Computation results demonstrate that the MMSE detector can achieve good performance under severe conditions, including a large number of users and large channel length. Yi Wang 0018, Jun Wu 0006, Zhimin Du, Weiling Wu |
ICC (2) | 2 |