Wen Ji 0003

dblp:74/4353-3 · DBLP profile ↗
← Back
54ranked-venue papers
20as first author
15since 2021 · last 2026
0000-0001-6895-3404ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 8 since 2021Systems, architecture and hardware · 15 · 2 first-author · 2 since 2021Computer networks · 14 · 8 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Toward Generation-Centric Coding: Compressing Latents representation for TI2V Synthesis
abstract
Text-Image-to-Video (TI2V) generation can be sensitive to compression distortions in the reference image. Small pixel-space artifacts may be amplified through iterative generation, leading to temporal inconsistency and semantic drift. One contributing factor is that conventional codecs are optimized for human-oriented pixel fidelity, without explicitly accounting for the distribution shift induced in the generator’s conditioning latents. Motivated by this mismatch, we propose a generation-centric latent-domain coding framework that compresses generator-facing deep latents instead of raw pixels, and jointly optimizes rate and latent-space distortion to preserve conditioning statistics. Experiments across four bitrate levels show that our approach mitigates distortion accumulation over iterations. It improves temporal coherence and content fidelity for both short and extended sequences, consistently reducing the pixel-level discrepancy relative to videos generated with lossless conditioning.
Jianran Liu, Wen Ji 0003, Xiaokai Meng, Wancai Zhang
ICMR2
2026 CLAP: Cross-Layer Adaptive Pipelining Inference Scheduling for Resource-Efficient Edge-Cloud Vision Systems
abstract
With the rapid growth of video-based applications, edge-cloud collaboration has become a mainstream paradigm for large-scale visual inference. However, existing edge-cloud systems primarily emphasize task offloading and static resource allocation, often overlooking the dynamic and heterogeneous nature of real-world scenarios. The significant variability in scene complexity across tasks leads to inefficient system performance. In this article, we propose CLAP, a cross-layer adaptive pipelining inference scheduling framework for edge-cloud vision systems. First, CLAP introduces a lightweight multiscale scene-aware module that accurately characterizes the visual complexity of incoming tasks at different granularities with minimal overhead. Based on this complexity profile, we design an adaptive multi-stage pipeline scheduling strategy, which dynamically adjusts processing granularity and selectively activates stages across edge and cloud nodes. Furthermore, we formulate the resource allocation as a multi-agent decision-making problem and employ cross-layer reinforcement learning to optimize task distribution under complex objectives, efficiently balancing accuracy, delay, and energy consumption. Extensive evaluations on public datasets demonstrate that CLAP can improve the throughput by more than 2.1x compared to traditional cloud-only and edge-only solutions while meeting accuracy requirements. Compared to state-of-the-art edge-cloud methods, CLAP achieves a 3% improvement in inference accuracy while simultaneously reducing end-to-end resource overhead, delay, and energy consumption by over 35%, proving its effectiveness in dynamic, large-scale vision applications.
Zheming Yang, Wen Ji 0003, Qi Guo 0009, Jian Zhao 0006, Xingzhou Zhang, Yangyu Zhang, Yang You 0001
ACM Trans. Archit. Code Optim.2
2026 End-to-End Coding Based on ViT Feature for Compression Distortion
abstract
Standard image codecs are misaligned with the feature space of vision-language models like CLIP, because they are designed to exploit spatial redundancies in pixel grids. In contrast, these features are dominated by channel-wise redundancy and contain only limited spatial structure. We empirically find that ViT-B/32 and ViT-B/16 features exhibit weak yet non-negligible spatial redundancy, sufficient to benefit from light local modeling but not from the aggressive spatial downsampling commonly used in learned image codecs. Guided by this, we develop an end-to-end feature compression framework for ViT-based CLIP. A compact transform maintains essential local structure while focusing modeling capacity on channel dependencies and entropy efficiency. Beyond the codec, we present a distortion-aware post-training strategy that adjusts CLIP to compression-related perturbations. A min–max optimization objective enhances robustness against worst-case feature distortions, and a comparison with an MSE-based post-training variant highlights the benefit of the robust formulation. Experiments on Flickr30k retrieval with \(224\times 224\) inputs show that our method consistently outperforms both traditional codecs (e.g., BPG and VVC intra) and learning based codecs (e.g., MLIC++ and STF), even without post-training. Additional post-training further increases this advantage across bitrates. Additional results with ViT-B/16, latent-space visualizations, kernel-size and stride ablations, and CPU inference measurements support the effectiveness and practicality of the proposed framework.
Jianran Liu, Wen Ji 0003
ACM Trans. Multim. Comput. Commun. Appl.4
2025 AODMS: Adaptive Online Edge-Cloud Collaborative Inference with Dynamic Model Switching and Resource Allocation
abstract
Cloud inference consumes massive amounts of resources and carbon emissions, which has prompted edge-cloud inference to become a new paradigm. However, constrained by limited edge resources, existing approaches predominantly adopt periodic model switching strategies, overlooking the heterogeneity of application requirements and the highly dynamic nature of task demands in large-scale edge-cloud environments. To address these challenges, this paper proposes an Adaptive Online Dynamic Model Switching (AODMS) framework, which jointly optimizes real-time model switching, cross-layer resource allocation, and edge-cloud collaborative inference to minimize long-term system cost, including task execution time and carbon emissions. First, an average saving cost evaluation mechanism is developed to determine optimal model switching points, coupled with a matrix-constrained pairwise search strategy for efficient exploration of feasible edge model loading. And then a distributed iterative optimization scheme is proposed for joint resource allocation and edge-cloud task scheduling. Extensive experimental evaluations demonstrate that AODMS can significantly reduce model switching frequency, average inference time, and Carbon Footprint (CF), offering a scalable and adaptive solution for the sustainable edge-cloud system.
Lulu Zuo, Zheming Yang, Wen Ji 0003
ICPADS4
2025 A Stackelberg Evolutionary Game Theoretic Framework for Dynamical Data Trading in Artificial Intelligence of Things
abstract
Artificial Intelligence of Things (AIoT) aims to build a self-learning, self-adaptive, and self-evolving Internet of Things ecosystem, which has facilitated many promising intelligent services. Data is an important foundational element for many applications. Establishing a well-designed trading mechanism to collect the necessary data from various sources is essential to realize the vision of AIoT. In this article we investigate the data trading incentive mechanism between multiproviders and multibuyers for AIoT. To address the two-sided dilemma, we develop a joint optimization game to maximize the payoff of all market participants. A two-layer Stackelberg evolutionary game theoretic framework is developed to divide the optimization problem into two subproblems: one for data pricing by providers and the other for purchasing decisions by buyers. The subproblem of optimal data pricing for providers is modeled as a noncooperative game. Providers utilize the game's equilibrium solution to dynamically modify their pricing strategies in response to a changing competitive environment and demanding. This is because buyers have limited information, their behaviors are modeled via evolutionary game. By encouraging data providers to take the buyers' evolutionary dynamics into account has the potential to overcome the myopia behaviors. The equilibrium solution is obtained via replicator dynamics. Extensive experiments demonstrate the efficacy and efficiency of the proposed hierarchical interaction framework. Overall, our results show the proposed Stackelberg evolutionary game framework establishes a desired data market and achieves higher long-term revenue for both sides of participants in the market. The hierarchical framework can effectively prompts data trading in the market.
Bo Shen 0006, Gang Yang 0008, Wen Ji 0003
IEEE Internet Things J.5
2025 Screen Content-Aware Video Coding Through Non-Local Model Embedded With Intra-Inter In-Loop Filtering
abstract
Many studies have focused on utilizing convolutional neural networks (CNNs) to enhance loop filter performance in video encoding. However, existing methods primarily concentrate on improving the natural sequence quality rather than addressing the specific needs of screen content sequences, which have gained increased attention due to the growing demands of remote desktops and online meetings. This paper proposed to understand machine behavior from the machine’s point of view, and adopts the machine intelligence to screen content coding. It presents a novel loop filter specifically tailored for screen content coding (SCC), referred to as video coding-SCC (VC-SCC). It employs a multiscale feature extraction structure and introduces two innovative non-local models to address distortions in different frame types across various coding setups. Specifically, considering regions of text and graphic textures in screen content, three types of prior maps, including screen content maps, coding configuration maps, and traditional filtering maps, are designed as auxiliary information in the model, promoting distortion pattern learning under different configurations. Two novel non-local models are proposed to enhance the model’s ability to capture global features in intra- and inter-frames while keeping low computational complexity. Finally, the VC-SCC is proposed for parallel implementation with the standard in-loop filter, and the optimal results are selected in each patch. Experimental results demonstrate significant performance improvements, with average BD-rate savings of 9.93%, 11.05%, and 10.73% for the all-intra(AI), low-delay(LD), and random-access(RA) configurations, respectively, outperforming other state-of-the-art approaches.
Mingxuan Li 0002, Wen Ji 0003
IEEE Trans. Circuits Syst. Video Technol.2
2023 Multi-stream Adaptive Offloading of Joint Compressed Video Streams, Feature Streams, and Semantic Streams in Edge Computing Systems
abstract
Edge computing (EC) is a promising paradigm for serving latency-sensitive video applications. However, massive compressed video transmission and analysis require considerable bandwidth and computing resources, posing enormous challenges for current multimedia frameworks. Novel multi-stream frameworks that incorporate feature streams are more practical. The reason is that feature streams containing compact video frame feature data have a lower bitrate and better serve machine vision tasks. Nevertheless, feature extraction by devices increases the latency and energy consumption of local computing. Therefore, how to offload suitable streams according to video task requirements and system resources is a challenging issue. This paper studies EC-based multi-stream adaptive offloading. We model the multi-stream offloading and computation problem to maximize system utility by jointly optimizing offloading decisions, computation resource allocation, and video frame sampling rates. Frame sampling rates, processing latency, and energy consumption are considered in system utility modeling. The formulated optimization problem is a mixed-integer programming (MIP) problem. We propose an efficient algorithm to address this MIP problem. The proposed algorithm relies on the Hungarian algorithm and improved greedy Markov approximation. The simulation results validate our proposed algorithm’s superior performance.
Dieli Hu 0001, Wen Ji 0003, Zhi Wang 0001
ICME2
2023 JAVP: Joint-Aware Video Processing with Edge-Cloud Collaboration for DNN Inference
abstract
Currently, massive video inference tasks are processed through edge-cloud collaboration. However, the diverse scenarios make it difficult to allocate the inference tasks efficiently, resulting in many wasted resources. In this paper, we propose a joint-aware video processing (JAVP) architecture for edge-cloud collaboration. First, we develop a multiscale complexity-aware model for predicting task complexity and determining its suitability for edge or cloud servers. The task is subsequently efficiently scheduled to the appropriate servers by integrating complexity with an adaptive resource-aware optimization algorithm. For input tasks, JAVP can dynamically and intelligently select the most appropriate server. The evaluation results on public datasets show that JAVP can improve the through-put by more than 70% compared to traditional cloud-only solutions while meeting accuracy requirements. And JAVP can improve the accuracy by 3%-5% and reduce delay and energy consumption by 16%-50% compared to state-of-the-art edge-cloud solutions.
Zheming Yang, Wen Ji 0003, Qi Guo 0009, Zhi Wang 0001
ACM Multimedia2
2023 JVAP: A Joint Video Acceleration Processing Architecture for Online Edge Systems
abstract
In visual intelligent scenarios, large amounts of real-time video data are generated at the end. During the optimization process for different video tasks, frequent data copying between devices and hosts can be limited by data bandwidth, resulting in high system latency. We investigate computing bottlenecks in online video processing to reduce processing latency and improve efficiency. In this paper, we propose a joint video acceleration processing (JVAP) architecture for online edge systems. First, video-compressed streams are transmitted to the GPU for decoding and conversion of data content. Second, we design data pre-processing and post-processing modules to achieve specific functional operators and separately complete operator combinations and stitching. Different computing tasks can reuse the implemented operator library. Third, we modify the data interface of the inference task model to maintain the consistent flow of data in the GPU. We conduct experiments using videos of different qualities and model frameworks of varying scales. The results indicate that the proposed method enhances the average processing efficiency over 11% with respect to existing representative acceleration frameworks and extends the potential application of online intelligent inference algorithms.
Ningzhou Li, Zheming Yang, Mingxuan Li 0002, Wen Ji 0003
SMC4
2023 Lightweight Multiattention Recursive Residual CNN-Based In-Loop Filter Driven by Neuron Diversity
abstract
Many convolutional neural network (CNN)-based in-loop filters have been proposed to improve coding performance. However, considering the single perception scale, high parameter complexity, and the need to train multiple models for various quantization parameters (QPs), the performance and practicability of most existing methods are limited. Inspired by neuron diversity, this paper proposes a lightweight multiattention recursive residual CNN-based in-loop filter that can handle encoded frames with various QP values, frame types (FTs), and temporal layers (TLs) via a single model. First, multiscale features are learned in the neural network and fused with the proposed multidensity block (MDB) and multiscale fusion attention group (MFAG). Second, a recursive structure is adopted to improve the model depth while saving many parameters. The proposed auxiliary parameter fusion attention (APFA) and long-short-term skip connection (LSTSC) models integrate QPs, FTs and TLs into the model while accelerating training. Finally, we propose implementing LMA-RRCNN in parallel with the standard in-loop filter and select the optimal enhanced result in each patch. The experimental results on standard test sequences show that the proposed method achieves on average 13.70% and 11.87% BD rate savings under all-intra and random-access configurations, respectively, outperforming other state-of-the-art approaches.
Mingxuan Li 0002, Wen Ji 0003
IEEE Trans. Circuits Syst. Video Technol.2
2022 Design of Personalization Warehouse Management Platform Based on SaaS Model
abstract
SaaS is widely used in the fields of operation management, business process outsourcing, data analysis, and information security. However, with the improvement of the degree of information, the generalized SaaS platform is unable to satisfy the requirements of enterprise personalization. SaaS products face great challenges in personalized customization technology due to the application of multi-tenant architecture. The challenges include multi-tenant customization in SaaS model, data isolation during customization, and mapping of multi-tenant virtual warehouse locations to actual locations. Therefore, we decompose personalization technology into metadata driver, cloud data placement and mapping mechanism for research. We decompose the personalized technologies into metadata drive, cloud data placement, and mapping mechanism. In order to solve the problem that traditional personalization customization is unable to be applied in the SaaS field, we propose a multi-tenant personalized warehousing mode architecture, designed a personalization warehouse management platform based on SaaS model, and realized SaaS-based data isolation, interface customization, rapid positioning, and on-demand customization services. Experimental results show that the proposed multi-tenant personalized warehouse model can achieve data isolation and customization, and reflect the advantages of highly automated warehousing, shared storage resources, and on-demand customization in terms of warehouse management, data security, and user experience.
Qi Guo 0009, Hongbo Sun 0004, Wen Ji 0003
CSCWD3
2022 Mixed-Precision Neural Network Quantization via Learned Layer-Wise Importance
Kai Ouyang, Zhi Wang 0001, Yifei Zhu 0001, Wen Ji 0003, Yaowei Wang 0001, Wenwu Zhu 0001
ECCV (11)5
2022 Astute Video Transmission for Geographically Dispersed Devices in Visual IoT Systems
abstract
Visual IoT (VIoT) is a promising IoT paradigm that visualizes sensing data from massive numbers of dispersed devices. A key objective in VIoT is to efficiently manage the devices to perform complex task-related visual data processing. Prior multimedia IoT systems have mainly focused on the delivery of captured video to remote servers, without considering the video tasks’ characteristics and the devices’ heterogeneous capabilities. In this work, we propose an astute video transmission framework for such a VIoT system composed of heterogeneous visual devices. First, we formulate the problem of joint video task allocation and heterogeneous device management by constructing a device hypergraph (DH) structure, which enables devices with different capabilities to perform complex video tasks cooperatively. Second, we model the video transmission within a VIoT system by applying fractal theory considering the NP-hardness of the optimization. In particular, we construct a comprehensive fractal submodular optimization framework through a DH and explore the inner submodular property to effectively leverage both video-task complexity and device heterogeneity. Third, we consider the geographically dispersed characteristic of massive numbers of VIoT devices and propose a multi-hop dispersed transmission mechanism for achieving globally cooperative optimality. The proposed architecture has been evaluated under diverse parameter settings. Numerical results are provided to validate the proposed algorithm in terms of delay, computational efficiency, and bandwidth utilization. Simulation results confirm the effectiveness and superiority of the proposed method.
Wen Ji 0003, Ling-Yu Duan, Xi Huang 0002, Yueting Chai
IEEE Trans. Mob. Comput.1
2021 An Intelligent End-Edge-Cloud Architecture for Visual IoT-Assisted Healthcare Systems
abstract
Recently, the Internet of Things (IoT) has played a powerful role in healthcare. However, the rapid growth of healthcare devices has produced many heterogeneous data and most of them are visual. It brings great difficulties to the calculation, cache, and transmission of data. The geographical dispersion and the dynamicity of nodes also challenge the development of healthcare IoT (HIoT). In this article, we propose an intelligent end–edge–cloud architecture for visual IoT-assisted healthcare systems (intelligent V-HIoT) to improve the end-to-end performance of next-generation smart healthcare. First, we systemically analyze the characteristics of human–machine–things in end side from the perspective of data processing, then define the end intelligence, solving the problem of intelligence measurement of heterogeneous devices. Second, we propose an efficiency intelligence measurement model in the edge side and cloud side, which provides a theoretical basis for the dynamic management of edge nodes. Third, we present an end–edge–cloud framework that optimizes the efficiency of data processing and node deployment. The intelligence level of HIoT is maximized as well as intelligent management of nodes is implemented. To verify the effectiveness, we perform the experiments for different approaches. The simulation results demonstrate that the intelligent V-HIoT significantly outperforms existing approaches because the proposed method can achieve maximum intelligence level of both in many heterogeneous devices and an emergency medical situation.
Zheming Yang, Wen Ji 0003
IEEE Internet Things J.3
2021 Risk Optimization for Revenue-Driven Wireless Video Broadcasting Systems: A Copula-Based Framework
abstract
The revenue of wireless service providers (WSPs) relies on their ability to efficiently satisfy the variable demands from end users (EUs). However, emerging video services create new risks owing to the diverse content requirements of heterogeneous EUs operating in uncertain markets. It is challenging as the risks created by demands and prices are highly uncertain. In this work, a risk problem of high revenues is studied in which multiple video content items with different prices are broadcast to wireless EUs. The video content prices differ in their value functions in terms of popularity, ratings, and types. The objective is to maximize the revenue of WSPs under a certain Value-at-Risk (VaR) by adjusting the content prices and allocation of bandwidth resources. Furthermore, a VaR-based optimization framework for wireless video broadcasting systems is presented. First, the content characteristics are analyzed, and a copula model is then used to build the content value structure. The copula of a multivariate distribution corresponds to the description of the price-dependent structure. Second, a risk analysis for the effects of price fluctuations on revenues caused by uncertainty is conducted. A VaR model is associated with changes in the prices and allocated rates. Copulas are used to derive a bound on the VaR for functions of dependent risks. Subsequently, detailed representations are provided to identify the distributional bounds for revenue functions of dependent risks. Lastly, a VaR-based framework that optimizes pricing and bandwidth provision is presented. For the solution, the risk regions of WSPs are modeled as polymatroidal structures to minimize the risk caused by different service demands and variable market prices. Experiments on different price markets demonstrated that the proposed method is effective, thereby verifying the feasibility of the proposed method.
Wen Ji 0003, H. Vincent Poor
IEEE Trans. Multim.1
2020 Reinforcement Learning-Based Mobile Offloading for Edge Computing Against Jamming and Interference
abstract
Mobile edge computing systems help improve the performance of computational-intensive applications on mobile devices and have to resist jamming attacks and heavy interference. In this paper, we present a reinforcement learning based mobile offloading scheme for edge computing against jamming attacks and interference, which uses safe reinforcement learning to avoid choosing the risky offloading policy that fails to meet the computational latency requirements of the tasks. This scheme enables the mobile device to choose the edge device, the transmit power and the offloading rate to improve its utility including the sharing gain, the computational latency, the energy consumption and the signal-to-interference-plus-noise ratio of the offloading signals without knowing the task generation model, the edge computing model, and the jamming/interference model. We also design a deep reinforcement learning based mobile offloading for edge computing that uses an actor network to choose the offloading policy and a critic network to update the actor network weights to improve the computational performance. We discuss the computational complexity and provide the performance bound that consists of the computational latency and the energy consumption based on the Nash equilibrium of the mobile offloading game. Simulation results show that this scheme can reduce the computational latency and save energy consumption.
Liang Xiao 0003, Xiaozhen Lu, Tangwei Xu, Xiaoyue Wan, Wen Ji 0003, Yanyong Zhang
IEEE Trans. Commun.5
2020 Profit Maximization for Sponsored Data in Wireless Video Transmission Systems
abstract
Recently, sponsored data started to gain wide-spread use in wireless networks. Existing approaches for regular data services when applying to video transmission in sponsored data faces tremendous challenges. In two-sided markets, internal and external competitions among service providers (SPs), content provider (CPs), and end-users (EUs) become more complex, because larger video-bandwidth significantly increases the cost. When the three parties benefit from sponsored data, pricing and transmission design become much more complicated and difficult in parameterized systems by profit decomposing models, because maximizing the total profits among the three parties is a NP-hard problem. In this paper, we address the problem of maximizing the total profits for sponsored video data. We propose a submodular-based optimization framework for wireless video transmission systems. First, we formulate the sponsored data models in wireless video transmission respective to SP, CPs and EUs sides. The proposed profit models are used in wireless broadcasting networks with a specific attention to sponsored video data. Second, we construct a layered video transmission model in two-sided market through incorporating submodularity. We approximate the NP-hard layered-video broadcasting problem by applying submodular theory. Third, we present a profit maximization framework that optimizes the number of sponsored users, pricing and bandwidth provision. The profit of SP is maximized as well as the profit of CPs is improved and the cost-effective QoE of EUs is guaranteed. To verify the efficiency, we perform the experiments for different pricing schemes. The simulation results demonstrate that the proposed method significantly outperforms existing approaches because the proposed method can achieve maximum profits of both SP and CPs in a wide range of transmission rates.
Wen Ji 0003, Wenwu Zhu 0001
IEEE Trans. Mob. Comput.1
2019 Profit optimization in service-oriented data market: A Stackelberg game approach
Bo Shen 0006, Yulong Shen 0001, Wen Ji 0003
Future Gener. Comput. Syst.3
2019 Guest Editorial Multimedia Economics for Future Networks: Theory, Methods, and Applications
abstract
With the growing integration of telecommunication networks, Internet of Things (IoT), and 5G networks, there is a tremendous demand for multimedia services over heterogeneous networks. According to recent survey reports, mobile video traffic accounted for 60 percent of total mobile data traffic in 2016, and it will reach up to 78 percent by the end of 2021. Users’ daily lives are inundated with multimedia services, such as online video streaming (e.g., YouTube and Netflix), social networks (e.g., Facebook, Instagram, and Twitter), IoT and machine generated video (e.g, surveillance cameras), and multimedia service providers (e.g., Over-the-Top (OTT) services). Multimedia data is thus becoming the dominant traffic in the near future for both wired and wireless networks.
Wen Ji 0003, Zhu Li 0001, H. Vincent Poor, Christian Timmerer, Wenwu Zhu 0001
IEEE J. Sel. Areas Commun.1
2018 AIEM: AI-enabled affective experience management
Yongfeng Qian, Yiming Miao, Wen Ji 0003, Renchao Jin, Enmin Song
Future Gener. Comput. Syst.4
2018 Deadline-aware rate allocation for IoT services in data center network
Bo Shen 0006, Naveen K. Chilamkurti, Xingshe Zhou 0001, Wen Ji 0003
J. Parallel Distributed Comput.6
2017 Green Video Transmission in the Mobile Cloud Networks
abstract
Video transmission is an indispensable component of most applications related to the mobile cloud networks (MCNs). However, because of the complexity of the communication environment and the limitation of resources, attempts to develop an effective solution for video transmission in the MCN face certain difficulties. In this paper, we propose a novel green video transmission (GVT) algorithm that uses video clustering and channel assignment to assist in video transmission. A video clustering model is designed based on game theory to classify the different video parts stored in mobile devices. Using the results of video clustering, the GVT algorithm provides the function of channel assignment, and its assignment process depends on the content of the video to improve channel utilization in the MCN. Extensive simulations are carried out to evaluate the GVT with several performance criteria. Our analysis and simulations show that the proposed GTV demonstrates a superior video transmission performance compared with the existing methods.
Jeungeun Song 0001, Jiming Luo, Wen Ji 0003, M. Shamim Hossain, Ahmed Ghoneim
IEEE Trans. Circuits Syst. Video Technol.4
2017 Guest Editorial for ACM TECS Special Issue on Effective Divide-and-Conquer, Incremental, or Distributed Mechanisms of Embedded Designs for Extremely Big Data in Large-Scale Devices
abstract
No abstract available.
Bo-Wei Chen, Wen Ji 0003, Zhu Li 0001
ACM Trans. Embed. Comput. Syst.2
2017 Feedback-Free Binning Design for Mobile Wyner-Ziv Video Coding: An Operational Duality between Source Distortion and Channel Capacity
abstract
Most mobile video applications require the encoder to have low complexity. Wyner-Ziv (WZ) video coding removes complex motion estimation from the encoder, and provides error resilience from the embedded channel coding module. WZ video coding is regarded as a promising encoder for wireless video systems. Most WZ video coding based on channel codes is a practical implementation of the binning schemes. In this work, we present a novel two-tier binning scheme that consists of the inner and outer structure, for improving rate-distortion performance. First, we develop a Raptor coding with side information to construct the inner binning structure, which provides a lower rate. Second, for the outer binning, we model the WZ video coding architecture as a multiaccess channel, so that we can exploit the property of channel capacity. Third, we exploit the duality property of WZ video coding. Based on such a property, both the primal and dual solutions are subsequently provided in this study. For the primal problem of distortion minimization, we develop dynamic programming to find the optimal binning policy, whereas for the dual problem of capacity maximization, we devise a near sum-capacity binning algorithm. The objective is to lower the coding rate with lower complexity. Experimental results showed that when compared with the state-of-the-art coding, the decoding performance and the quality of our proposed method were respectively enhanced. Besides, we observed that the decoding distortion was reduced through the proposed outer binning, while the proposed inner binning based on Raptor coding by jointly considering side information (SI) lead to a low bitrate when a target decoding quality was specified. Such findings have substantiated the effectiveness of our method.
Wen Ji 0003, Xiangyang Ji, Yiqiang Chen 0001
IEEE Trans. Mob. Comput.1
2016 Divide-and-conquer signal processing, feature extraction, and machine learning for big data
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho
Neurocomputing2
2016 Evaluate mobile video quality in hybrid spatial and temporal domain
Wen Ji 0003, Seungmin Rho, Bo-Wei Chen, Yiqiang Chen 0001
Multim. Tools Appl.2
2016 A semi-supervised privacy-preserving clustering algorithm for healthcare
Meiyu Huang, Yiqiang Chen 0001, Bo-Wei Chen, Junfa Liu, Seungmin Rho, Wen Ji 0003
Peer-to-Peer Netw. Appl.6
2016 Cross-Layer Opportunistic Scheduling for Device-to-Device Video Multicast Services
abstract
In this article, we address the problem of how to make the wireless device-to-device (D2D) video multicast systems have better quality provision with consideration of internet-of-things (IoT) applications. We propose an opportunistic transmission and fair resource allocation framework, including joint application-layer and physical-layer transmission and optimization. First, we use a parallel subchannels structure by concatenating the Fountain codes and diversity-embedded space-time block codes to provide reliable and flexible transmission in heterogeneous circumstances. Second, we exploit the quality of heterogeneous user experience (quality of experience) metric under D2D video multicast systems, with consideration of various channel states, device capability, video content urgency, and the number of demanding users. Third, we formulate reliable multiple video streams broadcasting to heterogeneous devices as an aggregate maximum utility achieving problem, and we use opportunistic scheduling to select suitable users in each transmission interval to improve the broadcasting utility. Fourth, we use the utility fair scheme to guide rate allocation among multicontent video multicast. Extensive performance comparison and analysis are presented to demonstrate efficiency of the proposed solution.
Wen Ji 0003, Bo-Wei Chen, Haiyong Luo, Mucheol Kim, Yiqiang Chen 0001
ACM Trans. Embed. Comput. Syst.1
2016 Large-scale image colorization based on divide-and-conquer support vector machines
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung
J. Supercomput.3
2016 Support vector analysis of large-scale data based on kernels with iteratively increasing order
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung
J. Supercomput.3
2016 Erratum to: Large-scale image colorization based on divide-and-conquer support vector machines
Bo-Wei Chen, Wen Ji 0003, Seungmin Rho, Sun-Yuan Kung
J. Supercomput.3
2016 Optimal filter based on scale-invariance generation of natural images
Feng Jiang 0001, Bo-Wei Chen, Seungmin Rho, Wen Ji 0003, Liqiang Pan, Debin Zhao
J. Supercomput.4
2016 Profit Maximization through Online Advertising Scheduling for a Wireless Video Broadcast Network
abstract
In this paper, we address the problem of how to make the wireless service provider (WSP) earn profits in a wireless video broadcast network with consideration of advertisement insertion. At the beginning, this study examines the profit components by analyzing traffic provision and advertisement insertion. This study considers using two components for profit maximization-one is the function for allocating video rates, and the other is the function for inserting advertisement duration. The maximum achievable profit depends on joint optimization of optimal video-rate vectors and advertisement-duration vectors, which are usually computationally intensive. To resolve such a complexity problem, this work also proposes an effective algorithm for joint optimization. First, the overall profit is formulated as the solution of four local optimization problems through horizontal and vertical decomposition. Second, a theoretic polymatroidal framework is introduced in our work for optimization as this framework is proved effective in profit maximization of multiuser systems. Third, this study shows that the overall profit can be maximized by finding the optimal profit points on the boundary of the rate and duration regions. As a result, the optimum points and the total profit can be obtained through a hierarchical greedy algorithm. Experimental results demonstrate that the proposed method is capable of making maximum profits for WSPs in a wide range of broadcasting rates.
Wen Ji 0003, Yingying Chen 0001, Min Chen 0003, Bo-Wei Chen, Yiqiang Chen 0001, Sun-Yuan Kung
IEEE Trans. Mob. Comput.1
2016 Fixed-Point Computing Element Design for Transcendental Functions and Primary Operations in Speech Processing
abstract
This brief presents a fixed-point architecture based on a reconfigurable scheme for integrating several commonly used mathematical operations of speech signal processing. The proposed design can perform two transcendental mathematical operations called logarithm and powering, and three commonly used computations with similar operations named polynomial calculation, filtering, and windowing. By analyzing the adopted algorithms of the above five operations, a simplified computing unit is designed. This unit can combine six types of operations by reconfiguring the data paths, and the same multiply-add architecture can be reused for reducing the redundant usage of logic gates. The experimental results reveal that the proposed design can work at a 200-MHz clock rate, and its gate count only has 11.9k. Compared with the results of the floating-point function, the median errors of the proposed design for computing the powering and logarithmic functions are 0.57% and 0.11%, respectively. Such results indicate that this simple architecture can be effectively used in most speech processing applications.
Chung-Hsien Chang, Shi-Huang Chen, Bo-Wei Chen, Wen Ji 0003, K. Bharanitharan, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.4
2016 A New Binary-Halved Clustering Method and ERT Processor for ASSR System
abstract
This paper presents an automatic speech–speaker recognition (ASSR) system implemented in a chip which includes a built-in extraction, recognition, and training (ERT) core. For VLSI design (here, ASSR system), the hardware cost and time complexity are always the important issues which are improved in this proposed design in two levels: 1) algorithmic and 2) architecture. At the algorithm level, a newly binary-halved clustering (BHC) is proposed to achieve low time complexity and low memory requirement. In addition, at the architecture level, a new ERT core is proposed and implemented based on data dependence and reuse mechanism to reduce the time and hardware cost as well. Finally, the chip implementation is synthesized, placed, and routed using TSMC 90-nm technology library. To verify the performance of the proposed BHC method, a case study is performed based on nine speakers. Moreover, the validation of the ASSR system is examined in two parts: 1) speech recognition and 2) speaker recognition. The results show that the proposed system can achieve 93.38% and 87.56% of recognition rates during speech and speaker recognition, respectively. Furthermore, the proposed ASSR chip includes 396k gate counts, and consumes power in 8.74 mW. Such results demonstrate that the performance of the proposed ASSR system is superior to the conventional systems.
Chih-Hung Chou, Ta-Wen Kuan, Shovan Barma, Bo-Wei Chen, Wen Ji 0003, Chih-Hsiang Peng, Jhing-Fa Wang
IEEE Trans. Very Large Scale Integr. Syst.5
2015 Accurate and Robust Moving-Object Segmentation for Telepresence Systems
abstract
Moving-object segmentation is the key issue of Telepresence systems. With monocular camera--based segmentation methods, desirable segmentation results are hard to obtain in challenging scenes with ambiguous color, illumination changes, and shadows. Approaches based on depth sensors often cause holes inside the object and missegmentations on the object boundary due to inaccurate and unstable estimation of depth data. This work proposes an adaptive multi-cue decision fusion method based on Kinect (which integrates a depth sensor with an RGB camera). First, the algorithm obtains an initial foreground mask based on the depth cue. Second, the algorithm introduces a postprocessing framework to refine the segmentation results, which consists of two main steps: (1) automatically adjusting the weight of two weak decisions to identify foreground holes based on the color and contrast cue separately; and (2) refining the object boundary by integrating the motion probability weighted temporal prior, color likelihood, and smoothness constraint. The extensive experiments we conducted demonstrate that our method can segment moving objects accurately and robustly in various situations in real time.
Meiyu Huang, Yiqiang Chen 0001, Wen Ji 0003, Chunyan Miao
ACM Trans. Intell. Syst. Technol.3
2015 Game theoretic analysis for large-scale networks and traffic data
Daniel Bo-Wei Chen, Wen Ji 0003, Yong Liu 0013
J. Supercomput.2
2015 Profit Improvement in Wireless Video Broadcasting System: A Marginal Principle Approach
abstract
In this paper, we address the problem of how to make the wireless service provider have better profits with consideration of user experience provision in wireless video broadcasting systems. We propose a marginal-based pricing and a resource-allocation framework to achieve better resource utilization and profit improvement. The marginal principle includes 1) marginal user principle, in which a pricing mechanism is established on the basis of marginal users, such that the WSP can seek its own maximum profit of each content with a QoE guarantee; 2) marginal profit principle, in which a WSP can earn the maximum profit through multicontent-service provision by regulating rate allocation in limited available bandwidth. Furthermore, we present a two-tier framework consisting of the inner and outer loops. The inner loop focuses on pricing-based service provision based on the notion of marginal user principle. The outer loop concentrates on allocating bandwidth among multiple video contents according to marginal profit principle. For the solution, we model the profit regions of WSPs and end-users as the polymatroid structures and model the corresponding allocated rate regions as the contra-polymatroid structures. Through exploiting the properties of polymatroid and contra-polymatroid structures, the broadcasting profit problem is solved by finding the optimal rate vector on the sum-rate facet which satisfies the maximal achievable profit. Extensive performance comparison and analysis are presented to demonstrate efficiency of the proposed solution.
Wen Ji 0003, Bo-Wei Chen, Yiqiang Chen 0001, Sun-Yuan Kung
IEEE Trans. Mob. Comput.1
2015 Profit Optimization for Wireless Video Broadcasting Systems Based on Polymatroidal Analysis
abstract
This study addresses the problem of profit maximization between wireless service providers (WSPs) and content providers (CPs) in wireless broadcasting systems , while simultaneously providing high quality of experience for end-users (EUs). We first study the profit model in wireless broadcasting networks with a particular attention to the heterogeneous requirements of EUs, e.g., different display sizes and variable channel conditions. Then, we propose a profit formulation that describes the requirements of wireless service providers and content providers, as well as the satisfaction of EUs that essentially depends on video quality and service charges. We propose a new polymatroidal theoretic framework for maximizing the resulting three-side achievable profit through proper bandwidth allocation. Our framework exploits two particular structures, namely the underlying polymatroidal structure of the profit region and the contra- polymatroidal structure of the rate region. We then propose a profit maximization solution by finding a rate allocation vector on the sum-rate facet that satisfies the maximal achievable profit among the WSP, CPs, and EUs. Experiments on different broadcasting scenarios demonstrate the effectiveness of the proposed method. The WSP is capable of generating more revenues by applying the proposed approach to their marketing strategies while satisfying the demands from CPs and EUs.
Wen Ji 0003, Pascal Frossard, Bo-Wei Chen, Yiqiang Chen 0001
IEEE Trans. Multim.1
2014 EXIT-Based Side Information Refinement in Wyner-Ziv Video Coding
abstract
The accuracy of the side information (SI) is critical in the performance of distributed video coding algorithms. The SI is typically built at a decoder based on the reconstructed data and on channel coding parity bits transmitted by the encoder. The optimal encoding rate is generally difficult to compute precisely due to the dynamics of video content with varying correlation. Effective methods for the refinement of imprecise SI are therefore important for improved decoding quality. In this paper, we propose to exploit the intrinsic property of channel coding algorithms in Wyner-Ziv video coding. The SI is refined via both the information-plane and the parity-plane bits, which rapidly increases the accuracy of refined SI. We use extrinsic information transfer chart analysis in order to estimate the variations of the mutual information in the iterative decoding. In particular, we characterize mutual information variations for punctured regular and irregular rate-compatible low-density parity-check codes. Tracking the mutual information changes permits to decrease the coding rate of the information and parity bitstreams, while preserving the decoding quality. Simulation results confirm that our method improves on the decoding quality of recent distributed video coding algorithms, especially for high-motion sequences or at high-coding rate regimes.
Wen Ji 0003, Pascal Frossard, Yiqiang Chen 0001
IEEE Trans. Circuits Syst. Video Technol.1
2013 A Binning Design for Wyner-Ziv Video Coding
abstract
In this work, we proposes a two-tier binning scheme. First, we develop a Fountain coding with side information to construct the inner binning structure. Second, for the the outer binning, we model the WZ video coding architecture as a multi-access channel and exploit the duality property between the WZ coding and channel coding techniques. Third, we provide both the primal and dual solutions. For the primal distortion minimization problem, we use dynamic programming approach to find the optimal binning policy, and for the dual capacity maximization problem, we give a near sum-capacity binning algorithm. The objective is to lower the coding rate under same video reconstruction quality.
Wen Ji 0003, Yiqiang Chen 0001
DCC1
2013 Modeling the hybrid temporal and spatial resolutions effect for web video quality evaluation
abstract
Understanding and modelling the users' perceptual quality of a video are the key steps towards improving multimedia service provision. In this paper, we take an analytical approach to study the joint impact of spatial and temporal resolution on the perceptual video quality. First, we map the subjective quality data as an evaluation function in terms of the spatial resolution and frame rate, which reflect the major features of web video quality. Second, we use ε-Support Vector Regression (ε-SVR) with Radial Basis Function (RBF) kernel to give a more accurate hybrid temporal-spatial quality evaluation model so as to predict the perceptual video quality. The comprehensive subjective quality experiments were carried out to construct and validate this model. The experimental results demonstrated the effectiveness of the proposed model. Besides, the quality model can be easily deployed in a practical multimedia delivery system.
Wen Ji 0003, Min Chen 0003, Yiqiang Chen 0001
GLOBECOM2
2013 TiPS: A lightweight Tele-immersive photograph system
abstract
This paper presents a lightweight Tele-immersive photograph system named TiPS. TiPS allows remote users to take a photo as a souvenir in a virtual space with natural social behavior, through adaptively merging two participants of remote video interaction into the same shared background. We address two key technologies in this paper, first, we propose an adaptive multi-cue decision fusion algorithm for accurate and robust foreground segmentation of live video. Second, we present a user behavioral intention driven video composition algorithm, aiming at obtaining life-like photographs in real time. Experimental results show TiPS achieves high performance in both foreground segmentation and video composition, which well satisfies remote users' requirement of taking photos together.
Meiyu Huang, Yiqiang Chen 0001, Wen Ji 0003
ICME3
2012 EXIT Chart-Based Side Information Refinement for Wyner-Ziv Video Coding
abstract
This paper focuses on side information (SI) refinement in Wyner-Ziv video coding and proposes to exploit the intrinsic property of channel coding for improving the joint decoding performance. In this paper, we propose to use syndrome and information bits from the encoder to help the decoder in refining the SI. We use extrinsic information transfer (EXIT) chart analysis to deduce the mutual information variation in LDPC iterative decoding during the SI refinement process. The objective is to obtain the same decoding quality under lower coding rates. Simulation results demonstrate the effectiveness of the proposed solution.
Wen Ji 0003, Pascal Frossard, Yiqiang Chen 0001
DCC1
2012 QoE-based opportunistic transmission for video broadcasting in heterogeneous circumstance
abstract
This paper presents an opportunistic transmission scheme for layered video broadcasting to multiple heterogeneous devices. In contrast to conventional wireless video broadcasting system, the main ideas proposed here include: (i) exploit the quality of heterogeneous user experience (QoE) metric under wireless broadcasting scenario, with consideration of various channel state, device capability, video content urgency and the number of demanding users. (ii) formulate reliable multiple video streams broadcasting to heterogeneous devices as an aggregate maximum utility achieving problem. (iii) use opportunistic scheduling to select suitable users in each transmission interval so as to improve the broadcasting utility. (iv) use parallel-pipe structure to transmit the layered video with Fountain coding protection, which provide reliable and low-latency transmission in heterogeneous circumstance. Numerical experiments demonstrate that the proposed scheme outperforms conventional methods.
Wen Ji 0003, Zhu Li 0001, Yiqiang Chen 0001
ACM Multimedia1
2012 Power-efficient video encoding on resource-limited systems: A game-theoretic approach
Wen Ji 0003, Jiangchuan Liu, Min Chen 0003, Yiqiang Chen 0001
Future Gener. Comput. Syst.1
2012 Joint Source-Channel Coding and Optimization for Layered Video Broadcasting to Heterogeneous Devices
abstract
Heterogeneous quality-of-service (QoS) video broadcast over wireless network is a challenging problem, where the demand for better video quality needs to be reconciled with different display size, variable channel condition requirements. In this paper, we present a framework for broadcasting scalable video to heterogeneous QoS mobile users with diverse display devices and different channel conditions. The framework includes joint video source-channel coding and optimization. First, we model the problem of broadcasting a layered video to heterogeneous devices as an aggregate utility achieving problem. Second, based on scalable video coding, we introduce the temporal-spatial content distortion metric to build adaptive layer structure, so as to serve mobile users with heterogeneous QoS requirements. Third, joint Fountain coding protection is introduced so as to provide flexible and reliable video stream. Finally, we use dynamic programming approach to obtain optimal layer broadcasting policy, so as to achieve maximum broadcasting utility. The objective is to achieve maximum overall receiving quality of the heterogeneous QoS receivers. Experimental results demonstrate the effectiveness of the solution.
Wen Ji 0003, Zhu Li 0001, Yiqiang Chen 0001
IEEE Trans. Multim.1
2011 Heterogeneous QoS Video Broadcasting with Optimal Joint Layered Video and Digital Fountain Coding
abstract
Heterogeneous QoS video broadcast over wireless network is a challenging problem, where the demand for better video quality needs to be reconciled with different display size, channel condition and QoS requirements. In this paper, we present a framework for broadcasting scalable video to heterogeneous mobile users with diverse display devices and different channel conditions, which includes joint spatial-temporal layered video and Fountain coding optimization. First, we develop an elastic rate video adaptation method so as to serve mobile users with heterogeneous QoS requirements, which includes a hybrid temporal-spatial quality metric. Joint Fountain coding protection is introduced so as to provide adaptive and reliable video streams. Second, we use dynamic programming approach to obtain optimal video layer structure, so as to achieve maximum broadcasting utility. The objective is to achieve maximum overall receiving quality of the heterogeneous QoS users. Experimental results demonstrate the effectiveness of the solution.
Wen Ji 0003, Zhu Li 0001
ICC1
2011 Content-aware utility-fair video streaming in wireless broadcasting networks
abstract
In wireless multi-content video broadcasting system, a critical problem is fairness among contents with respect to heterogeneous characteristics. To address this problem, we propose an approach of content-aware utility-fair streaming control scheme, which aims at heterogeneous QoS video provision and ensures max-min utility-fair sharing among video streams. First, we introduce a hybrid temporal-spatial quality metric to model content-aware utility so as to serve mobile users with heterogeneous QoS requirements. Second, we use max-min utility-fair scheme to guide rate allocation and video content generation among multi-content video broadcasting. Simulation results demonstrate the proposed approach can achieve utility fair among multiple video contents, and provide better quality of service to all broadcasting users especially when available bandwidth is limited.
Wen Ji 0003, Zhu Li 0001, Yiqiang Chen 0001
ICIP1
2011 Joint scalable video and digital fountain coding for heterogeneous QoS video broadcasting
abstract
Serving broadcasting video to heterogeneous mobile devices with different display size and channel conditions has many valuable applications and technical challenges. In this work, we present a cross-layer, multi-time scale approach solution to this problem. At the outer loop of control, a physical layer diversity-embedded space-time coding scheme is employed to maximize the total effective data rates among mobiles. New spatial-temporal adaptation scheme and metrics are developed to enhance the elasticity and robustness of the broadcasting video, then a joint layered video and digital fountain coding optimization is performed at the inner loop to deliver the best possible QoE among the heterogeneous receivers. Simulation results demonstrated the benefits and effectiveness of the proposed solutions.
Wen Ji 0003, Zhu Li 0001, Yiqiang Chen 0001
ICME1
2011 A perceptual macroblock layer power control for energy scalable video encoder based on just noticeable distortion principle
Wen Ji 0003, Min Chen 0003, Xiaohu Ge, Yiqiang Chen 0001
J. Netw. Comput. Appl.1
2011 Design and integration of the OpenCore-based mobile TV framework for DVB-H/T wireless network
Chin-Feng Lai, Yueh-Min Huang, Jiann-Liang Chen, Wen Ji 0003, Min Chen 0003
Multim. Syst.4
2010 Joint layered video and digital fountain coding for multi-channel video broadcasting
abstract
In this paper, we consider a scenario where multiple video content channels are broadcasted to a set of heterogeneous mobile users with diverse display devices and different channel conditions. The objective is to design a joint coding and rate allocation algorithm which achieves maximum overall receiving quality of the heterogeneous users, measured by broadcasting utility. We use hierarchy optimization to solve this problem, and decompose the problem into a two-tier solution: the inner loop aims at single content broadcasting, solved with a joint coding algorithm; while the outer loop focus on multiple contents broadcasting, using dynamic programming approach to find the optimal rate allocation policy. Numerical experiments demonstrate the effectiveness of the solution.
Wen Ji 0003, Zhu Li 0001
ACM Multimedia1
2008 A complexity scalable decoder in an AVS video codec
abstract
The control of energy consumption in mobile devices is very important and is being paid more and more attentions to. An effective way to reduce the energy consumption of video decoding is to reduce the computational complexity of the video decoder. This paper aims at the tradeoff between complexity and video quality, and proposes a complexity scalable AVS video decoder. AVS is the recent video coding standard developed by the Audio and Video Coding Standard Workgroup of China, which has similar performance with H.264/AVC but a more succinct technical plan. With little information of the characteristics of video sequences added to the bitstream by the encoder, the decoder provides complexity scalable output. Given a percentage K, the decoder can accurately reduce K% computational complexity with slight video quality degradation. Our experimental studies show that, for typical video, using the complexity scalable technology, the quality of video sequences remains acceptable.
Chen Lei, Yiqiang Chen 0001, Wen Ji 0003
MoMM3