VLDB 2026 Research / reviewers in the wild / expert
Jie Ren 0007
dblp:r/JieRen-7
· DBLP profile ↗
33ranked-venue papers
4as first author
26since 2021 · last 2027
0000-0003-3183-7228ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 14 · 4 first-author · 8 since 2021Systems, architecture and hardware · 7 · 7 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | An empirical study of neural network based graph representations for software defect prediction
Zhiqiang Li 0003, Yanwei Xiang, Jie Ren 0007, Hongyu Zhang 0002, Fumin Qi, Xiaoyuan Jing |
Sci. Comput. Program. | 3 |
| 2026 | Lifting Optimized Binaries to Canonical Compiler IR via Structure-Aware Retrieval and Iterative VerificationabstractLifting stripped and highly optimized binaries to the canonical compiler intermediate representation (IR) enables program analysis when source code is unavailable.However, compiler optimizations severely distort controlflow and data-flow structure, making existing rule-based and LLM-based decompilation approaches brittle.We present BRIDGE, a system that reliably lifts optimized binaries to analysis-friendly compiler IR.BRIDGE combines control-flow-aware retrieval-augmented generation with feedback-driven verification.It uses pseudo-probe instrumentation to align optimized binary fragments with normalized IR semantics, and then employs an iterative refinement loop guided by static analysis and runtime feedback to improve executability and semantic consistency.We evaluate BRIDGE on HumanEval-Decompile and MBPP, lifting x86-64 and ARM64 binaries to LLVM IR.BRIDGE outperforms seven baselines, achieving an average of over 30% higher re-executability than the strongest general-purpose LLM baseline.Void func (){ PROBE(1); If else branch …… PROBE(2); PROBE(3); for Loop … PROBE(4);} Xiaoao Zhu, Jie Ren 0007, Zhiqiang Li 0003, Jie Zheng 0005, Zhanyong Tang, Zheng Wang 0001 |
ACL (1) | 2 |
| 2026 | GRASP: Optimizing VLIW Instruction Scheduling via Graph Reinforcement LearningabstractVery Long Instruction Word (VLIW) processors can expose substantial instruction-level parallelism (ILP), yet compiler-generated schedules often fail to fully utilize available functional units, leaving performance on the table and forcing costly manual tuning. We present GRASP (Graph Reinforcement Assembly Scheduling Platform), an assembly-level post-pass optimizer that improves VLIW instruction scheduling and packing via graph reinforcement learning. GRASP represents each basic block as an Instruction Dependency Graph (IDG) that explicitly encodes dependence, latency, and machine resource constraints, and trains a GNN-based reinforcement learning agent to iteratively select legal scheduling decisions and construct high-throughput issue packets. By operating after compilation, GRASP requires no source-code changes and does not modify compiler internals, making it easy to integrate into existing toolchains. Across a suite of AI and HPC kernels, GRASP accelerates execution by up to 1.38 × (1.22 × on average), narrowing the gap between compiler output and expert-tuned VLIW code. Weiyuan Tong, Jianbin Fang, Wei Wang 0056, Jie Ren 0007, Zhanyong Tang |
ICS | 6 |
| 2025 | An empirical comparison of data transformation techniques for clustering-based unsupervised software defect predictionabstractAs a key pre-processing step of software defect prediction, data transformation techniques aim to eliminate dimensional inconsistencies among software metrics and improve model performance. However, the impact of data transformation techniques on the performance of clustering-based unsupervised defect prediction (UDP) has not been fully explored. To this end, our aim is to empirically investigate the effect of data transformation techniques on the performance of clustering-based UDP models. In this paper, we systematically evaluated 23 unsupervised clustering models under 4 data conditions: (1) raw untransformed data (original), (2) logarithmic transformation, (3) z-score standardization, and (4) min-max normalization. Our experimental design consists of two dimensions: comparative analysis of individual models’ performance before and after data transformation and evaluation of relative performance across different models under various transformations. Extensive experiments conducted on 22 software projects indicate that: (1) Data transformation techniques have significant potential to improve the performance of unsupervised clustering models. They achieve improvements of 6.9% -22.9% in AUC,10.7% -41.6% in MCC, and 3.1% -12.27% in F_measure@20%. However, these transformations lead to a degradation in IFA performance, which inevitably increases code inspection efforts but within tolerable operational thresholds. (2) The performance of unsupervised clustering models varies depending on the data transformation techniques used, with no single clustering model showing consistent superiority. Empirical findings indicate that data transformation techniques exhibit general effectiveness in clustering-based UDP models, and the choice of an unsupervised clustering model should align with the specific data transformation technique employed. Zhengxiang Chen, Zhiqiang Li 0003, Hongyu Zhang 0002, Jie Ren 0007, Feng Tian 0005 |
APSEC | 4 |
| 2025 | LLVMTuner: Predictive Compiler Optimization for LLVM IR Across Heterogeneous PlatformsabstractThe Low Level Virtual Machine Intermediate Representation (LLVM IR) is a key component of modern compilers, valued for its universality and cross-platform adaptability. However, identifying optimal optimization passes across platforms remains challenging, as current autotuning frameworks struggle on mobile devices due to resource constraints and network variability. This paper introduces a novel LLVM IR performance optimization framework, LL VMTUNER. At its core is a deep neural network-based predictive model specifically designed to forecast the execution time of input passes on various platforms by analyzing IR features. By integrating this predictive model with the advanced autotuning framework, we enable a rapid and precise search for optimal pass list, removing the need for real-time latency measurements on target platforms and reducing data transmission between devices and cloud-based autotuning frameworks. LLVMTuner significantly enhances the performance of LLVM IR, providing a robust solution for efficient compiler optimizations across a spectrum of computing environments, from high-performance laptops to resource-constrained mobile devices. We evaluate LL VMTuNER on three heterogeneous platforms using over 600 LLVM IR benchmarks. The results show that LLVMTuner achieves an average speedup of 2.41x over the -03 optimization configuration. Additionally, LL VMTUNER reduces search overhead by 83.68% compared to OpenTuner. Xiaoao Zhu, Jie Ren 0007, Zhiqiang Li 0003, Feng Tian 0005, Jie Zheng 0005 |
CSCWD | 2 |
| 2025 | Optimizing Direct Convolutions on High-Performance Multi-Core DSPsabstractConvolution operations form the computational backbone of deep learning inference but often become performance bottlenecks on conventional architectures. While multi-core Digital Signal Processors (DSPs) offer energy-efficient alternatives through long vector units and software-managed memory hierarchies, existing convolution optimizations designed for CPUs/GPUs underperform due to architectural mismatches in memory systems and execution pipelines. We present mtConv, an optimizing convolution method for multi-core DSPs. mtConv achieves high performance by exploiting data reuse, managing on-chip memory, and designing efficient micro-kernels. It maximizes the overlap between computation and communication to hide data transfer latency, leveraging DSPs’ long vector units and hierarchical scratchpad memories. We evaluate mtConv against state-of-the-art convolution optimizations on DSPs. Experimental results show that mtConv delivers the best overall performance across various convolution layers, achieving up to 93.25% of the hardware’s peak performance on a single DSP core and 92.31% when using all 8 cores of a single DSP cluster. Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Chun Huang 0006, Jie Ren 0007 |
ICPP | 7 |
| 2025 | Optimizing Personalized Federated Learning Through Adaptive Layer-Wise LearningabstractReal-life deployment of federated Learning (FL) often faces non-IID data, which leads to poor accuracy and slow convergence. Personalized FL (pFL) tackles these issues by tailoring local models to individual data sources and using weighted aggregation methods for client-specific learning. However, existing pFL methods often fail to provide each local model with global knowledge on demand while maintaining low computational overhead. Additionally, local models tend to over-personalize their data during the training process, potentially dropping previously acquired global information. We propose FLAYER, a novel layer-wise learning method for pFL that optimizes local model personalization performance. FLAYER considers the different roles and learning abilities of neural network layers of individual local models. It incorporates global information for each local model as needed to initialize the local model cost-effectively. It then dynamically adjusts learning rates for each layer during local training, optimizing the personalized learning process for each local model while preserving global knowledge. Additionally, to enhance global representation in pFL, FLAYER selectively uploads parameters for global aggregation in a layer-wise manner. We evaluate FLAYER on four representative datasets in computer vision and natural language processing domains. Compared to eight state-of-the-art pFL methods, FLAYER improves the inference accuracy, on average, by 5.20% (up to 14.29%). Code is available at https://github.com/lancasterJie/FLAYER/. Weihang Chen, Jie Ren 0007, Zhiqiang Li 0003, Zheng Wang 0001 |
IJCAI | 3 |
| 2025 | StrongLive: Adaptive Offloading and Scene-Aware SR Learning for 4 K Live Streaming on MobileabstractThe demand for 4 K live streaming has grown rapidly, driven by the desire for ultra-high-definition viewing experiences. However, delivering seamless 4 K live streams on mobile devices presents significant challenges due to mobile network limitations, including restricted bandwidth and upload capacity, which can lead to increased latency in broadcasting high-resolution video. Additionally, encoding 4 K video on mobile devices requires substantial computational resources, straining their capabilities. This paper presents StrongLive, a novel computation offloading framework designed to optimize$\mathbf{4 K}$live video streaming on mobile platforms. Specifically, StrongLive leverages advanced SR models to process the selected frames and utilizes decoding techniques to restore non-key frames by combining residual information with previously SR-processed frames. Additionally, StrongLive harnesses a GPU-accelerated video processing pipeline that seamlessly integrates decoding, upscaling, and encoding tasks, significantly enhancing processing speed without compromising 4 K quality. To adeptly handle scene changes, the StrongLive incorporates an online learning mechanism that continuously refines the SR model in real time, ensuring consistent perceptual quality across diverse broadcasting scenarios. We evaluate StrongLive on over 30004 K video clips under three typical network environments. Results show that StrongLive outperforms state-of-the-art methods, achieving an average reduction of 89.5 % in FPS violations and delivering high-quality$\mathbf{4 K}$live video streaming that meets user perceptual expectations. Rongqing Liu, Jie Ren 0007, Jie Zheng 0005 |
IWQoS | 3 |
| 2025 | Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core ProcessorabstractOptimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries. Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Jie Ren 0007 |
SC | 5 |
| 2025 | The impact of unsupervised feature selection techniques on the performance and interpretation of defect prediction models
Zhiqiang Li 0003, Wenzhi Zhu, Hongyu Zhang 0002, Yuantian Miao, Jie Ren 0007 |
Autom. Softw. Eng. | 5 |
| 2025 | FMCC-RT: a scalable and fine-grained all-reduce algorithm for large-scale SMP clusters
Jintao Peng, Jie Liu 0002, Jianbin Fang, Zhiquan Lai, Bo Yang 0023, Chunye Gong, Xinjun Mao, Guo Mao, Jie Ren 0007 |
Sci. China Inf. Sci. | 11 |
| 2025 | Generative AI-Aided Multimodal Parallel Offloading for AIGC Metaverse Service in IoT NetworksabstractMobile edge computing (MEC) enabled artificial intelligence-generated content (AIGC) has garnered considerable attention. To support AIGC metaverse applications within MEC in Internet of Things (IoT) networks, it is effective to offload computation tasks, particularly those involving neural networks generative in AIGC, from mobile devices to edge clouds. Existing solutions typically assume the availability of a dedicated and powerful edge server for each user with single modal data, which can handle the entire AIGC service offloading. However, the practical availability of such dedicated and powerful servers may be limited, necessitating the utilization of less capable alternatives. Thus, we propose the multimodal parallel offloading AIGC framework which partitions multimodal content and offloads partial diffusion tasks to multiple servers. Our proposed scheme accelerates mobile deep vision multimodal metaverse applications through parallel offloading provided by multiple servers. We further utilize the generative AI scheme to solve offloading problems to adapt the dynamic and available communication and computing resource in wireless IoT network. Our framework proposed a multimodal parallel diffusion offloading scheme with integrating the recurrent region proposal prediction algorithm to optimize communication and computing resources while minimizing delay. Simulation results show that our approach can significantly reduce delay compared to conventional algorithms. Weizhe Zeng, Jie Zheng 0005, Jinping Niu, Jie Ren 0007, Hai Wang 0010, Rui Cao 0003 |
IEEE Internet Things J. | 5 |
| 2025 | VIKCSE: Visual-knowledge enhanced contrastive learning with prompts for sentence embedding
Rui Cao 0003, Yihao Wang 0010, Jie Ren 0007, Jie Zheng 0005, Jianfen Yang |
Knowl. Based Syst. | 4 |
| 2025 | nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUsabstractConvolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures. Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007 |
IEEE Trans. Computers | 11 |
| 2025 | Trust Online Over-the-Air Computation for Wireless Federated LearningabstractUsing the wireless waveform superposition property, over-the-air computation (OAC) enables federated learning (FL) to achieve fast model aggregation. However, this computing paradigm is vulnerable to poisoning attacks due to the openness of a wireless channel over time, where malicious mobile devices can introduce cumulative errors for the global FL model in a time-varying wireless environment for each communication round. This article presents a trust online OAC (TO-OAC) scheme to minimize impacts on the global model introduced by malicious devices adjusting to dynamic attack and wireless channel fluctuations over time. TO-OAC achieves this by utilizing trustworthy security quantification of OAC for each FL training round. To optimize the cumulative training loss at the aggregation node with the long-term power and trust constraints of mobile devices, we propose a joint trust, power, and channel-aware algorithm to flexibly update local and global models in response to the dynamic changes in the wireless and secure environment. We analyze the performance limits for the aggregation of trust models, considering metrics for computation and communication through time. We then propose another trust online regularization over-the-air computation (TOR-OAC) as an improved version of the TO-OAC scheme to decrease convergence time while ensuring long-term trust and power limitation. Experimental results performed on real-life datasets show that the two proposed schemes (TO-OAC and TOR-OAC) outperform prior works, especially in noisy, time-varying wireless channels and malicious attacks. Mingjie Sun, Jie Zheng 0005, Hongyang Du 0001, Haijun Zhang 0001, Dusit Niyato, Jiawen Kang 0001, Jiacheng Wang 0001, Jie Ren 0007, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | Optimizing Stencil Computation on Multi-core DSPsabstractStencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart. Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007 |
ICPP | 10 |
| 2024 | Online Learning to parallel offloading in heterogeneous wireless networks
Yulin Qin, Jie Zheng 0005, Hai Wang 0010, Yuhui Ma, Jie Ren 0007, Rui Cao 0003, Yongxing Zheng |
Comput. Commun. | 8 |
| 2024 | JavaScript Performance Tuning as a Crowdsourced ServiceabstractJavaScript (JS) is one of the most used programming languages for mobile applications. As JS is increasingly used in computation-intensive and latency-sensitive components, JS application performance can significantly impact user experience. While compilers play a crucial role in optimizing JS performance on mobile systems, their optimizations must be simple due to the computation and battery usage limitations of the underlying hardware platforms. We presentJSTuner, a machine-learning system to leverage compiler-based autotuning techniques to optimize JS performance by finding a good compiler optimization sequence.JSTuneris designed to reduce the cost of autotuning by using prior knowledge of JS programs collected through a crowdsourcing framework to bootstrap the search process. It allows the user to seamlessly utilize the computation resources of a cloud server to perform the heavy-lifting autotuning process for repeatedly running JS components. This enables aggressive search-based optimizations that are too expensive to run on the user's device. We evaluateJSTunerby applying it to 60 JS benchmarks across three distinct mobile devices and comparing it against four search-based techniques. Experimental results show that JSTuner consistently outperforms prior techniques and improves JS performance by 1.62x on average (up to 3.33x) over the default compiler setting used by the Chrome V8 JS engine. Jie Ren 0007, Zheng Wang 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Noise processing and multitask learning for far-field dialect classificationabstractSummary Deep learning has made great achievements in the field of speech recognition. With the popularization of embedded devices such as intelligent speaker and the demand for dialect interaction scenes, it poses great challenges to far‐field speech recognition and dialect language recognition. In order to solve the dialect language recognition of embedded devices in far‐field speech recognition, we propose a deep learning neural network model with multitask learning. First, the audio is passed through the end‐to‐end noise reduction model to improve the effect of audio recognition. Then we define dialect recognition as the main task and dialect area as the auxiliary task, using the multitask learning method to improve the accuracy of dialect classification. The experimental results show that the end‐to‐end noise reduction model can improve the accuracy of audio recognition, and the best effect can be 7.54% higher than the baseline, and the accuracy of dialect language recognition can be improved by about 5% through multi task learning model. Hai Wang 0010, Yuhui Ma, Chenguang Qin, Jie Ren 0007 |
Concurr. Comput. Pract. Exp. | 7 |
| 2023 | Covert Federated Learning via Intelligent Reflecting SurfacesabstractOver-the-air computation (OAC) is a promising technology that can achieve rapid model aggregation by utilizing the wireless waveform superposition feature to harness the interference of multiple-access channel for wireless federated learning (FL). However, OAC-based aggregation for OAC faces critical security challenges due to unfavorable and wireless broadcast properties, such as privacy leaks and eavesdropping attacks. In this paper, we propose to utilize an intelligent reflecting surface (IRS) to support covert OAC-based FL. We first derive the optimal condition for covertness in OAC with IRS and formulate a joint optimization problem to select the maximum covert devices participating in the model aggregation while satisfying the mean squared error (MSE) requirement. We then design a covert difference-of-convex-functions program (CDC) to efficiently determine the transmission power of the device, aggregation beamforming of base station (BS), phase shifts, and reflection amplitudes at the IRS. Simulation results demonstrate that our proposed approach can achieve significant performance gain compared to the baseline algorithms by deploying IRS into covert OAC-based FL. Jie Zheng 0005, Haijun Zhang 0001, Jiawen Kang 0001, Jie Ren 0007, Dusit Niyato |
IEEE Trans. Commun. | 5 |
| 2023 | DSSDPP: Data Selection and Sampling Based Domain Programming Predictor for Cross-Project Defect PredictionabstractCross-project defect prediction (CPDP) refers to recognizing defective software modules in one project (i.e., target) using historical data collected from other projects (i.e., source), which can help developers find defects and prioritize their testing efforts. Unfortunately, there often exists large distribution difference between the source and target data. Most CPDP methods neglect to select the appropriate source data for a given target at the project level. More importantly, existing CPDP models are parametric methods, which usually require intensive parameter selection and tuning to achieve better prediction performance. This would hinder wide applicability of CPDP in practice. Moreover, most CPDP methods do not address the cross-project class imbalance problem. These limitations lead to suboptimal CPDP results. In this paper, we propose a novel data selection and sampling based domain programming predictor (DSSDPP) for CPDP, which addresses the above limitations. DSSDPP is a non-parametric CPDP method, which can perform knowledge transfer across projects without the need for parameter selection and tuning. By exploiting the structures of source and target data, DSSDPP can learn a discriminative transfer classifier for identifying defects of the target project. Extensive experiments on 22 projects from four datasets indicate that DSSDPP achieves betterMCCandAUCresults against a range of competing methods both in the single-source and multi-source scenarios. Since DSSDPP is easy, effective, extensible, and efficient, we suggest that future work can use it with the well-chosen source data to conduct CPDP especially for the projects with limited computational budget. Zhiqiang Li 0003, Hongyu Zhang 0002, Xiaoyuan Jing, Juanying Xie, Jie Ren 0007 |
IEEE Trans. Software Eng. | 6 |
| 2021 | Learning to Remove: Towards Isotropic Pre-trained BERT Embedding
Rui Cao 0003, Jie Zheng 0005, Jie Ren 0007 |
ICANN (5) | 4 |
| 2021 | Timeception Single Shot Action Detector: A Single-Stage Method for Temporal Action Detection
Xiaoqiu Chen, Miao Ma, Zhuoyu Tian, Jie Ren 0007 |
ICIG (1) | 4 |
| 2021 | ATO-EDGE: Adaptive Task Offloading for Deep Learning in Resource-Constrained Edge Computing SystemsabstractOn-device deep learning enables mobile devices to perform complex tasks, such as object detection and voice translation, regardless of the network condition. The advanced deep learning model gives an excellent performance, also leads to a heavy burden on resource-limited devices (i.e., mobile devices). To speed up the on-device deep learning. Prior studies focus on developing lightweight network architecture for real-time inference by sacrificing model accuracy. This paper presents ATO-EDGE: adaptive task offloading for deep learning based on edge computing. Considering three optimization goals, energy consumption, accuracy, and latency, ATO-EDGE leverages an offline pre-trained model to select a suitable deep learning model on a specific device to process the given task. We apply our approach to object detection and evaluate it on Jetson TX2, Xilinx ZYNQ 7020, and Raspberry 3B+. The deep learning model candidates contain ten typical object detection models trained on Microsoft COCO 2017 dataset. We obtain, on average, 28.25%, 35.44%, and 0.9 improvements respectively for latency, energy consumption, and mAP (mean average precision) when compared to the SOTA DETR model on the Raspberry Pi. Yihao Wang 0010, Jie Ren 0007, Rui Cao 0003, Hai Wang 0010, Jie Zheng 0005, Quanli Gao |
ICPADS | 3 |
| 2021 | Towards practical 3D ultrasound sensing on commercial-off-the-shelf mobile devices
Shuangjiao Zhai, Guixin Ye, Zhanyong Tang, Jie Ren 0007, Dingyi Fang, Baoying Liu, Zheng Wang 0001 |
Comput. Networks | 4 |
| 2021 | eICIC Configuration of Downlink and Uplink Decoupling With SWIPT in 5G Dense IoT HetNetsabstractInterference management and power transfer can provide a significant improvement over the 5th generation mobile networks (5G) dense Internet of Things (IoT) heterogeneous networks (HetNets). In this paper, we present a novel approach to simultaneously manage inferences at the downlink (DL) and uplink (UL), and to identify opportunities for power transfer and additional UL transmissions integrated with existing protocols and infrastructures for enhanced inter-cell interference coordination (eICIC) protocol in dense IoT HetNets, while considering practical non-linear energy harvesting (EH) model. The design is formulated as the joint optimization of interference aware UL/DL decoupling, airtime resource allocation and energy transfer. The key insight of our algorithm is to translate the original, intractable joint-optimization problem into a problem space where a good approximate solution can be quickly found. We evaluate our scheme through theoretical analysis and simulation. The evaluation shows that our approach improves the system utility by over 20% compared to start-of-the-art in dense IoT HetNets. Compared to alternative schemes, our approach maintains the best user fairness and rate experience and can solve the problem in a fast and scalable way. Jie Zheng 0005, Haijun Zhang 0001, Dusit Niyato, Jie Ren 0007, Hai Wang 0010, Zheng Wang 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2020 | Camel: Smart, Adaptive Energy Optimization for Mobile Web InteractionsabstractWeb technology underpins many interactive mobile applications. However, energy-efficient mobile web interactions is an outstanding challenge. Given the increasing diversity and complexity of mobile hardware, any practical optimization scheme must work for a wide range of users, mobile platforms and web workloads. This paper presents CAMEL, a novel energy optimization system for mobile web interactions. CAMEL leverages machine learning techniques to develop a smart, adaptive scheme to judiciously trade performance for reduced power consumption. Unlike prior work, CAMEL directly models how a given web content affects the user expectation and uses this to guide energy optimization. It goes further by employing transfer learning and conformal predictions to tune a previously learned model in the end-user environment and improve it over time. We apply CAMEL to Chromium and evaluate it on four distinct mobile systems involving 1,000 testing webpages and 30 users. Compared to four state-of-the-art web-event optimizers, CAMEL delivers 22% more energy savings, but with 49% fewer violations on the quality of user experience, and exhibits orders of magnitudes less overhead when targeting a new computing environment. Jie Ren 0007, Petteri Nurmi, Miao Ma, Zhanyong Tang, Jie Zheng 0005, Zheng Wang 0001 |
INFOCOM | 1 |
| 2020 | Smart Edge Caching-Aided Partial Opportunistic Interference Alignment in HetNets
Jie Zheng 0005, Hai Wang 0010, Jinping Niu, Jie Ren 0007 |
Mob. Networks Appl. | 5 |
| 2018 | Proteus: network-aware web browsing on heterogeneous mobile systemsabstractWe present Proteus, a novel network-aware approach for optimizing web browsing on heterogeneous multi-core mobile systems. It employs machine learning techniques to predict which of the heterogeneous cores to use to render a given webpage and the operating frequencies of the processors. It achieves this by first learning offline a set of predictive models for a range of typical networking environments. A learnt model is then chosen at runtime to predict the optimal processor configuration, based on the web content, the network status and the optimization goal. We evaluate Proteus by implementing it into the open-source Chromium browser and testing it on two representative ARM big.LITTLE mobile multi-core platforms. We apply Proteus to the top 1,000 popular websites across seven typical network environments. Proteus achieves over 80% of best available performance. It obtains, on average, over 17% (up to 63%), 31% (up to 88%), and 30% (up to 91%) improvement respectively for load time, energy consumption and the energy delay product, when compared to two state-of-the-art approaches. Jie Ren 0007, Jianbin Fang, Yansong Feng 0002, Dongxiao Zhu, Zhunchen Luo, Jie Zheng 0005, Zheng Wang 0001 |
CoNEXT | 1 |
| 2018 | Max-Min Energy-Efficient eICIC Configuration in Heterogeneous NetworkabstractThe adaptive enhanced inter-cell interference coordination (eICIC) configuration is critical for interference management. This problem is challenging especially from energy efficiency perspective and taking individual user fairness into account. Therefore, we formulate a max-min energy efficiency eICIC configuration problem, i.e., determining the number of almost blank subframes (ABS) and user associates with macro or pico while considering fairness jointly. Since the mixed combinatorial and non-smooth features of the problem, an iterative- distributed algorithm is proposed with using fractional programming and Lagrangian dual theory. Numerical results demonstrate the effectiveness of the proposed algorithm and verify fairness achieved among users, and validate the tradeoff between energy efficiency and fairness for eICIC in HetNets comparing with the existing algorithms. Jie Zheng 0005, Haijun Zhang 0001, Hai Wang 0010, Jinping Niu, Xiaoya Li 0003, Jie Ren 0007 |
ICC | 7 |
| 2017 | Optimise web browsing on heterogeneous mobile platforms: A machine learning based approachabstractWeb browsing is an activity that billions of mobile users perform on a daily basis. Battery life is a primary concern to many mobile users who often find their phone has died at most inconvenient times. The heterogeneous multi-core architecture is a solution for energy-efficient processing. However, the current mobile web browsers rely on the operating system to exploit the underlying hardware, which has no knowledge of individual web contents and often leads to poor energy efficiency. This paper describes an automatic approach to render mobile web workloads for performance and energy efficiency. It achieves this by developing a machine learning based approach to predict which processor to use to run the web rendering engine and at what frequencies the processors should operate. Our predictor learns offline from a set of training web workloads. The built predictor is then integrated into the browser to predict the optimal processor configuration at runtime, taking into account the web workload characteristics and the optimisation goal: whether it is load time, energy consumption or a trade-off between them. We evaluate our approach on a representative ARM big.LITTLE mobile architecture using the hottest 500 webpages. Our approach achieves 80% of the performance delivered by an ideal predictor. We obtain, on average, 45%, 63.5% and 81% improvement respectively for load time, energy consumption and the energy delay product, when compared to the Linux heterogeneous multi-processing scheduler. Jie Ren 0007, Hai Wang 0010, Zheng Wang 0001 |
INFOCOM | 1 |
| 2017 | A preference elicitation method based on bipartite graphical correlation and implicit trust
Quanli Gao, Jianping Fan 0001, Jie Ren 0007 |
Neurocomputing | 4 |
| 2017 | EE-eICIC: Energy-Efficient Optimization of Joint User Association and ABS for eICIC in Heterogeneous Cellular NetworksabstractThe densification and expansion of heterogeneous cellular networks (HetNets) pose new challenges on interference management and reduction of energy consumption. The 3GPP has proposed enhanced intercell interference coordination (eICIC) by making a macrocell silent in almost blank subframes (ABSs) to mitigate interference for low power base stations (BSs) in HetNets. However, energy efficiency (EE) is very crucial for the deployment of a large number of low power nodes as they consume a lot of energy. In this work, we develop a novel EE-eICIC algorithm to determine the amount of ABSs and user equipment (UE) that should associate with picocells or macrocells from energy efficiency perspective. Due to the nonsmooth and mixed combinatorial features of this formulation, we focus on a suboptimal algorithm design. Using generalized fractional programming and the convex programming theory, we propose an iterative and relaxed-rounding algorithm to solve the problem. Numerical results illustrate that the proposed EE-eICIC algorithm achieves superior performance in comparison with state-of-the-art methods in terms of energy efficiency of both system and user. Jie Zheng 0005, Hai Wang 0010, Jinping Niu, Xiaoya Li 0003, Jie Ren 0007 |
Wirel. Commun. Mob. Comput. | 6 |