VLDB 2026 Research / reviewers in the wild / expert
Xiaoli Gong
dblp:45/1397
· DBLP profile ↗
33ranked-venue papers
3as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JSQKV: Joint Sparsification and Quantization for KV-Cache Compression and Decode Acceleration
Xiaoli Gong, Huayou Su, Qingxia Chen, Jin Zhang 0003 |
APPT | 2 |
| 2026 | Visual Question Explainable Reasoning on Hypothesis Agent Interaction with Scene
Baoyu Fan, Cong Xu 0001, Lu Liu 0009, Xiaoli Gong, Jin Zhang 0003 |
Signal Process. | 5 |
| 2025 | So Far Yet So Near: Time Series Data Augmentation with Exploring non-Semantic Boundaries based on Reinforcement LearningabstractData augmentation effectively expands feature distribution in time series classification, enhancing downstream task performance. However, existing techniques often fail to maintain semantic consistency between augmented and original time series data, causing label noise and thereby degrading downstream task performance. We argue that data augmentation should preserve time series semantic consistency and expand the non-semantic information space. In this paper, we reformulate data augmentation as a semantic path planning problem between original data and augmented data, modeled as a Markov Decision Process (MDP). We propose a reinforcement learning-based algorithm (RL) named FreqSYN, where the action space is defined by a set of learnable Gaussian kernels that perturbs the frequency domain of the original data to generate augmented samples. The confidence coefficients of augmented data in semantically relevant classification tasks are used as a reward to iteratively refine the FreqSYN. Our method is validated across four datasets, achieving state-of-the-art performance, with a 2% improvement in F1 score over the SimPSI method. The code and models are available at https://github.com/NKU-EmbeddedSystem/FreqSYN. Haoran Li 0014, Jiarong Kang, Xun Jiang 0001, Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Andrzej Cichocki |
ICASSP | 5 |
| 2025 | Essentia: Boosting Artifact Removal from EEG through Semantic Guidance Utilizing Diffusion ModelabstractElectroencephalography (EEG) is a time-series signal containing semantic information that can be used to determine human brain activities. Artifacts within EEG data can interfere with the intrinsic distribution of this semantic information, so removing artifacts is crucial for improving EEG analysis performance on downstream tasks. In this paper, we redefine the efficacy of the artifact removal model by evaluating the performance of the noisy EEG data in downstream tasks before and after artifact removal. Currently, most artifact removal models fail to ensure semantic consistency, rendering them ineffective. To solve it, we propose an artifact removal model based on the 1-dimensional diffusion model utilizing the U-Net, referred to as Essentia. Moreover, we find that the skip-connection layer in U-Net contains mid-to-high-frequency information that interferes with the semantic representation. We introduce a semantic guidance module (SGM) that leverages contrastive learning to generate semantic distribution weights, boosting semantic representation. We evaluate Essentia on three datasets with six solutions. The accuracy of downstream tasks from the denoised EEG data increased by 4% compared with the DeepSeparetor. The code and models are available at https://github.com/NKU-EmbeddedSystem/Essentia. Haoran Li 0014, Xiaoli Gong, Jin Zhang 0003, Tingjuan Lu, Zhe Sun 0009, Andrzej Cichocki |
ICASSP | 4 |
| 2025 | Dropletvideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
Guoguang Du 0001, Xiaochuan Li 0001, Qi Jia 0004, Lu Liu 0009, Cong Xu 0001, Zhenhua Guo 0003, Yaqian Zhao, Xiaoli Gong, RenGang Li, Baoyu Fan |
ICCV | 11 |
| 2024 | G2G: Generalized Learning by Cross-Domain Knowledge Transfer for Federated Domain GeneralizationabstractWe propose G2G, based on the global model of Generalized learning to solve the Federated Domain Generalization (FedDG) task. FedDG aims to collaboratively train a global model that can directly generalize to the unseen target domain without data sharing. Existing methods face challenges from both data heterogeneity, arising from imbalanced as well as non-independent and identical distributions (non-IID) among all domains, and model heterogeneity due to personalized requirements for client models. Also, these methods suffer from unnecessary time cost due to aggregation and distribution. G2G addresses these issues by making the global model acquire in-domain classification knowledge during local knowledge transfer and gain extensive knowledge in cross-domain training. Moreover, G2G eliminates waiting time by allowing clients to train independently when not trained with the global model. G2G outperforms state-of-the-art (SOTA) methods by 1.52%, 2.35%, and 1.11% on three datasets, respectively. Xinqian Chen, Xiaoli Gong |
ICASSP | 3 |
| 2024 | Prism: Decomposing Program Semantics for Code Clone Detection through CompilationabstractCode clone detection (CCD) is of critical importance in software engineering, while semantic similarity is a key evaluation factor for CCD. The embedding technique, which represents an object using a numerical vector, is utilized to generate code representations, where code snippets with similar semantics (clone pairs) should have similar vectors. However, due to the diversity and flexibility of high-level program languages, the code representation of clone pairs may be inconsistent. Assembly code provides the program execution trace and can normalize the diversity of high-level languages in terms of the program behavior semantics. After revisiting the assembly language, we find that different assembly codes can align with the computational logic and memory access patterns of cloned pairs. Therefore, the use of multiple assembly languages can capture the behavior semantics to enhance the understanding of programs. Thus, we propose Prism, a new method for code clone detection fusing behavior semantics from multiple architecture assembly code, which directly captures multilingual domains' syntax and semantic information. Additionally, we introduce a multi-feature fusion strategy that leverages global information interaction to expand the representation space. This fusion process allows us to capture the complementary information from each feature and leverage the relationships between them to create a more expressive representation of the code. After testing the OJClone dataset, the Prism model exhibited exceptional performance with precision and recall scores of 0.999 and 0.999, respectively. Haoran Li 0014, Siqian Wang, Weihong Quan, Xiaoli Gong, Huayou Su, Jin Zhang 0003 |
ICSE | 4 |
| 2024 | SyncIntellects: Orchestrating LLM Inference with Progressive Prediction and QoS-Friendly ControlabstractLarge Language Models (LLMs) have shown impressive capabilities, especially in the realm of Human-Machine Chat Systems. Nevertheless, these models entail significant computational expenses, particularly when generating tokens. As a remedy to enhance system throughput and hardware utilization, batch scheduling is commonly adopted. This method involves initiating a batch of inference requests concurrently and then waiting for their completion. A significant challenge encountered with task-batching is the need to group requests with similar response lengths. However, accurately predicting response length proves to be a daunting task, and the inherent variability in response length leads to suboptimal resource utilization.In this paper, we introduce SyncIntellects, a framework designed to orchestrate Large Language Model (LLM) Inference with fine-grained response length prediction and Quality of Service (QoS)-Friendly length control. Specifically, SyncIntellects enhances response length prediction by leveraging embedding information during token generation through a transformer-based model. Subsequently, a dynamic response length controller based on Prompt Engineering techniques is employed to ensure alignment of response lengths without compromising the QoS of the responses. We have implemented SyncIntellects and seamlessly integrated it with a chatbot engine based on the llama2 7B model. We conduct comprehensive experiments on an NVIDIA A100-based testbed, and the results demonstrate a significant reduction in latency by 17.76% on average, along with an increase in throughput by 9.34%. Xue Lin 0006, Peining Yue, Haoran Li 0014, Jin Zhang 0003, Baoyu Fan, Huayou Su, Xiaoli Gong |
IWQoS | 8 |
| 2024 | OneGraph: a cross-architecture framework for large-scale graph computing on GPUs based on oneAPI
Jiaxun Han, Xiaoli Gong, Gang Wang 0001, Jin Zhang 0003, Xuqiang Wang |
CCF Trans. High Perform. Comput. | 6 |
| 2024 | Distance-based feature repack algorithm for video coding for machines
Yuan Zhang 0023, Xiaoli Gong, Hualong Yu, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | JiuJITsu: Removing Gadgets with Safe Register Allocation for JIT Code GenerationabstractCode-reuse attacks have the capability to craft malicious instructions from small code fragments, commonly referred to as “gadgets.” These gadgets are generated by JIT (Just-In-Time) engines as integral components of native instructions, with the flexibility to be embedded in various fields, including Displacement . In this article, we introduce a novel approach for potential gadget insertion, achieved through the manipulation of ModR/M and SIB bytes via JavaScript code. This manipulation influences a JIT engine’s register allocation and code generation algorithms. These newly generated gadgets do not rely on constants and thus evade existing constant blinding schemes. Furthermore, they can be combined with 1-byte constants, a combination that proves to be challenging to defend against using conventional constant blinding techniques. To showcase the feasibility of our approach, we provide proof-of-concept (POC) code for three distinct types of gadgets. Our research underscores the potential for attackers to exploit ModR/M and SIB bytes within JIT-generated native instructions. In response, we propose a practical defense mechanism to mitigate such attacks. We introduce JiuJITsu , a security-enhanced register allocation scheme designed to prevent harmful register assignments during the JIT code generation phase, thereby thwarting the generation of these malicious gadgets. We conduct a comprehensive analysis of JiuJITsu ’s effectiveness in defending against code-reuse attacks. Our findings demonstrate that it incurs a runtime overhead of under 1% when evaluated using JetStream2 benchmarks and real-world websites. Zhang Jiang, Ying Chen 0034, Xiaoli Gong, Jin Zhang 0003, Wenwen Wang 0001, Pen-Chung Yew |
ACM Trans. Archit. Code Optim. | 3 |
| 2024 | Hybrid-Memcached: A Novel Approach for Memcached Persistence Optimization With Hybrid MemoryabstractMemcached is a widely adopted, high-performance, in-memory key-value object caching system utilized in data centers. Nonetheless, its data is stored in volatile DRAM, making the cached data susceptible to loss during system shutdowns. Consequently, cold restarts experience significant delays. Persistent memory is a byte-addressable, large-capacity, and non-volatility storage media, which can be employed to avoid the cold restart problem. However, deploying Memcached on persistent memory requires consideration of issues such as write endurance, asymmetric read/write latency and bandwidth, and write granularity of persistent memory. In this paper, we propose Hybrid-Memcached, an optimized Memcached framework based on a hybrid combination of DRAM and persistent memory. Hybrid-Memcached includes three key components: (1) a DRAM-based data aggregation buffer to avoid multiple fine-grained writes, which extends the write endurance of persistent memory, (2) a data-object alignment mechanism to avoid write amplification, and (3) a non-temporal store instruction-based writing strategy to improve the bandwidth utilization. We have implemented Hybrid-Memcached on the Intel Optane persistent memory. Several micros-benchmarks are designed to evaluate Hybrid-Memcached by varying read/write ratios, access distributions, and key-value item sizes. Additionally, we evaluated it with the YCSB benchmark, showing a 21.2% performance improvement for fully write-intensive workloads and 11.8% for read-write balanced workloads. Zhang Jiang, Xianduo Li, Tianxiang Peng, Haoran Li 0014, Jingxuan Hong, Jin Zhang 0003, Xiaoli Gong |
IEEE Trans. Computers | 7 |
| 2023 | KylinArm: An Arm Gesture Recognition System for Mobile Devices
Shikun Zhao, Jingxuan Hong, Xuqiang Wang, Xiaoli Gong |
ICA3PP (2) | 6 |
| 2023 | Privacy-Preserving Multi-Source Domain Adaptation for Medical DataabstractGreat progress has been made in diagnosing medical diseases based on deep learning. Large-scale medical data are expected to improve deep learning performance further. It is almost impossible for a single institution to collect so much data due to the time-consuming and costly collection and labeling of medical data. Many studies have turned attention to data sharing among multiple medical institutions. However, due to different data acquiring and processing procedures, multiple institutions' medical data is characterized by distribution heterogeneity. Besides, the protection of patient privacy in medical data sharing has also been a common concern. To simultaneously address the problems of heterogeneous data distribution and privacy protection, we propose a novel multi-source source free domain adaptation. When aligning distributed heterogeneous data, our method only require to transfer the pre-trained source models rather than the direct source domain data, thus protecting patients' privacy. In addition, it has the advantages of being efficient and less costly in network resources. The proposed method is evaluated on the multi-site fMRI database Autism Brain Imaging Data Exchange (ABIDE) and yields an average accuracy of 69.37%. We also analyzed its effectiveness on network resource-saving and conducted additional experiments on Camelyon17 to validate the generalization. Xiaoli Gong, Jin Zhang 0003, Zhe Sun 0009, Yu Zhang 0009 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | Liberator: A Data Reuse Framework for Out-of-Memory Graph Computing on GPUsabstractGraph analytics are widely used including recommender systems, scientific computing, and data mining. Meanwhile, GPU has become the major accelerator for such applications. However, the graph size increases rapidly and often exceeds the GPU memory, incurring severe performance degradation due to frequent data transfers between the main memory and GPUs. To relieve this problem, we focus on the utilization of data in GPUs by taking advantage of the data reuse across iterations. In our studies, we deeply analyze the memory access patterns of graph applications at different granularities. We have found that the memory footprint is accessed with a roughly sequential scan without a hotspot, which infers an extremely long reuse distance. Based on our observation, we propose a novel framework, calledLiberator, to exploit the data reuse within GPU memory. InLiberator, GPU memory is reserved for the data potentially accessed across iterations to avoid excessive data transfer between the main memory and GPUs. For the data not existing in GPU memory, a Merged and Aligned memory access manner is employed to improve the transmission efficiency. We also further optimize the framework by parallel processing of data in GPU memory and data in the main memory. We have implemented a prototype of theLiberatorframework and conducted a series of experiments on performance evaluation. The experimental results show thatLiberatorcan significantly reduce the data transfer overhead, which achieves an average of 2.7x speedup over a state-of-the-art approach. Ruiqi Tang, Xiaoli Gong, Wenwen Wang 0001, Jin Zhang 0003, Pen-Chung Yew |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2022 | An Efficient Transformer Inference Engine on DSP
Kangkang Chen, Huayou Su, Chaorun Liu, Xiaoli Gong |
ICA3PP | 4 |
| 2021 | Enhancing Atomic Instruction Emulation for Cross-ISA Dynamic Binary TranslationabstractDynamic Binary Translation (DBT) is a key enabler for cross-ISA emulation, system virtualization, runtime instrumentation, and many other important applications. Among several critical requirements for DBT, it is important to provide equivalent semantics for atomic synchronization instructions such as Load - Link / Store - Conditional (LL/SC), which are mostly included in the reduced-instruction set architectures (RISC) and Compare-and-Swap(CAS), which is mostly in the complex instruction set architectures (CISC). However, the state-of-the-art DBT tools often do not provide a fully correct translation of these atomic instructions, in particular, from RISC atomic instructions (i.e. LL/SC) to CISC atomic instructions (i.e. CAS), due to performance concerns. As a result, some may cause the well-known ABA problem, which could lead to wrong results or program crashes. In our experimental studies on QEMU, a state-of-the-art DBT, that runs multi-threaded lock-free stack operations implemented with ARM instruction set (i.e. using LL/SC) on Intel x86 platforms (i.e. using CAS), it often crashes within 2 seconds. Although attempts have been made to provide correct emulation for such atomic instructions, they either result in heavy execution overheads or require additional hardware support. In this paper, we propose several schemes to address those issues and implement them on QEMU to evaluate their performance overheads. The results show that all of the proposed schemes can provide correct emulation and, for the best solution, can achieve a min, max, geomean speedup of 1.25x, 3.21x, 2.03x respectively, over the best existing software-based scheme. Zhang Jiang, Ying Chen 0034, Xiaoli Gong, Wenwen Wang 0001, Pen-Chung Yew |
CGO | 4 |
| 2021 | Ascetic: Enhancing Cross-Iterations Data Efficiency in Out-of-Memory Graph Processing on GPUsabstractGraph analytics are widely used in real-world applications, and GPUs are major accelerators for such applications. However, as graph sizes become significantly larger than the capacity of GPU memory, the performance can degrade significantly due to the heavy overhead required in moving a large amount of graph data between CPU main memory and GPU memory. Ruiqi Tang, Kailun Wang, Xiaoli Gong, Jin Zhang 0003, Wenwen Wang 0001, Pen-Chung Yew |
ICPP | 4 |
| 2021 | Effective exploitation of SIMD resources in cross-ISA virtualizationabstractSystem virtualization is a fundamental technology that enables many important applications. However, existing virtualization techniques suffer from a critical limitation due to their limited exploitation of host SIMD hardware resources, especially when a guest application does not have inherently fine-grained data-level parallelism. To bridge this utilization gap and unleash the full potential of host SIMD resources, this paper proposes an effective and unconventional SIMD exploitation technique. The proposed exploitation takes advantage of ample host SIMD registers and powerful host SIMD instructions to generate more efficient host binary code for guest applications even without any fine-grained data-level parallelism. It also mitigates the shortage of general-purpose registers on the host platform, as well as improves the efficiency of accessing guest registers. We have implemented the exploitation in an extensively-used virtualization platform, QEMU. Experimental results on a comprehensive list of benchmarks from PARSEC, SPEC-CPU2017, and Google Octane JavaScript benchmark suite show that an average of 2.2X performance speedup can be achieved for AArch64 binaries on an x86-64 host machine. We believe the proposed technique will provide a new perspective for our community to rethink the exploitation of SIMD hardware resources. Jian Dong 0010, Ruili Fang, Xiaoli Gong, Wenwen Wang 0001, De-Cheng Zuo |
VEE | 5 |
| 2021 | Efficient attention based deep fusion CNN for smoke detection in fog environment
Lijun He 0001, Xiaoli Gong, Sirou Zhang, Fan Li 0003 |
Neurocomputing | 2 |
| 2021 | A Thread Level SLO-Aware I/O Framework for Embedded VirtualizationabstractWith the development of virtualization technology, it is practical and necessary to integrate virtual machine software into embedded systems. I/O scheduling is important for embedded systems, because embedded systems always face different situations and their requests have more diversity on the requirement of real-time and importance. However, the semantic information associated with the I/O data is completely lost when crossing the virtualized I/O software stack. Here, we present an I/O scheduling framework to connect the semantic gap between the application threads in virtual machines and hardware schedulers in the host machine. Therefore, the details for the I/O request can be passed through the layers of the software stack and each layer can get the specific information about the device environment. Also, various scheduling points have been provided to implement different I/O strategies. Our framework was implemented based on Linux operating system, KVM, QEMU and virtio protocol. A prototype scheduler, Orthrus, was implemented to evaluate the effectiveness of the framework. Comprehensive experiments were conducted and the results show that our framework can guarantee the real-time requirements, and reserve more system resources for critical tasks, with negligible memory consumption and throughput overhead. Xiaoli Gong, Dingyuan Cao 0001, Yusen Li, Jin Zhang 0003, Tao Li 0022 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | Towards Minimizing Resource Usage With QoS Guarantee in Cloud GamingabstractCloud gaming has been very popular recently, but providing satisfactory gaming experiences to players at a modest cost is still challenging. Colocating several games onto one server could improve server utilization. However, prior work regarding colocating games either ignores the performance interference between games or uses simple performance model to charaterize it, which may make inefficient game colocation decisions and cause QoS violations. In this article, we address the resource allocation issues for colocating games in cloud gaming. We first propose a novel machine learning-based performance model, which is able to capture the complex relationship among the performance interference, the contention features of colocated games and resource partition. Guided by the performance model, we then propose efficient and effective algorithms for two resource allocation scenarios in cloud gaming. We evaluate the proposed solutions through extensive experiments using a large number of real popular games. The results show that our performance model is able to identify whether a colocated game satisfies QoS requirement within an average error of 5 percent, which significantly outperforms the alternatives. Our resource allocation algorithms are able to increase the resource utilization by up to 60 percent compared to the state-of-the-art solutions. Yusen Li, Changjian Zhao, Xueyan Tang, Wentong Cai 0001, Xiaoguang Liu 0001, Gang Wang 0001, Xiaoli Gong |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2020 | DQEMU: A Scalable Emulator with Retargetable DBT on Distributed PlatformsabstractThe scalability of a dynamic binary translation (DBT) system has become important due to the prevalence of multicore systems and large multi-threaded applications. Several recent efforts have addressed some critical issues in extending a DBT system to run on multicore platforms for better scalability. In this paper, we present a distributed DBT framework, called DQEMU, that goes beyond a single-node multicore processor and can be scaled up to a cluster of multi-node servers. Zhang Jiang, Xiaoli Gong, Wenwen Wang 0001, Pen-Chung Yew |
ICPP | 4 |
| 2020 | Regaining Lost Seconds: Efficient Page Preloading for SGX EnclavesabstractIntel SGX is already here, with a strong emphasis on security and privacy. However, it is not free. Studies have shown that it incurs a significant performance overhead to take advantage of the security and privacy enhancement offered by SGX. In particular, it only provides limited physical memory for applications to use SGX. As a result, page faults can be frequently triggered during program execution, especially for memory-intensive applications with a large memory footprint. Therefore, it is imperative to look into possible optimization opportunities to enhance the efficiency of SGX. Wenwen Wang 0001, Xiaoli Gong, Pen-Chung Yew |
Middleware | 4 |
| 2020 | Monitoring Memory Behaviors and Mitigating NUMA Drawbacks on Tiered NVM Systems
Shengjie Yang, Xinglei Dou, Xiaoli Gong, Hao Liu 0107, Lei Liu 0037 |
NPC | 4 |
| 2019 | GAugur: Quantifying Performance Interference of Colocated Games for Improving Resource Utilization in Cloud GamingabstractCloud gaming has been very popular recently, but providing satisfactory gaming experiences to players at a modest cost is still challenging. Colocating several games onto one server could improve server utilization. To enable efficient colocations while providing Quality of Service (QoS) guarantees, a precise quantification of performance interference among colocated games is required. However, achieving such precise interference prediction is very challenging for games due to the complexity introduced by the contention on many shared resources across CPU and GPU. Moreover, the distinctive properties of cloud gaming require that the prediction model should be constructed beforehand and the prediction should be made instantaneously at request arrivals, which further increases the difficulty. The existing solutions are either not applicable or not effective due to many limitations. In this paper, we present GAugur, a novel methodology that enables highly accurate prediction of the performance interference among games arbitrarily colocated. By leveraging machine learning technologies, GAugur is able to capture the complex relationship between the interference and the contention features of colocated games. We evaluate GAugur through extensive experiments using a large number of real popular games. The results show that GAugur is able to identify whether a colocated game satisfies QoS requirement within an average error of 5%, and is able to quantify the performance degradation of a colocated game within an average error of 7.9%, which significantly outperforms the alternatives. Moreover, GAugur incurs an offline profiling cost linear to the number of games, and negligible overhead for online prediction. We apply GAugur to guiding efficient game colocations for cloud gaming. Experimental results show that GAugur is able to increase the resource utilization by 20% to 60%, and improve the overall performance by up to 15%, compared to the state-of-the-art solutions. Yusen Li, Chuxu Shan, Ruobing Chen 0002, Xueyan Tang, Wentong Cai 0001, Shanjiang Tang, Xiaoguang Liu 0001, Gang Wang 0001, Xiaoli Gong, Ying Zhang 0015 |
HPDC | 9 |
| 2019 | AliISA: Creating an Interactive Search Experience in E-commerce PlatformsabstractOnline shopping has been a habit of more and more people, while most users are unable to craft an informative query, and thus it often takes a long search session to satisfy their purchase intents. We present AliISA - a shopping assistant which offers users some tips to further specify their queries during a search session. With such an interactive search, users tend to find targeted items with fewer page requests, which often means a better user experience. Currently, AliISA assists tens of millions of users per day, earns more usage than existing systems, and consequently brings in a 5% improvement in CVR. In this paper, we present our system, describe the underlying techniques, and discuss our experience in stabilizing reinforcement learning under an E-commerce environment. Fei Xiao 0023, Zhen Wang 0036, Haikuan Huang, Jun Huang 0007, Hongbo Deng, Minghui Qiu, Xiaoli Gong |
SIGIR | 8 |
| 2019 | Dual buffer rotation four-stage pipeline for CPU-GPU cooperative computing
Tao Li 0022, Qiankun Dong, Xiaoli Gong, Yulu Yang |
Soft Comput. | 4 |
| 2018 | Improving Dynamically-Generated Code Performance on Dynamic Binary TranslatorsabstractThe recent transition in the software industry toward dynamically generated code poses a new challenge to existing dynamic binary translation (DBT) systems. A significant re-translation overhead could be introduced due to the maintenance of the consistency between the dynamically-generated guest code and the corresponding translated host code. To address this issue, this paper presents a novel approach to optimize DBT systems for guest applications with dynamically-generated code. The proposed approach can maximize the reuse of previously translated host code to mitigate the re-translation overhead. A prototype based on such an approach has been implemented on an existing DBT system HQEMU. Experimental results on a set of JavaScript applications show that it can achieve a 1.24X performance speedup on average compared to the original HQEMU. Wenwen Wang 0001, Jiacheng Wu 0001, Xiaoli Gong, Tao Li 0022, Pen-Chung Yew |
VEE | 3 |
| 2018 | A Webpage Offloading Framework for Smart Devices
Jin Zhang 0003, Weilai Liu, Wenjian Zhao, Haocong Xu, Xiaoli Gong |
Mob. Networks Appl. | 6 |
| 2017 | Automatically Difficulty Grading Method Based on Knowledge Tree
Jin Zhang 0003, Haoxiang Yang, Xiaoli Gong |
KSEM | 5 |
| 2016 | WWOF: An Energy Efficient Offloading Framework for Mobile WebpageabstractCurrently, the smart-phone has become a significant part for many people. As the major role in providing excellent surfing experiences for users, web browser can not only serve web sites visiting, but also support mobile web applications in smart-phones. In the meantime, in order to attract users, mobile web applications provide more and more diverse choices, which, however, results in the increase of CPU usage, time and energy consumption. Hence, how to improve efficiency of these applications without affecting user experiences has became a significant project. One of the solutions is to migrate the heavy computing tasks of mobile web applications from browser to cloud, which is called Offloading. This paper presents the design and implementation of a generic framework to realize the offloading from local browsers to cloud--Web Worker Offloading Framework (WWOF). The designed framework is easy to apply and is fully compatible with HTML5 API. On client side, the native Web Worker API is replaced by a library to offload seamlessly. On cloud side, corresponding interfaces are implemented to run workers on the server. To prove the effect of WWOF, evaluation of several benchmarks has been made, which shows 85% energy saving and 2-4 times execution speed-up on average for some mobile web applications. Xiaoli Gong, Weilai Liu, Jin Zhang 0003, Haocong Xu, Wenjian Zhao |
MobiQuitous | 1 |
| 2011 | Automatic Model Building and Verification of Embedded Software with UPPAALabstractEmbedded systems are becoming ubiquitous and taking more and more important part in our daily life. Increasingly complex functionality leads to higher develop cost and lower software quality. Model checking has the potential of alleviating these problems. In this paper, we present an approach to construct model directly from the source code. An embedded system design language, Virgil, is selected as the target. Without losing any information, the UPPAAL model is generated based on the typed intermediate language. The timing information and stack behavior are estimated and after merging the hardware platform model, the whole system can be simulated on the model checker and some safety and aliveness properties of the program are verified. Xiaoli Gong, Qingcheng Li, Jin Zhang 0003 |
TrustCom | 1 |