VLDB 2026 Research / reviewers in the wild / expert
Chen Yu 0003
dblp:98/4839-3
· DBLP profile ↗
53ranked-venue papers
4as first author
36since 2021 · last 2026
0000-0003-0782-0450ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 34 · 2 first-author · 23 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sonnet: A Workflow-Aware Serverless Platform for Time-Sensitive Edge Computing With WebAssemblyabstractThe serverless computing paradigm has emerged as a promising solution to address the resource underutilization and inflexible service scaling in edge environments by decoupling the monolithic application into a serverless workflow. However, existing serverless platforms are primarily designed for cloud centers, relying on heavyweight isolation mechanisms that are illsuited for resource-constrained edge computing. These limitations result in high latency, low deployment density, and restricted parallelism. In this paper, we proposeSonnet, a serverless platform tailored for edge computing, capable of rapidly responding to user requests and supporting efficient and elastic service scaling. Sonnet offers these features by (i) employing lightweight WebAssembly as the execution environment for functions, (ii) leveraging serverless workflow information to optimize function deployment on resource-constrained edge environments, and (iii) designing a function deployment algorithm that achieves dynamic load balancing within the cluster. An extensive evaluation ofSonnetwith real-world serverless workflows demonstrates its effectiveness and practical applicability. Compared with SOTA and commonly used edge computing serverless solutions, our experiments show that Sonnet can reduce end-to-end latency by 27% and improve throughput by 2.83×. Quanfeng Deng, Jing Wu 0024, Qiangyu Pei, Chuangxun Lin, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 5 |
| 2026 | KGNAS: Prior-Knowledge-Guided NAS for Real-Time Object Detection on Edge DevicesabstractAlthough deep neural networks have significantly improved object detection accuracy, manually designing efficient architectures is costly, and NAS networks incur high computational overhead, limiting their applicability in resource-constrained or real-time scenarios. This study aims to develop an efficient neural architecture search method capable of rapidly generating low-latency, high-performance detection networks on edge devices. We propose Knowledge-Guided Neural Architecture Search (KGNAS). First, explicit knowledge from a hardware universal deployment framework is used to enrich the search space; second, implicit knowledge characterizes the impact of memory interactions in hierarchical operator blocks on latency; finally, a zero-cost proxy metric evaluates the feature extraction capability of the backbone without training. Experiments show that KGNAS can design efficient networks for various edge devices within one hour, reducing inference latency by 88% compared to DetNAS with only a 3% accuracy loss. Yijun Mo, Xingao Tu, Chen Yu 0003 |
IEEE Trans. Computers | 3 |
| 2026 | Optimized Scheduling of Dependent Tasks and Idle Computational Resources for Edge IntelligenceabstractDue to the dependencies among different computing tasks, edge devices must await the completion of preceding computing tasks in order to continue with the current computing task. As a result, there is a long waiting time. In the past, reducing waiting time was often achieved by optimizing the scheduling of edge devices to execute computing tasks more efficiently. However, this approach incurs a certain level of communication overhead. Additionally, the waiting process for edge devices results in a waste of their computational resources. In order to tackle the challenge of excessive waiting time, we propose a wait-time compression scheme (WCS) based on model early-exit. The WCS selects early exit points for computing tasks based on the user-tolerated latency and reduces the waiting time without scheduling computing tasks or edge devices. Furthermore, to optimize the use of idle computational resources of edge devices during the waiting process, we introduce an idle-resource-based computing task scheduling scheme (ICTS). A significant number of experiments demonstrate that compared to other computing tasks scheduling schemes, WCS achieves a latency reduction of up to 10.7%, while ICTS enhances computational resource utilization by as much as 28.8%. Xin Niu 0001, Wang Chen 0004, Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 4 |
| 2026 | Puffer: A Serverless Platform Based on Vertical Memory ScalingabstractThis paper quantitatively analyses the potential of vertical scaling MicroVMs in serverless computing. Our analysis shows that under real-world serverless workloads, vertical scaling can significantly improve execution performance and resource utilization. However, we also find that the memory scaling of MicroVMs is the bottleneck that hinders vertical scaling from reaching the performance ceiling. We propose Faascale, a novel mechanism that efficiently scales the memory of MicroVMs for serverless applications. Faascale employs a series of techniques to tackle this bottleneck: 1) it sizes up/down the memory for a MicroVM by blocks that bind with a function instance instead of general pages; and 2) it pre-populates physical memory for function instances to reduce the delays introduced by the lazy-population. Compared with existing memory scaling mechanisms, Faascale improves the memory scaling efficiency by 2 to 3 orders of magnitude. Based on Faascale, we realize a serverless platform, named Puffer. Experiments conducted on eight serverless benchmark functions demonstrate that compared with horizontal scaling strategies, Puffer reduces time for cold-starting MicroVMs by 89.01%, improves memory utilization by 17.66%, and decreases functions execution time by 23.93% on average. Hao Fan 0006, Kun Wang 0059, Haibo Mi, Song Wu 0001, Chen Yu 0003 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2026 | Efficient Cluster-Based Knowledge Distillation for Deep Face RecognitionabstractKnowledge distillation has been widely used to improve the performance of small compact models for face recognition. However, selecting key knowledge and effectively transferring it from teacher to student remains a challenging problem. In this work, we propose an efficient Cluster-based Knowledge Distillation (CKD) dedicated to aligning the student model with the teacher model in terms of both sample relations and class centers. Specifically, CKD first determines the key sample relations based on the similarities between the sample features extracted by the teacher and their cluster centers generated by existing clustering algorithms. Then, CKD effectively transfers the knowledge of the above relations from the teacher to the student by designing a cluster-based relation distillation loss. Finally, CKD further improves the quality of the student's class centers by constructing a center loss between the above representative cluster centers and the student's class centers. We validate the proposed CKD on multiple face benchmarks. For example, CKD improves the baseline student performance from 91.95% to 94.20% on MegaFace and consistently outperforms recent competitive distillation methods on multiple benchmarks. These results demonstrate the effectiveness and superiority of CKD. Xianwei Lv 0001, Haibo Mi, Xin Niu 0001, Wang Chen 0004, Kun Wang 0059, Chen Yu 0003 |
IEEE Trans. Sustain. Comput. | 6 |
| 2025 | WAF: An Efficient WebAssembly-Based Execution Environment for User-Defined FunctionsabstractUser-Defined Functions (UDFs) have long served as the standard method for extending the capabilities of data management systems. With the advent of WebAssembly (WASM), UDFs' dependencies, such as language runtimes and libraries, can be compiled into a WASM module, which is then instantiated to execute the UDF. This approach offers several key advantages: 1) it allows developers to write UDFs in their preferred programming language, rather than being limited to those natively supported by the database engine; 2) it isolates UDFs' dependencies within the WASM module, mitigating the risk of errors caused by conflicting dependencies on the same host; and 3) it promotes cross-platform compatibility, enabling seamless execution of UDFs across different engines, operating systems, and architectures. However, our analysis reveals that executing a WASM-based UDF incurs overhead due to data transfer between the database engine and the WASM runtime. This process involves data copying and data layout adjustments, which can significantly impact performance. To address these challenges, we present WAF, a WASM-based UDF execution environment. WAF leverages shared memory to eliminate data copying and shifts data layout adjustments from the execution phase to the compilation phase. Experimental results show that WAF reduces the execution overhead of WASM-based UDFs by 3.1x and achieves an 18.1x speedup compared to the container-based approach, eliminating nearly all data transfer delays. Hao Fan 0006, Junhui Peng, Song Wu 0001, Chen Yu 0003, Hai Jin 0001, Wei Yang 0013 |
ICDE | 6 |
| 2025 | It Takes Two to Tango: Serverless Workflow Serving via Bilaterally Engaged Resource AdaptationabstractServerless platforms typically adopt an earlybinding approach for function sizing, requiring developers to specify an immutable size for each function within a workflow beforehand. Accounting for potential runtime variability, developers must size functions for worst-case scenarios to ensure service-level objectives (SLOs), resulting in significant resource inefficiency. To address this issue, we propose Janus, a novel resource adaptation framework for serverless platforms. Janus employs a late-binding approach, allowing function sizes to be dynamically adapted based on runtime conditions. The main challenge lies in the information barrier between the developer and the provider: developers lack access to runtime information, while providers lack domain knowledge about the workflow. To bridge this gap, Janus allows developers to provide hints containing rules and options for resource adaptation. Providers then follow these hints to dynamically adjust resource allocation at runtime based on real-time function execution information, ensuring compliance with SLOs. We implement Janus and conduct extensive experiments with real-world serverless workflows. Our results demonstrate that Janus enhances resource efficiency by up to 34.7% compared to the state-of-the-art. Jing Wu 0024, Lin Wang 0015, Quanfeng Deng, Chen Yu 0003, Bingheng Yan, Fangming Liu |
IPDPS | 4 |
| 2025 | Falcon: Advancing Asynchronous BFT Consensus for Lower Latency and Enhanced ThroughputabstractAsynchronous Byzantine Fault Tolerant (BFT) consensus protocols have garnered significant attention with the rise of blockchain technology. A typical asynchronous protocol is designed by executing sequential instances of the Asynchronous Common Sub-seQuence (ACSQ). The ACSQ protocol consists of two primary components: the Asynchronous Common Subset (ACS) protocol and a block sorting mechanism, with the ACS protocol comprising two stages: broadcast and agreement. However, current protocols encounter three critical issues: high latency arising from the execution of the agreement stage, latency instability due to the integral-sorting mechanism, and reduced throughput caused by block discarding. To address these issues, we propose Falcon, an asynchronous BFT protocol that achieves low latency and enhanced throughput. Falcon introduces a novel broadcast protocol, Graded Broadcast (GBC), which enables a block to be included in the ACS set directly, bypassing the agreement stage and thereby reducing latency. To ensure safety, Falcon incorporates a new binary agreement protocol called Asymmetrical Asynchronous Binary Agreement (AABA), designed to complement GBC. Additionally, Falcon employs a partial-sorting mechanism, allowing continuous rather than simultaneous block committing, enhancing latency stability. Finally, we incorporate an agreement trigger that, before its activation, enables nodes to wait for more blocks to be delivered and committed, thereby boosting throughput. We conduct a series of experiments to evaluate Falcon, demonstrating its superior performance. Xiaohai Dai, Chaozheng Ding, Wei Li 0058, Jiang Xiao 0001, Chen Yu 0003, Albert Y. Zomaya, Hai Jin 0001 |
Proc. VLDB Endow. | 6 |
| 2025 | Pako: Multi-Valued Byzantine Agreement Comparable to Partially-Synchronous BFTabstractAsynchronousByzantine Fault Tolerance(BFT) consensus protocols are gaining attention for their resilience against network attacks. Among them,Multi-valued Byzantine Agreement(MVBA) protocols play a critical role, which accepts input values from each replica and returns a consistent output. The state-of-the-art MVBA protocol, sMVBA, has a good-case latency of$6\delta$and an expected bad-case latency of$12\delta$, with$\delta$representing the network delay. Additionally, sMVBA exhibits a communication of$O(n^{2})$in both good and bad cases. Although it outperforms other MVBA protocols, sMVBA still lags behind partially-synchronous counterparts. For instance, PBFT achieves a good-case latency of$3\delta$, and HotStuff boasts a good-case communication of$O(n)$. This paper introduces a novel MVBA protocol, Pako, aiming for performance comparable to partially-synchronous protocols. Pako leverages an existing MVBA protocol as a black box and introduces an additional view with an optimistic path to commit values efficiently. Two Pako variants, Pako1 and Pako2, provide a trade-off between latency and communication. To be more specific, Pako1 achieves a good-case latency of$3\delta$with$O(n^{2})$communication, while Pako2 reduces the communication to$O(n)$with a slightly higher good-case latency of$5\delta$. A series of experiments demonstrate Pako's significant outperformance of counterparts. Xiaohai Dai, Zhengxuan Guo, Jiang Xiao 0001, Guanxiong Wang, Yifei Liang, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 6 |
| 2025 | CBuild: Cluster-Oriented Collaborative Image Building for ContainersabstractStarting a container needs to build a container image layer-by-layer if the required image is not available. However, the image building involves downloading a large amount of data, which significantly delays the development and deployment of containerized services. To reduce data downloads and accelerate image building, current methods typically focus on improving data sharing through reconstructing images. Unfortunately, these approaches show limited performance improvement in clusters as they only improve data sharing on a single node. In this paper, we find that there are significant duplicated remote file downloads between nodes in a cluster. Accordingly, we propose cBuild, a distributed file cache to minimize costly image data downloads in cluster environments. Specifically, to enable inter-node image data sharing, cBuild designs a non-intrusive interception mechanism based on network namespace, instead of directly detecting building instructions that dirty images. Based on the distribution characteristics of duplicated files in layers, cBuild places image files among nodes in a balanced manner to prevent transfer bottlenecks caused by hotspot nodes and employs a layer-aware searching strategy to quickly locate the desired files. We implement cBuild on the basis of Docker. Experiments show that cBuild improves building speed by up to 15.3 × and reduces the data downloading by 80%. Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 5 |
| 2025 | Computing Tasks Saving Schemes Through Early Exit in Edge Intelligence-Assisted SystemsabstractEdge intelligence (EI) is a promising paradigm where end devices collaborate with edge servers to provide artificial intelligence services to users. In most realistic scenarios, end devices often move unconsciously, resulting in frequent computing migrations. Moreover, a surge in computing tasks offloaded to edge servers significantly prolongs queuing latency. These two issues obstruct the timely completion of computing tasks in EI-assisted systems. In this paper, we formulate an optimization problem aiming to maximize computing task completion under latency constraints. To address this issue, we first categorize computing tasks into new computing tasks (NCTs) and partially completed computing tasks (PCTs). Subsequently, based on model partitioning, we design a new computing task saving scheme (NSS) to optimize early exit points for NCTs and computing tasks in the queuing queue. Furthermore, we propose a partially completed computing task saving scheme (PSS) to set early exit points for PCTs during computing migrations. Numerous experiments show that computing saving schemes can achieve at least 90% computing task completion rate and up to 61.81% latency reduction compared to other methods. Xin Niu 0001, Xianwei Lv 0001, Wang Chen 0004, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Efficient Distributed Sparse Relative Similarity LearningabstractLearning a good similarity measure for large-scale high-dimensional data is a crucial task in machine learning applications, yet it poses a significant challenge. Distributed minibatch Stochastic Gradient Descent (SGD) serves as an efficient optimization method in large-scale distributed training, allowing linear speedup in proportion to the number of workers. However, communication efficiency in distributed SGD requires a sufficiently large minibatch size, presenting two distinct challenges. Firstly, a large minibatch size leads to high memory usage and computational complexity during parallel training of high-dimensional models. Second, a larger batch size of data reduces the convergence rate. To overcome these challenges, we propose an Efficient Distributed Sparse Relative Similarity Learning ( \(\mathbf{\mathsf{EDSRSL}}\) ) framework. This framework integrates two strategies: local minibatch SGD and sparse relative similarity learning. By effectively reducing the number of updates through synchronous delay while maintaining a large batch size, we address the issue of high computational cost. Additionally, we incorporate sparse model learning into the training process, significantly reducing computational cost. This article also provides theoretical proof that the convergence rate does not decrease significantly with increasing batch size. Various experiments on six high-dimensional real-world datasets demonstrate the efficacy and efficiency of the proposed algorithms, with a communication cost reduction of up to \(90.89\%\) and a maximum wall time speedup of \(5.66\times\) compared to the baseline methods. Dezhong Yao 0002, Sanmu Li, Peilin Zhao, Chen Yu 0003, Hai Jin 0001 |
ACM Trans. Knowl. Discov. Data | 6 |
| 2025 | End-to-End Steady-State Adaptive Slicing Method for Dynamic Network State and LoadabstractNetwork slicing has become a primary function of 5G/6G network resource management. However, the existing slicing schemes have not sufficiently discussed the reconfiguration optimization schemes brought by user behavior changes and mobile network environment fluctuations, leading to excessive service interruption rates and slice reconfiguration costs in dynamic environments. To address this problem, this paper proposes an End-to-end Steady-state Adaptive slicing method for Dynamic network state and load (ESAD). To realize the steady-state slicing decisions, ESAD takes the steady-state degree of network slicing and reconfiguration cost as the objective and constructs the slicing reconfiguration probability evaluation function based on the service load dynamics function and the time-varying function of the network channel conditions. To improve the predictability and steady-state degree of the slicing decision, ESAD introduces an ensemble deep learning method to predict the load service fluctuation based on the user behavior model and employs reinforcement learning to compute the channel dynamics boundary, which guides the slicing decision to balance the network dynamics factors. Experiments on quality of service assurance for 5G cloud game rendering class prove that ESAD can reduce reconfiguration probability and long-term reconfiguration cost by 49.45%–58.50% while improving system QoS assurance and capacity. Boyi Tang, Yijun Mo, Chen Yu 0003 |
IEEE Trans. Mob. Comput. | 3 |
| 2025 | KubeSPT: Stateful Pod Teleportation for Service Resilience With Live MigrationabstractContainer orchestration systems, such as Kubernetes, streamline containerized application deployment. As more and more applications are being deployed in Kubernetes, there is an increasing need for rescheduling - relocating a running pod to different nodes - due to system upgrades, node failures, and load-balancing optimizations. Live migration, which transfers services from source nodes to target nodes with minimal downtime, is the ideal support for rescheduling. However, implementing live migration for pods that run stateful services is challenging, because Kubernetes manages pods as stateless. First, the current pod's network namespace initialization process causes a mismatch in the network state between the migrated pod and internal containers. Second, migrating the memory state results in extended downtime. Third, Kubernetes operations on pods do not consider preserving the state of the pods. Therefore, we propose KubeSPT to achieve live migration of stateful pods in rescheduling scenarios. Firstly, we synchronize the network state of pods and internal containers by controlling packet flow and implement fast service redirection. Secondly, we introduce a Hot Data and Lazy-Restore method for memory restoration to reduce migration downtime. Finally, we decouple pod migration operations from other Kubernetes operations to ensure compatibility with live migration. Experimental results show that KubeSPT reduces downtime by 86%-93% compared to current rescheduling methods. Hansheng Zhang, Song Wu 0001, Hao Fan 0006, Weibin Xue, Chen Yu 0003, Shadi Ibrahim, Hai Jin 0001 |
IEEE Trans. Serv. Comput. | 6 |
| 2025 | Working Smarter Not Harder: Hybrid Cooling for Deep Learning in Edge DatacentersabstractThe proliferation of deep-learning-based mobile and IoT applications has driven the increasing deployment of edge datacenters equipped with domain-specific accelerators. The unprecedented computing power offered by these accelerators puts a heavy burden on the cooling system, motivating more potent cooling techniques like cold water cooling. However, we observe that cold water cooling results in significant energy waste in edge datacenters due to the fluctuating resource utilization both spatially and temporally. To tackle this issue, we propose the concept of “working smarter” by slowing down accelerators deliberately whenever possible and enabling warm water cooling during these times to achieve cooling efficiency. Based on this concept, we develop Hyco—a hybrid water cooling system tailored for edge datacenters running deep learning workloads. First, Hyco features a zone-based cooling architecture enabling dynamic switching between cold water and warm water cooling. Then, based on a lightweight latency estimation method, Hyco incorporates a learning-based scheduling scheme to determine “which” accelerator workers and “when” to slow down through an adaptive and intelligent power-latency trade-off for deep learning models. The simulation with real-world traces shows that Hyco reduces the cooling energy consumption by up to 34.74× while satisfying latency constraints more than 99% of the time for deep-learning-based applications. Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu |
IEEE Trans. Sustain. Comput. | 7 |
| 2025 | RLA: A Low-Latency and High-Smoothness Path Planning System Based on Interpolation and Velocity ControlabstractLocal planning is a key issue in the field of unmanned delivery. Unmanned delivery requires high delivery efficiency and lower equipment maintenance costs, which pose challenges to the latency and smoothness of path planning algorithms. After investigation of the work on optimizing the latency and smoothness of local planning, we proposed a Robotic-Look-Ahead approach based on Look Ahead approach. It consists of four parts: calculating the conjunction speed, circular arc interpolation, the Look-Ahead method, and path modification. The experiment showed that with different paths, different running memory, and different maximum running speeds, latency decreased by an average of 90% compared to the benchmark, and smoothness improved by an average of 40%. Under different loads, the average energy consumption decreases by 4%. Xupeng Zhu, Xin Niu 0001, Wang Chen 0004, Chen Yu 0003 |
IEEE Trans. Sustain. Comput. | 4 |
| 2024 | InferCool: Enhancing AI Inference Cooling through Transparent, Non-Intrusive Task ReassignmentabstractThe increasing power consumption of AI inference in modern datacenters has escalated cooling demands significantly, necessitating the adoption of potent cooling approaches like water cooling. Unlike traditional cloud workloads, AI inference has unique characteristics that create substantial gaps in achieving optimal cooling efficiency. In this work, we present the first comprehensive measurement study of AI inference cooling across various models within an industrial-ready scheduling framework, highlighting significant inefficiencies and their causes. To fill the gap while following the fundamental requirements of cooling systems, we explore a new opportunity presented by modern Multi-Instance GPU-enabled inference serving, where the scheduling dimension is naturally orthogonal to the cooling dimension. Building on this insight, we develop InferCool, a cooling middleware designed to enhance cooling efficiency for inference serving through transparent, non-intrusive task reassignment. It includes a streamlined power and temperature prediction approach and a thermal-aware, adaptive application deployment and request scheduling mechanism. Real-world experiments on a water-cooled testbed and a three-node cluster demonstrate that InferCool can reduce the maximum GPU temperature by 5°C across eight A100 GPUs, equivalent to cooling energy savings of about 20%. Importantly, InferCool requires no modifications to existing cooling infrastructures and is compatible with existing scheduling systems. Qiangyu Pei, Lin Wang 0015, Bingheng Yan, Chen Yu 0003, Fangming Liu |
SoCC | 5 |
| 2024 | Harnessing the Power of Large Language Model for Uncertainty Aware Graph ProcessingabstractHandling graph data is one of the most difficult tasks. Traditional techniques, such as those based on geometry and matrix factorization, rely on assumptions about the data relations that become inadequate when handling large and complex graph data. On the other hand, deep learning approaches demonstrate promising results in handling large graph data, but they often fall short of providing interpretable explanations. To equip the graph processing with both high accuracy and explainability, we introduce a novel approach that harnesses the power of a large language model (LLM), enhanced by an uncertainty-aware module to provide a confidence score on the generated answer. We experiment with our approach on two graph processing tasks: few-shot knowledge graph completion and graph classification. Our results demonstrate that through parameter efficient fine-tuning, the LLM surpasses state-of-the-art algorithms by a substantial margin across ten diverse benchmark datasets. Moreover, to address the challenge of explainability, we propose an uncertainty estimation based on perturbation, along with a calibration scheme to quantify the confidence scores of the generated answers. Our confidence measure achieves an AUC of 0.8 or higher on seven out of the ten datasets in predicting the correctness of the answer generated by LLM. Yiming Qian, Yuting Song, Hai Jin 0001, Chen Yu 0003 |
LREC/COLING | 6 |
| 2024 | Precise control of page cache for containers
Kun Wang 0005, Song Wu 0001, Shengbang Li, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001 |
Frontiers Comput. Sci. | 6 |
| 2024 | QoS-pro: A QoS-enhanced Transaction Processing Framework for Shared SSDsabstractSolid State Drives (SSDs) are widely used in data-intensive scenarios due to their high performance and decreasing cost. However, in shared environments, concurrent workloads can interfere with each other, leading to a violation of Quality of Service (QoS). While QoS mechanisms like fairness guarantees and latency constraints have been integrated into SSDs, existing transaction processing frameworks offer limited QoS guarantees and can significantly degrade overall performance in a shared environment. The reason is that the internal components of an SSD, originally designed to exploit parallelism, struggle to coordinate effectively when QoS mechanisms are applied to them. This article proposes a novel QoS -enhanced transaction pro cessing framework, called QoS-pro, which enhances QoS guarantees for concurrent workloads while maintaining high parallelism for SSDs. QoS-pro achieves this by redesigning transaction processing procedures to fully exploit the parallelism of shared SSDs and enhancing QoS-oriented transaction translation and scheduling with parallelism features in mind. In terms of fairness guarantees, QoS-pro outperforms state-of-the-art methods by achieving 96% fairness improvement and 64% maximum latency reduction. QoS-pro also shows almost no loss in throughput when compared with parallelism-oriented methods. Additionally, QoS-pro triggers the fewest Garbage Collection (GC) operations and minimally affects concurrently running workloads during GC operations. Hao Fan 0006, Yiliang Ye, Shadi Ibrahim, Xingru Li, Weibin Xue, Song Wu 0001, Chen Yu 0003, Xuanhua Shi, Hai Jin 0001 |
ACM Trans. Archit. Code Optim. | 8 |
| 2024 | vKernel: Enhancing Container Isolation via Private Code and DataabstractContainer technology is increasingly adopted in cloud environments. However, the lack of isolation in the shared kernel becomes a significant barrier to the wide adoption of containers. The challenges lie in how to simultaneously attain high performance and isolation. On the one hand, kernel-level isolation mechanisms, such asseccomp,capabilities, andapparmor, achieve good performance without much overhead, but lack the support for per-container customization. On the other hand, user-level and VM-based isolation offer superior security guarantees and allow for customization since a container is assigned a dedicated kernel, however, at the cost of high overhead. We presentvKernel, a kernel isolation framework. It maintains a minimal set of code and data that are either sensitive or are prone to interference in a virtual kernel instance (vKI). vKernel relies on inline hooks to intercept and redirect requests sent to the host kernel to a vKI, where container-specific security rules, functions, and data are implemented. Through case studies, we demonstrate that under vKernel user-defined data isolation and kernel customization can be supported with a reasonable engineering effort. An evaluation of vKernel with micro-benchmarks, cloud services, real-world applications show that vKernel achieves good security guarantees, but with much less overhead. Hang Huang, Jia Rao, Song Wu 0001, Hao Fan 0006, Chen Yu 0003, Hai Jin 0001, Kun Suo, Lisong Pan |
IEEE Trans. Computers | 6 |
| 2024 | Multi-Grained Trace Collection, Analysis, and Management of Diverse Container ImagesabstractContainer technology is getting popular in cloud environments due to its lightweight feature and convenient deployment. Container Registry plays a critical role in container-based clouds, as many container startups involve downloading layer-structured container images from Container Registry. However, Container Registry is struggling to efficiently manage images (i.e., transfer and store) with the emergence of diverse services and new image formats. The reason is that Container Registry manages images uniformly at layer granularity. On the one hand, such uniform layer-level management probably cannot fit the various requirements of different kinds of containerized services well. On the other hand, new image formats organizing data in blocks or files cannot benefit from such uniform layer-level image management. In this paper, we perform the first analysis of image traces at multiple granularities (i.e., image-, layer-, and file-level) for various services and provide an in-depth comparison of different image formats. The traces were collected from a production-level Container Registry, amounting to 24 million requests and involving more than 184 TB of transferred data. We provide a number of valuable insights, including request patterns of services, file-level access patterns, and bottlenecks associated with different image formats. Based on these insights, we propose two optimizations to improve image transfer. Both the traces and toolkit for trace collection will be open-sourced. Qi Zhang 0009, Hao Fan 0006, Song Wu 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 5 |
| 2024 | Game-Based Adaptive FLOPs and Partition Point Decision Mechanism With Latency and Energy-Efficient Tradeoff for Edge IntelligenceabstractAs the product of the combination of edge computing and artificial intelligence, edge intelligence (EI) not only solves the problem of insufficient computing capacity of the end device, but also can provide users with various types of intelligent services. However, offline and online model partitioning methods respectively have problems of poor adaptability to the real computing environment and delayed feedback. In addition, previous work on optimizing energy consumption through model partitioning often ignores the latency of intelligent services. Similarly, the energy consumption of end devices and edge servers is usually not considered when optimizing latency. Therefore, we propose game-based adaptive floating-point operations and partition point decision mechanism (GAFPD) to efficiently find the optimal partition point that reduces latency and improves energy efficiency simultaneously in a dynamically changing computing environment. Numerous simulation experiments and robot-based EI system experiments show that GAFPD can simultaneously reduce the latency of intelligent services and improve the energy efficiency of edge devices, while exhibiting strong adaptability to bandwidth changes. Xin Niu 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 4 |
| 2024 | FedGKD: Toward Heterogeneous Federated Learning via Global Knowledge DistillationabstractFederated learning, as one enabling technology of edge intelligence, has gained substantial attention due to its efficacy in training deep learning models without data privacy and network bandwidth concerns. However, due to the heterogeneity of the edge computing system and data, many methods suffer from the“client-drift”issue that could considerably impede the convergence of global model training: local models on clients can drift apart, and the aggregated model can be different from the global optimum. To tackle this issue, one intuitive idea is to guide the local model training by global teachers,i.e.,past global models, where each client learns the global knowledge from past global models via adaptive knowledge distillation techniques. Inspired by these insights, we propose a novel approach for heterogeneous federated learning,FedGKD, which fuses the knowledge from historical global models and guides local training to alleviate the“client-drift”issue. In this paper, we evaluateFedGKDthrough extensive experiments across various CV and NLP datasets (i.e.,CIFAR-10/100, Tiny-ImageNet, AG News, SST5) under different heterogeneous settings. The proposed method is guaranteed to converge under common assumptions and outperforms the state-of-the-art baselines in the non-IID federated setting. Dezhong Yao 0002, Wanning Pan, Yutong Dai 0002, Yao Wan 0001, Xiaofeng Ding 0001, Chen Yu 0003, Hai Jin 0001, Zheng Xu 0002, Lichao Sun 0001 |
IEEE Trans. Computers | 6 |
| 2024 | Un-IOV: Achieving Bare-Metal Level I/O Virtualization Performance for Cloud Usage With Migratability, Scalability and TransparencyabstractI/O virtualization is utilized by cloud platforms to provide tenants with efficient, scalable, and manageable network and storage services. The de-facto industrial standard, paravirtualization, offers rich cloud functionality by introducing split front-end and back-end drivers in the guest and host operating systems, respectively. Given this fact, paravirtualization incurs host inefficiency and performance overhead. Thus, emerging hardware virtio accelerators (i.e., SRIOV-capable devices that conform to virtio specification) with device passthrough technologies mitigate the performance issue. However, adopting these devices presents the challenge of insufficient support for live migration.This paper proposes Un-IOV, a novel I/O virtualization system that simultaneously achieves bare-metal level I/O performance and migratability. The key idea is to develop a new hybrid virtualization stack with: (1) a host-bypassed direct data path for virtio accelerators, and (2) a relayed control path guaranteeing seamless live migration support. Un-IOV achieves high scalability by consuming minimum host resources. Extensive experiment results demonstrate that Un-IOV achieves superior network and storage virtualization performance than software implementations with comparable performance of direct passthrough I/O virtualization, while imposing zero guest modification (i.e., guest transparency). Zongpu Zhang, Chenbo Xia, Cunming Liang, Jian Li 0021, Chen Yu 0003, Tiwei Bie, Roberts Martin, Dan Daly, Xiao Wang 0084, Haibing Guan |
IEEE Trans. Computers | 5 |
| 2023 | Graph-Reinforcement-Learning-Based Task Offloading for Multiaccess Edge ComputingabstractNetwork applications involve massive heterogeneous data fusion and analysis. Artificial intelligence can significantly improve the convenience and user experience, but it requires a lot of storage, bandwidth, and computing resources. Multiaccess edge computing (MEC) extends intelligence services to IoT devices through offloading approaches and joint processing, which solves the resource bottleneck. However, designing advanced collaboration technology to offload tasks to MEC servers is still challenging. Heuristic algorithms and deep reinforcement learning (DRL)-based approaches have been proposed to offload tasks and minimize application latency. However, heuristic algorithms heavily depend on accurate mathematical models for the MEC system, and DRL does not make fair use of the relationship between devices in the MEC graph. To solve this, we propose a task offloading mechanism based on graph neural network (GNN), which can directly learn on graph data with messages passing and aggregation. We propose a graph reinforcement learning-based offloading (GRLO) framework, which models MEC as an acyclic graph and the offloading policy by graph state migration. GRLO combines GNN with the actor-critic network and trains offloading decision makers without labels. To efficiently train the GRLO, we propose a method that quickly explores action space and approaches the optimal solution. The numerical results show that the GRLO has lower latency compared to baselines while having generalization ability to new environments and topologies. Moreover, we verified the effectiveness of GRLO on a prototype. Zhenchuan Sun, Yijun Mo, Chen Yu 0003 |
IEEE Internet Things J. | 3 |
| 2023 | A Feedback-Driven DNN Inference Acceleration System for Edge-Assisted Video AnalyticsabstractWith the proposal of edge computing, lots of intelligence applications have made significant progress. For enormous video analysis, how to further accelerate the process is still a major challenge. To overcome the challenge, researchers propose various video frame filtering systems to reduce the data transmission. In this work, we propose a Feedback-Driven DNN Inference Acceleration system (FDDIA). FDDIA is committed to further reducing the latency of DNN inference and the transmission according to the feedback information. Specifically, on the device side, FDDIA first uses the detection results of prior frames as feedback to determine the key candidate regions and uses the inter-frame difference to determine the new object regions. Those regions containing large objects are down-sampled to further reduce the frame. On the edge side, FDDIA first integrates all candidate regions into a smaller new image. Then it remaps the detection result for the new image back to the original frame and returns the result to the device. We evaluate FDDIA on different video benchmarks for three object detection tasks. The results show FDDIA improves the average end-to-end latency by 44% and average bandwidth usage by 41% than the existing advanced method while maintaining a high accuracy. Xianwei Lv 0001, Qianqian Wang 0017, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 3 |
| 2023 | EventTube: An Artificial Intelligent Edge Computing Based Event Aware System to Collaborate With Individual Devices in Logistics SystemsabstractArtificial intelligence has been adopted to facilitate monitoring, operation, and decision in the logistics field. Logistics robots with environment perception capability have been used to improve warehousing efficiency in logistics systems. However, autonomous mobile robots face computationally intensive and real-time demanding tasks such as navigation, localization, and obstacle avoidance. In this article, we present EventTube, an edge computing based event-aware system that can efficiently discover events from the video data captured by RGB-Monoculars and collaborate with individual devices to make timely decisions. EventTube deploys a semantic context extraction pipeline on edge servers to aggregate video streams from mobile robots and feed a few keyframes, including the start and end of the specific events to the successive perception pods, accelerating logistics robots’ response speed. The event-related model parameters are trained and updated online on a server. The video data collected at the warehouse site for our mobile robots show that EventTube significantly improves parcel delivery efficiency without affecting regular deliveries. Yijun Mo, Zhenchuan Sun, Chen Yu 0003 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Decentralized Online Learning: Take Benefits from Others' Data without Sharing Your Own to Track Global TrendabstractDecentralized online learning (online learning in decentralized networks) has been attracting more and more attention, since it is believed that decentralized online learning can help data providers cooperatively better solve their online problems without sharing their private data to a third party or other providers. Typically, the cooperation is achieved by letting the data providers exchange their models between neighbors, e.g., recommendation model. However, the best regret bound for a decentralized online learning algorithm is 𝒪( n √ T ), where n is the number of nodes (or users) and T is the number of iterations. This is clearly insignificant, since this bound can be achieved without any communication in the networks. This reminds us to ask a fundamental question: Can people really get benefit from the decentralized online learning by exchanging information? In this article, we studied when and why the communication can help the decentralized online learning to reduce the regret. Specifically, each loss function is characterized by two components: the adversarial component and the stochastic component. Under this characterization, we show that decentralized online gradient enjoys a regret bound \( {\mathcal {O}(\sqrt {n^2TG^2 + n T \sigma ^2})} \) , where G measures the magnitude of the adversarial component in the private data (or equivalently the local loss function) and σ measures the randomness within the private data. This regret suggests that people can get benefits from the randomness in the private data by exchanging private information. Another important contribution of this article is to consider the dynamic regret—a more practical regret to track users’ interest dynamics. Empirical studies are also conducted to validate our analysis. Wendi Wu, Zongren Li, Chen Yu 0003, Peilin Zhao, Ji Liu 0002, Kunlun He |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2023 | Personalized Edge Intelligence via Federated Self-Knowledge DistillationabstractFederated Learning(FL) is an emerging approach in edge computing for collaboratively training machine learning models among multiple devices, which aims to address limited bandwidth, system heterogeneity, and privacy issues in traditional centralized training. However, the existing federated learning methods focus on learning a shared global model for all devices, which may not always be ideal for different devices. Such situations become even worse when each edge device has its own data distribution or task. In this paper, we study personalized federated learning in which our goal is to train models to perform well for individual clients. We observe that the initialization in each communication round causes the forgetting of historical personalized knowledge. Based on this observation, we propose a novelPersonalized Federated Learning(PFL) framework via self-knowledge distillation, named pFedSD. By allowing clients to distill the knowledge of previous personalized models to current local models, pFedSD accelerates the process of recalling the personalized knowledge for the latest initialized clients. Moreover, self-knowledge distillation provides different views of data in feature space to realize an implicit ensemble of local models. Extensive experiments on various datasets and settings demonstrate the effectiveness and robustness of pFedSD. Hai Jin 0001, Dongshan Bai, Dezhong Yao 0002, Yutong Dai 0002, Lin Gu 0002, Chen Yu 0003, Lichao Sun 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Container lifecycle-aware scheduling for serverless computingabstractAbstract Elastic scaling in response to changes on demand is a main benefit of serverless computing. When bursty workloads arrive, a serverless platform launches many new containers and initializes function environments (known as cold starts), which incurs significant startup latency. To reduce cold starts, platforms usually pause a container after it serves a request, and reuse this container for subsequent requests. However, this reuse strategy cannot efficiently reduce cold starts because the schedulers are agnostic of container lifecycle. For example, it may ignore soon available containers or evict soon needed containers. We propose a container lifecycle‐aware scheduling strategy for serverless computing, CAS. The key idea is to control distribution of requests and determine creation or eviction of containers according to different lifecycle phases of containers. We implement a prototype of CAS on OpenWhisk. Our evaluation shows that CAS reduces 81% cold starts and therefore brings a 63% reduction at 95th percentile latency compared with native scheduling strategy in OpenWhisk when there is worker contention between workloads, and does not add significant performance overhead. Song Wu 0001, Zhiheng Tao, Hao Fan 0006, Hai Jin 0001, Chen Yu 0003, Chun Cao |
Softw. Pract. Exp. | 7 |
| 2022 | CRSM: Computation Reloading Driven by Spatial-Temporal Mobility in Edge-Assisted Automated Industrial Cyber-Physical SystemsabstractEdge Computing, as an emerging computing mode, transfers computing capacity from the cloud center to the edge of new generation automation sensor networks. However, the surge in the number of sensor devices has resulted in the edge server overload. Fortunately, with the development of hardware technology, the computing capacity of sensor devices has improved. Therefore, we introduce a new concept of computation reloading, that is, resource-rich servers allocate tasks to sensor devices with stronger computing capacity. Most previous real-time EC researches have not considered sensor devices’ higher spatial–temporal mobility, which results in continuous computation migration and high delay. In this article, we expose a computation reloading scheme driven by spatial–temporal mobility (CRSM), for single time slice and multiple time slices cases, computation tasks are allocated in advance by predicting sensor devices locations. The experiments demonstrate that our scheme outperforms other methods in terms of shorting industrial cyber-physical systems delays. Xin Niu 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Ind. Informatics | 2 |
| 2022 | HQ2CL: A High-Quality Class Center Learning System for Deep Face RecognitionabstractBenefited from the proposals of function losses margin-based, face recognition has achieved significant improvements in recent years. Those losses aim to increase the margin between the different identities to enhance the discriminability. Ideally, the class center of different identities is far from each other, and face samples are compact around the corresponding class center. Hence, it's very vital to produce a high-quality class center. However, the distribution of training sets determines the class center. With low-quality samples being in the majority, the class center would be close to the samples with little identity information. As a result, it would impair the discriminability of the learned model for those unseen samples. In this work, we propose a High-Quality Class Center Learning system (HQ2CL). This is an effective system and guides the class center to approach the high-quality samples to keep the discriminability. Specifically, HQ2CL introduces a quality-aware scale and margin layer for the identification loss and constructs a new high-quality center loss. We implement the proposed system without additional burden. And we present the experimental evaluation over different face benchmarks. The experimental results show the superiority of our proposed HQ2CL over the state-of-the-arts. Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | Cost Efficient Sensor Positions Determination For Human Activity RecognitionabstractHuman activity recognition(HAR) is one of the most active topics in the field of ubiquitous computing. Multi-sensor based HAR has attracted extensive interest because of its high recognition accuracy. Correspondingly, the number of body sensors in terms of hardware cost, the overload on communications, the storage and the computational complexity will be very high. In this paper, we propose a novel approach which can optimize cost-efficient sensors positions to save all the costs while maintaining high recognition performance. The tradeoff among the sensor positions, the target category, and the redundancy is considered. We also propose a data set D that contains acceleration sensor data for seventeen positions of the human body by simulating the worker actions in a factory assembly lines. The experimental results show that only six sensors can maintain high activity recognition accuracy out of seventeen sensors by using the proposed method. Xianwei Lv 0001, Chen Yu 0003, Hai Jin 0001, Ruiguo Zhang |
IEEE Trans. Sustain. Comput. | 2 |
| 2021 | Accelerating Parallel Applications in Cloud Platforms via Adaptive Time-Slice ControlabstractCloud platforms can provide flexible and cost-effective environments for parallel applications. However, the resource over-commitment issues, i.e., cloud providers often provide much more executable virtual CPUs than available physical CPUs, still impede the synchronization operations of parallel applications, causing severe performance degradation. Existing methods optimize parallel applications by promoting the priorities of involved VMs. They cannot fully explore the performance of parallel applications, because they ignore the time-slice requirements of different phases of parallel applications. Furthermore, non-parallel applications experience unsatisfied performance because of low scheduling priorities. Given empirical analysis on time-slices of virtual machines (VMs), we find that shortening time-slices can mitigate synchronization overhead which incurs during communication phases, while over-short time-slices cause frequent cache misses in computation phases. Accordingly, we propose an Adaptive Time-slice Control (ATC) mechanism. ATC first detects the phases of parallel applications based on lock latency or cache misses. Then, ATC shortens time-slices during communication phases and prolongs time-slices during computation phases for parallel applications, and sets a uniform time-slice for non-parallel applications. We evaluate ATC using seven well-known benchmarks with 25+ applications. Experiments show that ATC obtains 1.5-75× performance gain for running parallel applications than state-of-the-art solutions, with nearly unaffected impact on non-parallel applications. Hao Fan 0006, Song Wu 0001, Zhenjiang Xie, Sheng Di, Jiang Xiao 0001, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 7 |
| 2021 | A3C-DO: A Regional Resource Scheduling Framework Based on Deep Reinforcement Learning in Edge ScenarioabstractCurrently, huge amounts of data are produced by edge device. Considering the heavy burden of network bandwidth and the service delay requirements of delay-sensitive applications, processing the data at network edge is a great choice. However, edge devices such as smart wearables, connected and autonomous vehicles usually have several limitations on computational capacity and energy which will influence the quality of service. As an effective and efficient strategy, offloading is widely used to address this issue. But when facing device heterogeneity problem and task complexity increase, service quality degradation and resource utility decrease often occur due to unreasonable task distribution. Since conventional simplex offloading strategies show limited performance in complex environment, we are motivated to design a dynamic regional resource scheduling framework which is able to work effectively taking different indexes into consideration. Thus, in this article we first propose a double offloading framework to simulate the offloading process in real edge scenario which consists of different edge servers and devices. Then we formulate the offloading as a Markov Decision Process (MDP) and utilize a deep reinforcement learning (DRL) algorithm named asynchronous advantage actor-critic (A3C) as the offloading decision making strategy to balance the workload of edge servers and finally reduce the overhead in terms of energy and time. Comparison experiments for local computing and wide-used DRL algorithm DQN are conducted in a comprehensive benchmark and the results show that our work performs much better on self-adjusting and overhead reduction. Junfeng Zou, Tongbo Hao, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Computers | 3 |
| 2020 | A parameter-level parallel optimization algorithm for large-scale spatio-temporal data mining
Xuanhua Shi, Ligang He, Dongxiao Yu, Hai Jin 0001, Chen Yu 0003, Hulin Dai, Zezhao Feng |
Distributed Parallel Databases | 6 |
| 2019 | Evaluation model for business sites planning based on online and offline datasets
Chen Yu 0003, Hai Jin 0001 |
Future Gener. Comput. Syst. | 2 |
| 2019 | Using Crowdsourcing to Provide QoS for Mobile Cloud ComputingabstractQuality of cloud service (QoS) is one of the crucial factors for the success of cloud providers in mobile cloud computing. Context-awareness is a popular method for automatic awareness of the mobile environment and choosing the most suitable cloud provider. Lack of context information may harm the users' confidence in the application rendering it useless. Thus, mobile devices need to be constantly aware of the environment and to test the performance of each cloud provider, which is inefficient and wastes energy. Crowdsourcing is a considerable technology to discover and select cloud services in order to provide intelligent, efficient, and stable discovering of services for mobile users based on group choice. This article introduces a crowdsourcing-based QoS supported mobile cloud service framework that fulfills mobile users' satisfaction by sensing their context information and providing appropriate services to each of the users. Based on user's activity context, social context, service context, and device context, our framework dynamically adapts cloud service for the requests in different kinds of scenarios. The context-awareness based management approach efficiency achieves a reliable cloud service supported platform to supply the Quality of Service on mobile device. Dezhong Yao 0002, Chen Yu 0003, Laurence T. Yang, Hai Jin 0001 |
IEEE Trans. Cloud Comput. | 2 |
| 2017 | SSDUP: a traffic-aware ssd burst buffer for HPC systemsabstractMany high performance computing (HPC) applications are highly data intensive. Current HPC storage systems still use hard disk drives (HDDs) as their dominant storage devices, which suffer from disk head thrashing when accessing random data. New storage devices such as solid state drives (SSDs), which can handle random data access much more efficiently, have been widely deployed as the buffer to HDDs in many production HPC systems. Burst buffer has also been proposed to manage the SSD buffering of bursty write requests. Although burst buffer can improve I/O performance in many cases, we find that it has some limitations such as requiring large SSD capacity and harmonious overlapping between computation phase and data flushing stage. Xuanhua Shi, Wei Liu 0004, Hai Jin 0001, Chen Yu 0003, Yong Chen 0001 |
ICS | 5 |
| 2017 | A semi-supervised social relationships inferred model based on mobile phone data
Chen Yu 0003, Namin Wang, Laurence T. Yang, Dezhong Yao 0002, Ching-Hsien Hsu, Hai Jin 0001 |
Future Gener. Comput. Syst. | 1 |
| 2017 | Predicting Transportation Carbon Emission with Urban Big DataabstractTransportation carbon emission is a significant contributor to the increase of greenhouse gases, which directly threatens the change of climate and human health. Under the pressure of the environment, it is very important to master the information of transportation carbon emission in real time. In the traditional way, we get the information of the transportation carbon emission by calculating the combustion of fossil fuel in the transportation sector. However, it is very difficult to obtain the real-time and accurate fossil fuel combustion in the transportation field. In this paper, we predict the real-time and fine-grained transportation carbon emission information in the whole city, based on the spatio-temporal datasets we observed in the city, that is taxi GPS data, transportation carbon emission data, road networks, points of interests (POIs), and meteorological data. We propose a three-layer perceptron neural network (3-layerPNN) to learn the characteristics of collected data and infer the transportation carbon emission. We evaluate our method with extensive experiments based on five real data sources obtained in Zhuhai, China. The results show that our method has advantages over the well-known three machine learning methods (Gaussian Naive Bayes, Linear Regression, and Logistic Regression) and two deep learning methods (Stacked Denoising Autoencoder and Deep Belief Networks). Xiangyong Lu, Kaoru Ota, Mianxiong Dong, Chen Yu 0003, Hai Jin 0001 |
IEEE Trans. Sustain. Comput. | 4 |
| 2015 | Sparse Online Relative Similarity LearningabstractFor many data mining and machine learning tasks, the quality of a similarity measure is the key for their performance. To automatically find a good similarity measure from datasets, metric learning and similarity learning are proposed and studied extensively. Metric learning will learn a Mahalanobis distance based on positive semi-definite (PSD) matrix, to measure the distances between objectives, while similarity learning aims to directly learn a similarity function without PSD constraint so that it is more attractive. Most of the existing similarity learning algorithms are online similarity learning method, since online learning is more scalable than offline learning. However, most existing online similarity learning algorithms learn a full matrix with d2parameters, where d is the dimension of the instances. This is clearly inefficient for high dimensional tasks due to its high memory and computational complexity. To solve this issue, we introduce several Sparse Online Relative Similarity (SORS) learning algorithms, which learn a sparse model during the learning process, so that the memory and computational cost can be significantly reduced. We theoretically analyze the proposed algorithms, and evaluate them on some real-world high dimensional datasets. Encouraging empirical results demonstrate the advantages of our approach in terms of efficiency and efficacy. Dezhong Yao 0002, Peilin Zhao, Chen Yu 0003, Hai Jin 0001, Bin Li 0027 |
ICDM | 3 |
| 2015 | Mining user check-in features for location classification in location-based social networksabstractWith the increasing popularity of location-based social networks, a large number of users have been involved in the check-ins. The venues where the user frequently repeats check-ins tend to play a very important role in his daily life, as they not only dominate the user's mobility behavior but also imply the user's personal preferences. Therefore, fast discerning of such check-in venues could enable us to improve a wide range of location-based services. In this paper, we propose a new location classification problem for users of location-based social networks, in which we aim to discern, given the observation that a user makes a "new" check-in at a venue, whether he will frequently repeat check-ins at this venue. To solve the problem, we first extract 16 features attached to the user's "new" check-ins. With the publicly available check-in dataset, we then train a location classifier based on Support Vector Machine and compare it with two baselines based on majority voting. The comparison results demonstrate the practicability of the trained location classifier. Chen Yu 0003, Yang Liu 0082, Dezhong Yao 0002, Hai Jin 0001, Feng Lu 0003, Hanhua Chen |
ISCC | 1 |
| 2014 | Temporal-Based Ranking in Heterogeneous Networks
Chen Yu 0003, Ruidan Li, Dezhong Yao 0002, Feng Lu 0003, Hai Jin 0001 |
NPC | 1 |
| 2014 | Energy efficient indoor tracking on smartphones
Dezhong Yao 0002, Chen Yu 0003, Anind K. Dey, Christian Koehler 0002, Geyong Min, Laurence T. Yang, Hai Jin 0001 |
Future Gener. Comput. Syst. | 2 |
| 2013 | CloudThings: A common architecture for integrating the Internet of Things with Cloud ComputingabstractThe Internet of Things presents the user with a novel means of communicating with the Web world through ubiquitous object-enabled networks. Cloud Computing enables a convenient, on demand and scalable network access to a shared pool of configurable computing resources. This paper mainly focuses on a common approach to integrate the Internet of Things (IoT) and Cloud Computing under the name of CloudThings architecture. We review the state of the art for integrating Cloud Computing and the Internet of Things. We examine an IoT-enabled smart home scenario to analyze the IoT application requirements. We also propose the CloudThings architecture, a Cloud-based Internet of Things platform which accommodates CloudThings IaaS, PaaS, and SaaS for accelerating IoT application, development, and management. Moreover, we present our progress in developing the CloudThings architecture, followed by a conclusion. Jiehan Zhou, Teemu Leppänen, Erkki Harjula, Mika Ylianttila, Timo Ojala, Chen Yu 0003, Hai Jin 0001 |
CSCWD | 6 |
| 2013 | Energy Efficient Task Scheduling in Mobile Cloud Computing
Dezhong Yao 0002, Chen Yu 0003, Hai Jin 0001, Jiehan Zhou |
NPC | 2 |
| 2013 | Location-aware private service discovery in pervasive computing environment
Chen Yu 0003, Dezhong Yao 0002, Xi Li 0003, Yan Zhang 0002, Laurence T. Yang, Naixue Xiong, Hai Jin 0001 |
Inf. Sci. | 1 |
| 2013 | Design and application of the stereo vision manipulator with novel scheduling policies control
Kuei-Shu Hsu, Limei Peng, Chen Yu 0003 |
Multim. Tools Appl. | 3 |
| 2012 | Dynamic Spray and Wait Routing Protocol for Delay Tolerant Networks
Longbo Zhang, Chen Yu 0003, Hai Jin 0001 |
NPC | 2 |
| 2011 | Live Virtual Machine Migration via Asynchronous Replication and State SynchronizationabstractLive migration of virtual machines (VM) across physical hosts provides a significant new benefit for administrators of data centers and clusters. Previous memory-to-memory approaches demonstrate the effectiveness of live VM migration in local area networks (LAN), but they would cause a long period of downtime in a wide area network (WAN) environment. This paper describes the design and implementation of a novel approach, namely, CR/TR-Motion, which adopts checkpointing/recovery and trace/replay technologies to provide fast, transparent VM migration for both LAN and WAN environments. With execution trace logged on the source host, a synchronization algorithm is performed to orchestrate the running source and target VMs until they reach a consistent state. CR/TR-Motion can greatly reduce the migration downtime and network bandwidth consumption. Experimental results show that the approach can drastically reduce migration overheads compared with memory-to-memory approach in a LAN: up to 72.4 percent on application observed downtime, up to 31.5 percent on total migration time, and up to 95.9 percent on the data to synchronize the VM state. The application performance overhead due to migration is kept within 8.54 percent on average. The results also show that for a variety of workloads migrated across WANs, the migration downtime is less than 300 milliseconds. Haikun Liu, Hai Jin 0001, Xiaofei Liao, Chen Yu 0003, Cheng-Zhong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2009 | Live migration of virtual machine based on full system trace and replayabstractLive migration of virtual machines (VM) across distinct physical hosts provides a significant new benefit for administrators of data centers and clusters. Previous migration schemes focused on transferring the runtime memory state of the VM. Those approaches employed memory pre-copy algorithm to synchronize the migrating VM states, which make VM live migration cost much network traffic and application downtime, especially for memory intensive workloads. This paper describes the design and implementation of a novel approach CR/TR-Motion that adopts checkpointing/recovery and trace/replay technology to provide fast, transparent VM migration. With execution trace logged on the source host, a synchronization algorithm is performed to orchestrate the running source and target VM until they get a consistent state. We also give a formalized characterization about the migration evaluation metrics and make a mathematical analysis about our algorithm. Our scheme can greatly reduce the migration downtime and network bandwidth consumption. Experimental measurements show that our approach can drastically reduce migration overheads compared with pre-copy algorithm: up to 72.4% on application observed downtime, up to 31.5% on total migration time and up to 95.9% on the data to synchronize the VM state, while the application performance overhead due to migration is less than 8.54% on average. Haikun Liu, Hai Jin 0001, Xiaofei Liao, Liting Hu, Chen Yu 0003 |
HPDC | 5 |