VLDB 2026 Research / reviewers in the wild / expert
Hao Tian 0012
dblp:29/1374-12
· DBLP profile ↗
15ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-5335-5898ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 5 · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsabstractLarge Language Models (LLMs) have revolutionized intelligent interactions, enabling mobile applications such as personal assistants on edge devices for local execution. Speculative decoding (SD) has emerged as a promising paradigm to accelerate LLM inference without compromising generation quality, employing a draft-then-verify manner. However, due to the constrained computing and memory resources on edge devices, existing SD works heavily rely on an auxiliary draft model that incurs additional memory burden and hinders the adaptability, as well as static token trees that yield suboptimal inference performance. To this end, we propose DIAA, a Decoding-efficient Inference Acceleration Approach for on-device LLMs. DIAA achieves plug-and-play and model-agnostic inference speedup with memory and computation efficiency for edge devices. Specifically, a pair of lightweight look-up tables (LUTs) is constructed by Top-K token sampling to cache historical tokens and probabilities for rapid candidate drafting. DIAA integrates a dynamic token tree with prior LUTs enabling paralleled verification, updated during decoding process, to adapt the online context. A computation overlap is then employed to pipeline the update operations of token tree, LUTs, and KV cache to improve the computational efficiency. Finally, through extensive experiments implemented on edge platform NVIDIA Jetson, DIAA outperforms existing baselines in generation speed and inference wall-clock time, while incurring minimal memory overhead. Hao Tian 0012, Fuwen Tian, Guangming Cui, Zheng Li 0026, Xuyun Zhang, Quan Z. Sheng, Wan-Chun Dou |
AAAI | 1 |
| 2025 | A Cost-Aware Approach for Collaborating Large Language Models and Small Language ModelsabstractThe emerging reasoning ability of large language models (LLMs) and accompanying commercial applications offer a promising path for service providers to deploy intelligent agents on their own products through API calls. However, the black-box nature of LLMs has driven providers to try prompt tuning to improve reasoning quality for competitiveness, while the generated reasoning logic results in additional service costs. Although some works have proposed collaborating LLMs and Small Language Models (SLMs) to reduce the frequency of LLM calls, most overlook the actual number of tokens interacting with the LLMs, which results in a potentially high cost still. Furthermore, directly compressing the prompt to reduce tokens often leads to a significant accuracy loss. To address the above challenges, we propose a cost-aware approach for collaborating LLMs and SLMs, named Coco. In our method, a confidence-based task assignment method is designed which leverages the result confidence of SLMs to assess task complexity and determine whether LLM involvement is necessary. For complex tasks, the SLM adapts the input by compressing unnecessary information according to confidence. Considering the potential loss of accuracy, prompt tuning-based reasoning optimization methods are introduced to guide the LLM in generating both the reasoning logic sketch and the final result. Finally, logic alignment is applied to fuse sketches from both models, ensuring the rationality of the reasoning logic. Experimental results on three open-source datasets demonstrate that our approach effectively reduces the cost of API calls to LLMs while ensuring the reasoning accuracy and the reasonableness of generated logic. Zheng Li 0026, Xuyun Zhang, Hao Tian 0012, Wan-Chun Dou |
CIKM | 5 |
| 2025 | A Dynamic Inference Method for Autoregressive Transformer ModelsabstractAutoregressive transformer models have achieved state-of-the-art performance in advanced services such as text generation and machine translation. Given the significant computational bottlenecks of model inference, layer-wise skipping has emerged as a promising method to accelerate inference by bypassing redundant layers. However, existing methods face challenges, including sub-optimal performance resulting from the premature skipping of critical layers and an unbalanced focus on either multi-head attention or feed-forward sub-blocks, ultimately leading to global performance degradation. In light of the above challenges, we propose a Dynamic Inference Method, named DIM, for autoregressive transformer models. DIM dy-namically selects sub-blocks from both multi-head attention and feed-forward networks through the importance score alignment, ensuring a balanced selection that optimizes both efficiency and model performance. To further mitigate the potential performance loss of skipped sub-blocks, a lightweight adjustment is developed to approximate the computations of skipped sub-blocks. Finally, extensive experiments using several benchmarks validate that DIM outperforms existing inference methods. Hao Tian 0012, Zheng Li 0026, Wan-Chun Dou |
ICWS | 1 |
| 2025 | A Consortium Blockchain-Based Edge Task Offloading Method for Connected Autonomous VehiclesabstractIn recent years, the proliferation of Connected Autonomous Vehicles (CAV) has revolutionized the transportation industry. However, these vehicles often face limitations in terms of local computing resources, leading to the need for offloading interactive-intensive application tasks to servers for processing. Traditional paradigm has its limitations in meeting the demands of massive task processing. The combination of Web3.0 and edge computing offers users high-reliable, low-latency, and highly flexible services. Nevertheless, the new paradigm also presents its own challenges such as ensuring privacy data protection, and reducing the time and energy costs associated with task offloading. To tackle these challenges, an edge task offloading framework based on consortium blockchain for CAVs has been developed. Within this framework, a consortium blockchain-based interaction-intensive task offloading method, called CBIToMe, has been designed. CBIToMe specifically addresses the multi-stage nature of interactive-intensive CAV tasks and aims to minimize task completion time and offloading costs, particularly when the waiting time for interaction is uncertain. Additionally, CBIToMe effectively utilizes consortium blockchain technology to safeguard the CAV privacy data. Results from experiments conducted in various scenarios demonstrate that CBIToMe outperforms three representative methods, showcasing its superior performance. Bowen Liu 0002, Hao Tian 0012, Zhijie Shen, Yueyue Xu, Wan-Chun Dou |
ACM Trans. Auton. Adapt. Syst. | 2 |
| 2024 | An Inference Acceleration Approach for Boosting DNN Cold Start in Cloud-Edge Computing
Hao Tian 0012, Haolong Xiang, Tingtong Zhu, Siyuan Wu 0002, Zheng Li 0026, Mingxu Jiang, Wan-Chun Dou |
ADMA (1) | 1 |
| 2024 | A Heterogeneous Federated Learning Method Based on Dual Teachers Knowledge Distillation
Siyuan Wu 0002, Hao Tian 0012, Weiran Zhang, Tingtong Zhu, Fuwen Tian, Zhehong Wang, Wan-Chun Dou |
ADMA (2) | 2 |
| 2024 | SFSM: A Serverless Function Scheduling Method for FaaS Applications over Edge ComputingabstractServerless edge computing is emerging as an enabler to provision scalable and flexible Function-as-a-Service (FaaS) applications with lightweight function instances at network edge. In serverless edge computing, the function instances with inter-dependencies are scheduled to proximate edge nodes in a distributed manner. However, the heterogeneity and unpredictability of edge networks bring significant challenge in realizing optimal scheduling decision to guarantee execution performance of applications without any prior. In view of this challenge, a Serverless Function Scheduling Method, named SFSM, is proposed in this paper for FaaS applications over edge computing. First, a long-term optimization problem is formulated to reduce completion time and decoupled into time-slot sub-problems via Lyapunov optimization. To avoid the cross-edge redundant data transmission overhead of inter-functions, a two-level graph optimization is designed with vertical and horizontal data merging. Then, SFSM incorporates an online multi-armed bandit-based scheduling algorithm that only requires the context of requests without complete information of edge networks. Finally, extensive experimental results based on real-world datasets demonstrate the effectiveness and superiority of SFSM. Hao Tian 0012, Fei Dai 0002, Wan-Chun Dou |
ICWS | 1 |
| 2024 | A Contrastive Collaborative Filtering Method for Personalized Recommendation with Self-AttentionabstractRecently, sequential and Collaborative Filtering (CF) based recommender systems have shown their research popularity in both academia and industry. Graph-based CF and Transformer-based sequential models have independently shown state-of-art recommendation performance respectively. However, each approach has its own limitation and retrieves different potential factors from users and items. CF aims to retrieve collaborative signals between users and items while Transformer models the temporal dependencies between user historical actions. The under utilization of latent factors can limit the performance and generalization of recommender systems. To address the above limitation, in this paper, we proposed a novel multi-view recommender method called Contrastive Attentive Collaborative Filtering (CACF) that combines collaborative signals and temporal dependencies. Our method leverages the strengths of both approaches by combining the temporal dependecies captured by the Transformer with the user-item collaborative signals retrieved by graph convolutional network (GCN). The Transformer learns temporal dependencies and generates embeddings of items while GCN component capture collaborative signals. By contrastivly optimizing output embeddings from both methods, our approach aims to provide more accurate and personalized recommendations. Finally, we evaluate our method on several benchmark datasets and compare its performance with state-of-the-art recommendation systems. The experimental results show that our method outperforms in several evaluation metrices. Hao Tian 0012, Wan-Chun Dou |
ISPA | 2 |
| 2024 | A Crowdsensing Service Pricing Method in Vehicular Edge ComputingabstractThe rapid advancements in vehicular networking technology have enabled onboard users to access a variety of emerging services like precise navigation and real-time hazard avoidance. However, as the diversity and volume of required data expand and demands intensify, most of vehicular services are unable to satisfy user expectations. Although crowdsensing architectures based on edge computing have been proposed in vehicular networks, how to optimally allocate the sensing service time of roadside units and set appropriate prices for services to maximize user benefits still remain significant challenges. To solve the above issue, in this paper, a crowdsensing service pricing method is proposed in vehicular edge computing. Specifically, a vehicular networking crowdsensing framework based on the edge computing is designed. Then, the optimization problem of perception time pricing during the crowdsensing process is modeled into a Stackelberg game and further the multi-agent deep deterministic policy gradient-based algorithm is employed to obtain the optimal pricing strategy with user profits maximization. Finally, simulation results demonstrate that the proposed method significantly increases the payoff for all participants during the crowdsensing process. Zheng Li 0026, Sizhe Tang, Hao Tian 0012, Haolong Xiang, Xiaolong Xu 0001, Wan-Chun Dou |
ISPA | 3 |
| 2024 | A Blockchain-Assisted Federated Learning Method for Recommendation SystemsabstractDue to the privacy advantages of federated learning (FL), federated recommendation systems (FedRSs) are gaining popularity for improving recommendation performance through training on local data. However, FedRSs frequently face the significant challenge of high communication costs between the server and clients. Most FedRSs utilize a client-server communication architecture, leading to heavy communication loads and single points of failure due to dependence on a central server. Clients may also encounter problems due to limited communication resources. In view of this challenge, in this paper, we propose a blockchain-assisted federated learning method at edge for communication-efficient recommendation systems, named BFedRec. Specifically, BFedRec reduces reliance on the central server by utilizing blockchain systems on edge servers to aggregate and distribute the recommendation model. To mitigate the high communication costs between clients and blockchain in each iteration, a communication-efficient training algorithm is used that trains the recommendation model directly on low-rank compressed parameters. Finally, we conduct extensive experiments on real-world datasets to verify the communication efficiency of BFedRec compared to existing methods. The experimental results show that BFedRec effectively improves communication efficiency without compromising recommendation performance. Hao Tian 0012, Chen Tian 0001, Wan-Chun Dou |
ISPA | 2 |
| 2023 | A Blockchain-enabled Secure Access Management Method in Edge ComputingabstractEdge computing has emerged as a transformative paradigm with applications ranging from IoT to smart homes and transportation systems. However, its decentralized nature presents significant challenges in securing sensitive data, particularly in scenarios involving data sharing. To address these issues, this paper introduces a secure data access control system for edge computing by leveraging blockchain to foster mutual trust among edge servers and employing attribute-based encryption for secure data sharing. Our proposed solution seeks to enhance data security in distributed environments. We validate our method through practical experiments conducted on multiple edge servers, confirming its effectiveness in safeguarding data integrity and privacy. Xutong Jiang, Hao Tian 0012, Haolong Xiang, Xuyun Zhang, Wan-Chun Dou |
ICPADS | 3 |
| 2023 | Real-time COVID-19 detection over chest x-ray images in edge computingabstractSevere Coronavirus Disease 2019 (COVID-19) has been a global pandemic which provokes massive devastation to the society, economy, and culture since January 2020. The pandemic demonstrates the inefficiency of superannuated manual detection approaches and inspires novel approaches that detect COVID-19 by classifying chest x-ray (CXR) images with deep learning technology. Although a wide range of researches about bran-new COVID-19 detection methods that classify CXR images with centralized convolutional neural network (CNN) models have been proposed, the latency, privacy, and cost of information transmission between the data resources and the centralized data center will make the detection inefficient. Hence, in this article, a COVID-19 detection scheme via CXR images classification with a lightweight CNN model called MobileNet in edge computing is proposed to alleviate the computing pressure of centralized data center and ameliorate detection efficiency. Specifically, the general framework is introduced first to manifest the overall arrangement of the computing and information services ecosystem. Then, an unsupervised model DCGAN is employed to make up for the small scale of data set. Moreover, the implementation of the MobileNet for CXR images classification is presented at great length. The specific distribution strategy of MobileNet models is followed. The extensive evaluations of the experiments demonstrate the efficiency and accuracy of the proposed scheme for detecting COVID-19 over CXR images in edge computing. Weijie Xu, Beijing Chen, Haoyang Shi, Hao Tian 0012, Xiaolong Xu 0001 |
Comput. Intell. | 4 |
| 2022 | DisCOV: Distributed COVID-19 Detection on X-Ray Images with Edge-Cloud Collaborationabstract[J1C2 Presentation Abstract at IEEE SERVICES 2022 for IEEE Transactions on Services Computing DOI 10.1109/TSC.2022.3142265] Xiaolong Xu 0001, Hao Tian 0012, Xuyun Zhang, Lianyong Qi, Qiang He 0001, Wan-Chun Dou |
SERVICES | 2 |
| 2022 | DisCOV: Distributed COVID-19 Detection on X-Ray Images With Edge-Cloud CollaborationabstractCurrently, the world is experiencing the rapid spread of Coronavirus Disease 2019 (COVID-19). Since the epidemic continues to take a devastating impact on the society, economy, and healthcare, the real-time detection of COVID-19 is essential for fast and cost-effective diagnosis services. Fortunately, deep learning (DL), as a promising technology, enables the COVID-19 diagnosis services on chest X-ray (CXR) images. The training task of DL model is generally implemented at the centralized cloud. However, due to the geo-distributed data sources and the transmission of large amounts of raw data to the centralized cloud, the transmission latency becomes a bottleneck of the COVID-19 diagnosis model training. In this paper, we propose aDistributedCOVID-19 detection model training method on CXR images with edge-cloud collaboration, named DisCOV. Specifically, to improve the training efficiency and guarantee the model accuracy, a distributed lightweight model-based training algorithm is designed with the cooperation of edge computing and cloud computing. In addition, a resource allocation algorithm is developed during the training to jointly minimize the time cost and energy consumption. Extensive experiments based on real-world CXR image datasets demonstrate that DisCOV is better performed and more promising than the existing baselines. Xiaolong Xu 0001, Hao Tian 0012, Xuyun Zhang, Lianyong Qi, Qiang He 0001, Wan-Chun Dou |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | DIMA: Distributed cooperative microservice caching for internet of things in edge computing by deep reinforcement learning
Hao Tian 0012, Xiaolong Xu 0001, Tingyu Lin 0001, Yong Cheng 0002, Lei Ren 0001, Muhammad Bilal 0003 |
World Wide Web | 1 |