VLDB 2026 Research / reviewers in the wild / expert
Fang Liu 0009
dblp:67/5807-9
· DBLP profile ↗
18ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0002-2165-6127ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Learning for Fake News Detection in Short Videos Using Linguistically Verified Data and Heterogeneous Modality FusionabstractThe rapid proliferation of short video platforms has necessitated advanced methods for detecting fake news. This need arises from the widespread influence and ease of sharing misinformation, which can lead to significant societal harm. Current methods often struggle with the dynamic and multi-modal nature of short video content. This paper presents HFN, Heterogeneous Fusion Net, a novel multimodal framework that integrates video, audio, and text data to evaluate the authenticity of short video content. HFN introduces a Decision Network that dynamically adjusts modality weights during inference and a Weighted Multi-Modal Feature Fusion module to ensure robust performance even with incomplete data. Additionally, we contribute a comprehensive dataset VESV (VEracity on Short Videos) specifically designed for short video fake news detection. Experiments conducted on the FakeTT and newly collected VESV datasets demonstrate improvements of 2.71% and 4.14% in Marco F1 over state-of-the-art methods. This work establishes a robust solution capable of effectively identifying fake news in the complex landscape of short video platforms, paving the way for more reliable and comprehensive approaches in combating misinformation. Shanghong Li, Chiam Wen Qi Ruth, Fang Liu 0009 |
GLOBECOM | 4 |
| 2023 | It Is About Weather: Explainable Machine Learning for Traffic Accident UnderstandingabstractRoad traffic accidents cause injuries, claim lives, and disrupt economic activities. It is among the key problems for intelligent transportation and smart cities. Cities, especially the mega ones, must strive for reducing accidents for public safety and sustainable growth, and the first task is to understand accidents. In this paper, we build up such understanding with the emerging explainable machine learning (ML) technique. We prepare a huge dataset with over two million accident records and use it to deliver ML models for accident modelling. Given ML models of high fidelity for mapping accident features and conditions to the accident severity, we apply several explainable ML techniques to explain the models and understand accidents. We first consider coarse granularity to capture the overall feature importance. Then we consider fine granularity methods including partial dependence plot and Shapley additive explanations. The former shows that feature impact varies at different feature values and uses a plot to reflect the impact changes. The latter puts attention to individual accident and quantifies the feature impact for the specific accident. Our core observation is that traffic accidents are often about weather, followed by location and road type. Extensive experimental study is performed to support our discussion and justify our conclusion. The deliverable of this paper offers an advanced way of understanding traffic accidents accurately in a quantitative manner and has great potential to be used for intelligent transportation and smart city applications. Syabil Soedirman, Fang Liu 0009, Zengyan Fan, Wei Zhang 0082 |
SMC | 2 |
| 2023 | TransLine: transfer learning for accurate and explainable power line anomaly detection with insufficient data
Fang Liu 0009, Wei Zhang 0082, Indriyati Atmosukarto, Teck Wei Low |
CCF Trans. Pervasive Comput. Interact. | 1 |
| 2022 | TransLine: Transfer Learning for Accurate Power Line Anomaly Detection with Insufficient DataabstractAccurate and automatic power line anomaly detection is critical to smart grid. However, effective solutions are yet available due to the insufficiency of anomaly data. In this paper, we first collect a dataset from various sources consisting both normal and abnormal power line images. With this dataset, anomaly detection becomes feasible though with limited accuracy due to the limited size of the dataset. As such, we propose TransLine, an approach based on transfer learning to apply the existing knowledge extracted from large-scale datasets to complement the data insufficiency of power line anomaly detection. TransLine customizes and optimizes the knowledge to automate the power line anomaly detection with high accuracy. The experiment results show that TransLine can achieve superb accuracy of 96.1% on average and up to 98.1% given only a hundred abnormal images for model training. TransLine can be a key enabler of smart grid for great stability and efficiency and can inspire the other industrial applications facing data insufficiency issues. Fang Liu 0009, Teck Wei Low, Wei Zhang 0082, Indriyati Atmosukarto |
ICC | 1 |
| 2022 | Utility Optimal Thread Assignment and Resource Allocation in Multi-Server SystemsabstractAchieving high performance in many multi-server systems (e.g., web hosting center, cloud) requires finding a good assignment of worker threads to servers and also effectively allocating each server’s resources to its assigned threads. The assignment and allocation components of this problem have been studied extensively but largely separately in the literature. In this paper, we introduce theassign and allocate (AA)problem, which seeks to simultaneously find an assignment and allocation that maximizes the total utility of the threads. Assigning and allocating the threads together can result in substantially better overall utility than performing the steps separately, as is traditionally done. We model each thread by a utility function giving its performance as a function of its assigned resources. We first prove that the AA problem is NP-hard. We then present a$2 (\sqrt {2}-1) > 0.828$factor approximation algorithm for concave utility functions, which runs in$O(mn^{2} + n (\log mC)^{2})$time for$n$threads and$m$servers with$C$amount of resources each. We also give a faster algorithm with the same approximation ratio and$O(n (\log mC)^{2})$time complexity. We then extend the problem to two more general settings. First, we consider threads with nonconcave utility functions, and give a 1/2 factor approximation algorithm. Next, we give an algorithm for threads using multiple types of resources, and show the algorithm achieves good empirical performance. We conduct extensive experiments to test the performance of our algorithms on threads with both synthetic and realistic utility functions, and find that they achieve over 92% of the optimal utility on average. We also compare our algorithms with a number of practical heuristics, and find that our algorithms achieve up to 9 times higher total utility. Pan Lai, Rui Fan 0004, Xiao Zhang 0006, Wei Zhang 0082, Fang Liu 0009, Joey Tianyi Zhou |
IEEE/ACM Trans. Netw. | 5 |
| 2021 | Towards Cost-Optimal Energy Procurement for Cooling as a Service: A Data-Driven ApproachabstractCoolingas a Service (CaaS) is an emerging business that provides air conditioning services for buildings. With the rapid development of the business and the continuous increase of energy load, CaaS providers need cost-effective energy procurement to meet the service requirements. In this paper, we propose a data-driven approach for energy procurement for CaaS providers. First, we focus on two dominant variables of cooling energy cost, including outdoor temperature and electricity price. We predict their trends in the next day and accordingly, we estimate the energy usage and purchase the energy in the day-ahead energy market, one day before the actual usage. During the real-time operation, we can obtain the actual temperature and price, and we use this information to adjust the quality of service without violating the service standards. The adjustment serves as the demand response to the real-time energy market and can be cost beneficial. We conducted experimental studies to verify the performance of the proposed solution. The results show that our solution provides high-quality cooling services with minimal energy expenditure and helps improve the stability of the power grid. Wei Zhang 0082, Yonggang Wen 0001, Fang Liu 0009 |
GLOBECOM | 3 |
| 2021 | Cost Optimal Data Center Servers: A Voltage Scaling ApproachabstractData centers have experienced dramatic growth in recent years in order to meet the ever-increasing demand for computing. As a result, minimizing the electrical cost to operate data centers has become a crucial issue. In this paper, we observe that electricity prices change over time, and that we can take advantage of periods with low prices by scaling up processor speeds to perform more work, while scaling down speeds during high price periods to reduce cost. We apply this observation to several settings. First, we consider an offline setting which assumes future electricity prices are given, and propose an efficient algorithm for optimally scaling a processor's speed in order to minimize the total electrical cost for completing a task by a deadline. We then consider a more realistic stochastic setting in which future prices are not known, but vary according to a Markov model. We present another efficient algorithm for minimizing the expected cost to meet a deadline. We performed a number of experiments using real electricity price traces to test the performance of our algorithms. We show that our stochastic algorithm is light-weight and relies only on easily obtainable price data, but that it achieves excellent performance, with only a 1 percent cost difference on average from the optimal offline algorithm. In addition, the stochastic algorithm significantly reduced costs compared to several candidate algorithms. Wei Zhang 0082, Yonggang Wen 0001, Loi Lei Lai, Fang Liu 0009, Rui Fan 0004 |
IEEE Trans. Cloud Comput. | 4 |
| 2020 | Electricity Cost Minimization for Interruptible Workload in Datacenter ServersabstractDatacenters have experienced dramatic growth in recent years, and the cost for powering them has become a significant problem. This paper proposes methods to minimize the energy cost for performing a task on a datacenter server before a deadline. We observe that energy prices fluctuate over time, and schedule the task to execute in periods of relatively low cost, despite not having knowledge of future costs during the execution. This problem is studied in several models, starting with an online setting where electricity prices can change arbitrarily. A$\sqrt{\varphi }$-competitive algorithm is proposed, where$\varphi$is the ratio between the maximum and minimum electricity prices, and this algorithm is also shown to be optimal by proving a matching lower bound. Next, we consider a stochastic setting in which prices vary in a Markovian fashion and propose an optimal algorithm based on dynamic programming. We then study the performance of our algorithms in practice using prices derived from real world data. The results show that the stochastic algorithm is very effective, and achieves cost that is within 3.4 percent of the optimum. Moreover, it performs well compared to several heuristics used in practice. Wei Zhang 0082, Yonggang Wen 0001, Loi Lei Lai, Fang Liu 0009, Rui Fan 0004 |
IEEE Trans. Serv. Comput. | 4 |
| 2019 | QoE-Driven Mobile Streaming: A Location-Aware ApproachabstractIn this paper, we maximize the quality of experience (QoE) for mobile video streaming. QoE is modeled to capture user's video quality assessment as well as the freezing and bitrate variation during playback. Based on an observation that network bandwidth is correlated with location, we predict the future locations and accordingly the bandwidth along a trip. Then, the predicted information is utilized to dynamically adapt the video version and bitrate with maximized QoE. We show that our proposed solution well approximates the offline optimal performance by almost 98% on average. The proposed solution is also competitive compared to several popular streaming algorithms. Fang Liu 0009, Wei Zhang 0082, Yonggang Wen 0001 |
ICME | 1 |
| 2018 | Fast media caching for geo-distributed data centers
Wei Zhang 0082, Yonggang Wen 0001, Fang Liu 0009, Yiqiang Chen 0001, Rui Fan 0004 |
Comput. Commun. | 3 |
| 2017 | Energy-Efficient Mobile Video Streaming: A Location-Aware ApproachabstractVideo streaming is one of the most widely used mobile applications today, and it also accounts for a large fraction of mobile battery usage. Much of the energy consumption is for wireless data transmission and is highly correlated to network bandwidth conditions. In periods of poor connectivity, up to 90% of mobile energy can be used for wireless data transfer. In this article, we study the problem of energy-efficient mobile video streaming. We make use of the observed correlation between bandwidth and user location , and also observe that a user’s location is predictable in many situations, such as when commuting to a known destination. Based on the user’s predicted locations and bandwidth conditions, we optimize wireless transmission times to achieve high quality video playback while minimizing energy use. We propose an optimal offline algorithm for this problem, which runs in O ( Tk ) time, where T is the duration of the video and k is the size of the video buffer. We also propose LAWS, a Location AWare Streaming algorithm. LAWS learns from historical location-aware bandwidth conditions and predicts future bandwidths along a planned route to make online wireless download decisions. We evaluate LAWS using real bandwidth traces, and show that LAWS closely approximates the performance of the optimal offline algorithm, achieving 90.6% of the optimal performance on average, and 97% in certain cases. LAWS also outperforms three popular strategies used in practice by, on average, 69%, 63%, and 38%, respectively. Lastly, we show that LAWS is able to deal with noisy data and can attain the stated performance after sampling bandwidth conditions only five times. Wei Zhang 0082, Rui Fan 0004, Yonggang Wen 0001, Fang Liu 0009 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2016 | Utility Maximizing Thread Assignment and Resource AllocationabstractAchieving high performance in many distributed systems requires finding a good assignment of threads to servers as well as effectively allocating each server's resources to its assigned threads. The assignment and allocation components of this problem have both been studied extensively, but separately in the literature. In this paper, we introduce the assign and allocate (AA) problem, which seeks to simultaneously find an assignment and allocations that maximize the total utility of the threads. Assigning and allocating the threads together can result in substantially better overall utility than performing the steps separately, as is traditionally done. We model each thread by a concave utility function giving its throughput as a function of its assigned resources. We first show that the AA problem is NP-hard, even when there are only two servers. We then present a 2(√2-1) > 0.828 factor approximation algorithm, which runs in O(mn2 + n (log mC)2) time for n threads and m servers with C amount of resources each. We also present a faster algorithm with the same approximation ratio and O(n(log mC)2) running time. We conducted experiments to test the performance of our algorithm on threads with different types of utility functions, and found that it achieves over 99% of the optimal utility on average. We also compared our algorithm against several other assignment and allocation algorithms, and found that it achieves up to 5.7 times better total utility. Pan Lai, Rui Fan 0004, Wei Zhang 0082, Fang Liu 0009 |
IPDPS | 4 |
| 2015 | Encrypted SVM for Outsourced Data MiningabstractIndividuals and companies, taking advantage of cloud computing which affords both resource and compute scalability, are willing to outsource their exploding data to save the storage and managing cost, however, users often do not fully trust the cloud and therefore outsource their private data after encryption to protect the data privacy. Here, as the data are both encrypted and outsourced in the cloud, how to securely and efficiently store and process such data becomes a challenging task and a primary concern. Support vector machine (SVM) classification, among different data mining and machine learning algorithms, has been very widely used in practical applications, which however, does not have a corresponding solution for such outsourced and encrypted data. Also, existing secure methods only assume that the data is locally stored by users rather than outsourced. To address this problem, we propose a novel Protocol for Outsourced SVM (POS) in this paper. POS lets cloud and users perform collaborative operations on encrypted and outsourced data without violating the data privacy contributed by each user. We formally verified that POS is correct and secure. We also conducted experimental analysis. Fang Liu 0009, Wee Keong Ng, Wei Zhang 0082 |
CLOUD | 1 |
| 2015 | Encrypted Gradient Descent Protocol for Outsourced Data MiningabstractWith the push of cloud computing which has both resource and compute scalability, data, which has been exploding in the past years, are often outsourced to a server. To this end, secure and efficient data processing and mining on outsourced private database becomes a primary concern for users. Among different secure data mining and machine learning algorithms, gradient descent method, as a widely used optimization paradigm, aims at approximating a target function to reach a local minimum, which is always deemed as a decision model to be discovered. In existing methods, users are assumed to hold and process their own data, and all users follow a secure protocol to perform gradient descent algorithm. However, such methods are not applicable to a cloud platform since that data is outsourced to a centralized server after encryption. To address this problem, we propose an Encrypted Gradient Descent Protocol (EGDP) in this paper. In EGDP, both users and server perform collaborative operations to learn and approximate the target function without violating data privacy. We formally proved that EGDP is secure and can return correct result. Fang Liu 0009, Wee Keong Ng, Wei Zhang 0082 |
AINA | 1 |
| 2015 | Encrypted Association Rule Mining for Outsourced Data MiningabstractRule mining, for discovering valuable relations between items in large databases, has been a popular and well researched method for years. However, such old but important technique faces huge challenges and difficulties in the era of cloud computing although which affords both storage and computing scalability: 1) data are outsourced to a cloud due to data explosion and high storage and management cost, 2) moreover, data are usually encrypted first before being outsourced for privacy's sake. Existing privacy-preserving rule mining methods only assume a distributed model where every data owner holds the self data without encryption and together follow a secure protocol to perform rule mining. To address this limitation, we propose a novel Protocol for Outsourced Rule Mining (PORM) in this paper. PORM performs rule mining in a cloud environment where data are both encrypted and outsourced. We formally proved that PORM is both correct and secure, and we also extended PORM to the multiple-user scenario. Fang Liu 0009, Wee Keong Ng, Wei Zhang 0082 |
AINA | 1 |
| 2015 | Energy-Aware CachingabstractTo achieve higher performance, cache sizes have been steadily increasing in computer processors and network systems. But caches are often over-provisioned for peak demand and underutilized in typical non-peak workloads. As caches consume substantial power, this results in significant amounts of wasted energy. To address this, existing works turn off parts of the cache when they do not contribute to higher performance. However, while these methods are effective empirically, they lack provable performance bounds. In addition, existing works focus on processor caches and are not applicable to network caches where data size and cost can vary. In this paper, we study the energy-aware caching (EAC) problem, and seek to minimize the total cost incurred due to cache misses and energy consumption. We propose three algorithms to solve different variants of this problem. The first is an optimal offline algorithm that runs in O(kn log n) time for a size k cache and n cache accesses. Then, we propose a simple online algorithm for uniform data size and cost that is $2 + {{h} \over {h-h+1}}$ competitive compared to an optimal algorithm with a size h ≤ k cache. Lastly, we propose a $2 + {{h-1} \over {h-h+1}}$ competitive online algorithm that allows arbitrary data sizes and costs. We give an efficient implementation of the algorithm that takes O(log k) amortized time per cache access, and also present an adaptive version that reacts to workload patterns to achieve better real-world performance. Using trace driven simulations, we show our algorithm has substantially lower cost than algorithms focused on maximizing cache hit rates or minimizing energy usage alone. Wei Zhang 0082, Rui Fan 0004, Fang Liu 0009, Pan Lai |
ICPADS | 3 |
| 2014 | Encrypted Scalar Product Protocol for Outsourced Data MiningabstractOrganizations and individuals nowadays face increasing daily operations closely rely on a huge amount of private data which is outsourced to a centralized server. Secure and efficient data processing and mining on such outsourced private data becomes a primary concern for users, especially with the push of cloud computing which has both resource and compute scalability. Among the building blocks of secure data mining algorithms, secure scalar product is used to calculate the sum of the products of the corresponding values of two vectors. Existing privacy preserving methods assume data is stored at the user side, and users follow a protocol to perform privacy preserving scalar product. However, such methods are not applicable as data now is outsourced to a centralized server in its encrypted form. To solve this problem, in this paper, we design a novel Protocol for Outsourced Scalar Product (POSP) that performs collaborative operations between server and users to produce the scalar product result without violating each user's data privacy. We proved that POSP can return the correct result and is secure. We also analysed that POSP has linear complexity in terms of space, computation, and communication with respect to the vector length. Fang Liu 0009, Wee Keong Ng, Wei Zhang 0082 |
IEEE CLOUD | 1 |
| 2014 | Encrypted Set Intersection Protocol for Outsourced DatasetsabstractSecure and efficient data storage and computation for an outsourced database is a primary concern for users, especially with the push for cloud computing that affords both compute and resource scalability. Among the diverse secure building blocks for secure analytical computations on outsourced databases, the encrypted set intersection operation extracts common sensitive information from datasets belonging to different users. In existing methods, each user holds their sensitive data and all users follow a secure protocol to perform set intersection. This approach is not applicable to the cloud platform, where data resides in the cloud platform in encrypted form and not at each user site. To address this limitation, in this paper, we design the Encrypted Set Intersection Protocol (ESIP) that allows server and users to perform collaborative operations to obtain the correct set intersection result without violating privacy of data contributed by each user at the server. Fang Liu 0009, Wee Keong Ng, Wei Zhang 0082, Hoang Giang Do, Shuguo Han |
IC2E | 1 |