Fan Zhang 0003

dblp:21/3626-3 · DBLP profile ↗
← Back
27ranked-venue papers
12as first author
8since 2021 · last 2025
0000-0002-5576-7271ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 8 first-authorArtificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Trustworthy machine learning · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Smart cities and intelligent transportation · 80% Bioinformatics and computational biology · 20%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Parallel and multicore computing · 85% Cloud and datacenter computing · 15%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%
Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 79% Machine learning and data management · 21%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › interpretability
example-based explanation
0.512021
On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation · ACL/IJCNLP (1) 2021
Machine learning › Trustworthy machine learning › interpretability › explanation evaluation
explanation faithfulness
0.512021
On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation · ACL/IJCNLP (1) 2021
Machine learning › Trustworthy machine learning
interpretability
0.512021
On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation · ACL/IJCNLP (1) 2021
Smart cities and intelligent transportation › disaster management
emergency evacuation
0.312017
MacroServ: A Route Recommendation Service for Large-Scale Evacuations · IEEE Trans. Serv. Comput. 2017
Smart cities and intelligent transportation › route planning
route recommendation
0.312017
MacroServ: A Route Recommendation Service for Large-Scale Evacuations · IEEE Trans. Serv. Comput. 2017
Smart cities and intelligent transportation
traffic management
0.312017
MacroServ: A Route Recommendation Service for Large-Scale Evacuations · IEEE Trans. Serv. Comput. 2017
Query processing and optimization › preference query
skyline query
0.212016
Skyline Discovery and Composition of Multi-Cloud Mashup Services · IEEE Trans. Serv. Comput. 2016
Services computing and microservices › service selection
qos-aware service selection
0.212016
Skyline Discovery and Composition of Multi-Cloud Mashup Services · IEEE Trans. Serv. Comput. 2016
Services computing and microservices
service composition
0.212016
Skyline Discovery and Composition of Multi-Cloud Mashup Services · IEEE Trans. Serv. Comput. 2016
Bioinformatics and computational biology › biomedical text mining
biomedical named entity recognition
0.212015
Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › data-parallel programming
mapreduce
0.212015
Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields · IEEE Trans. Parallel Distributed Syst. 2015
Parallel and multicore computing › parallel algorithms
parallel algorithm design
0.212015
Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields · IEEE Trans. Parallel Distributed Syst. 2015
Cloud and datacenter computing › cloud deployment
multi-cloud
0.112016
Skyline Discovery and Composition of Multi-Cloud Mashup Services · IEEE Trans. Serv. Comput. 2016
Machine learning and data management
data management for machine learning
0.112015
Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields · IEEE Trans. Parallel Distributed Syst. 2015

Methods — techniques the papers use, named apart from their topics

skyline query processing · 0.8similarity pruning · 0.8viterbi algorithm · 0.7mapreduce · 0.7sampling-based explanation · 0.5conditional random field · 0.4L-BFGS · 0.4probability distribution · 0.3multi-objective optimization · 0.3conditional random fields · 0.2LBFGS · 0.2
YearPublicationVenuePosition
2025 Deception detection in videos using the facial action coding system
Hammad Ud Din Ahmed, Usama Ijaz Bajwa, Naeem Iqbal Ratyal, Fan Zhang 0003, Muhammad Waqas Anwar
Multim. Tools Appl.4
2024 Leveraging coverless image steganography to hide secret information by generating anime characters using GAN
Hafiz Abdul Rehman, Usama Ijaz Bajwa, Rana Hammad Raza, Sultan Alfarhood, Mejdl S. Safran, Fan Zhang 0003
Expert Syst. Appl.6
2023 Selecting Distinctive-Variant Training Samples Base on Intra-class Similarity
Hang Diao, Zhengchang Liu, Fan Zhang 0003, Jiaqing Huang, Feiyu Zhou, Samee Ullah Khan
ICANN (9)3
2023 Find Important Training Dataset by Observing the Training Sequence Similarity
Zhengchang Liu, Hang Diao, Fan Zhang 0003, Samee Ullah Khan
ICANN (3)3
2022 Mining Influential Training Data by Tracing Influence on Hard Validation Samples
abstract
The ever-growing deep learning model size is constantly driven by the ever-growing dataset size. Mining the influential training data has significant payoff of either reducing the training time, model complexity as well as potentially increasing the model accuracy. In this paper, we propose a few approaches, e.g. classifying the validation dataset into easy, medium and hard levels, introducing influence value by calculating each training data on the hard validation data, to co-prune the validation dataset and the training dataset. Empirically we conclude that the portion of the hard validation data could be used to mine the most influential training data, whereby reducing the training dataset size by 50% without losing accuracy in our experiments.
Qikai Zhang, Fan Zhang 0003, Samee Ullah Khan
ICTAI2
2022 HCA Operator: A Hybrid Cloud Auto-scaling Tooling for Microservice Workloads
abstract
Elastic cloud platform, e.g. Kubernetes, enables dy-namically scale in or out computing resources in accordance with the workloads fluctuation. As the cloud evolves to hybrid, where public and private clouds co-exist as the underline substrate, autoscaling applications within a hybrid cloud is no longer straightforward. The difficulty lies in all aspects, e.g. global load balancing, hybrid-cloud monitoring and alerting, storage sharing and replication, security and privacy, etc. However, it will significantly pay off if hybrid-cloud autoscaling is supported and boundless computing resources can be utilized per request. In this paper, we design Hybrid Cloud Autoscaler Operator (HCA Operator), a customized Kubernetes Controller that leverages the Kubernetes Custom Resource to auto-scale microservice applications across hybrid clouds. HCA Operator load balances across hybrid clouds, monitors metrics, and autoscales to des-tination clusters that exist in other clouds. We discuss the implementation details and perform experiments in a hybrid cloud environment. The experimental results demonstrate that if the workload changes quickly, our Operator can properly auto-scale the microservice applications across hybrid cloud in order to meet the Service Level Agreement (SLA) requirements.
Fan Zhang 0003, Samee Ullah Khan
MSN2
2021 On Sample Based Explanation Methods for NLP: Faithfulness, Efficiency and Semantic Evaluation
abstract
Wei Zhang, Ziming Huang, Yada Zhu, Guangnan Ye, Xiaodong Cui, Fan Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Zhang 0057, Ziming Huang, Yada Zhu, Guangnan Ye, Fan Zhang 0003
ACL/IJCNLP (1)6
2021 Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation
abstract
In Speech Emotion Recognition (SER), emotional characteristics often appear in diverse forms of energy patterns in spectrograms. Typical attention neural network classifiers of SER are usually optimized on a fixed attention granularity. In this paper, we apply multiscale area attention in a deep convolutional neural network to attend emotional characteristics with varied granularities and therefore the classifier can benefit from an ensemble of attentions with different scales. To deal with data sparsity, we conduct data augmentation with vocal tract length perturbation (VTLP) to improve the generalization capability of the classifier. Experiments are carried out on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) dataset. We achieved 79.34% weighted accuracy (WA) and 77.54% unweighted accuracy (UA), which, to the best of our knowledge, is the state of the art on this dataset.
Mingke Xu, Fan Zhang 0003, Wei Zhang 0057
ICASSP2
2019 Facial Expression Recognition Based on Edge Computing
abstract
Facial action unit (AU) detection recognizes facial expressions by analyzing cues about the movement of certain atomic muscles in the local facial area. According to the detection of facial feature points, we can calculate the value of AU, and then classification algorithms are performed on these AU values for emotion detection in realtime. When this advanced system is in the actual production process, large network bandwidth is usually required to transmit a large number of frames from cameras to the backend servers. To overcome this issue, we propose to offload the computing to edge devices in which the raw image data from each camera is directly processed with our optimized and customized algorithms, and then the detected emotions are transmitted to the end-user more easily compared with less data.
Tiantian Qian, Fan Zhang 0003, Samee Ullah Khan
MSN2
2019 Quantifying cloud elasticity with container-based autoscaling
Fan Zhang 0003, Xuxin Tang, Xiu Li 0001, Samee Ullah Khan, Zhijiang Li
Future Gener. Comput. Syst.1
2019 Empirical Discovery of Power-Law Distribution in MapReduce Scalability
abstract
Understanding the scalability of MapReduce applications is a challenging problem. The difficulty lies in the distributed mapping of the input big data. The distribution of data and compute resources must match with fluctuating network substrates. User-defined Map and Reduce functions over application parameters further complicate the issue. Therefore, it offers great payoff to use small datasets and limited test runs to reveal the behavior of MapReduce applications over big-data. In this paper, we analyze the scaling effects of server cluster-size over varieties of Map- and Reduce-intensive applications. In our study, we discover specific conditions which lead to the power-law conformity in representative MapReduce applications. We report four major discoveries: (1) Within a range of scaling parameters, MapReduce execution time follows the power-law distribution. (2) Power-law scalability for Map-intensive applications work well even with a small cluster size. (3) Shuffle-intensive applications exhibit power-law behavior starting from larger cluster size. (4) The scaling effects may depart from power-law distribution, if the cloud resources are heavily overprovisioned than the workload demands. The above findings enable users to use bounded test runs to allocate and configure virtual and physical resources in large-scale MapReduce applications. These results can be also applied in generating business models for providing cost-effective cloud computing services.
Fan Zhang 0003, Majd F. Sakr, Kai Hwang 0001, Samee Ullah Khan
IEEE Trans. Cloud Comput.1
2017 MobiContext: A Context-Aware Cloud-Based Venue Recommendation Framework
abstract
In recent years, recommendation systems have seen significant evolution in the field of knowledge engineering. Most of the existing recommendation systems based their models on collaborative filtering approaches that make them simple to implement. However, performance of most of the existing collaborative filtering-based recommendation system suffers due to the challenges, such as: (a) cold start, (b) data sparseness, and (c) scalability. Moreover, recommendation problem is often characterized by the presence of many conflicting objectives or decision variables, such as users' preferences and venue closeness. In this paper, we proposed MobiContext, a hybrid cloud-based bi-objective recommendation framework (BORF) for mobile social networks. The MobiContext utilizes multi-objective optimization techniques to generate personalized recommendations. To address the issues pertaining to cold start and data sparseness, the BORF performs data preprocessing by using the Hub-Average (HA) inference model. Moreover, the Weighted Sum Approach (WSA) is implemented for scalar optimization and an evolutionary algorithm (NSGA-II) is applied for vector optimization to provide optimal suggestions to the users about a venue. The results of comprehensive experiments on a large-scale real dataset confirm the accuracy of the proposed recommendation framework.
Rizwana Irfan, Osman Khalid, Muhammad Usman Shahid Khan, Camelia Chira, Rajiv Ranjan 0001, Fan Zhang 0003, Samee Ullah Khan, Bharadwaj Veeravalli, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Cloud Comput.6
2017 MacroServ: A Route Recommendation Service for Large-Scale Evacuations
abstract
To respond to emergencies in a fast and an effective manner, it is of critical importance to have efficient evacuation plans that lead to minimum road congestions. Although emergency evacuation systems have been studied in the past, the existing approaches, mostly based on multi-objective optimizations, are not scalable enough when involve numerous time varying parameters, such as traffic volume, safety status, and weather conditions. In this paper, we propose a scalable emergency evacuation service, termed the MacroServ that recommends the evacuees with the most preferred routes towards safe locations during a disaster. Unlike many existing approaches that model systems with static network characteristics, our approach considers real-time road conditions to compute the maximum flow capacity of routes in the transportation network. The evacuees are directed towards those routes that are safe and have least congestion resulting in decreased evacuation time. We utilized probability distributions to model the real-life stochastic behaviors of evacuees during emergency scenarios. The results indicate that recommendation of appropriate routes during emergency scenarios play a critical role in quicker and safe evacuation of the population.
Muhammad Usman Shahid Khan, Osman Khalid, Rajiv Ranjan 0001, Fan Zhang 0003, Bharadwaj Veeravalli, Samee Ullah Khan, Keqin Li 0001, Albert Y. Zomaya
IEEE Trans. Serv. Comput.5
2016 MapReduce-based fast fuzzy c-means algorithm for large-scale underwater image segmentation
Xiu Li 0001, Jingdong Song, Fan Zhang 0003, Xiaogang Ouyang, Samee Ullah Khan
Future Gener. Comput. Syst.3
2016 Skyline Discovery and Composition of Multi-Cloud Mashup Services
abstract
A cloud mashup is composed of multiple services with shared datasets and integrated functionalities. For example, the elastic compute cloud (EC2) provided by Amazon Web Service (AWS), the authentication and authorization services provided by Facebook, and the Map service provided by Google can all be mashed up to deliver real-time, personalized driving route recommendation service. To discover qualified services and compose them with guaranteed quality of service (QoS), we propose an integrated skyline query processing method for building up cloud mashup applications. We use a similarity test to achieve optimal localized skyline. This mashup method scales well with the growing number of cloud sites involved in the mashup applications. Faster skyline selection, reduced composition time, dataset sharing, and resources integration assure the QoS over multiple clouds. We experiment with the quality of Web service (QWS) benchmark over 10,000 Web services along six QoS dimensions. By utilizing block-elimination, data-space partitioning, and service similarity pruning, the skyline process is shortened by three times, when compared with two state-of-the-art methods.
Fan Zhang 0003, Kai Hwang 0001, Samee Ullah Khan, Qutaibah M. Malluhi
IEEE Trans. Serv. Comput.1
2015 A task-level adaptive MapReduce framework for real-time streaming data in healthcare applications
Fan Zhang 0003, Samee Ullah Khan, Keqin Li 0001, Kai Hwang 0001
Future Gener. Comput. Syst.1
2015 CloudFlow: A data-aware programming model for cloud workflow applications on modern HPC systems
Fan Zhang 0003, Qutaibah M. Malluhi, Tamer Elsayed, Samee Ullah Khan, Keqin Li 0001, Albert Y. Zomaya
Future Gener. Comput. Syst.1
2015 Maximizing reliability with energy conservation for parallel task scheduling in a heterogeneous cluster
Longxin Zhang, Kenli Li 0001, Yuming Xu, Jing Mei, Fan Zhang 0003, Keqin Li 0001
Inf. Sci.5
2015 Adaptive Workflow Scheduling on Cloud Computing Platforms with IterativeOrdinal Optimization
abstract
The scheduling of multitask jobs on clouds is an NP-hard problem. The problem becomes even worse when complex workflows are executed on elastic clouds, such as Amazon EC2 or IBM RC2. The main difficulty lies in the large search space and high overhead of generating optimal schedules, especially for real-time applications with dynamic workloads. In this work, a new iterative ordinal optimization (IOO) method is proposed. The ordinal optimization method is applied in each iteration to achieve sub-optimal schedules. IOO aims at generating more efficient schedules from a global perspective over a long period. We prove through overhead analysis the advantages in time and space efficiency in using the IOO method. The IOO method is designed to adapt to system dynamism to yield suboptimal performance. In cloud experiments on IBM RC2 cloud, we execute 20,000 tasks in LIGO (Laser Interferometer Gravitational-wave Observatory) verification workflow on 128 virtual machines. The IOO schedule is generated in less than 1,000 seconds, while using the Monte Carlo simulation takes 27.6 hours, 100 times longer to yield an optimal schedule. The IOO-optimized schedule results in a throughput of 1,100 tasks/sec with 7 GB memory demand, compared with 60 percent decrease in throughput and 70 percent increase in memory demand in using the Monte Carlo method. Our LIGO experimental results clearly demonstrate the advantage of using the IOO-based workflow scheduling over the traditional blind-pick, ordinal optimization, or Monte Carlo methods. These numerical results are also validated by the theoretical complexity and overhead analysis provided.
Fan Zhang 0003, Kai Hwang 0001, Keqin Li 0001, Samee Ullah Khan
IEEE Trans. Cloud Comput.1
2015 Hadoop Recognition of Biomedical Named Entity Using Conditional Random Fields
abstract
Processing large volumes of data has presented a challenging issue, particularly in data-redundant systems. As one of the most recognized models, the conditional random fields (CRF) model has been widely applied in biomedical named entity recognition (Bio-NER). Due to the internally sequential feature, performance improvement of the CRF model is nontrivial, which requires new parallelized solutions. By combining and parallelizing the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) and Viterbi algorithms, we propose a parallel CRF algorithm called MapReduce CRF (MRCRF) in this paper, which contains two parallel sub-algorithms to handle two time-consuming steps of the CRF model. The MapReduce L-BFGS (MRLB) algorithm leverages the MapReduce framework to enhance the capability of estimating parameters. Furthermore, the MapReduce Viterbi (MRVtb) algorithm infers the most likely state sequence by extending the Viterbi algorithm with another MapReduce job. Experimental results show that the MRCRF algorithm outperforms other competing methods by exhibiting significant performance improvement in terms of time efficiency as well as preserving a guaranteed level of correctness.
Kenli Li 0001, Wei Ai 0001, Zhuo Tang, Fan Zhang 0003, Lingang Jiang, Keqin Li 0001, Kai Hwang 0001
IEEE Trans. Parallel Distributed Syst.4
2014 Performance Variations in Resource Scaling for MapReduce Applications on Private and Public Clouds
abstract
In this paper, we delineate the causes of performance variations when scaling provisioned virtual resources for a variety of MapReduce applications. Hadoop MapReduce facilitates the development and execution processes of large-scale batch applications on big data. However, provisioning suitable resources to achieve desired performance at an affordable cost requires expertise into the execution model of MapReduce, the resources available for provisioning and the execution behavior of the application at hand. As an initial step towards automating this process, we characterize the difference in execution response for different MapReduce applications while varying the number of virtualized CPUs and memory resources, number of map slots as well as cluster size on a private cloud. This characterization helps illustrate the performance variation, 5x compared to 36x speedup, of Reduce-intensive and Map-intensive applications at effectively utilizing provisioned resources at different scales (1-64 VMs). By comparing the scalability efficiency, we clearly indicate the under-provisioning or over-provisioning of resources for different MapReduce applications at large scale.
Fan Zhang 0003, Majd F. Sakr
IEEE CLOUD1
2014 Multi-objective scheduling of many tasks in cloud platforms
Fan Zhang 0003, Keqin Li 0001, Samee Ullah Khan, Kai Hwang 0001
Future Gener. Comput. Syst.1
2013 Cluster-Size Scaling and MapReduce Execution Times
abstract
Understanding performance scalability in MapReduce applications presents a challenging problem. The difficulty lies in the distributed locations of input data and the distributed compute resources that utilize varied network substrates. User-defined Map and Reduce stages, with numerous application parameters, further complicate the problem. Using small datasets and limited test runs to understand how MapReduce applications will behave with "big data" can have a significant payoff. In this paper, we evaluate the impact of cluster-size scaling on execution time for a set of Map- and Reduce-intensive applications. We model the MapReduce framework, specify conditions and implications of power-law conformity, and verify our model with data from benchmark MapReduce applications. Empirical results indicate that: (1) within a range of scaling parameters, MapReduce execution times follow a power-law distribution. (2) Power-law scalability for Map-intensive applications starts from a small cluster size. (3) Shuffle-intensive applications exhibit power-law behavior starting from larger clusters. (4) Cluster-scaling performance gains fail to show power-law behavior when computing resources far exceed those needed. Our findings will facilitate using small-scale test runs to allocate and configure virtual and physical computing resources in large scale clouds.
Fan Zhang 0003, Majd F. Sakr
CloudCom (1)1
2013 ConMR: Concurrent MapReduce Programming Model for Large Scale Shared-Data Applications
abstract
The rapid growth of large-data processing has brought in the MapReduce programming model as a widely accepted solution. However, MapReduce limits itself to a one map-to-one-reduce framework. Meanwhile, it lacks built-in support and optimization when the input datasets are shared among concurrent applications and/or jobs. The performance might be improved when the shared and frequently accessed data is read from local instead of distributed file system.To enhance the performance of big data applications, this paper presents Concurrent MapReduce, a new programming model built on top of MapReduce that deals with large amount of shared data items. Concurrent MapReduce provides support for processing heterogeneous sources of input datasets and offers optimization when the datasets are partially or fully shared. Experimental evaluation has shown an execution runtime speedup of 4X compared to traditional nonconcurrent MapReduce implementation with a manageable time overhead.
Fan Zhang 0003, Qutaibah M. Malluhi, Tamer M. Elsyed
ICPP1
2011 Ordinal Optimized Scheduling of Scientific Workflows in Elastic Compute Clouds
abstract
Elastic compute clouds are best represented by the virtual clusters in Amazon EC2 or in IBM RC2. This paper proposes a simulation based approach to scheduling scientific workflows onto elastic clouds. Scheduling multitask workflows in virtual clusters is a NP-hard problem. Excessive simulations in months of time may be needed to produce the optimal schedule using Monte Carlo simulations. To reduce this scheduling overhead is necessary in real-time cloud computing. We present a new workflow scheduling method based on iterative ordinal optimization (IOO). This new method outperforms the Monte Carlo and Blind-Pick methods to yield higher performance against rapid workflow variations. For example, to execute 20,000 tasks on 128 virtual machines for gravitational wave analysis, an ordinal optimized schedule can be generated in a few minutes, which is O(103)~O(104) faster than using Monte Carlo simulations. The ordinal optimized schedule results in higher throughput with lower memory demand. The cloud experimental results being reported verified our theoretical findings on the relative performance of three workflow scheduling methods studied in this paper.
Fan Zhang 0003, Kai Hwang 0001, Cheng Wu 0002
CloudCom1
2011 Adaptive Virtual Machine Provisioning in Elastic Multi-tier Cloud Platforms
abstract
Virtual machines are allocated on demand in virtualized cloud platforms to provide flexible and reliable services. The major difficulty lies in satisfying the conflicting objectives of reducing response time while lowering resource costs. In this paper, a mathematical multi-tier framework for virtual machine allocation is proposed, which can be used to capture the performance of the cloud platform. We first use simulations to derive virtual resource allocation policies, and later use real benchmarking applications to verify the effectiveness of this framework. Experimental results show that the model can be simply and effectively used to satisfy the response time requirement as well as lowering the cost of using the virtual machine resources.
Fan Zhang 0003, James J. Mulcahy, Cheng Wu 0002
NAS1
2011 Formal Verification of Temporal Properties for Reduced Overhead in Grid Scientific Workflows
Fan Zhang 0003, Lianchen Liu, Cheng Wu 0002
J. Comput. Sci. Technol.2