Fangfei Chen

dblp:58/1240 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 8 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 95% Data integration and cleaning · 5%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 36% High-performance computing · 35% Parallel and multicore computing · 12%
Computer networks
3 papers
Content delivery and video streaming · 58% Wireless networking · 17% Internet architecture and protocols · 15%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 100%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
process mining
0.722020
Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020
Scientific Workflow Mining in Clouds · IEEE Trans. Parallel Distributed Syst. 2017
Services computing and microservices
workflow management
0.612022
Identifying a Minimum Sequence of High-Level Changes Between Workflows · IEEE Trans. Serv. Comput. 2022
Data mining › process mining
event log analysis
0.412020
Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020
Data mining › process mining
process discovery
0.412020
Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020
High-performance computing › scientific workflow
scientific workflow management
0.312017
Scientific Workflow Mining in Clouds · IEEE Trans. Parallel Distributed Syst. 2017
Content delivery and video streaming
content delivery network
0.112012
Intra-cloud lightning: Building CDNs in the cloud · INFOCOM 2012
Content delivery and video streaming
content placement
0.112012
Intra-cloud lightning: Building CDNs in the cloud · INFOCOM 2012
Wireless networking › wireless data broadcast › wireless data dissemination
mobile data dissemination
0.112012
Who, When, Where: Timeslot Assignment to Mobile Clients · IEEE Trans. Mob. Comput. 2012
Cloud and datacenter computing › datacenter services › online service systems › internet services
cloud-based content delivery
0.112012
Intra-cloud lightning: Building CDNs in the cloud · INFOCOM 2012
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.112012
Joint scheduling of processing and Shuffle phases in MapReduce systems · INFOCOM 2012
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
mapreduce scheduling
0.112012
Joint scheduling of processing and Shuffle phases in MapReduce systems · INFOCOM 2012
Electronic design automation › high-level synthesis
scheduling
0.112012
Who, When, Where: Timeslot Assignment to Mobile Clients · IEEE Trans. Mob. Comput. 2012
Parallel and multicore computing
task scheduling
0.112012
Joint scheduling of processing and Shuffle phases in MapReduce systems · INFOCOM 2012
High-performance computing
scientific workflow
0.112020
Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020
Data integration and cleaning
data provenance
0.112017
Scientific Workflow Mining in Clouds · IEEE Trans. Parallel Distributed Syst. 2017
Internet architecture and protocols
domain name system
0.112015
End-User Mapping: Next Generation Request Routing for Content Delivery · SIGCOMM 2015
Internet architecture and protocols › naming and addressing
mapping system
0.112015
End-User Mapping: Next Generation Request Routing for Content Delivery · SIGCOMM 2015

Methods — techniques the papers use, named apart from their topics

prom plug-in · 0.9prom · 0.6process mining · 0.6heuristics · 0.6edit distance · 0.6a* search · 0.6simulation · 0.3online and offline scheduling algorithms · 0.3heuristic algorithm · 0.3measurement study · 0.2heuristic · 0.1approximation algorithm · 0.1
YearPublicationVenuePosition
2022 Identifying a Minimum Sequence of High-Level Changes Between Workflows
abstract
Adaptive workflow management systems allow workflows to be changed in both the modeling and runtime stages, resulting in many workflow variants. Identifying a minimum sequence of high-level changes between two workflows represents a fundamental yet critical issue. The state-of-the-art approach utilizes digital logic to seek the optimal solution; however, this approach may face difficulties when advanced workflow patterns (e.g., loops) are involved, and it does not scale well. To address this problem, we first propose a naive approach that applies all valid changes to one workflow until the other workflow is found. Then, the approach is optimized from two aspects. First, we present advanced heuristics that significantly reduce the search space without pruning the optimal solution. Second, we employ the A$^\ast$search algorithm to direct the search procedure. Because the heuristic function used in the A$^\ast$algorithm is problem specific, we devise a consistent heuristic function to approximate the edit distance between two workflows, thereby accelerating the search. We implement our approach in a prototype tool and conduct extensive experiments on two data sets to evaluate its effectiveness and efficiency. The experimental results demonstrate that our approach outperforms the state of the art in terms of both application scope and scalability.
Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Chengzhen Zhang
IEEE Trans. Serv. Comput.2
2020 Scientific Workflow Protocol Discovery from Public Event Logs in Clouds
abstract
With the advancement of cloud computing, many challenging scientific problems can be solved using scientific workflow technology which integrates geo-distributed instruments, applications, and big data effectively and efficiently. For workflow collaboration, the workflow protocols of all participants are needed. However, workflow protocols are not always available and are often outdated as the workflow evolve frequently. To address this problem, we propose a novel workflow discovery approach which can extract up-to-date scientific workflow protocols from public event logs in clouds, without the need to access the full-fledged event logs involving private events. Our approach leverages transitive precedence relations between events to achieve this. We implement our approach as a ProM plug-in, and evaluate it through extensive experiments on event logs of real-world scientific workflows. The experimental results demonstrate that our approach requires a weaker completeness notion of event logs than the state-of-the-art do, and our approach derives the same workflow protocol from the public event log as that discovered from the original event log, and thus the private events can be protected.
Wei Song 0003, Hans-Arno Jacobsen, Fangfei Chen
IEEE Trans. Knowl. Data Eng.3
2017 Scientific Workflow Mining in Clouds
abstract
Computing clouds have become the platform of choice for the deployment and execution of scientific workflows. Due to the uncertainty and unpredictability of scientific exploration, the execution plan for a scientific workflow may vary from the definition. It is therefore of great significance to be able to discover actual workflows from execution histories (event logs) to reproduce experimental results and to establish provenance. However, most existing process mining techniques focus on discovering control flow-oriented business processes in a centralized environment, and thus, they are mostly inapplicable to the discovery of data flow-oriented, unstructured scientific workflows in distributed cloud environments. In this paper, we present Scientific Workflow Mining as a Service (SWMaaS) to support both intra-cloud and inter-cloud scientific workflow mining. The approach is implemented as a ProM plug-in and is evaluated on event logs derived from real-world scientific workflows. Through experimental results, we demonstrate the effectiveness and efficiency of our approach.
Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Xiaoxu Xia, Chunyang Ye, Xiaoxing Ma
IEEE Trans. Parallel Distributed Syst.2
2016 Effa: a proM plugin for recovering event logs
abstract
While event logs generated by business processes play an increasingly significant role in business analysis, the quality of data remains a serious problem. Automatic recovery of dirty event logs is desirable and thus receives more attention. However, existing methods only focus on missing event recovery, or fall short of efficiency. To this end, we present Effa, a ProM plugin, to automatically recover event logs in the light of process specifications. Based on advanced heuristics including process decomposition and trace replaying to search the minimum recovery, Effa achieves a balance between repairing accuracy and efficiency.
Xiaoxu Xia, Wei Song 0003, Fangfei Chen, Xuansong Li, Pengcheng Zhang 0001
Internetware3
2015 End-User Mapping: Next Generation Request Routing for Content Delivery
abstract
Content Delivery Networks (CDNs) deliver much of the world's web, video, and application content on the Internet today. A key component of a CDN is the mapping system that uses the DNS protocol to route each client's request to a ``proximal'' server that serves the requested content. While traditional mapping systems identify a client using the IP of its name server, we describe our experience in building and rolling-out a novel system called end-user mapping that identifies the client directly by using a prefix of the client's IP address. Using measurements from Akamai's production network during the roll-out, we show that end-user mapping provides significant performance benefits for clients who use public resolvers, including an eight-fold decrease in mapping distance, a two-fold decrease in RTT and content download time, and a 30% improvement in the time-to-first byte. We also quantify the scaling challenges in implementing end-user mapping such as the 8-fold increase in DNS queries. Finally, we show that a CDN with a larger number of deployment locations is likely to benefit more from end-user mapping than a CDN with a smaller number of deployments.
Fangfei Chen, Ramesh K. Sitaraman, Marcelo Torres
SIGCOMM1
2014 Throughput Maximization in Mobile WSN Scheduling With Power Control and Rate Selection
abstract
We study a data dissemination scenario in which data items are to be transmitted to mobile clients via one of the stationary data access points (APs) that the clients pass by en route to their destinations. The scheduler dedicates sequences of consecutive timeslots of an AP to downloading a data item to a client during the time window in which it is in range, which corresponds to assigning a job (the client's download) to a machine (the AP) among many. The transmission rate chosen for each assignment partly corresponds to setting a machine's speed, but it also has subtler effects. The APs may control transmission power to tune its transmission range making sure that no interference occurs with neighboring APs' transmissions. The problem is a generalization of an already NP-hard parallel-machine scheduling problem in which jobs' release times and deadlines depend on the machine to which they are assigned. We define this joint timeslot, power control, and rate assignment problem formally and apply both new algorithms and adaptations of existing algorithms to it. We evaluate these algorithms through simulations which show that our proposed algorithms achieve near-optimal throughput.
Yosef Alayev, Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta, Kin K. Leung
IEEE Trans. Wirel. Commun.2
2012 Throughput Maximization in Mobile WSN Scheduling with Power Control and Rate Selection
abstract
We study a data dissemination scenario in which data items are to be transmitted to mobile clients via one of the stationary data access points (APs) that the clients pass by en route to their destinations. The scheduler dedicates sequences of consecutive timeslots of an AP to downloading a data item to a client during the time window in which it is in range, which corresponds to assigning a job (the client's download) to a machine (the AP) among many. The transmission rate chosen for each assignment partly corresponds to setting a machine's speed, but it also has subtler effects. The APs may control transmission power to tune its transmission range making sure that no interference occurs with neighboring APs' transmissions. The problem is a generalization of an already NP-hard parallel-machine scheduling problem in which jobs' release times and deadlines depend on the machine to which they are assigned. We define this joint timeslot, power control, and rate assignment problem formally and apply both new algorithms and adaptations of existing algorithms to it. We evaluate these algorithms through simulations which show that our proposed algorithms achieve near-optimal throughput.
Yosef Alayev, Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta, Kin K. Leung
DCOSS2
2012 Resource Allocation with Stochastic Demands
abstract
Resources in modern computer systems include not only CPU, but also memory, hard disk, bandwidth, etc. To serve multiple users simultaneously, we need to satisfy their requirements in all resource dimensions. Meanwhile, their demands follow a certain distribution and may change over time. Our goal is then to admit as many users as possible to the system without violating the resource capacity more often than a predefined overflow probability. In this paper, we study the problem of allocating multiple resources among a group of users/tasks with stochastic demands. We model it as a stochastic multi-dimensional knapsack problem. We extend and apply the concept of effective bandwidth in order to solve this problem efficiently. Via numerical experiments, we show that our algorithms achieve near-optimal performance with specified overflow probability.
Fangfei Chen, Thomas La Porta, Mani Srivastava 0001
DCOSS1
2012 Intra-cloud lightning: Building CDNs in the cloud
abstract
Content distribution networks (CDNs) using storage clouds have recently started to emerge. Compared to traditional CDNs, storage cloud-based CDNs have the advantage of cost effectively offering hosting services to Web content providers without owning infrastructure. However, existing work on replica placement in CDNs does not readily apply in the cloud. In this paper, we investigated the joint problem of building distribution paths and placing Web server replicas in cloud CDNs to minimize the cost incurred on the CDN providers while satisfying QoS requirements for user requests. We formulate the cost optimization problem with accurate cost models and QoS requirements and show that the monthly cost can be as low as 2.62 US Dollars for a small Web site. We develop a suite of offline, online-static and online-dynamic heuristic algorithms that take as input network topology and work load information such as user location and request rates. We then evaluate the heuristics via Web trace-based simulation, and show that our heuristics behave very close to optimal under various network conditions.
Fangfei Chen, Katherine Guo, John Lin, Thomas La Porta
INFOCOM1
2012 Joint scheduling of processing and Shuffle phases in MapReduce systems
abstract
MapReduce has emerged as an important paradigm for processing data in large data centers. MapReduce is a three phase algorithm comprising of Map, Shuffle and Reduce phases. Due to its widespread deployment, there have been several recent papers outlining practical schemes to improve the performance of MapReduce systems. All these efforts focus on one of the three phases to obtain performance improvement. In this paper, we consider the problem of jointly scheduling all three phases of the MapReduce process with a view of understanding the theoretical complexity of the joint scheduling and working towards practical heuristics for scheduling the tasks. We give guaranteed approximation algorithms and outline several heuristics to solve the joint scheduling problem.
Fangfei Chen, Murali S. Kodialam, T. V. Lakshman
INFOCOM1
2012 Convergecast with aggregatable data classes
abstract
Data-gathering or convergecast problems have traditionally been studied in two combinations of settings: one-shot scheduling of data items with no aggregation, and periodic scheduling of data items with full aggregation meaning that any number of unit-size data items can, if available, be aggregated into a single (unit-size) data item (e.g., by summing or averaging values). In this paper, we extend beyond these problem settings in two ways. First, we study a) one-shot throughput maximization in settings with aggregation and b) periodic scheduling in settings without aggregation. Second, we generalize the notion of aggregatability in both one-shot and periodic scheduling beyond the binary choice of either all sets of items being aggregatable or none being so. Modeling the presence of multiple semantic data types (e.g., target counts to be summed and temperature readings to be averaged), we partition data items into classes, whereby items are aggregatable if they belong to the same class, in both periodic and non-periodic settings. For these two problems we provide guaranteed approximations and heuristics, for a variety of general and special cases. We then evaluate the algorithms in a systematic simulation study, both under the conditions in which our provable guarantees apply and in more general settings, where we find the algorithms continue to perform well on typical problem inputs.
Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta
SECON1
2012 Who, When, Where: Timeslot Assignment to Mobile Clients
abstract
We consider variations of a problem in which data must be delivered to mobile clients en route, as they travel toward their destinations. The data can only be delivered to the mobile clients as they pass within range of wireless base stations. Example scenarios include the delivery of building maps to firefighters responding to multiple alarms. We cast this scenario as a parallel-machine scheduling problem with the little-studied property that jobs may have different release times and deadlines when assigned to different machines. We present new algorithms and also adapt existing algorithms, for both online and offline settings. We evaluate these algorithms on a variety of problem instance types, using both synthetic and real-world data, including several geographical scenarios, and show that our algorithms produce schedules achieving near-optimal throughput.
Fangfei Chen, Matthew P. Johnson 0001, Yosef Alayev, Amotz Bar-Noy, Thomas La Porta
IEEE Trans. Mob. Comput.1
2012 Proactive data dissemination to mission sites
Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta
Wirel. Networks1
2009 Who, When, Where: Timeslot Assignment to Mobile Clients
abstract
We consider variations of a problem in which data must be delivered to mobile clients en-route, as they travel towards their destinations. The data can only be delivered to the mobile clients as they pass within range of wireless base stations. Example scenarios include the delivery of building maps to firefighters responding to multiple alarms, and the in-transit ldquoilluminationrdquo of simultaneous surface-to-air missiles. We cast this scenario as a parallel-machine scheduling problem with the little-studied property that jobs may have different release times and deadlines when assigned to different machines. We present new algorithms and also adapt existing algorithms, for both online and offline settings. We evaluate these algorithms on a variety of problem instance types, using both synthetic and real-world data, including several geographical scenarios, and show that our algorithms produce schedules achieving near-optimal throughput.
Fangfei Chen, Matthew P. Johnson 0001, Yosef Alayev, Amotz Bar-Noy, Thomas La Porta
MASS1
2009 Proactive Data Dissemination to Mission Sites
abstract
In many situations it is important to deliver information to personnel as they work in the field. We consider such a specialized content distribution application in wireless mesh networks. When a new mission arrives-for example, when an alarm for a fire is reported-data is pushed to storage nodes at the mission site where it may be retrieved locally by responding personnel (e.g., police, firefighters, paramedics, government officials, and the media). It is important that information is available at low latency, when requested or pulled by the personnel. The total latency experienced will be a combination of the push delay (if the personnel arrive at the mission site before all the data can be pushed), and the pull delay. Each delay component will in turn be a function of 1) the hop distance traveled by the data when pushed or pulled and 2) the congestion on the links. In this paper, we define algorithms and protocols that trade-off the push and pull latencies depending on the type of application. Our goal is to choose a storage node assignment minimizing the total latency-based cost. We start with a simple model in which cost is a function of distance, and then extend the model explicitly taking congestion into account. Since the problem is NP-hard to approximate, our focus is on developing efficient algorithms and distributed protocols that can be easily deployed in wireless mesh networks. In NS2 simulations, we find that our heuristic algorithms achieve on average a cost within at most 15 % of the optimum.
Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Iris Fermin, Thomas La Porta
SECON1
2008 Multiple Backhaul Mobile Access Router: Design and Experimentation
abstract
The multiple backhaul mobile access router aims to provide high capacity and high performance Internet access for emerging mobile wireless applications. In this paper we describe the framework and implementation of a modular mobile access router system. A set of backhaul interface monitoring APIs are provided to support flexible handover policies. We present the experimental results to validate the performance of our mobile access router and illustrate tradeoffs when designing handover policies.
Yan Sun 0007, Fangfei Chen, Thomas La Porta
ICC2