VLDB 2026 Research / reviewers in the wild / expert
Fangfei Chen
dblp:58/1240
· DBLP profile ↗
16ranked-venue papers
9as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 8 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 95% Data integration and cleaning · 5% | |
| Computer architecture, parallel and distributed computing, and storage systems
5 papers |
Cloud and datacenter computing · 36% High-performance computing · 35% Parallel and multicore computing · 12% | |
| Computer networks
3 papers |
Content delivery and video streaming · 58% Wireless networking · 17% Internet architecture and protocols · 15% | |
| Software engineering, system software, and programming languages
1 paper |
Services computing and microservices · 100% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
process mining |
0.7 | 2 | 2020 | Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020 Scientific Workflow Mining in Clouds · IEEE Trans. Parallel Distributed Syst. 2017 |
Services computing and microservices
workflow management |
0.6 | 1 | 2022 | Identifying a Minimum Sequence of High-Level Changes Between Workflows · IEEE Trans. Serv. Comput. 2022 |
Data mining › process mining
event log analysis |
0.4 | 1 | 2020 | Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020 |
Data mining › process mining
process discovery |
0.4 | 1 | 2020 | Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020 |
High-performance computing › scientific workflow
scientific workflow management |
0.3 | 1 | 2017 | Scientific Workflow Mining in Clouds · IEEE Trans. Parallel Distributed Syst. 2017 |
Content delivery and video streaming
content delivery network |
0.1 | 1 | 2012 | Intra-cloud lightning: Building CDNs in the cloud · INFOCOM 2012 |
Content delivery and video streaming
content placement |
0.1 | 1 | 2012 | Intra-cloud lightning: Building CDNs in the cloud · INFOCOM 2012 |
Wireless networking › wireless data broadcast › wireless data dissemination
mobile data dissemination |
0.1 | 1 | 2012 | Who, When, Where: Timeslot Assignment to Mobile Clients · IEEE Trans. Mob. Comput. 2012 |
Cloud and datacenter computing › datacenter services › online service systems › internet services
cloud-based content delivery |
0.1 | 1 | 2012 | Intra-cloud lightning: Building CDNs in the cloud · INFOCOM 2012 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling |
0.1 | 1 | 2012 | Joint scheduling of processing and Shuffle phases in MapReduce systems · INFOCOM 2012 |
Cloud and datacenter computing › cluster resource management and scheduling › cluster scheduling
mapreduce scheduling |
0.1 | 1 | 2012 | Joint scheduling of processing and Shuffle phases in MapReduce systems · INFOCOM 2012 |
Electronic design automation › high-level synthesis
scheduling |
0.1 | 1 | 2012 | Who, When, Where: Timeslot Assignment to Mobile Clients · IEEE Trans. Mob. Comput. 2012 |
Parallel and multicore computing
task scheduling |
0.1 | 1 | 2012 | Joint scheduling of processing and Shuffle phases in MapReduce systems · INFOCOM 2012 |
High-performance computing
scientific workflow |
0.1 | 1 | 2020 | Scientific Workflow Protocol Discovery from Public Event Logs in Clouds · IEEE Trans. Knowl. Data Eng. 2020 |
Data integration and cleaning
data provenance |
0.1 | 1 | 2017 | Scientific Workflow Mining in Clouds · IEEE Trans. Parallel Distributed Syst. 2017 |
Internet architecture and protocols
domain name system |
0.1 | 1 | 2015 | End-User Mapping: Next Generation Request Routing for Content Delivery · SIGCOMM 2015 |
Internet architecture and protocols › naming and addressing
mapping system |
0.1 | 1 | 2015 | End-User Mapping: Next Generation Request Routing for Content Delivery · SIGCOMM 2015 |
Methods — techniques the papers use, named apart from their topics
prom plug-in · 0.9prom · 0.6process mining · 0.6heuristics · 0.6edit distance · 0.6a* search · 0.6simulation · 0.3online and offline scheduling algorithms · 0.3heuristic algorithm · 0.3measurement study · 0.2heuristic · 0.1approximation algorithm · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Identifying a Minimum Sequence of High-Level Changes Between WorkflowsabstractAdaptive workflow management systems allow workflows to be changed in both the modeling and runtime stages, resulting in many workflow variants. Identifying a minimum sequence of high-level changes between two workflows represents a fundamental yet critical issue. The state-of-the-art approach utilizes digital logic to seek the optimal solution; however, this approach may face difficulties when advanced workflow patterns (e.g., loops) are involved, and it does not scale well. To address this problem, we first propose a naive approach that applies all valid changes to one workflow until the other workflow is found. Then, the approach is optimized from two aspects. First, we present advanced heuristics that significantly reduce the search space without pruning the optimal solution. Second, we employ the A$^\ast$search algorithm to direct the search procedure. Because the heuristic function used in the A$^\ast$algorithm is problem specific, we devise a consistent heuristic function to approximate the edit distance between two workflows, thereby accelerating the search. We implement our approach in a prototype tool and conduct extensive experiments on two data sets to evaluate its effectiveness and efficiency. The experimental results demonstrate that our approach outperforms the state of the art in terms of both application scope and scalability. Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Chengzhen Zhang |
IEEE Trans. Serv. Comput. | 2 |
| 2020 | Scientific Workflow Protocol Discovery from Public Event Logs in CloudsabstractWith the advancement of cloud computing, many challenging scientific problems can be solved using scientific workflow technology which integrates geo-distributed instruments, applications, and big data effectively and efficiently. For workflow collaboration, the workflow protocols of all participants are needed. However, workflow protocols are not always available and are often outdated as the workflow evolve frequently. To address this problem, we propose a novel workflow discovery approach which can extract up-to-date scientific workflow protocols from public event logs in clouds, without the need to access the full-fledged event logs involving private events. Our approach leverages transitive precedence relations between events to achieve this. We implement our approach as a ProM plug-in, and evaluate it through extensive experiments on event logs of real-world scientific workflows. The experimental results demonstrate that our approach requires a weaker completeness notion of event logs than the state-of-the-art do, and our approach derives the same workflow protocol from the public event log as that discovered from the original event log, and thus the private events can be protected. Wei Song 0003, Hans-Arno Jacobsen, Fangfei Chen |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | Scientific Workflow Mining in CloudsabstractComputing clouds have become the platform of choice for the deployment and execution of scientific workflows. Due to the uncertainty and unpredictability of scientific exploration, the execution plan for a scientific workflow may vary from the definition. It is therefore of great significance to be able to discover actual workflows from execution histories (event logs) to reproduce experimental results and to establish provenance. However, most existing process mining techniques focus on discovering control flow-oriented business processes in a centralized environment, and thus, they are mostly inapplicable to the discovery of data flow-oriented, unstructured scientific workflows in distributed cloud environments. In this paper, we present Scientific Workflow Mining as a Service (SWMaaS) to support both intra-cloud and inter-cloud scientific workflow mining. The approach is implemented as a ProM plug-in and is evaluated on event logs derived from real-world scientific workflows. Through experimental results, we demonstrate the effectiveness and efficiency of our approach. Wei Song 0003, Fangfei Chen, Hans-Arno Jacobsen, Xiaoxu Xia, Chunyang Ye, Xiaoxing Ma |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2016 | Effa: a proM plugin for recovering event logsabstractWhile event logs generated by business processes play an increasingly significant role in business analysis, the quality of data remains a serious problem. Automatic recovery of dirty event logs is desirable and thus receives more attention. However, existing methods only focus on missing event recovery, or fall short of efficiency. To this end, we present Effa, a ProM plugin, to automatically recover event logs in the light of process specifications. Based on advanced heuristics including process decomposition and trace replaying to search the minimum recovery, Effa achieves a balance between repairing accuracy and efficiency. Xiaoxu Xia, Wei Song 0003, Fangfei Chen, Xuansong Li, Pengcheng Zhang 0001 |
Internetware | 3 |
| 2015 | End-User Mapping: Next Generation Request Routing for Content DeliveryabstractContent Delivery Networks (CDNs) deliver much of the world's web, video, and application content on the Internet today. A key component of a CDN is the mapping system that uses the DNS protocol to route each client's request to a ``proximal'' server that serves the requested content. While traditional mapping systems identify a client using the IP of its name server, we describe our experience in building and rolling-out a novel system called end-user mapping that identifies the client directly by using a prefix of the client's IP address. Using measurements from Akamai's production network during the roll-out, we show that end-user mapping provides significant performance benefits for clients who use public resolvers, including an eight-fold decrease in mapping distance, a two-fold decrease in RTT and content download time, and a 30% improvement in the time-to-first byte. We also quantify the scaling challenges in implementing end-user mapping such as the 8-fold increase in DNS queries. Finally, we show that a CDN with a larger number of deployment locations is likely to benefit more from end-user mapping than a CDN with a smaller number of deployments. Fangfei Chen, Ramesh K. Sitaraman, Marcelo Torres |
SIGCOMM | 1 |
| 2014 | Throughput Maximization in Mobile WSN Scheduling With Power Control and Rate SelectionabstractWe study a data dissemination scenario in which data items are to be transmitted to mobile clients via one of the stationary data access points (APs) that the clients pass by en route to their destinations. The scheduler dedicates sequences of consecutive timeslots of an AP to downloading a data item to a client during the time window in which it is in range, which corresponds to assigning a job (the client's download) to a machine (the AP) among many. The transmission rate chosen for each assignment partly corresponds to setting a machine's speed, but it also has subtler effects. The APs may control transmission power to tune its transmission range making sure that no interference occurs with neighboring APs' transmissions. The problem is a generalization of an already NP-hard parallel-machine scheduling problem in which jobs' release times and deadlines depend on the machine to which they are assigned. We define this joint timeslot, power control, and rate assignment problem formally and apply both new algorithms and adaptations of existing algorithms to it. We evaluate these algorithms through simulations which show that our proposed algorithms achieve near-optimal throughput. Yosef Alayev, Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta, Kin K. Leung |
IEEE Trans. Wirel. Commun. | 2 |
| 2012 | Throughput Maximization in Mobile WSN Scheduling with Power Control and Rate SelectionabstractWe study a data dissemination scenario in which data items are to be transmitted to mobile clients via one of the stationary data access points (APs) that the clients pass by en route to their destinations. The scheduler dedicates sequences of consecutive timeslots of an AP to downloading a data item to a client during the time window in which it is in range, which corresponds to assigning a job (the client's download) to a machine (the AP) among many. The transmission rate chosen for each assignment partly corresponds to setting a machine's speed, but it also has subtler effects. The APs may control transmission power to tune its transmission range making sure that no interference occurs with neighboring APs' transmissions. The problem is a generalization of an already NP-hard parallel-machine scheduling problem in which jobs' release times and deadlines depend on the machine to which they are assigned. We define this joint timeslot, power control, and rate assignment problem formally and apply both new algorithms and adaptations of existing algorithms to it. We evaluate these algorithms through simulations which show that our proposed algorithms achieve near-optimal throughput. Yosef Alayev, Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta, Kin K. Leung |
DCOSS | 2 |
| 2012 | Resource Allocation with Stochastic DemandsabstractResources in modern computer systems include not only CPU, but also memory, hard disk, bandwidth, etc. To serve multiple users simultaneously, we need to satisfy their requirements in all resource dimensions. Meanwhile, their demands follow a certain distribution and may change over time. Our goal is then to admit as many users as possible to the system without violating the resource capacity more often than a predefined overflow probability. In this paper, we study the problem of allocating multiple resources among a group of users/tasks with stochastic demands. We model it as a stochastic multi-dimensional knapsack problem. We extend and apply the concept of effective bandwidth in order to solve this problem efficiently. Via numerical experiments, we show that our algorithms achieve near-optimal performance with specified overflow probability. Fangfei Chen, Thomas La Porta, Mani Srivastava 0001 |
DCOSS | 1 |
| 2012 | Intra-cloud lightning: Building CDNs in the cloudabstractContent distribution networks (CDNs) using storage clouds have recently started to emerge. Compared to traditional CDNs, storage cloud-based CDNs have the advantage of cost effectively offering hosting services to Web content providers without owning infrastructure. However, existing work on replica placement in CDNs does not readily apply in the cloud. In this paper, we investigated the joint problem of building distribution paths and placing Web server replicas in cloud CDNs to minimize the cost incurred on the CDN providers while satisfying QoS requirements for user requests. We formulate the cost optimization problem with accurate cost models and QoS requirements and show that the monthly cost can be as low as 2.62 US Dollars for a small Web site. We develop a suite of offline, online-static and online-dynamic heuristic algorithms that take as input network topology and work load information such as user location and request rates. We then evaluate the heuristics via Web trace-based simulation, and show that our heuristics behave very close to optimal under various network conditions. Fangfei Chen, Katherine Guo, John Lin, Thomas La Porta |
INFOCOM | 1 |
| 2012 | Joint scheduling of processing and Shuffle phases in MapReduce systemsabstractMapReduce has emerged as an important paradigm for processing data in large data centers. MapReduce is a three phase algorithm comprising of Map, Shuffle and Reduce phases. Due to its widespread deployment, there have been several recent papers outlining practical schemes to improve the performance of MapReduce systems. All these efforts focus on one of the three phases to obtain performance improvement. In this paper, we consider the problem of jointly scheduling all three phases of the MapReduce process with a view of understanding the theoretical complexity of the joint scheduling and working towards practical heuristics for scheduling the tasks. We give guaranteed approximation algorithms and outline several heuristics to solve the joint scheduling problem. Fangfei Chen, Murali S. Kodialam, T. V. Lakshman |
INFOCOM | 1 |
| 2012 | Convergecast with aggregatable data classesabstractData-gathering or convergecast problems have traditionally been studied in two combinations of settings: one-shot scheduling of data items with no aggregation, and periodic scheduling of data items with full aggregation meaning that any number of unit-size data items can, if available, be aggregated into a single (unit-size) data item (e.g., by summing or averaging values). In this paper, we extend beyond these problem settings in two ways. First, we study a) one-shot throughput maximization in settings with aggregation and b) periodic scheduling in settings without aggregation. Second, we generalize the notion of aggregatability in both one-shot and periodic scheduling beyond the binary choice of either all sets of items being aggregatable or none being so. Modeling the presence of multiple semantic data types (e.g., target counts to be summed and temperature readings to be averaged), we partition data items into classes, whereby items are aggregatable if they belong to the same class, in both periodic and non-periodic settings. For these two problems we provide guaranteed approximations and heuristics, for a variety of general and special cases. We then evaluate the algorithms in a systematic simulation study, both under the conditions in which our provable guarantees apply and in more general settings, where we find the algorithms continue to perform well on typical problem inputs. Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta |
SECON | 1 |
| 2012 | Who, When, Where: Timeslot Assignment to Mobile ClientsabstractWe consider variations of a problem in which data must be delivered to mobile clients en route, as they travel toward their destinations. The data can only be delivered to the mobile clients as they pass within range of wireless base stations. Example scenarios include the delivery of building maps to firefighters responding to multiple alarms. We cast this scenario as a parallel-machine scheduling problem with the little-studied property that jobs may have different release times and deadlines when assigned to different machines. We present new algorithms and also adapt existing algorithms, for both online and offline settings. We evaluate these algorithms on a variety of problem instance types, using both synthetic and real-world data, including several geographical scenarios, and show that our algorithms produce schedules achieving near-optimal throughput. Fangfei Chen, Matthew P. Johnson 0001, Yosef Alayev, Amotz Bar-Noy, Thomas La Porta |
IEEE Trans. Mob. Comput. | 1 |
| 2012 | Proactive data dissemination to mission sites
Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Thomas La Porta |
Wirel. Networks | 1 |
| 2009 | Who, When, Where: Timeslot Assignment to Mobile ClientsabstractWe consider variations of a problem in which data must be delivered to mobile clients en-route, as they travel towards their destinations. The data can only be delivered to the mobile clients as they pass within range of wireless base stations. Example scenarios include the delivery of building maps to firefighters responding to multiple alarms, and the in-transit ldquoilluminationrdquo of simultaneous surface-to-air missiles. We cast this scenario as a parallel-machine scheduling problem with the little-studied property that jobs may have different release times and deadlines when assigned to different machines. We present new algorithms and also adapt existing algorithms, for both online and offline settings. We evaluate these algorithms on a variety of problem instance types, using both synthetic and real-world data, including several geographical scenarios, and show that our algorithms produce schedules achieving near-optimal throughput. Fangfei Chen, Matthew P. Johnson 0001, Yosef Alayev, Amotz Bar-Noy, Thomas La Porta |
MASS | 1 |
| 2009 | Proactive Data Dissemination to Mission SitesabstractIn many situations it is important to deliver information to personnel as they work in the field. We consider such a specialized content distribution application in wireless mesh networks. When a new mission arrives-for example, when an alarm for a fire is reported-data is pushed to storage nodes at the mission site where it may be retrieved locally by responding personnel (e.g., police, firefighters, paramedics, government officials, and the media). It is important that information is available at low latency, when requested or pulled by the personnel. The total latency experienced will be a combination of the push delay (if the personnel arrive at the mission site before all the data can be pushed), and the pull delay. Each delay component will in turn be a function of 1) the hop distance traveled by the data when pushed or pulled and 2) the congestion on the links. In this paper, we define algorithms and protocols that trade-off the push and pull latencies depending on the type of application. Our goal is to choose a storage node assignment minimizing the total latency-based cost. We start with a simple model in which cost is a function of distance, and then extend the model explicitly taking congestion into account. Since the problem is NP-hard to approximate, our focus is on developing efficient algorithms and distributed protocols that can be easily deployed in wireless mesh networks. In NS2 simulations, we find that our heuristic algorithms achieve on average a cost within at most 15 % of the optimum. Fangfei Chen, Matthew P. Johnson 0001, Amotz Bar-Noy, Iris Fermin, Thomas La Porta |
SECON | 1 |
| 2008 | Multiple Backhaul Mobile Access Router: Design and ExperimentationabstractThe multiple backhaul mobile access router aims to provide high capacity and high performance Internet access for emerging mobile wireless applications. In this paper we describe the framework and implementation of a modular mobile access router system. A set of backhaul interface monitoring APIs are provided to support flexible handover policies. We present the experimental results to validate the performance of our mobile access router and illustrate tradeoffs when designing handover policies. Yan Sun 0007, Fangfei Chen, Thomas La Porta |
ICC | 2 |