VLDB 2026 Research / reviewers in the wild / expert
Xiao-Feng Li
dblp:81/3637
· DBLP profile ↗
19ranked-venue papers
3as first author
1since 2021 · last 2023
0000-0002-1799-844XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9Artificial intelligence and machine learning · 3 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
3 papers |
Runtime systems and virtual machines · 74% Compilers and program optimization · 26% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 45% Memory systems · 41% Performance modeling and evaluation · 14% |
Topics — the 7 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Runtime systems and virtual machines › garbage collection
compaction |
0.1 | 1 | 2012 | Packer: Parallel Garbage Collection Based on Virtual Spaces · IEEE Trans. Computers 2012 |
Runtime systems and virtual machines
garbage collection |
0.1 | 1 | 2012 | Packer: Parallel Garbage Collection Based on Virtual Spaces · IEEE Trans. Computers 2012 |
Runtime systems and virtual machines › garbage collection
parallel garbage collection |
0.1 | 1 | 2012 | Packer: Parallel Garbage Collection Based on Virtual Spaces · IEEE Trans. Computers 2012 |
Compilers and program optimization › parallelization
speculative parallelization |
0.0 | 1 | 2004 | A cost-driven compilation framework for speculative parallelization of sequential programs · PLDI 2004 |
Compilers and program optimization › parallelization
thread-level speculation |
0.0 | 1 | 2004 | A cost-driven compilation framework for speculative parallelization of sequential programs · PLDI 2004 |
Memory systems
memory management |
0.0 | 1 | 2012 | Packer: Parallel Garbage Collection Based on Virtual Spaces · IEEE Trans. Computers 2012 |
Routing and switching
packet forwarding |
0.0 | 1 | 2005 | Shangri-La: achieving high performance from compiled network applications while enabling ease of programming · PLDI 2005 |
Methods — techniques the papers use, named apart from their topics
DAG traversal parallelization · 0.3software-controlled caching · 0.1scalar optimization · 0.1custom stack model · 0.1runtime speculation · 0.1cost-driven compilation · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | NEOP: A Framework for Distributed Mobile Apps on Heterogeneous DevicesabstractToday’s apps on a mobile device, such as a smartphone and a tablet, need to access various resources to deliver quality service to users’ satisfaction. These resources may include cameras, microphones, screens, processors, various specialized sensors, and data. In today’s client-server framework, resources accessible to an app are limited to those available in the device running the app, on the cloud, and likely in a few statically connected devices. However, there can be abundant resources on devices near the app-running one with desirable functionalities that can enable or empower the app’s new features and services, but cannot be easily accessed and leveraged. The NEOP (Neutron Operation Platform) framework is an app development and execution environment that removes the barrier across the devices. Heterogeneous IoT devices make the capabilities in their hardware and service software available after security and privacy authentication. An app is developed as a composition of capabilities distributed across various end devices and the cloud. Its constituent computing tasks can be dynamically created and scheduled. Different device capabilities can be selectively and dynamically recruited into the app for the optimal user experience. In this paper we describe example scenarios that motivate the next-generation app framework, the framework’s architecture, design principles, technical challenges, and details on its design and implementation. We also compare this work with related efforts on distributed mobile computing to highlight the unique contributions made by the NEOP platform. Song Jiang 0001, Weidong Zhong, Lizhong Wang, Xiao-Feng Li |
ISADS | 5 |
| 2018 | Development and Validation of Empirical Wave Retrieval Algorithms for Sentinel-1 Synthetic Aperture Radar in HH-PolarizationabstractIn our work, three empirical algorithms for significant wave height (SWH) retrieval have been tuned for horizontal-horizontal (HH) polarization Sentinel-1 SAR. We extracted more than ten thousand sub-scenes from available 200 images, which were treated as a dataset with collocated SWH from European Centre for Medium-Range Weather Forecasts (ECMWF) reanalysis wave data at a 0.125 ° grid. Two algorithms are based on the relation between SWH and azimuthal cutoff wavelength named CSAR_WAVEs and the other is called XWAVE which was originally explored for SWH retrieval from X-band SAR data. The three empirical algorithms have been tuned for HH-polarization Sentinel-1 SAR through the dataset. Additional 83 HH-polarization Sentinel-1 SAR images were carried out to retrieve SWH by using the four algorithms. Although the three algorithms allow estimating SWH from HH-polarization Sentinel-1 SAR, the algorithm herein called CSAR_WAVE_S is recommended for use due to less comparative error of a 0.57m STD of SWH. Weizeng Shao, Xiao-Feng Li, Zhanfeng Sun, Juncheng Zuo |
IGARSS | 2 |
| 2017 | Using Stacked Auto-encoder to Get Feature with Continuity and Distinguishability in Multi-object Tracking
Haoyang Feng, Xiao-Feng Li, Peixin Liu |
ICIG (1) | 2 |
| 2014 | Application of the fuzzy gain scheduling IMC-PID for the boiler pressure controlabstractIn this paper, the use of a Fuzzy Gain Scheduling IMC-PID (FGS+IMC-PID) scheme has been presented based on fuzzy performance degree coefficient n self-adjusting controller for the improvement of IMC-PID control. It is shown that the IMC-PID controller with the Fuzzy PID parameters Gain Scheduler provides satisfactory closed-loop responses with less overshoot and shorter rising times in case of both set point disturbance and the plant/model mismatch. Simulations are given, in which the proposed method is compared to other PID tuning methods (IMC, SPMG). The proposed scheme is suitable to implement in the complex process control system in power generation, since it does not demand significant computing resources. It has already been implemented through Function Code in many typical DCS (EDPU, XDPS, Ovation etc). The industrial applications show that the scheme achieves better performance in specific load variation range. Xiao-Feng Li, Shi-He Chen, Ruiyuan Wu |
FUZZ-IEEE | 1 |
| 2014 | The IMC-PID controller design for TITO process using closed-loop identification methodabstractA new auto-tuning method is presented in this paper for the TITO process with time delay. The first step is to identify the process model and the second step is to design the IMC-PID controller. In the identifying procedure, both the bias relay test and the idea relay test are involved to identify the model. The important feature of the proposed method is that it does not require the prior information about the process and the steady-state process gain can be directly obtained. In the second step, this method is used to design the IMC-PID controller of the TITO control system. The new auto-tuning scheme is implemented through the function code in typical DCSs (EDPF-NT, OVATION, Symphony, etc) and is used in 300MW fossil-fuel units. The industrial applications show that the scheme achieves better performance in specific load variation range. Xiao-Feng Li, Ruiyuan Wu, Weidong Zhang 0004 |
ICARCV | 1 |
| 2012 | A novel Gaussian Scale Space-based joint MGRF framework for precise lung segmentationabstractA new framework for the precise segmentation of lung tissues from Computed Tomography (CT) is proposed. The CT images, Gaussian Scale Space (GSS) data generation using Gaussian Kernels (GKs), and desired maps of regions (lung and the other chest tissues) are described by a joint Markov-Gibbs Random Field Model (MGRF) of independent image signals and interdependent region labels. We focus on the most accurate model identification of the joint MGRF models. To better specify region borders, each empirical distribution of signals is rigorously approximated by a Linear Combination of Discrete Gaussians (LCDG) with positive and negative components. The classical Expectation-Maximization (EM) algorithm has been adapted for the LCDG model. The initial segmentations from the original and the generated GSS CT images are based on the LCDG-models; then they are iteratively refined using an MGRF model with analytically estimated potentials. Finally, these initial segmentations are fused together using a Bayesian fusion approach to get the final segmentation of the lung region. Experiments on eleven real data sets based on Dice Similarity Coefficient (DSC) metric confirms the high accuracy of the proposed approach. Behnoush Abdollahi, Ahmed Soliman 0001, Ali Cahid Civelek, Xiao-Feng Li, Georgy L. Gimel'farb, Ayman El-Baz |
ICIP | 4 |
| 2012 | Packer: Parallel Garbage Collection Based on Virtual SpacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and normal object space (non-LOS). When either space is full, garbage collection is triggered even though the other space may still have plenty of room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection inefficient on multiple spaces. To address these problems, we propose Packer, a parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physical space. With multiple virtual spaces, Packer offers efficient memory management. With one physical space, Packer avoids the problem of an inefficient space utilization. To reduce the garbage collection pause time, we also propose a novel parallelization method that is applicable to multiple virtual spaces. Specifically, we reduce the compacting GC parallelization problem into a discreted acyclic graph (DAG) traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Jie Tang 0003, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IEEE Trans. Computers | 4 |
| 2012 | Achieving middleware execution efficiency: hardware-assisted garbage collection operationsabstractAlthough virtualization technologies bring many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we achieved an order of magnitude speedup and significant improvement on energy efficiency. In addition, the results of our performance estimation study indicate that the hardware-assisted GC instructions can reduce the GC execution time by half and lead to a 7% improvement on the overall execution time. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
J. Supercomput. | 4 |
| 2011 | TypeCastor: demystify dynamic typing of JavaScript applicationsabstractDynamic typing is a barrier for JavaScript applications to achieve high performance. Compared with statically typed languages, the major overhead of dynamic typing comes from runtime type resolution and runtime property lookup. Common folks' belief is that the traditional static compilation techniques are no longer effective for dynamic languages. The best known JavaScript engines such as Mozilla TraceMonkey and Chrome V8 have developed non-traditional techniques to reduce the runtime overhead. This paper describes TypeCastor, a new JavaScript engine that tries to investigate where and how much the dynamism really is in JavaScript applications, thus to demystify their dynamic typing behavior. To verify our findings, we evaluate TypeCastor with SunSpider benchmark. For type resolution, we find 99% of all the primitive type instances can be statically identified before the program execution. For object property lookup, more than 97% of all runtime property accesses can be satisfied by inline cache. These data mean that the representative JavaScript applications are not that dynamic as people expect, although the language provides the flexible dynamism supports. Though not developed for pure performance, TypeCastor achieves 5.6% and 12.7% higher scores compared to current Chrome V8 and Mozilla TraceMonkey engines respectively. Buqi Cheng, Xiao-Feng Li |
HiPEAC | 3 |
| 2011 | Improve Google Android User Experience with Regional Garbage Collection
Yunan He, Xiao-Feng Li |
NPC | 3 |
| 2010 | Hardware-assisted middleware: Acceleration of garbage collection operationsabstractAlthough the virtualization technology brings many benefits to cloud computing environments, as the virtual machines provide more features, the middleware layer has become bloated, introducing a high overhead. Our ultimate goal is to provide hardware-assisted solutions to improve the middleware performance in cloud computing environments. As a starting point, in this paper, we design, implement, and evaluate specialized hardware instructions to accelerate GC operations. We select GC because it is a common component in virtual machine designs and it incurs high performance and energy consumption overheads. We performed a profiling study on various GC algorithms to identify the GC performance hotspots, which contribute to more than 50% of the total GC execution time. By moving these hotspot functions into hardware, we managed to achieve an order of magnitude speedup. Jie Tang 0003, Shaoshan Liu, Zhimin Gu, Xiao-Feng Li, Jean-Luc Gaudiot |
ASAP | 4 |
| 2010 | Vectorization for Java
Jiutao Nie, Buqi Cheng, Ligang Wang 0001, Xiao-Feng Li |
NPC | 5 |
| 2009 | Packer: An innovative space-time-efficient parallel garbage collection algorithm based on virtual spacesabstractThe fundamental challenge of garbage collector (GC) design is to maximize the recycled space with minimal time overhead. For efficient memory management, in many GC designs the heap is divided into large object space (LOS) and non-large object space (non-LOS). When one of the spaces is full, garbage collection is triggered even though the other space may still have a lot of free room, thus leading to inefficient space utilization. Also, space partitioning in existing GC designs implies different GC algorithms for different spaces. This not only prolongs the pause time of garbage collection, but also makes collection not efficient on multiple spaces. To address these problems, we propose Packer, a space-and-time-efficient parallel garbage collection algorithm based on the novel concept of virtual spaces. Instead of physically dividing the heap into multiple spaces, Packer manages multiple virtual spaces in one physically shared space. With multiple virtual spaces, Packer offers the advantage of efficient memory management. At the same time, with one physically shared space, Packer avoids the problem of inefficient space utilization. To reduce the garbage collection pause time of Packer, we also propose a novel parallelization method that is applicable to multiple virtual spaces. We reduce the compacting GC parallelization problem into a tree traversal parallelization problem, and apply it to both normal and large object compaction. Shaoshan Liu, Ligang Wang 0001, Xiao-Feng Li, Jean-Luc Gaudiot |
IPDPS | 3 |
| 2008 | Runtime Engine for Dynamic Profile Guided Stride Prefetching
Qiong Zou, Xiao-Feng Li, Long-Bing Zhang |
J. Comput. Sci. Technol. | 2 |
| 2007 | A Throughput-Driven Task Creation and Mapping for Network Processors
Xiao-Feng Li, Michael K. Chen 0001, Roy Dz-Ching Ju |
HiPEAC | 2 |
| 2007 | Task-pushing: a Scalable Parallel GC Marking Algorithm without Synchronization OperationsabstractThis paper describes a scalable parallel marking technique for garbage collection that does not employ any synchronization operation. To achieve good scalability, two major design issues have to be resolved in parallel marking algorithm, i.e., the overhead of synchronization operations and load balance. This paper presents task-pushing, a novel parallel marking algorithm where each thread proactively gives up its spare tasks to other threads. Enlightened by the idea of communicating sequential process (CSP), task-pushing arranges the computation into a process network, eliminating synchronization operations in the whole marking process. Load balance is achieved by dripping tasks from thread local mark-stack for other threads to execute. To the best of our knowledge, this is the first parallel marking algorithm that completely avoids the synchronization primitives. We evaluated task-pushing in aspects of queuing efficiency, load balancing strategy, synchronization overhead, and overall scalability. The results on a 16-way Intel Xeon machine showed that task-pushing has better scalability than work-stealing technique with pseudojbb and GCOld server-kind Java benchmarks. Xiao-Feng Li |
IPDPS | 2 |
| 2006 | Fuzzy Optimization Control System and its Application in Ball Mill Pulverizing SystemabstractThe pulverizing process of a ball mill is rather complex. It is difficult to describe the pulverizing process of a ball mill with precise mathematics model for its complex characteristics. A fuzzy control method including the data process, characteristic parameter detection, operating mode recognition, fuzzy optimization control, Fuzzy PID controller was put forward in this paper aimed to design a Fuzzy optimized control system for the pulverizing system. The operation result shows that it is quite suitable for the control of mill and a large amount of electric energy is saved. Xiao-Feng Li, Yu-Xing Zeng, Hui-Yan Wu |
FUZZ-IEEE | 1 |
| 2005 | Shangri-La: achieving high performance from compiled network applications while enabling ease of programmingabstractProgramming network processors is challenging. To sustain high line rates, network processors have extremely tight memory access and instruction budgets. Achieving desired performance has traditionally required hand-coded assembly. Researchers have recently proposed high-level programming languages for packet processing, but the challenges of compiling these languages into code that is competitive with hand-tuned assembly remain unanswered.This paper describes the Shangri-La compiler, which accepts a packet program written in a C-like high-level language and applies scalar and specialized optimizations to generate a highly optimized binary. Hot code paths identified by profiling are mapped across processing elements to maximize processor utilization. Since our compilation target has no hardware caches, software-controlled caches are generated for frequently accessed application data structures. Packet handling optimizations significantly reduce per-packet memory access and instruction counts. Finally, a custom stack model maps stack frames to the fastest levels of the target processor's heterogeneous memory hierarchy.Binaries generated by the compiler were evaluated on the Intel IXP2400 network processor with eight packet processing cores and eight threads per core. Our results show the importance of both traditional and specialized optimization techniques for achieving the maximum forwarding rates on three network applications, L3-Switch, MPLS and Firewall. Michael K. Chen 0001, Xiao-Feng Li, Ruiqi Lian, Jason H. Lin, Roy Dz-Ching Ju |
PLDI | 2 |
| 2004 | A cost-driven compilation framework for speculative parallelization of sequential programsabstractThe emerging hardware support for thread-level speculation opens new opportunities to parallelize sequential programs beyond the traditional limits. By speculating that many data dependences are unlikely during runtime, consecutive iterations of a sequential loop can be executed speculatively in parallel. Runtime parallelism is obtained when the speculation is correct. To take full advantage of this new execution model, a program needs to be programmed or compiled in such a way that it exhibits high degree of speculative thread-level parallelism. We propose a comprehensive cost-driven compilation framework to perform speculative parallelization. Based on a misspeculation cost model, the compiler aggressively transforms loops into optimal speculative parallel loops and selects only those loops whose speculative parallel execution is likely to improve program Zhao-Hui Du, Chu-Cheow Lim, Xiao-Feng Li, Qingyu Zhao, Tin-Fook Ngai |
PLDI | 3 |