VLDB 2026 Research / reviewers in the wild / expert
Jianfeng Mao
dblp:37/203
· DBLP profile ↗
20ranked-venue papers
2as first author
10since 2021 · last 2025
0000-0002-8969-0284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Data mining · 45% Information retrieval · 40% Data integration and cleaning · 15% | |
| Theoretical computer science
1 paper |
Algorithms and data structures · 100% |
Topics — the 9 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › similarity search › nearest neighbor search
maximum inner product search |
0.9 | 1 | 2025 | A Theory-Driven Approach to Inner Product Matrix Estimation for Incomplete Data: An Eigenvalue Perspective · WWW 2025 |
Information retrieval
similarity search |
0.9 | 1 | 2025 | A Theory-Driven Approach to Inner Product Matrix Estimation for Incomplete Data: An Eigenvalue Perspective · WWW 2025 |
Algorithms and data structures › similarity search
nearest neighbor search |
0.9 | 1 | 2025 | A Theory-Driven Approach to Inner Product Matrix Estimation for Incomplete Data: An Eigenvalue Perspective · WWW 2025 |
Data mining › clustering
affinity learning |
0.7 | 1 | 2023 | Boosting Spectral Clustering on Incomplete Data via Kernel Correction and Affinity Learning · NeurIPS 2023 |
Data mining
clustering |
0.7 | 1 | 2023 | Boosting Spectral Clustering on Incomplete Data via Kernel Correction and Affinity Learning · NeurIPS 2023 |
Data integration and cleaning
missing data |
0.7 | 1 | 2023 | Boosting Spectral Clustering on Incomplete Data via Kernel Correction and Affinity Learning · NeurIPS 2023 |
Data mining › clustering
spectral clustering |
0.7 | 1 | 2023 | Boosting Spectral Clustering on Incomplete Data via Kernel Correction and Affinity Learning · NeurIPS 2023 |
Energy-efficient computing › voltage scaling
dynamic voltage scaling |
0.1 | 1 | 2007 | Optimal Dynamic Voltage Scaling in Energy-Limited Nonpreemptive Systems with Real-Time Constraints · IEEE Trans. Mob. Comput. 2007 |
Embedded and real-time systems
real-time scheduling |
0.1 | 1 | 2007 | Optimal Dynamic Voltage Scaling in Energy-Limited Nonpreemptive Systems with Real-Time Constraints · IEEE Trans. Mob. Comput. 2007 |
Methods — techniques the papers use, named apart from their topics
random matrix theory · 1.7eigenvalue correction · 1.7spectral clustering · 0.7self-expressive framework · 0.7kernel correction · 0.7online optimization · 0.1critical task decomposition · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UltraTWD: Optimizing Ultrametric Trees for Tree-Wasserstein DistanceabstractThe Wasserstein distance is a widely used metric for measuring differences between distributions, but its super-cubic time complexity introduces substantial computational burdens. To mitigate this, the tree-Wasserstein distance (TWD) offers a linear-time approximation by leveraging a tree structure; however, existing TWD methods often compromise accuracy due to suboptimal tree structures and edge weights. To address it, we introduce UltraTWD, a novel unsupervised framework that simultaneously optimizes both ultrametric tree structures and edge weights to more faithfully approximate the cost matrix. Specifically, we develop algorithms based on minimum spanning trees, iterative projection, and gradient descent to efficiently learn high-quality ultrametric trees. Empirical results across document retrieval, ranking, and classification tasks demonstrate that UltraTWD achieves superior approximation accuracy and competitive downstream performance. Code is available at: https://github.com/NeXAIS/UltraTWD. Fangchen Yu, Yanzhen Chen, Jiaxing Wei, Jianfeng Mao, Wenye Li 0001, Qiang Sun 0007 |
ICML | 4 |
| 2025 | A Theory-Driven Approach to Inner Product Matrix Estimation for Incomplete Data: An Eigenvalue PerspectiveabstractAddressing the critical challenge of data incompleteness in inner product matrix estimation, we introduce a novel eigenvalue correction method designed to precisely reconstruct true inner product matrices from incomplete data. Utilizing random matrix theory, our method adjusts the eigenvalue distribution of the estimated inner product matrix to align with the ground truth. This approach significantly reduces estimation errors for both inner product matrices and the associated Euclidean distance matrices, thereby enhancing the effectiveness of similarity searches on incomplete data. Our method surpasses traditional data imputation and similarity calibration techniques in both maximum inner product search and nearest neighbor search tasks, demonstrating marked advancements in managing incomplete data. Fangchen Yu, Yicheng Zeng, Jianfeng Mao, Wenye Li 0001 |
WWW | 3 |
| 2025 | Semantic Encoding Algorithm for Classification and Retrieval of Aviation Safety ReportsabstractAutomated analysis of aviation safety reports is helpful in effectively preventing future accidents and improving emergency response capabilities. To date, there are no publicly available large-scale aviation text similarity datasets, which hinders the successful application of NLP techniques in the aviation domain. We present an automatically created aviation text similarity dataset consisting of more than 500,000 pairs for fine-tuning pretrained language models. Since technical terms have specialized meanings that differ from everyday language, we propose an efficient semantic encoding algorithm to improve the ability of embeddings to adequately represent aviation terms. We provide new solutions and revised evaluation metrics for the classification and the retrieval of safety reports, confirming the reliability of our dataset and the superiority of our algorithm.Note to Practitioners—Text representation is an essential task in natural language processing(NLP). A crucial step towards the successful application of NLP in safety reports analysis is to ensure that aviation texts are adequately encoded. Aiming at the problem of poor ability of current embeddings to represent technical terms, we automatically create an aviation text similarity dataset and propose a semantic encoding algorithm for aviation terms. It is clear that the proposed method has great potential in representation of technical terms, thus providing assistance for downstream tasks such as text classification, information retrieval and question answering. Yubing Gao, Guangyu Zhu 0001, Ya Duan, Jianfeng Mao |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Robustness Optimization of Air Transportation Network With Total Route Cost ConstraintabstractDesigning a robust air transportation network is critical for aviation activities to maintain highly efficient operations. Given certain total route cost, designing a strategy to improve the robustness of the network is a challenging issue. In this paper, first, a comparison experiment based on a small example shows the total effective resistance is a superior measure satisfying two critical criteria that other measures do not satisfy. Consequently, we explicitly formulate the network optimization problem as the minimization of total effective resistance under several constraints. The problem is often too time-consuming to solve to the optimal by exact methods. To achieve better efficiency, we propose a convex relaxation method for medium-scale networks making use of the problem properties. For large-scale networks, a clustering based convex relaxation method, reducing the dimensions of the network via selecting critical airports, is further proposed. To demonstrate the efficiency and efficacy of the proposed methods, simulations are performed on nine typical scale-free network and two real air transportation networks including Jetstar Asia Airway and domestic American Airlines.Note to Practitioners—The crucial need for designing robust air transportation networks in the aviation industry serves as the impetus for this work. The application scenarios range from the design of airport and flight operation networks for airlines, to air traffic networks with airways and waypoints, extending to drone cargo networks. An essential characteristic of a well-designed network is its ability to preserve connectivity, as much as possible, when it experiences disruptions such as flight cancellation or airway navigation equipment failure. The paper presents an argument for using total effective resistance as the key measure of network robustness, showing it to outperform other measures. We’ve developed specialized algorithms that leverage the unique properties of this optimization problem to minimize the total effective resistance of air transportation networks subject to certain total route cost to enhance the network structure. Although this study focuses on air transportation networks, the applicability of the algorithms extends beyond this field. Practitioners can apply them in other scenarios to design robust networks in various aviation operations. Changpeng Yang, Jianfeng Mao, Xiongwen Qian, Qi Wu 0003 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | DocReal: Robust Document Dewarping of Real-Life Images via Attention-Enhanced Control Point PredictionabstractDocument image dewarping is a crucial task in computer vision with numerous practical applications. The control point method, as a popular image dewarping approach, has attracted attention due to its simplicity and efficiency. However, inaccurate control point prediction due to varying background noises and deformation types can result in unsatisfactory performance. To address these issues, we propose a robust document dewarping approach for real-life images, namely DocReal, which utilizes Enet to effectively remove background noise and an attention-enhanced control point (AECP) module to better capture local deformations. Moreover, we augment the training data by synthesizing 2D images with 3D deformations and additional deformation types. Our proposed method achieves state-of-the-art performance on the DocUNet benchmark and a newly proposed benchmark of 200 Chinese distorted images, exhibiting superior dewarping accuracy, OCR performance, and robustness to various types of image distortion. Fangchen Yu, Yina Xie, Yafei Wen, Guozhi Wang, Shuai Ren 0002, Xiaoxin Chen 0001, Jianfeng Mao, Wenye Li 0001 |
WACV | 8 |
| 2024 | Reinforcement-Learning-Informed Prescriptive Analytics for Air Traffic Flow ManagementabstractAir Traffic Flow Management (ATFM) is a complex sequential decision-making problem that involves dynamically matching flights with sectors under changing environmental conditions. Finding an optimal solution for ATFM is challenging due to its dynamic nature and operational constraints. Reinforcement learning is a well-suited approach for sequential decision-making problems. However, ATFM poses three potential challenges: 1) large state space, 2) combinatorial action space, and 3) variational feasible action set, resulting from numerous agents with tightly-coupled constraints. These challenges can hinder the effectiveness of direct application of reinforcement learning methods. While prescriptive analytics can readily handle hard constraints via a mathematical optimization model, but it is computationally intractable for online sequential decision-making problems under changing environments. To address these challenges, we propose a novel framework, Reinforcement-Learning-Informed Prescriptive Analytics (RLIPA), in which an “informing” scheme is devised to integrate reinforcement learning and prescriptive analytics and leverage their strengths in predicting future reward and coping with hard constraints respectively. RLIPA is a general framework that can be adapted to other problems beyond ATFM, which typically involves many agents with tightly-coupled hard constraints. We demonstrate the usage and performance of RLIPA using numerical results and a real case study in comparison to two baseline approaches.Note to Practitioners—To improve Air Traffic Flow Management (ATFM) and reduce flight congestion, we propose a new method called reinforcement-learning-informed prescriptive analytics (RLIPA). RLIPA is a general framework that facilitates online sequential decision-making problems with multiple agents coupled with hard constraints. The approach consists of two stages: first, estimating future potential rewards for each agent via reinforcement learning, and second, informing the potential rewards to the following prescriptive analysis and using the information to construct and solve the downstream optimization problem dealing with hard coupling constraints among agents. Numerical experiments demonstrate the efficiency and effectiveness of RLIPA in the application of ATFM. In the most cases, RLIPA can offer more than 10x improvement in computational efficiency while maintaining or improving the level of optimality. The framework of RLIPA can be further extended to problems such as order dispatch in ride-hailing systems and food delivery. Yuan Wang 0069, Weilin Cai, Yilei Tu, Jianfeng Mao |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2023 | Highly-Efficient Robinson-Foulds Distance Estimation with Matrix CorrectionabstractPhylogenetic trees are essential in studying evolutionary relationships, and the Robinson-Foulds (RF) distance is a widely used metric to calculate pairwise dissimilarities between phylogenetic trees, with various applications in both the biology and computing communities. However, generating a precise RF distance matrix becomes difficult or even intractable when tree information is partially missing. To address this issue, we introduce a novel distance correction algorithm for estimating the RF distance matrix of incomplete phylogenetic trees. Our method innovatively harnesses the assumption of Euclidean embedding, correcting an approximate distance matrix into a valid distance metric, guaranteed to be closer to the unknown ground-truth. Despite its simplicity, our approach exhibits robust performance, efficiency, and scalability in empirical evaluations, outperforming classical distance correction algorithms and holding potential benefits in downstream applications. Our code is available at https://github.com/CUHKSZ-Yu/EMC. Fangchen Yu, Rui Bao, Jianfeng Mao, Wenye Li 0001 |
ECAI | 3 |
| 2023 | From Incompleteness to Unity: A Framework for Multi-view Clustering with Missing Values
Fangchen Yu, Jianfeng Mao, Wenye Li 0001 |
ICONIP (11) | 4 |
| 2023 | Boosting Spectral Clustering on Incomplete Data via Kernel Correction and Affinity LearningabstractSpectral clustering has gained popularity for clustering non-convex data due to its simplicity and effectiveness. It is essential to construct a similarity graph using a high-quality affinity measure that models the local neighborhood relations among the data samples. However, incomplete data can lead to inaccurate affinity measures, resulting in degraded clustering performance. To address these issues, we propose an imputation-free framework with two novel approaches to improve spectral clustering on incomplete data. Firstly, we introduce a new kernel correction method that enhances the quality of the kernel matrix estimated on incomplete data with a theoretical guarantee, benefiting classical spectral clustering on pre-defined kernels. Secondly, we develop a series of affinity learning methods that equip the self-expressive framework with $\ell_p$-norm to construct an intrinsic affinity matrix with an adaptive extension. Our methods outperform existing data imputation and distance calibration techniques on benchmark datasets, offering a promising solution to spectral clustering on incomplete data in various real-world applications. Fangchen Yu, Jicong Fan 0001, Yicheng Zeng, Jianfeng Mao, Wenye Li 0001 |
NeurIPS | 7 |
| 2023 | Online estimation of similarity matrices with incomplete dataabstractThe similarity matrix measures pairwise similarities between a set of data points and is an essential concept in data processing, routinely used in practical applications. Obtaining a similarity matrix is typically straightforward when data points are completely observed. However, incomplete observations can make it challenging to obtain a high-quality similarity matrix, which becomes even more complex in online data. To address this challenge, we propose matrix correction algorithms that leverage the positive semi-definiteness (PSD) of the similarity matrix to improve similarity estimation in both offline and online scenarios. Our approaches have a solid theoretical guarantee of performance and excellent potential for parallel execution on large-scale data. Empirical evaluations demonstrate their high effectiveness and efficiency with significantly improved results over classical imputation-based methods, benefiting downstream applications with superior performance. Our code is available at \url{https://github.com/CUHKSZ-Yu/OnMC}. Fangchen Yu, Yicheng Zeng, Jianfeng Mao, Wenye Li 0001 |
UAI | 3 |
| 2020 | Energy-Efficient Elevating Transfer Vehicle Routing for Automated Multi-Level Material Handling SystemsabstractWe investigate an energy-efficient elevating transfer vehicle routing problem (ETVRP), in which an elevating transfer vehicle (ETV) serves a multi-level freight handling system to transport cargo containers between airside and landside in an air cargo terminal. The problem can be regarded as a special case of stacker crane problem defined on a regular grid graph constructed by uniform rectangular tiles. Even with the special grid network structure, the ETVRP is still NP-hard in general. We manage to identify a subset of the ETVRP instances that are polynomially solvable based on the condition of free-permutation. For general ETVRPs, we propose a new and more efficient exact formulation whose dimensionality does not constantly increase with the number of requests and is bounded by the size of the underlying grid network. To further enhance computational efficiency, we develop two approximation algorithms, one of which is asymptotically optimal and has the time-complexity that grows linearly with the number of requests; the other has a bounded time-complexity and works better for instances with smaller arc lengths. Combining these two algorithms can guarantee an approximation ratio of 5/3. The performances of the proposed formulation and the approximation algorithms are further examined through numerical simulations. Note to Practitioners-Elevating transfer vehicles (ETVs) are widely utilized in automated multi-level material handling systems for transporting, storing, and retrieving units vertically and horizontally. We develop an efficient and effective method to reduce energy consumption of operating ETV, which is critical for improving both financial and environmental sustainability of these systems. The method can be adopted to realize online timely decision making for automatic vehicle routing when facing a huge number of pickup and delivery requests in multi-level material handling systems. A class of popular application scenarios, termed “Free-Permutation,” is identified and mathematically characterized, in which the energy consumption can be exactly minimized in a theoretical sense within tractable computational times. For general application scenarios, an efficient and robust method is designed to guarantee to save the energy consumption to a level lower than 67% above the theoretical minimum level. Preliminary numerical experiments suggest the efficiency and effectiveness of this approach, but it has not yet been incorporated into nor tested in a practical system. In the future research, we will approach the application scenarios with two or more ETVs simultaneously operated in a multi-level material handling system. Jianfeng Mao |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2019 | Designing Robust Air Transportation Networks via Minimizing Total Effective ResistanceabstractDesigning a robust air transportation network is an ongoing research effort that seeks to improve the extent of a network being connected against failures and attacks. The total effective resistance can be a promising measure for network robustness as demonstrated by case studies conducted in this paper. To enhance the robustness of air transportation networks, we consider to solve a flight route selection problem in which a set of routes is chosen from a candidate route set to minimize the utility function defined by the total effective resistance. Since it is an integer nonlinear programming problem, to balance the tradeoff between optimality performance and computational efficiency, two methods that implement the total effective resistance are developed to suit different network scales. For small/medium-scale networks, we develop an interior-point method based on convex relaxation and duality gap, which achieves a near optimality performance within acceptable time. For large-scale networks, we develop an accelerated greedy algorithm based on proved monotone submodularity, which can substantially reduce the computation time in deriving a good solution with a guaranteed optimality gap. Their optimality performance and computational efficiency are verified and compared through numerical results. Moreover, three case studies from real-world examples are also performed to demonstrate the application of the proposed methods for different network scales. Changpeng Yang, Jianfeng Mao, Xiongwen Qian |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Hierarchical Decentralized Optimization Architecture for Economic Dispatch: A New Approach for Large-Scale Power SystemabstractIn this paper, a new hierarchical decentralized optimization architecture is proposed to solve the economic dispatch problem for a large-scale power system. Conventionally, such a problem is solved in a centralized way, which is usually inflexible and costly in computation. In contrast to centralized algorithms, in this paper we decompose the centralized problem into local problems. Each local generator only solves its own problem iteratively, based on its own cost function and generation constraint. An extra coordinator agent is employed to coordinate all the local generator agents. Besides, it also takes responsibility to handle the global demand supply constraint based on a newly proposed concept named virtual agent. In this way, different from existing distributed algorithms, the global demand supply constraint and local generation constraints are handled separately, which would greatly reduce the computational complexity. In addition, as only local individual estimate is exchanged between the local agent and the coordinator agent, the communication burden is reduced and the information privacy is also protected. It is theoretically shown that under proposed hierarchical decentralized optimization architecture, each local generator agent can obtain the optimal solution in a decentralized fashion. Several case studies implemented on the IEEE 30-bus and the IEEE 118-bus are discussed and tested to validate the proposed method. Fanghong Guo, Changyun Wen, Jianfeng Mao, Jiawei Chen 0002, Yongduan Song 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2016 | Hybrid BF-PSO and fuzzy support vector machine for diagnosis of fatigue status using EMG signal features
Qi Wu 0003, Jianfeng Mao, Chuanfeng Wei, Shan Fu, Rob Law 0001, Bi-Ting Yu, Jia Bo, Changpeng Yang |
Neurocomputing | 2 |
| 2015 | Distributed Cooperative Secondary Control for Voltage Unbalance Compensation in an Islanded MicrogridabstractThis paper presents a distributed cooperative control scheme for voltage unbalance compensation (VUC) in an islanded microgrid (MG). By letting each distributed generator (DG) share the compensation effort cooperatively, unbalanced voltage in sensitive load bus (SLB) can be compensated. The concept of contribution level (CL) for compensation is first proposed for each local DG to indicate its compensation ability. A two-layer secondary compensation architecture consisting of a communication layer and a compensation layer is designed for each local DG. A totally distributed strategy involving information sharing and exchange is proposed, which is based on finite-time average consensus and newly developed graph discovery algorithm. This strategy does not require the whole system structure as a prior and can detect the structure automatically. The proposed scheme not only achieves similar VUC performance to the centralized one, but also brings some advantages, such as communication fault tolerance and plug-and-play property. Case studies including communication failure, CL variation, and DG plug-and-play are discussed and tested to validate the proposed method. Fanghong Guo, Changyun Wen, Jianfeng Mao, Jiawei Chen 0002, Yongduan Song 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2014 | Optimal Control of Multilayer Discrete Event Systems With Real-Time Constraint GuaranteesabstractWe consider discrete event systems (DESs) involving tasks with dependability requirements in the form of real-time constraints. We seek to control their processing times so as to satisfy these constraints while also minimizing a given cost function. When tasks are processed by a single resource, it has been shown that there are structural properties of the optimal state trajectory for this problem that lead to the critical task decomposition algorithm (CTDA) with a time complexity of O(N2). For a DES with multiple resources, we consider a multilayer network where each layer contains multiple nodes, each node may have multiple inputs and multiple outputs, and tasks are processed so that the real-time constraints apply on an end-to-end basis. Extending earlier results (where each layer contained a single node), we derive structural properties of the optimal solution that lead to the idea of introducing “virtual” deadlines at each node (except for the last layer) and decouple nodes so that the CTDA for single-node problems can be used. We prove that an appropriately constructed sequence of solutions of these simpler problems converges to the global optimum of the original problem and hence obtain an efficient scalable multilayer virtual deadline algorithm (MLVDA). We illustrate the efficiency of the MLVDA through numerical examples. Jianfeng Mao, Christos G. Cassandras |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2008 | Time separations of cyclic event rule systems with min-max timing constraints
Qianchuan Zhao, Jianfeng Mao |
Theor. Comput. Sci. | 2 |
| 2007 | Optimal Dynamic Voltage Scaling in Energy-Limited Nonpreemptive Systems with Real-Time ConstraintsabstractDynamic voltage scaling is used in energy-limited systems as a means of conserving energy and prolonging their life. We consider a setting in which the tasks performed by such a system are nonpreemptive and aperiodic. Our objective is to control the processing rate over different tasks so as to minimize energy subject to hard real-time processing constraints. Under any given task scheduling policy, we prove that the optimal solution to the offline version of the problem can be efficiently obtained by exploiting the structure of optimal sample paths, leading to a new dynamic voltage scaling algorithm termed the critical task decomposition algorithm (CTDA). The efficiency of the algorithm rests on the existence of a set of critical tasks that decompose the optimal sample path into decoupled segments within which optimal processing times are easily determined. The algorithm is readily extended to an online version of the problem as well. Its worst-case complexity of both offline and online problems is O(N2) Jianfeng Mao, Christos G. Cassandras, Qianchuan Zhao |
IEEE Trans. Mob. Comput. | 1 |
| 2001 | Spoken dialogue management as planning and acting under uncertaintyabstractSome stochastic models like Markov decision process (MDP) are used to model the dialogue manager. MDP-based system degrades fast when uncertainty about user’s intention increases. We propose a novel dialogue model based on the partially observable Markov decision process (POMDP). We use hidden system states and user intentions as the state set, parser results and low-level information as the observation set, domain actions and dialogue repair actions as the action set. Here the low-level information is extracted from different input modals using Bayesian networks. Because of the limitation of exact algorithms, we focus on heuristic methods and their applicability in dialogue management. Bo Zhang 0025, Qingsheng Cai, Jianfeng Mao, Eric Chang, Baining Guo |
INTERSPEECH | 3 |
| 2001 | Planning and Acting under Uncertainty: A New Model for Spoken Dialogue System
Bo Zhang 0025, Qingsheng Cai, Jianfeng Mao, Baining Guo |
UAI | 3 |