VLDB 2026 Research / reviewers in the wild / expert
Yuhong Feng
dblp:84/485
· DBLP profile ↗
23ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-7691-5587ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFDGraph: Privacy-Preserving Graph Processing for Large-Scale Collaborative Fraud DetectionabstractInternational audience Qiulin Wu, Amelie Chi Zhou, Tristan Allard, Shadi Ibrahim, Yuhong Feng, Lichun Li, Amr El Abbadi |
ICDE | 5 |
| 2025 | ORCAS: Obfuscation-Resilient Binary Code Similarity Analysis using Dominance Enhanced Semantic GraphabstractBinary code similarity analysis (BCSA) serves as a foundational technique for binary analysis tasks such as vulnerability detection and malware identification. Existing graph based BCSA approaches capture more binary code semantics and demonstrate remarkable performance. However, when code obfuscation is applied, the unstable control flow structure degrades their performance. To address this issue, we develop ORCAS, an Obfuscation-Resilient BCSA model based on Dominance Enhanced Semantic Graph (DESG). The DESG is an original binary code representation, capturing more binaries' implicit semantics without control flow structure, including inter-instruction relations (e.g., def-use), inter-basic block relations (i.e., dominance and post-dominance), and instruction-basic block relations. ORCAS takes binary functions from different obfuscation options, optimization levels, and instruction set architectures as input and scores their semantic similarity more robustly. Extensive experiments have been conducted on ORCAS against eight baseline approaches over the BinKit dataset. For example, ORCAS achieves an average 12.1% PR-AUC improvement when using combined three obfuscation options compared to the state-of-the-art approaches. In addition, an original obfuscated real-world vulnerability dataset has been constructed and released to facilitate a more comprehensive research on obfuscated binary code analysis. ORCAS outperforms the state-of-the-art approaches over this newly released real-world vulnerability dataset by up to a recall improvement of 43%. Yuhong Feng, Yixuan Cao 0002, Haiyue Feng |
CIKM | 2 |
| 2025 | TOPSIS-FGD: A Multidimensional Resource Scheduling Strategy for Fragmentation OptimizationabstractWith the widespread adoption of cloud computing and containerization technologies, large-scale clusters now host increasingly complex applications, yet the fragmentation of multidimensional resources (CPU, GPU, memory) has significantly degraded resource utilization efficiency. To address this challenge, this study investigates scheduling strategies for multidimensional resource fragmentation optimization. First, a unified quantitative metric is established by extending a task-statistics-based fragmentation measurement approach to threedimensional resource constraints (CPU, memory, GPU), enabling objective assessment of fragmentation levels. Subsequently, a multidimensional resource fragmentation optimization scheduling strategy is proposed based on the FGD framework. This approach employs TOPSIS to holistically address multidimensional resource fragmentation optimization. Building upon this methodology, we design and implement both two-dimensional and three-dimensional scheduling strategies. Comprehensive evaluation using an extended Kubernetes scheduler simulator with Alibaba production dataset reveals that: 1) All TOPSIS-based multi-dimensional strategies achieve performance comparable to or exceeding baseline approaches in their target dimensions while demonstrating synergistic effects.; 2) Two-dimensional strategies achieve optimal performance in their target dimensions but may degrade others 3) The three-dimensional optimization strategy TOPSIS-FGD-CGM can overcome the limitations of two-dimensional optimization strategies and demonstrates robust comprehensive performance. Haohan Chen, Bingjian Yao, Jieting Zhang, Zhijiao Xiao, Yuhong Feng, Shenghua Zhong |
HPCC | 6 |
| 2025 | HyMetricScaler: A Multi-Metric-Driven Hybrid Autoscaling Framework for KubernetesabstractExisting Kubernetes autoscaling solutions predominantly rely on single metrics with single-direction scaling (either horizontal or vertical), making it challenging to balance end-to-end latency and CPU utilization effectively. This paper introduces HyMetricScaler, a hybrid autoscaling framework. The framework employs a horizontal scaling component that makes decisions based on distributed tracing data and incorporates an exponential moving average mechanism to mitigate scaling oscillations, thereby optimizing end-to-end latency. Additionally, it utilizes a vertical scaling component based on PID control theory to adjust CPU quotas, improving CPU utilization. Experimental results demonstrate that HyMetricScaler significantly reduces oscillations in scaling decisions under high-load scenarios. In low-load scenarios, it explores the performance trade-off space by adjusting CPU target utilization. Through configuration adjustments of its components, the framework can achieve optimal end-to-end latency or CPU utilization, meeting various business requirements for different performance metrics. Comparative experiments against baseline algorithms validate the effectiveness of the proposed solution. Yuhong Feng, Jianming Li, Zhijiao Xiao |
HPCC | 1 |
| 2025 | Tech-ASan: Two-stage check for Address SanitizerabstractAddress Sanitizer (ASan) is a sharp weapon for detecting memory safety violations, including temporal and spatial errors hidden in C/C++ programs during execution.However, ASan incurs significant runtime overhead, which limits its efficiency in testing large software.The overhead mainly comes from sanitizer checks due to the frequent and expensive shadow memory access.Over the past decade, many methods have been developed to speed up ASan by eliminating and accelerating sanitizer checks, however, they either fail to adequately eliminate redundant checks or compromise detection capabilities.To address this issue, this paper presents Tech-ASan, a two-stage check based technique to accelerate ASan with safety assurance.First, we propose a novel two-stage check algorithm for ASan, which leverages magic value comparison to reduce most of the costly shadow memory accesses.Second, we design an efficient optimizer to eliminate redundant checks, which integrates a novel algorithm for removing checks in loops.Third, we implement Tech-ASan as a memory safety tool based on the LLVM compiler infrastructure.Our evaluation using the SPEC CPU2006 benchmark shows that Tech-ASan outperforms the state-of-theart methods with 33.70% and 17.89% less runtime overhead than ASan and ASan--, respectively.Moreover, Tech-ASan detects 56 fewer false negative cases than ASan and ASan--when testing on the Juliet Test Suite under the same redzone setting. Yixuan Cao 0002, Yuhong Feng, Chongyi Huang, Fangcao Jian, Xu Wang 0006 |
Internetware | 2 |
| 2025 | Flexible Computing: A New Framework for Improving Resource Allocation and Scheduling in Elastic ComputingabstractSince the advent of cloud computing, Elastic Computing (EC) has become the standard architecture for resource allocation and scheduling. EC typically allocates computing resources based on predefined specifications, such as virtual machine or container flavors. However, these flavors are often constrained by fixed CPU-to-memory ratios, which frequently fail to match the actual resource needs of applications. As a result, cloud providers experience high resource allocation rates nearing saturation ($> $80%) but with low utilization ($< $25%). This study introduces Flexible Computing (FC), a novel approach to resource allocation and scheduling. Unlike EC, FC allocates resources based on an application resource usage profile, derived from the historical resource consumption of workloads, rather than relying on fixed specifications. Additionally, FC incorporates a real-time performance degradation detection mechanism to address performance issues caused by the noisy-neighbor effect when colocated workloads interfere with each other. FC dynamically adjusts resource allocation according to actual usage, ensuring that application performance meets Service Level Agreements (SLAs), while preventing resource waste and performance degradation from improper resource over-commitment. Large-scale experimental validations conducted on the FC architecture within Huawei Cloud data centers demonstrate that, compared to EC, FC can reduce computing resource consumption by over 33% while managing the same workloads. Furthermore, FC's real-time performance degradation detection model achieves a prediction error of less than 5% across various testing environments, highlighting its commercial viability. Weipeng Cao, Jiongjiong Gu, Zhong Ming 0001, Zhiyuan Cai, Yuzhao Wang, Changping Ji, Zhijiao Xiao, Yuhong Feng, Liang-Jie Zhang |
IEEE Trans. Serv. Comput. | 8 |
| 2024 | PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive LearningabstractCounterfactually Augmented Data (CAD) involves creating new data samples by applying minimal yet sufficient modifications to flip the label of existing data samples to other classes.Training with CAD enhances model robustness against spurious features that happen to correlate with labels by spreading the casual relationships across different classes.Yet, recent research reveals that training with CAD may lead models to overly focus on modified features while ignoring other important contextual information, inadvertently introducing biases that may impair performance on out-ofdistribution (OOD) datasets.To mitigate this issue, we employ contrastive learning to promote global feature alignment in addition to learning counterfactual clues.We theoretically prove that contrastive loss can encourage models to leverage a broader range of features beyond those modified ones.Comprehensive experiments on two human-edited CAD datasets demonstrate that our proposed method outperforms the state-of-the-art on OOD datasets. Xiaoqi Qiu, Xu Guo 0002, Yu Yue, Yuhong Feng, Chunyan Miao |
ACL (1) | 6 |
| 2024 | Leveraging Large Language Models for Automated Chinese Essay Scoring
Haiyue Feng, Sixuan Du, Gaoxia Zhu, Yan Zou, Poh Boon Phua, Yuhong Feng, Haoming Zhong, Zhiqi Shen 0001, Siyuan Liu 0003 |
AIED (1) | 6 |
| 2024 | CRABS-former: CRoss-Architecture Binary Code Similarity Detection based on TransformerabstractBinary code similarity detection (BCSD) is widely used in software analysis such as vulnerability detection and malware identification. Among various forms of binary representation, assembly is particularly feasible for real-world applications due to its efficient preprocessing compared to graph and intermediate representation (IR). Existing assembly-based methods leverage the text embedding capabilities of pretrained language models such as BERT, which still encounter limitations in cross-architecture BCSD due to the characteristics of assembly code and the lack of cross-architecture vocabulary. In this paper, we first design several normalization strategies to preprocess assembly code from multiple instruction set architectures (ISAs), in order to decrease the token length of assembly code inputs and reduce the size of vocabulary, thereby improving processing efficiency and simplifying model structure. Then, we propose a method to collect token instances and construct a tokenizer capable of processing assembly code from multiple ISAs, enhancing the model’s ability to interpret such code. Based on this tokenizer, we develop a CRoss-Architecture Binary code Similarity detection model based on Transformer (CRABS-former). CRABS-former compares two binary functions from different ISAs, compilers or optimization options and computes their similarity score. Finally, we conduct experiments for two BCSD tasks (one-to-one and one-to-many) using CRABS-former, comparing its performance against four baselines: SAFE, Trex, jTrans, and TE3L. The results indicate that CRABS-former, with a pool size of 10,000, improves recall by 10.85%, 18.02%, and 3.33% across different ISAs, compilers, and optimizations, respectively, underscoring the effectiveness of our approach. Yuhong Feng, Yixuan Cao 0002, Haiyue Feng |
Internetware | 1 |
| 2024 | Robust Calibration of Vehicle Solid-State Lidar-Camera Perception System Using Line-Weighted Correspondences in Natural EnvironmentsabstractWith the rapid development of autonomous driving and SLAM technology, the perception system of a vehicle heavily relies on laser and image sensors to capture the real-world scenario and avoid obstacles autonomously. To achieve accurate and robust multi-sensor fusion computation, high-precision extrinsic calibration of camera and laser scanner is a necessary requirement. Traditional multi-sensor calibration methods based on manual features rely on specific scenarios and may not provide feature information over long distances. In this paper, we present a novel approach for robustly calibrating the extrinsic parameters of a solid-state(SS) lidar-camera system in a natural environment. Our proposed method begins with obtaining robust line feature information. we first innovatively employ a super-voxel clustering method to extract global 3D line features from the complete point cloud and then back-project these 3D line features into 2D space. Afterward, a transformer-based edge detection network, EDTER, is used to detect the edge features and estimate the probability pixel-by-pixel. To consider the uncertainty of two-dimensional line features and the inconsistency of residuals at different distances, we construct a line feature weight model for line feature residual calculation. Finally, we minimize the residual errors using least squares optimization to recover the relative pose of the camera and the lidar sensor. We conducted a performance study to compare our proposed method against existing targetless calibration methods on various natural scenarios. The experimental results demonstrate that our proposed method achieves higher robustness, accuracy, and consistency, making it suitable for real-world applications. Shengjun Tang, Xiaoming Li 0009, Zhihan Lyu, Yuhong Feng, Weixi Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | An Efficient Service-Aware Virtual Machine Scheduling Approach Based on Multi-Objective Evolutionary AlgorithmabstractService providers tend to deploy application services to several different virtual machines (VMs) to improve the scalability and manageability of the cloud data center (CDC). Therefore, high frequency communication traffic is always involved among those VMs that are deployed the same application service. In order to reduce the communication cost (CC) of CDC, all VMs running the same service should be redeployed in the same subnet as much as possible by using live migration technology, because CC between VMs in different subnets is much higher than that within the same subnet. On the other hand, the migration time (MT) to complete all migration tasks is also crucial for providers and customers, because a prolonged MT will lead to the increased maintenance cost and the deterioration of quality of service (QoS). To address the aforementioned issues, this paper proposes an efficient service-aware VM scheduling approach (ES-VSA) based on a multi-objective evolutionary algorithm to minimize CC and MT, simultaneously. Finally, experiments are conducted on four different scenarios, and the simulation results demonstrate that our proposed algorithm is superior to several state-of-the-art algorithms in terms of both CC and MT. Zhijiao Xiao, Qijie Qiu, Yuhong Feng, Qiuzhen Lin, Zhong Ming 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2022 | Latency-driven Optimization of Switching Pipeline Design in Network ChipsabstractA network switch implements multiple services and each service is formed by a number of match-action operations through several pipeline stages. These services running in the switch equipment are to process various packets based on standard internet protocols to decide the route of each packet. Data packets come in serial to a port, where each packet is processed by a service according to the contents of the packet headers and then send out via another port. Design of the switch, i.e., mapping services to physical resources in the pipeline stages, aims to achieve low switching latency with small chip area while respecting data-flow dependencies and hardware constraints. The current practice relies on expertise of engineers empirically, which is laborious and generates mediocre results. In this paper, we propose a switching pipeline design optimizatton technique, called SPOT. Our main contributions are as follows: (i) We first formulate the bi-objective (latency and chip area) constrained design optimization problem; (ii) SPOT quickly spots a feasible solution from a largely unfeasible design space using a dependency-aware greedy algorithm; (iii) Based on the above feasible seed, SPOT explores the design space with hundreds of decision dimensions towards Pareto optimal solutions using non-dominated sorting genetic algorithm II (NSGA-II) and multi-objective tabu search (MOTS), both adapted to be deployed in this problem setting. We apply SPOT on three sets of real-world network services. In comparison to the design sheets prepared by expert engineers, experiments show that SPOT offers 20.63% shorter service latency and 4.55% smaller chip area on average. As a by-product, the power consumption is lowered by 23.72% on average, which is correlated to the chip area. For hard real-time scenarios, the longest service latency a data packet may experience is the major concern. SPOT reduces the worst-case service latency by 12.65% on average. SPOT is the first automated optimization solution for switching pipeline design in network chips, being utilized in millions of network products of various kinds and saving manual efforts from days to minutes. Debayan Roy, Hui Chen 0016, Ping Xiang, Yuhong Feng, Wanli Chang 0001 |
RTSS | 7 |
| 2019 | Right Ventricle Segmentation of Cine MRI Using Residual U-net Convolutinal NetworksabstractRight ventricle (RV) segmentation is difficult due to the variable shape and ill-defined borders of the RV. In this paper, we propose a method to segment RV using a residual U-net convolutional network. A U-net shaped network structure is employed in our method to extract RV features in the encoding layers and make end-to-end decisions in the decoding layers. In the encoding layers, several residual blocks are cascaded extract RV features. In the decoding layers, convolutional layers are employed to make the RV predication. Our network is light with less parameters compared with state-of-art networks. Experiments on public datasets demonstrate that our network outperforms most existed automated segmentation method in respect of several commonly used evaluation measures. Zexiong Liu, Yuhong Feng |
PDCAT | 2 |
| 2017 | Optimize the FP-Tree Based Graph Edge Weight Computation on Multi-core MapReduce ClustersabstractThe FP-tree based edge weight computation (EWC for short) with MapReduce has demonstrated its remarkable performance for extracting weighted graphs from big data for data analysis. However, our investigation finds that existing algorithm includes unnecessary scan on the datasets as well as unnecessary information for the FP-tree construction, which prolong the runtime execution. In addition, applying inappropriate Reducers-to-cores mapping strategy may make it exhaust the resources and fail to complete the job execution. This paper designs, implements and evaluates an optimized FP-tree based graph EWC algorithm with MapReduce on Multi-core Clusters. First, we design a more compact FP-tree based EWC with 2-phase MapReduce, reducing one phase scan of the dataset. Second, we propose a reduced FP-tree data structure to reduce the FP-tree construction cost. Third, we examine two strategies for mapping Reducers to cores for EWC on each multi-core computer: {\em one-Reducer-one-core} and {\em one-Reducer-multiple-cores}. Finally, an empirical comparison performance study has been carried out on the optimized EWC algorithm against the existing one over a massive application dataset generated by a real social network. The results demonstrate that the optimized FP-tree based EWC algorithm obtains about 39\% to 55\% percentage improvement in execution time, and in the meantime achieves better scale-out and scale-up speedup. This paper's findings can also be applied to improve the scalability and efficiency of the parallel and distributed execution of applications involving large scale all-pairs set intersection computation over multi-core MapReduce clusters. Yuhong Feng, Meihong Guo, Kezhong Lu, Zhong Ming 0001, Haoming Zhong, Wentong Cai 0001, Zengxiang Li |
ICPADS | 1 |
| 2016 | The Edge Weight Computation with MapReduce for Extracting Weighted GraphsabstractAutomated weighted graph construction from massive data is essential to weighted graph theory based data mining processes, where the edge weight computation is time consuming or even fails to complete on a single machine when necessary resources are exhausted. In addition, existing work lacks of the measurement on the accuracy of the edge weights, which represents the graph accuracy and affects the following data mining results. This paper describes the classification, implementation and evaluation of edge weight computation algorithms with MapReduce Framework, which is a powerful parallel and distributed processing model. First, a classification of the edge weight computation algorithms is developed and how they can be applied on MapReduce is also discussed. Then we propose comprehensive measurements on the edge weight accuracy in terms of the number of edges, strength distribution, community structure, Hop-plot and effective diameters. Finally, a performance study has been conducted to evaluate these algorithms in terms of memory and disk usage, execution time and accuracy using a real massive social network application dataset. The results are presented and discussed. Our comparison results can help find out the most effective parallel and distributed edge weight computation algorithm for constructing a weighted graph for a given massive dataset. Yuhong Feng, Junpeng Wang 0008, Haoming Zhong, Zhong Ming 0001, Rui Mao 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | Comparing the learning effectiveness of BP, ELM, I-ELM, and SVM for corporate credit ratings
Haoming Zhong, Chunyan Miao, Zhiqi Shen 0001, Yuhong Feng |
Neurocomputing | 4 |
| 2014 | Fine-Grained Localization for Multiple Transceiver-Free Objects by using RF-Based TechnologiesabstractIn traditional radio-based localization methods, the target object has to carry a transmitter (e.g., active RFID), a receiver (e.g., 802.11 × detector), or a transceiver (e.g., sensor node). However, in some applications, such as safe guard systems, it is not possible to meet this precondition. In this paper, we propose a model of signal dynamics to allow the tracking of a transceiver-free object. Based on radio signal strength indicator (RSSI), which is readily available in wireless communication, three centralized tracking algorithms, and one distributed tracking algorithm are proposed to eliminate noise behaviors and improve accuracy. The midpoint and intersection algorithms can be applied to track a single object without calibration, while the best-cover algorithm has higher tracking accuracy but requires calibration. The probabilistic cover algorithm is based on distributed dynamic clustering. It can dramatically improve the localization accuracy when multiple objects are present. Our experimental test-bed is a grid sensor array based on MICA2 sensor nodes. The experimental results show that the localization accuracy for single object can reach about 0.8 m and for multiple objects is about 1 m. Dian Zhang 0001, Kezhong Lu, Rui Mao 0001, Yuhong Feng, Yunhuai Liu, Zhong Ming 0001, Lionel M. Ni |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2010 | Approximation algorithm for minimizing relay node placement in wireless sensor networks
Kezhong Lu, Guoliang Chen 0005, Yuhong Feng, Gang Liu 0028, Rui Mao 0001 |
Sci. China Inf. Sci. | 3 |
| 2008 | A Self-configuring Personal Agent Platform for Pervasive ComputingabstractMobile agent technologies have been widely used in distributed computing to take care of the task execution for the user. However, pervasive computing presents new challenges to existing mobile agent systems, especially the need for the context-aware self-configuring collaboration with the services provided by the physical objects. In order to address the problem, this paper presents a self-configuring personal agent platform to enable a mobile agent to adapt to the on-demand collaboration with the services. The platform consists of a Ubiquitous Intelligent Object (UIO) model the pervasive computing environment modeling, a code repository to provide executable codes which can be downloaded and instantiated as a mobile agent's capability at runtime, a service registry server for UIOs to publish and subscribe services, and personal agents, one for each individual user. A prototype of platform has been implemented as proof-of-concept and a preliminary performance study has also been carried out on it using a case study. Yuhong Feng, Jiannong Cao 0001, Ivan Lau, Xuan Liu 0001 |
EUC (1) | 1 |
| 2008 | Temporal fuzzy cognitive mapsabstractThis paper is concerned with the design, implementation, and evaluation of a novel extension of FCMs, temporal fuzzy cognitive maps (tFCMs). FCMs have advantages such as simplicity, supporting of inconsistent knowledge, and circle causalities for knowledge modeling and inference. However, the lack of the time dimension limits the usage of FCMs from the long term inference and the time related knowledge modeling. In order to narrow down the gap, a temporalized FCM is proposed to define a complete discrete temporal extension of the FCM. To reduce the complexities brought by the temporalization, a design approach is also introduced to construct the map, using simplified patterns and fuzzy logic based effect functions to capture fuzzy knowledge from domain experts. The paper also discusses how the errors in the fuzzy knowledge affect the result. Two different causality models are studied in theory and experiments to compare the error effects during the inferring. The result shows a significant difference in error accumulation in different causality models and gives a guideline to help users balance the error accumulation and sensitivities of their maps. Haoming Zhong, Chunyan Miao, Zhiqi Shen 0001, Yuhong Feng |
FUZZ-IEEE | 4 |
| 2008 | Execution coordination in mobile agent-based distributed job workflow execution
Yuhong Feng, Wentong Cai 0001 |
J. Syst. Archit. | 1 |
| 2007 | Dynamic partner identification in mobile agent-based distributed job workflow execution
Yuhong Feng, Wentong Cai 0001, Jiannong Cao 0001 |
J. Parallel Distributed Comput. | 1 |
| 2004 | MCCF: A Distributed Grid Job Workflow Execution Framework
Yuhong Feng, Wentong Cai 0001 |
ISPA | 1 |