Chaoyi Ma

dblp:219/6220 · DBLP profile ↗
← Back
23ranked-venue papers
4as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 11 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Dual-stream frequency-domain framework with contextual graph enhancer and consensus-difference fusion for cross-view geo-localization
Haitong Li, Chaoyi Ma, Yuehuan Wang, Ruonan Wei
J. Vis. Commun. Image Represent.3
2025 Towards Guaranteed Accuracy for Flow Spread Measurement with $(\epsilon, \beta)$-Nonduplicate Sampling
Haibo Wang 0004, Chaoyi Ma, Dimitrios Melissourgos, Guoju Gao, Shigang Chen
INFOCOM2
2025 AntAkso: Claims Management System for Health Insurance in Alipay
abstract
The rapid growth of health insurance and the rising incidence of fraudulent claims underscore the necessity for an efficient and professional claims management system. However, there is a noticeable lack of shared relevant experience from previous research in this field. In response to this challenge, we introduce AntAkso, a robust claims management system specifically designed for health insurance operations within Alipay. AntAkso incorporates a digital and professional management system, achieving a notable decrease in the volume of false claims, reduction in administrative costs, and heightened satisfaction among its policyholders. We begin by highlighting the core components of this system, including the case stratification, hospital recommendation, and case dispatch modules, along with the pivotal algorithms employed, i.e., the fraud detection, recommendation, and robust satisficing algorithms. We also detail the system's implementation and deployment. We substantiate the proposed system's effectiveness and efficiency with empirical evidence from experiments on a large set of real-world health insurance claims data.
Qitao Shi, Jun Zhou 0011, Ya-Lin Zhang 0001, Chaoyi Ma, Yifan Wu 0020, Xiaobo Qin
KDD (1)5
2024 Cost-Efficient Fraud Risk Optimization with Submodularity in Insurance Claim
abstract
The fraudulent insurance claim is critical for the insurance industry.Insurance companies or agency platforms aim to confidently estimate the fraud risk of claims by gathering data from various sources.Although more data sources can improve the estimation accuracy, they inevitably lead to increased costs.Therefore, a great challenge of fraud risk verification lies in well balancing these two aspects.To this end, this paper proposes a framework named cost-efficient fraud risk optimization with submodularity (CEROS) to optimize the process of fraud risk verification.CEROS efficiently allocates investigation resources across multiple information sources, balancing the trade-off between accuracy and cost.CEROS consists of two parts that we propose: a submodular set-wise classification model * Equal Contribution.
Zhibo Zhu, Chaoyi Ma, Hong Qian, Xingyu Lu 0004, Yangwenhui Zhang, Xiaobo Qin, Binjie Fei, Jun Zhou 0011, Aimin Zhou
KDD3
2023 Attention Weighted Mixture of Experts with Contrastive Learning for Personalized Ranking in E-commerce
abstract
Ranking model plays an essential role in e-commerce search and recommendation. An effective ranking model should give a personalized ranking list for each user according to the user preference. Existing algorithms usually extract a user representation vector from the user behavior sequence, then feed the vector into a feed-forward network (FFN) together with other features for feature interactions, and finally produce a personalized ranking score. Despite tremendous progress in the past, there is still room for improvement. Firstly, the personalized patterns of feature interactions for different users are not explicitly modeled. Secondly, most of existing algorithms have poor personalized ranking results for long-tail users with few historical behaviors due to the data sparsity.To overcome the two challenges, we propose Attention Weighted Mixture of Experts (AW-MoE) with contrastive learning for personalized ranking. Firstly, AW-MoE leverages the MoE framework to capture personalized feature interactions for different users. To model the user preference, the user behavior sequence is simultaneously fed into expert networks and the gate network. Within the gate network, one gate unit and one activation unit are designed to adaptively learn the fine-grained activation vector for experts using an attention mechanism. Secondly, a random masking strategy is applied to the user behavior sequence to simulate long-tail users, and an auxiliary contrastive loss is imposed to the output of the gate network to improve the model generalization for these users. This is validated by a higher performance gain on the long-tail user test set.Experiment results on a JD real production dataset and a public dataset demonstrate the effectiveness of AW-MoE, which significantly outperforms state-of-art methods. Notably, AW-MoE has been successfully deployed in the JD e-commerce search engine, serving the real traffic of hundreds of millions of active users.
Juan Gong, Zhenlin Chen, Chaoyi Ma, Zhuojian Xiao, Guoyu Tang, Sulong Xu, Bo Long, Yunjiang Jiang
ICDE3
2023 Policy enforcement in traditional non-SDN networks
Olufemi Odegbile, Chaoyi Ma, Shigang Chen, Yuanda Wang
J. Parallel Distributed Comput.2
2023 Single Update Sketch with Variable Counter Structure
abstract
Per-flow size measurement is key to many streaming applications and management systems, particularly in high-speed networks. Performing such measurement on the data plane of a network device at the line rate requires on-chip memory and computing resources that are shared by other key network functions. It leads to the need for very compact and fast data structures, called sketches, which trade off space for accuracy. Such a need also arises in other application context for extremely large data sets. The goal of sketch design is two-fold: to measure flow size as accurately as possible and to do so as efficiently as possible (for low overhead and thus high processing throughput). The existing sketches can be broadly categorized to multi-update sketches and single update sketches. The former are more accurate but carry larger overhead. The latter incur small overhead but their accuracy is poor. This paper proposes a Single update Sketch with a Variable counter Structure (SSVS), a new sketch design which is several times faster than the existing multi-update sketches with comparable accuracy, and is several times more accurate than the existing single update sketches with comparable overhead. The new sketch design embodies several technical contributions that integrate the enabling properties from both multi-update sketches and single update sketches in a novel structure that effectively controls the measurement error with minimum processing overhead.
Dimitrios Melissourgos, Haibo Wang 0004, Shigang Chen, Chaoyi Ma, Shiping Chen 0002
Proc. VLDB Endow.4
2023 Randomized Error Removal for Online Spread Estimation in High-Speed Networks
abstract
Flow spread measurement provides fundamental statistics that can help network operators better understand flow characteristics and traffic patterns with applications in traffic engineering, cybersecurity and quality of service. Past decades have witnessed tremendous performance improvement for single-flow spread estimation. However, when dealing with numerous flows in a packet stream, it remains a significant challenge to measure per-flow spread accurately while reducing memory footprint. The goal of this paper is to introduce new multi-flow spread estimation designs that incur much smaller processing overhead and query overhead than the state of the art, yet achieves significant accuracy improvement in spread estimation. We formally analyze the performance of these new designs. We implement them in both hardware and software, and use real-world data traces to evaluate their performance in comparison with the state of the art. The experimental results show that our best sketch significantly improves over the best existing work in terms of estimation accuracy, packet processing throughput, and online query throughput.
Haibo Wang 0004, Chaoyi Ma, Olufemi Odegbile, Shigang Chen, Jih-Kwon Peir
IEEE/ACM Trans. Netw.2
2022 Supporting Real-time Networkwide T-Queries in High-speed Networks
abstract
Traffic measurement is key to many important network functions. Supporting real-time queries at the individual flow level over networkwide traffic represents a major challenge that has not been successfully addressed yet. This paper provides the first solutions in supporting real-time networkwide queries and allowing a local network function (for performance, security or management purpose) to make queries at any measurement point at any time on any flow’s networkwide statistics, while the packets of the flow may traverse different paths in the network, some of which may not come across the point where the query is made. Our trace-based experiments demonstrate that the proposed solutions significantly outperform the baseline solutions derived from the existing techniques.
Yuanda Wang, Haibo Wang 0004, Chaoyi Ma, Shigang Chen
ICDCS3
2022 Online Cardinality Estimation by Self-morphing Bitmaps
abstract
Estimating the cardinality of a data stream is a fundamental problem underlying numerous applications such as traffic monitoring in a network or a datacenter, popularity tracking on social media, and cache optimization in proxy servers. Existing solutions suffer from high processing/query overhead or memory in-efficiency, which prevents them from operating online for data streams with very high arrival rates. This paper takes a new solution path different from the prior art and proposes a self-morphing bitmap, which combines operational simplicity with structural dynamics, allowing the bitmap to be morphed in a series of steps with an evolving sampling probability that automatically adapts to different stream sizes. We evaluate the self-morphing bitmap theoretically and experimentally. The results demonstrate that it significantly outperforms the prior art.
Haibo Wang 0004, Chaoyi Ma, Shigang Chen, Yuanda Wang
ICDE2
2022 Super Spreader Identification Using Geometric-Min Filter
abstract
Super spreader identification has a lot of applications in network management and security monitoring. It is a more difficult problem than heavy hitter identification because flow spread is harder to measure than flow size due to the requirement of duplicate removal. The prior work either incurs heavy memory overhead or requires heavy computations. This paper designs a new super-spreader monitor capable of identifying all flows whose spreads are greater than a user-specified threshold with a probability that can be arbitrarily set. It introduces a generalized geometric hash function, a generalized geometric counter, and a novel geometric-min filter that blocks out the vast majority of small/medium flows from being tracked, allowing us to focus on a small number of flows in which super spreaders are identified. We provide an analytical way of properly setting the system threshold to meet probabilistically guaranteed identification of super spreaders, and implement it on both hardware (FPGA) and software platforms. We perform extensive experiments based on real Internet traffic traces from CAIDA. The results show that with proper parameter settings, the new monitor can identify more than 99% super spreaders with a low memory requirement, better than the prior art.
Chaoyi Ma, Shigang Chen, Youlin Zhang, Qingjun Xiao, Olufemi Odegbile
IEEE/ACM Trans. Netw.1
2022 Virtual Filter for Non-Duplicate Sampling With Network Applications
abstract
Sampling is key to handling mismatch between the line rate and the throughput of a network traffic measurement module. Flow-spread measurement requires non-duplicate sampling, which only samples the elements (carried in packet header or payload) in each flow when they appear for the first time and blocks them for subsequent appearances. The only prior work for non-duplicate sampling incurs considerable overhead, and has two practical limitations: It lacks a mechanism to set an appropriate sampling probability under dynamic traffic conditions, and it cannot efficiently handle multiple concurrent sampling tasks. This paper proposes a virtual filter design for non-duplicate sampling, which reduces the processing overhead by about half and reduces the memory overhead by an order of magnitude or more under some practical settings. It has a mechanism to automatically adapt its sampling probability to the traffic dynamics. It can be modified to handle sampling for multiple independent tasks with different probabilities. We also enhance the virtual filter for flow spread measurement and super spreader detection with a large measurement period.
Chaoyi Ma, Haibo Wang 0004, Olufemi Odegbile, Shigang Chen, Dimitrios Melissourgos
IEEE/ACM Trans. Netw.1
2022 Fast and Accurate Cardinality Estimation by Self-Morphing Bitmaps
abstract
Estimating the cardinality of a data stream is a fundamental problem underlying numerous applications such as traffic monitoring in a network or a datacenter and query optimization of Internet-scale P2P data networks. Existing solutions suffer from high processing/query overhead or memory in-efficiency, which prevents them from operating online for data streams with very high arrival rates. This paper takes a new solution path different from the prior art and proposes a self-morphing bitmap, which combines operational simplicity with structural dynamics, allowing the bitmap to be morphed in a series of steps with an evolving sampling probability that automatically adapts to different stream sizes. We further generalize the design of self-morphing bitmap. We evaluate the self-morphing bitmap theoretically and experimentally. The results demonstrate that it significantly outperforms the prior art.
Haibo Wang 0004, Chaoyi Ma, Shigang Chen, Yuanda Wang
IEEE/ACM Trans. Netw.2
2021 Supporting Real-Time ${T}$-Queries on Network Traffic with a Cloud-Based Offloading Model
abstract
Traffic measurement provides fundamental statistics for network management functions. To implement the measurement modules on the data plane for real-time query response, modern sketches are designed to work with limited on-die memory allocation from network processors and collect traffic statistics in epochs of a preset length. To handle real-time queries at arbitrary times over traffic in a preceding period${T}$(called${T}$-queries), the prior art sets the epoch length to${T \over n}$and keeps the measurement results in a window of$n - 1$past epochs to support approximate${T}$-queries. Such an approach however drastically increases the memory cost or decreases the accuracy in the query results if the memory allocation is fixed. In this paper, motivated by the concept of offloading in today's edge-cloud computing, we propose a collaborative edge-center traffic measurement model, where the traffic measurement modules at all network devices form the edge, which offloads the traffic measurement results to a measurement center possibly hosted in a datacenter. The center synthesizes the measurements from the past epochs and sends the aggregate results back to the measurement modules to support T-queries. We conduct experiments using real traffic traces to evaluate the performance of the proposed edge-center measurement model. The experimental results demonstrate that the proposed designs significantly outperform the prior art.
Yuanda Wang, Haibo Wang 0004, Chaoyi Ma, Shigang Chen, Ye Xia 0001
CLOUD3
2021 Virtual Filter for Non-duplicate Sampling
abstract
Sampling is key to handling mismatch between the line rate and the throughput of a network traffic measurement module. Flow-spread measurement requires non-duplicate sampling, which only samples the elements (carried in packet header or payload) in each flow when they appear for the first time and blocks them for subsequent appearances. The only prior work for non-duplicate sampling incurs considerable overhead, and has two practical limitations: It lacks a mechanism to set an appropriate sampling probability under dynamic traffic conditions, and it cannot efficiently handle multiple concurrent sampling tasks. This paper proposes a virtual filter design for non-duplicate sampling, which reduces the processing overhead by about half and reduces the memory overhead by an order of magnitude or more under some practical settings. It has a mechanism to automatically adapt its sampling probability to the traffic dynamics. It can be extended to solve a new problem called non-duplicate distribution sampling, which samples packets based on a probability distribution to support multiple concurrent measurement tasks.
Chaoyi Ma, Haibo Wang 0004, Olufemi Odegbile, Shigang Chen
ICNP1
2021 On Outsourcing Artificial Neural Network Learning of Privacy-Sensitive Medical Data to the Cloud
abstract
Machine learning and artificial neural networks (ANNs) have been at the forefront of medical research in the last few years. It is well known that ANNs benefit from big data and the collection of the data is often decentralized, meaning that it is stored in different computer systems. There is a practical need to bring the distributed data together with the purpose of training a more accurate ANN. However, the privacy concern prevents medical institutes from sharing patient data freely. Federated learning and multi-party computation have been proposed to address this concern. However, they require the medical data collectors to participate in the deep-learning computations of the data users, which is inconvenient or even infeasible in practice. In this paper, we propose to use matrix masking for privacy protection of patient data. It allows the data collectors to outsource privacy-sensitive medical data to the cloud in a masked form, and allows the data users to outsource deep learning to the cloud as well, where the ANN models can be trained directly from the masked data. Our experimental results on deep-learning models for diagnosis of Alzheimer's disease and Parkinson's disease show that the diagnosis accuracy of the models trained from the masked data is similar to that of the models from the original patient data.
Dimitrios Melissourgos, Hanzhi Gao, Chaoyi Ma, Shigang Chen, Samuel S. Wu
ICTAI3
2021 Noise Measurement and Removal for Data Streaming Algorithms with Network Applications
abstract
Data streaming has multiple applications on the Internet including traffic measurement and intrusion detection. The bedrock underlying these applications is a set of data streaming algorithms that extract useful information from network packet stream, estimate the needed statistics such as the frequencies of TCP flows, and feed them to application software. Among such algorithms, counting sketches are most prevalent, which are very compact but do so at the cost of errors in their estimations. The dominant error-control method that has been widely accepted for more than a decade is to take the min error from multiple independent estimations. This method produces a positively-biased error and the error can grow large under stringent performance and resource conditions, but no existing work makes an intensive study of this error. This paper investigates the property of the error, which is also known as noise, and claims that it can be measured and removed so as to make the estimations unbiased. We introduce two new ideas, d-smallest noise and artificial data items for measuring the noise. Based on these two ideas, we propose four noise measurement methods. The mathematical analysis and experimental results based on real network traces show that by removing the measured noise, the error of estimations will be reduced to a much lower level than what the state of the art can do.
Chaoyi Ma, Haibo Wang 0004, Olufemi Odegbile, Shigang Chen
Networking1
2021 Randomized Error Removal for Online Spread Estimation in Data Streaming
abstract
Measuring flow spread in real time from large, high-rate data streams has numerous practical applications, where a data stream is modeled as a sequence of data items from different flows and the spread of a flow is the number of distinct items in the flow. Past decades have witnessed tremendous performance improvement for single-flow spread estimation. However, when dealing with numerous flows in a data stream, it remains a significant challenge to measure per-flow spread accurately while reducing memory footprint. The goal of this paper is to introduce new multi-flow spread estimation designs that incur much smaller processing overhead and query overhead than the state of the art, yet achieves significant accuracy improvement in spread estimation. We formally analyze the performance of these new designs. We implement them in both hardware and software, and use real-world data traces to evaluate their performance in comparison with the state of the art. The experimental results show that our best sketch significantly improves over the best existing work in terms of estimation accuracy, data item processing throughput, and online query throughput.
Haibo Wang 0004, Chaoyi Ma, Olufemi Odegbile, Shigang Chen, Jih-Kwon Peir
Proc. VLDB Endow.2
2021 Spread Estimation With Non-Duplicate Sampling in High-Speed Networks
abstract
Per-flow spread measurement in high-speed networks has many practical applications. It is a more difficult problem than the traditional per-flow size measurement. Most prior work is based on sketches, focusing on reducing their space requirements in order to fit in on-chip cache memory. This design allows the measurement to be performed at the line rate, but it suffers from expensive computation for spread queries (unsuitable for online operations) and large errors in spread estimation for small flows. This paper complements the prior art with a new spread estimator design based on an on-chip/off-chip model. By storing traffic statistics in off-chip memory, our new design faces a key technical challenge to design an efficient online module of non-duplicate sampling that cuts down the off-chip memory access. We first propose a two-stage solution for non-duplicate sampling, which is efficient but cannot handle well a sampling probability that is either too small or too big. We then address this limitation through a three-stage solution that is more space-efficient. Our analysis shows that the proposed spread estimator is highly configurable for a variety of probabilistic performance guarantees. We implement our spread estimator in hardware using FPGA. The experiment results based on real Internet traffic traces show that our estimator produces spread estimation with much better accuracy than the prior art, reducing the mean relative (absolute) error by about one order of magnitude. Moreover, it increases the query throughput by around three orders of magnitude, making it suitable for supporting online queries in real time.
He Huang 0001, Yu-e Sun, Chaoyi Ma, Shigang Chen, Yang Du 0006, Haibo Wang 0004, Qingjun Xiao
IEEE/ACM Trans. Netw.3
2021 A Novel Modified Sparrow Search Algorithm with Application in Side Lobe Level Reduction of Linear Antenna Array
abstract
Antenna arrays play an increasingly important role in modern wireless communication systems. However, how to effectively suppress and optimize the side lobe level (SLL) of antenna arrays is critical for communication performance and communication capabilities. To solve the antenna array optimization problem, a new intelligent optimization algorithm called sparrow search algorithm (SSA) and its modification are applied to the electromagnetics and antenna community for the first time in this paper. Firstly, aimed at the shortcomings of SSA, such as being easy to fall into local optimum and limited convergence speed, a novel modified algorithm combining a homogeneous chaotic system, adaptive inertia weight, and improved boundary constraint is proposed. Secondly, three types of benchmark test functions are calculated to verify the effectiveness of the modified algorithm. Then, the element positions and excitation amplitudes of three different design examples of the linear antenna array (LAA) are optimized. The numerical results indicate that, compared with the other six algorithms, the modified algorithm has more advantages in terms of convergence accuracy, convergence speed, and stability, whether it is calculating the benchmark test functions or reducing the maximum SLL of the LAA. Finally, the electromagnetic (EM) simulation results obtained by FEKO also show that it can achieve a satisfactory beam pattern performance in practical arrays.
Qiankun Liang, Huaning Wu, Chaoyi Ma, Senyou Li
Wirel. Commun. Mob. Comput.4
2020 Online Spread Estimation with Non-duplicate Sampling
abstract
Per-flow spread measurement in high-speed networks has many practical applications. It is a more difficult problem than the traditional per-flow size measurement. Most prior work is based on sketches, focusing on reducing their space requirements in order to fit in on-chip cache memory. This design allows measurement to be performed at the line rate, but it has to accept tradeoff with expensive computation for spread queries (unsuitable for online operations) and large errors in spread estimation for small flows. This paper complements the prior art with a new spread estimator design based on an on-chip/off-chip model which is common in practice. The new estimator supports online queries in real time and produces spread estimation with much better accuracy. By storing traffic data in off-chip memory, our new design faces a key technical challenge of efficient non-duplicate sampling. We propose a two-stage solution with on-chip/off-chip data structures and algorithms, which are not only efficient but also highly configurable for a variety of probabilistic performance guarantees. The experiment results based on real Internet traffic traces show that our estimator reduces the mean relative and absolute error by around one order of magnitude, and achieves both space-efficiency and accuracy-efficiency in flow classification for small flows compared to the prior art.
Yu-e Sun, He Huang 0001, Chaoyi Ma, Shigang Chen, Yang Du 0006, Qingjun Xiao
INFOCOM3
2020 An Efficient K-Persistent Spread Estimator for Traffic Measurement in High-Speed Networks
abstract
Traffic measurement in high-speed networks has many important functions in improving network performance, assisting resource allocation, and detecting anomalies. In this paper, we study a generalized problem called k-persistent spread estimation, which measures the volume of persist traffic elements in each flow that appear during at least k out of t measurement periods, where k and t are two positive integers that can be arbitrarily set in user queries, with k ≤ t. Solutions to this problem have interesting applications in network attack detection, popular content identification, user access profiling, etc. There is very limited prior art for this problem, only addressing the special case of k = t under a flawed assumption. Removing this assumption, we propose an efficient and accurate estimator for generalized k-persistent traffic measurement, with k ≤ t. Our method relies on bitwise SUM, instead of bitwise AND in the prior art, to combine the information collected from different periods. This change has fundamental impact on the probabilistic analysis that derives the estimator, particular over space-saving virtual bitmaps. Based on real network traces, we demonstrate experimentally the effectiveness of our new method in estimating the k-persistent spreads of all network flows. Our estimator performs much better than the prior art on its case of k = t. We also incorporate a sampling module to the estimator for improved flexibility, and give a use study on how to detect and find DDoS attackers using the proposed estimator.
He Huang 0001, Yu-e Sun, Chaoyi Ma, Shigang Chen, You Zhou 0003, Wenjian Yang, Shaojie Tang 0001, Hongli Xu 0001
IEEE/ACM Trans. Netw.3
2019 From Semantic Retrieval to Pairwise Ranking: Applying Deep Learning in E-commerce Search
abstract
We introduce deep learning models to the two most important stages in product search at JD.com, one of the largest e-commerce platforms in the world. Specifically, we outline the design of a deep learning system that retrieves semantically relevant items to a query within milliseconds, and a pairwise deep re-ranking system, which learns subtle user preferences. Compared to traditional search systems, the proposed approaches are better at semantic retrieval and personalized ranking, achieving significant improvements.
Yunjiang Jiang, Wenyun Yang, Guoyu Tang, Songlin Wang, Chaoyi Ma, Yihong Eric Zhao
SIGIR6