Yunxiao Ma

dblp:46/1709 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Blockchain-enabled dispersed computing paradigm in Web 3.0 metaverse
Zhonghui Wu, Changqiao Xu, Yunxiao Ma, Zicong Huang, Jingtian Liu, Lujie Zhong, Luigi Alfredo Grieco
Comput. Networks4
2024 Standard-Driven Software Component Reuse and Agile Testing
abstract
When developing software for the government, there is often a need to develop standard-driven software, which is convenient for management, beneficial for data unification and maintenance. However, during the development of standard-driven software, there may be many repetitive tasks, such as a piece of code that handles a specific logic independently or a form used to validate a particular standard. The reusability of software components can play a significant role in improving the efficiency of these repetitive tasks. Therefore, we take a standard-driven program developed by a team as an example, and propose a method for reusing software components based on this program, combined with agile testing techniques, to constantly verify whether it meets the requirements of the standard. This paper aims to provide a reference for the development of future standard-driven software based on the experience gained from this development.
Yunxiao Ma, Aliya Bao, Chengyuan Tian, Li Hua
COMPSAC1
2024 A Method of Network Attack Named Entity Recognition based on Deep Active Learning
abstract
In the face of data scarcity for network attack annotation and the possibility that static datasets may not anticipate future security threats, in this paper, we integrate active learning and deep learning techniques, and propose a method of network attack named entity recognition based on deep active learning. Considering that traditional active learning sampling strategies may ignore the inherent complexity of data, we propose a dual-dimension diversity sampling method that pays attention to both the internal and external diversity of unlabeled samples to promote the improvement of model generalization ability. Further, in order to fully mine valuable samples and achieve balanced sample selection in model training, we design a strategy that alternately applies experience-driven uncertainty sampling and dual-dimension diversity sampling. The advantages of the proposed method in improving the precision, recall and F1 value of named entity recognition tasks on self-built and public datasets are verified through ablation and comparison experiments.
Yunxiao Ma, Peilong Zhang
QRS2
2024 Patronus: Countering Model Poisoning Attacks in Edge Distributed DNN Training
abstract
As Deep Neural Networks (DNNs) are evolving in complexity to meet the demands of novel applications, a single device becomes insufficient for training, leading to the emergence of distributed DNN training. However, this evolution exposes a gap in research surrounding security vulnerabilities on model poisoning attacks, especially in model parallel setups, an area that has been scarcely studied. To bridge this gap, we introduce Patronus, an approach that counters model poisoning attacks in distributed DNN training, accommodating both data and model parallelism. With the employment of Loss-aware Credit Evaluation, Patronus scores each participating client. Based on the continuously updated credit, malicious clients are isolated and detected after multiple epochs by Shuffling-based Isolation Mechanism. Additionally, the training system is reinforced by Byzantine Fault-tolerant Aggregation to minimize malicious client impacts. Comprehensive experiments confirm Patronus's superior reliable and efficient performance over the existing methods under attack scenarios.
Zhonghui Wu, Changqiao Xu, Yunxiao Ma, Zhongrui Wu, Zhenyu Xiahou, Luigi Alfredo Grieco
WCNC4
2024 CA-Live360: Crowd-assisted transcoding and delivery for live 360-degree video streaming
Yunxiao Ma, Changqiao Xu, Zhonghui Wu, Renjie Ding, Lujie Zhong, Yirong Zhuang, Gabriel-Miro Muntean
Comput. Networks1
2024 Connectional-style-guided contextual representation learning for brain disease diagnosis
Gongshu Wang, Yunxiao Ma, Duanduan Chen, Tianyi Yan
Neural Networks3
2024 Adaptive and Robust Query Execution for Lakehouses At Scale
abstract
Many organizations have embraced the "Lakehouse" data management paradigm, which involves constructing structured data warehouses on top of open, unstructured data lakes. This approach stands in stark contrast to traditional, closed, relational databases and introduces challenges for performance and stability of distributed query processors. Firstly, in large-scale, open Lakehouses with uncurated data, high ingestion rates, external tables, or deeply nested schemas, it is often costly or wasteful to maintain perfect and up-to-date table and column statistics. Secondly, inherently imperfect cardinality estimates with conjunctive predicates, joins and user-defined functions can lead to bad query plans. Thirdly, for the sheer magnitude of data involved, strictly relying on static query plan decisions can result in performance and stability issues such as excessive data movement, substantial disk spillage, or high memory pressure. To address these challenges, this paper presents our design, implementation, evaluation and practice of the Adaptive Query Execution (AQE) framework, which exploits natural execution pipeline breakers in query plans to collect accurate statistics and re-optimize them at runtime for both performance and robustness. In the TPC-DS benchmark, the technique demonstrates up to 25× per query speedup. At Databricks, AQE has been successfully deployed in production for multiple years. It powers billions of queries and ETL jobs to process exabytes of data per day, through key enterprise products such as Databricks Runtime, Databricks SQL, and Delta Live Tables.
Maryann Xue, Yingyi Bu, Abhishek Somani, Wenchen Fan, Steven Chen, Herman Van Hövell, Bart Samwel, Mostafa Mokhtar, Rk Korlapati, Andy Lam, Yunxiao Ma, Vuk Ercegovac, Jiexing Li, Alexander Behm, Yuanjian Li, Xiao Li 0087, Sriram Krishnamurthy, Amit Shukla 0001, Michalis Petropoulos, Sameer Paranjpye, Reynold Xin, Matei Zaharia
Proc. VLDB Endow.12
2023 A Multi-User Cost-Efficient Crowd-Assisted VR Content Delivery Solution in 5G-and-Beyond Heterogeneous Networks
abstract
The latest evolution of wireless communications enables user access rich Virtual Reality (VR) services via the Internet, including while on the move. However, providing a premium immersive experience for massive number of concurrent users with various device configurations is a significant challenge due to the ultra-high data rate and ultra-low delay requirements of live VR services. This paper introduces an innovative multi-user cost-efficient crowd-assisted delivery and computing (MEC-DC) framework, which leverages mobile edge computing and end-user resources to support high performance VR content delivery over 5G-and-beyond heterogeneous networks (5G-HetNets). The proposed MEC-DC framework is based on three main solutions. First is a novel buffer-nadir-based multicast (BNM) mechanism for VR transmissions over 5G-HetNets. BNM ensures smooth and synchronized user viewing experience by maximizing the average playback buffer-nadir of all participants with stochastic optimization. Second and third are practical distributed algorithms: the cost-efficient multicast-aware transcoding offloading (MATO) and crowd-assisted delivery algorithm (CAD) which optimize jointly multicast delivery and video transcoding. The algorithms optimality and complexity were investigated. The proposed MATO-CAD solution was evaluated with real datasets, trace-driven numerical simulations, and prototype-based experiments. The trace-driven experimental results showed how the proposed solution provides 18% throughput improvement, lowest delay and best playback freeze ratio in comparison with three other state-of-the-art solutions.
Lujie Zhong, Xingyan Chen, Changqiao Xu, Yunxiao Ma, Yu Zhao 0019, Gabriel-Miro Muntean
IEEE Trans. Mob. Comput.4
2022 Edge Intelligence: A Computational Task Offloading Scheme for Dependent IoT Application
abstract
Computational offloading, as an effective way to extend the capability of resource-limited edge devices in Internet of Things (IoT), is considered as a promising emerging paradigm for coping with delay-sensitive services. However, on one hand, applications commonly include several subtasks with dependent relations and on the other hand, the dynamic changes in network environments make offloading decision-making become a coupling and complex NP-hard problem, difficult to address. This paper proposes an intelligent Computational Offloading scheme for Dependent IoT Application (CODIA), which decouples the performance enhancement problem into two processes: scheduling and offloading. First, a prioritized scheduling strategy is designed and its complexity is analyzed. Then, an offloading algorithm with offline training and online deployment is introduced. Due to the temporal continuity between subtasks, the dependency relation is transformed into a transition of device state, and the overhead for the whole application is considered to be the long-term benefit.CODIAleverages an Actor-Critic-based solution, where the IoT devices are able to deploy intelligent models and dynamically adjust the offloading strategy to achieve low latency, while controlling energy consumption. Finally, a series of experiments are conducted to verify the robustness and efficiency of the proposed solution in terms of convergence, latency, and energy consumption.
Changqiao Xu, Yunxiao Ma, Lujie Zhong, Gabriel-Miro Muntean
IEEE Trans. Wirel. Commun.3
2021 Edge Computing-Assisted Multimedia Service Energy Optimization based on Deep Reinforcement Learning
abstract
With the development of communication technology, emerging multimedia (e.g. virtual reality) can provide users with more immersive service experience. However, due to the ultra-high rendering and splicing requirements of multimedia content, the higher demand for computing resources is put forward for the playback device. The anomalies of energy consumption and latency caused by such computationally intensive tasks hinder the practical application of emerging multimedia technology in mobile networks. In this regard, this paper proposes an edge computing assisted multimedia service optimization scheme (ECMSO) to broaden the computing capacity of the viewer(i.e. requester), so as to ensure that content can be served in time and reduce the energy cost of computation from the perspective of executor and requester, respectively. First, a computational offloading scheme based on deep reinforcement learning is designed. It optimizes intelligently the energy consumption while meeting the latency requirements of the requester. Secondly, a heuristic algorithm to allocate power, bandwidth, and computing resources for candidate executors is proposed. Finally, a series of simulation experiments are conducted to demonstrate the effectiveness of our proposed scheme.
Changqiao Xu, Yunxiao Ma, Lujie Zhong, Gabriel-Miro Muntean
GLOBECOM3
2021 Fairness-Guaranteed Transcoding Task Assignment for Viewer-Assisted Crowdsourced Livecast Services
abstract
Recent years have witnessed an outstanding increase in popularity of Crowdsourced Livecast Services (CLS), which is the latest trend in social media. In CLS, transcoding enormous video contents from massive broadcasters and providing high-quality CLS for global viewers with heterogeneous devices are computation-intensive as well as time-consuming. There are some schemes that design viewer-assisted transcoding scheme, but it is challenging to achieve an efficient and fair task assignment due to the dynamic of computing and communication resources. This paper introduces a viewer-assisted CLS framework and focuses on proposing an innovative fairness-guaranteed task assignment scheme, which is a key challenge in this context. Considering the dynamic nature of viewers’ computing and communication resources and stability, a dynamic programming problem with fairness and QoS constraints is formulated. To solve the problem, we devise a Fair Bandit (FB) algorithm based on the Combinatorial Multi-Armed Bandit (CMAB). Finally, the effectiveness of proposed scheme is demonstrated by trace-driven simulations.
Yunxiao Ma, Changqiao Xu, Xingyan Chen, Lujie Zhong, Gabriel-Miro Muntean
ICC1
2009 Nonlinear static-rank computation
abstract
Mainstream link-based static-rank algorithms (e.g. PageRank and its variants) express the importance of a page as the linear combination of its in-links and compute page importance scores by solving a linear system in an iterative way. Such linear algorithms, however, may give apparently unreasonable static-rank results for some link structures. In this paper, we examine the static-rank computation problem from the viewpoint of evidence combination and build a probabilistic model for it. Based on the model, we argue that a nonlinear formula should be adopted, due to the correlation or dependence between links. We focus on examining some simple formulas which only consider the correlation between links in the same domain. Experiments conducted on 100 million web pages (with multiple static-rank quality evaluation metrics) show that higher quality static-rank could be yielded by the new nonlinear algorithms. The convergence of the new algorithms is also proved in this paper by nonlinear functional analysis.
Shuming Shi 0001, Yunxiao Ma, Ji-Rong Wen
CIKM3
2007 Web object retrieval
abstract
The primary function of current Web search engines is essentially relevance ranking at the document level. However, myriad structured information about real-world objects embedded in static Web pages and online Web databases. In this paper, we propose a paradigm shift to enable searching at the object level. In traditional information retrieval models, documents are taken as the retrieval units and the content of a document is considered reliable. However, this reliability assumption is no longer valid in the object retrieval context when multiple copies of information about the same object typically exist. These copies may be inconsistent because of diversity of Web site qualities and the limited performance of current information extraction techniques. In this paper, we propose several language models for Web object retrieval. We test these models on our academic search engine called Libra and compare their performances. 1.
Zaiqing Nie, Yunxiao Ma, Shuming Shi 0001, Ji-Rong Wen, Wei-Ying Ma
WWW2