Chao Huang 0028

dblp:18/4087-28 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 12 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Decoupled Split Learning via Auxiliary Loss
Anower Zihad, Felix Owino, Ming Tang 0006, Chao Huang 0028
INFOCOM4
2026 HOSL: Hybrid-Order Split Learning for Memory-Constrained Edge Training
Aakriti Lnu, Zhe Li 0083, Dandan Liang, Chao Huang 0028, Rui Li 0002, Haibo Yang 0001
WiOpt4
2026 Decentralized Information Elicitation Without Verification
abstract
Information Elicitation Without Verification (IEWV) refers to eliciting high-accuracy solutions from crowd members when the ground truth is unverifiable. While prior research on IEWV has focused on central entities providing incentives to motivate effort exertion, this work explores the less-studied decentralized setting, which is increasingly relevant in machine learning, crowd decision-making, and autonomous organization applications. We model members’ strategic interactions as a two-stage game, where each member decides her incentive contribution strategy in Stage I and her effort exertion strategy in Stage II. We examine two types of incentive allocation mechanisms: Equal Allocation (EA), where each member receives an equal proportion of the total incentives, and Output Agreement (OA), where a member receives incentives if her solution matches a reference solution generated by other members. This paper first analyzes the two-member case and provides closed-form equilibrium results. For more than two members, we use a binomial approximation to simplify the combinatorial computation of the majority voting problem and characterize the symmetric Nash equilibrium under EA. For OA, we derive equilibrium results for effort exertion and propose an algorithm for the incentive contribution game due to discontinuous payoffs. Our results show that OA outperforms EA in the aggregated team solution accuracy at equilibrium. Furthermore, we reveal that higher member ability beyond a certain threshold may lead to reduced effort exertion under EA, and that smaller teams achieve better accuracy when the effort cost is high due to less free-riding behavior. Numerical and empirical simulations validate our theory.
Chao Huang 0028, Jianwei Huang 0001
IEEE Trans. Netw.2
2025 An empirical study on impact of label noise on synthetic tabular data generation
abstract
Abstract Synthetic data has been actively used for various machine learning-based tasks due to its benefits such as massive reproducibility and privacy enhancement compared to using the original data. The quality of the generated synthetic dataset crucially depends on the quality of the original data, and the latter is often corrupted by label noise. While there have been studies on feature noise, how label noise affects synthetic data generation is under-explored. In this paper, we evaluate the impact of the noisy label on synthetic data generation with a focus on tabular data. One challenge is how to evaluate the quality of synthetic data under label noise. To this end, we design comprehensive experiments to measure the impact of label noise on synthetic data generation in different aspects: synthetic data quality, data utility, and convergence for training synthesizers and machine learning models for downstream tasks. The empirical results cover wide aspects of synthetic data generation under label noise and they show quality and utility degrades with higher noise levels while there is no significant effect on the synthesizer convergence observed.
Chao Huang 0028, Xin Liu 0002
Mach. Learn.2
2024 An Accuracy-Shaping Mechanism for Competitive Distributed Learning
Chao Huang 0028, Justin Dachille, Xin Liu 0002
ICANN (6)1
2024 Incentivizing Participation in SplitFed Learning: Convergence Analysis and Model Versioning
abstract
In SplitFed learning (SFL), a global model is split into two segments, where distributed clients train the first segment in a federated manner and a main server trains the other. Existing studies focus on algorithm development but ignore the important issue of incentives, without which self-interested clients may be unwilling to participate. We fill this gap by presenting a first incentive study in SFL. One challenge is that the design requires an understanding of how clients' participation affects the model performance. To this end, we provide a first convergence analysis for SFL considering partial client participation to guide the mechanism design. Another challenge is that monetary payment may not be viable for large distributed systems. To this end, we propose a model-versioning mechanism where the main server assigns different versions of models (of different qualities) to clients as incentives. The design is further complicated by clients' multi-dimensional private information. To this end, we design the model-versioning mechanism so that it decouples clients' decisions and admits a weakly dominant strategy at equilibrium. We prove that our mechanism is feasible, effective, and incentive compatible. Experimental results show that our mechanism greatly improves client participation and model accuracy compared to a benchmark.
Pengchao Han, Chao Huang 0028, Xingyan Shi, Jianwei Huang 0001, Xin Liu 0002
ICDCS2
2024 When Federated Learning Meets Oligopoly Competition: Stability and Model Differentiation
abstract
Federated learning (FL) is decentralized machine learning framework that finds various applications in health, finance, and the internet of things. This paper studies the under-explored business competition in FL, where organizations are both collaborators in training a shared model and competitors in providing model-based services to a continuum of customers. We focus on an oligopoly case with three organizations. To understand how competition affects FL collaboration, we start with a benchmark case where organizations are not competitors, and show that they have an incentive to collaborate. However, in the presence of competition, organizations may prefer to train local models instead of collaborating via FL (even if FL incurs zero training costs). The reason is that FL intensifies price competition by improving organizations’ model performance to a similar level. To address this issue, we devise a model differentiation mechanism in which organizations adaptively adjust their model performance, enabling differentiated model-based services to customers. We prove that the adaptive mechanism converges in polynomial time and is incentive compatible. Perhaps surprisingly, numerical experiments on CIFAR-10 show that the mechanism can simultaneously improve the model performance, organizations’ revenues, and social welfare. The improvement is up to 22.31%, 14.42%, and 19.50%, respectively.
Chao Huang 0028, Justin Dachille, Xin Liu 0002
IEEE Internet Things J.1
2024 Incentivizing Efficient Label Denoising in Federated Learning
abstract
Federated learning (FL) is a distributed machine learning scheme that enables clients to train a shared global model without exchanging local data. In FL, the presence of label noise can severely reduce the accuracy of the global model. Although some recent works have focused on designing algorithms for label denoising, they ignored the important issue that clients may not apply costly label denoising strategies due to them being self-interested and having heterogeneous valuations on the model accuracy. To fill this gap, we model the clients’ strategic interactions as a novel label denoising game and determine the clients’ equilibrium strategies. We prove that the equilibrium outcome always leads to a lower global model accuracy than the socially optimal solution does. To motivate the clients’ efficient label denoising behaviors, we propose a penalty-based incentive mechanism and design the degree of penalty for punishing the clients’ undesired denoising behaviors, addressing the inaccurate noise rate detection in FL. We prove that our mechanism can achieve social efficiency, individual rationality, and weak budget balance. Numerical experiments on MNIST and CIFAR-10 show that as clients’ data become noisier, the gap between the equilibrium outcome and the socially optimal solution increases, verifying the necessity of an incentive mechanism. We empirically show that our proposed mechanism improves the model accuracy by up to 4.4% and incentivizes clients to achieve equilibrium strategies that are close to the socially optimal solution.
Yizhou Yan, Chao Huang 0028, Ming Tang 0006
IEEE Internet Things J.3
2023 Incentive Mechanism Design for Distributed Ensemble Learning
abstract
Distributed ensemble learning (DEL) involves training multiple models at distributed learners, and then combining their predictions to improve performance. Existing related studies focus on algorithm development but ignore the important issue of incentives, without which self-interested learners may be unwilling to participate. We aim to fill this gap by presenting a first study on the incentive mechanism design in DEL. The mechanism specifies both the training data and the reward for learners with heterogeneous computation and communication costs. One challenge is that it is unclear how learners' diversity (in terms of training data) contributes to the ensemble accuracy. To this end, we decompose the ensemble accuracy into a diversity-precision tradeoff to guide the mechanism design. Another challenge is that the mechanism design is a mixed-integer program with a large search space. To this end, we propose an alternating algorithm that iteratively updates each learner's training data size and reward. We prove that the algorithm converges and is polynomial in the number of learners. Numerical results using MNIST dataset are consistent with our analysis. Interestingly, we show that the mechanism may prefer a lower level of learner diversity to achieve a higher ensemble accuracy. Our code is made publicly available.
Chao Huang 0028, Pengchao Han, Jianwei Huang 0001
GLOBECOM1
2023 Information Elicitation from Decentralized Crowd Without Verification
abstract
Information Elicitation Without Verification (IEWV) refers to the problem of eliciting high-accuracy solutions from crowd members when the ground truth is unverifiable. A high-accuracy team solution (aggregated from members' solutions) requires members‘ effort exertion, which should be incentivized properly. Previous research on IEWV mainly focused on scenarios where a central entity (e.g., the crowdsourcing platform) provides incentives to motivate crowd members. Still, the proposed designs do not apply to practical situations where no central entity exists. This paper studies the overlooked decentralized IEWV scenario, where crowd members act as both incentive contributors and task solvers. We model the interactions among members with heterogeneous team solution accuracy valuations as a two-stage game, where each member decides her incentive contribution strategy in Stage 1 and her effort exertion strategy in Stage 2. We analyze members‘ equilibrium behaviors under three incentive allocation mechanisms: Equal Allocation (EA), Output Agreement (OA), and Shapley Value (SV). We show that at an equilibrium under any allocation mechanism, a low-valuation member exerts no more effort than a high-valuation member. Counter-intuitively, a low-valuation member provides incentives to the collaboration while a high-valuation member does not at an equilibrium under SV. This is because a high-valuation member who values the aggregated team solution more needs fewer incentives to exert effort. In addition, when members‘ valuations are sufficiently heterogeneous, SV leads to team solution accuracy and social welfare no smaller than EA and OA.
Chao Huang 0028, Jianwei Huang 0001
WiOpt2
2023 On the Impact of Label Noise in Federated Learning
abstract
Federated Learning (FL) is a distributed machine learning paradigm where clients collaboratively train a model using their local datasets. While existing studies focus on FL algorithm development to tackle data heterogeneity across clients, the important issue of data quality (e.g., label noise) in FL is less explored. This paper aims to fill this gap by providing a quantitative study on the impact of label noise on FL. We derive an upper bound for the generalization error that is linear in the summation of clients' label noise levels. Then we conduct experiments on MNIST and CIFAR-10 datasets using various FL algorithms. Our empirical results show that the global model accuracy linearly decreases as the noise level increases, which is consistent with our theoretical analysis. We further find that label noise slows down the convergence of FL training, and the global model tends to overfit when the noise level is high.
Shuqi Ke, Chao Huang 0028, Xin Liu 0002
WiOpt2
2023 An Online Inference-Aided Incentive Framework for Information Elicitation Without Verification
abstract
We study the design of incentive mechanisms for the problem of information elicitation without verification (IEWV). In IEWV, a data requester seeks to design proper incentives to optimize the tradeoff between the quality of information (collected from distributed crowd workers) and the total cost of incentives (provided to crowd workers) without verifiable ground truth. While prior work often relies on sufficient knowledge of worker information, we study a scenario where the data requester cannot access workers’ heterogeneous information quality and costs ex-ante. We propose a continuum-armed bandit-based incentive mechanism that dynamically learns the optimal reward level from workers’ reported information. A key challenge is that the data requester cannot evaluate the workers’ information quality without verification, which motivates the design of an inference algorithm. The inference problem is non-convex, yet we reformulate it as a bi-convex problem and derive an approximate solution with a performance guarantee, which ensures the effectiveness of our online reward design. We further enhance the inference algorithm using part of the workers’ historical reports. We also propose a novel rule for the data requester to aggregate workers’ solutions more effectively. We show that our mechanism achieves a sub-linear regret$\tilde {O}(T^{1/2})$and outperforms several celebrated benchmarks.
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
IEEE J. Sel. Areas Commun.1
2023 Strategic Information Revelation Mechanism in Crowdsourcing Applications Without Verification
abstract
We study a crowdsourcing problem, where a platform aims to incentivize distributed workers to provide high-quality and truthful solutions that are not verifiable. We focus on a largely overlooked yet pratically important asymmetric information scenario, where the platform knows more information regarding workers’ average solution accuracy and can strategically reveal such information to workers. Workers will utilize the announced information to determine the likelihood of obtaining a reward. We first study the case where the platform and workers share the same prior regarding the average worker accuracy (but only the platform observes the realized value). We consider two types of workers: (1)naiveworkers who fully trust the platform's announcement, and (2)strategicworkers who update prior belief based on the announcement. For naive workers, we show that the platform should always announce a high average accuracy to maximize its payoff. However, this is not always optimal when facing strategic workers, and the platform may benefit from announcing an average accuracy lower than the actual value. We further study the more challenging non-common prior case, and show the counter-intuitive result that when the platform is uninformed of the workers’ prior, both the platform payoff and the social welfare may decrease as the high accuracy workers’ solutions become more accurate.
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
IEEE Trans. Mob. Comput.1
2023 Online Crowd Learning Through Strategic Worker Reports
abstract
When it is difficult to verify contributed solutions in mobile crowdsourcing, the majority voting mechanism is widely utilized to incentivize distributed workers to provide high-quality and truthful solutions. In the majority voting mechanism, a worker is rewarded based on whether his solution is consistent with the majority. However, most prior related work relies on a strong assumption that workers solution accuracy levels are public knowledge, which may not hold in many practical scenarios. We relax such an assumption and propose an online mechanism, which allows the platform to learn the distribution of the workers solution accuracy levels via asking workers to report their private accuracy levels (which do not need to be the true values), in addition to deciding their effort levels and solution reporting strategies. The mechanism design is challenging, as neither the workers task solutions nor their accuracy reports can be verified. We devise a randomized reward mechanism that computes the workers rewards based on their reported accuracy levels, under which the workers obtain rewards if their reported solutions match the majority. Our mechanism induces workers to truthfully report their solution accuracy levels in the long run, and the empirical accuracy distribution converges to the actual accuracy distribution.
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
IEEE Trans. Mob. Comput.1
2022 Using Truth Detection to Incentivize Workers in Mobile Crowdsourcing
abstract
Mobile crowdsourcing platforms often want to incentivize workers to finish tasks with high quality and truthfully report their solutions by providing proper rewards. Most existing incentive mechanisms reward workers based on the comparison among workers’ reported solutions. However, these mechanisms are vulnerable to worker collusion, i.e., workers coordinate to misreport their solutions. We address such an issue by proposing a novel rewarding mechanism based on a${truth detection}$technology, which relies on the independent verification of the correctness of each worker’s response to some question with animperfectaccuracy. We model the interactions between the platform and workers as a two-stage Stackelberg game. In Stage I, the platform optimizes the reward mechanism parameters associated withtruth detectionto maximize its payoff. In Stage II, the workers decide their effort levels and reporting strategies to maximize their payoffs (which depend on the output of the truth detector). We analyze the game’s equilibrium and show that our proposed mechanism can effectively mitigate worker collusion. We also propose a novel rule, namedfiltered majority, for the platform to more effectively aggregate the workers’ solutions. Our proposed aggregation rule utilizes truth detection and outperforms the conventional simple majority rule. We further characterize the impact of the truth detection accuracy on the platform’s decisions. Surprisingly, under the simple majority rule, we show that as the truth detection accuracy improves, the platform should always incentivize more workers to exert effort and truthfully report. However, under our proposed filtered majority rule, we show that as the truth detection accuracy improves, in some cases, the platform should incentivize fewer workers and save costs. We further examine the impact of the workers’ imperfect estimation of the truth detection accuracy on the platform’s decisions.
Chao Huang 0028, Haoran Yu 0001, Randall Berry, Jianwei Huang 0001
IEEE Trans. Mob. Comput.1
2022 Eliciting Information From Heterogeneous Mobile Crowdsourced Workers Without Verification
abstract
In mobile crowdsourcing, platforms seek to incentivize heterogeneous workers to complete tasks (e.g., road traffic sensing) and truthfully report their solutions. When platforms cannot verify the quality of the workers’ solutions, the crowdsourcing problem is known asinformation elicitation without verification(IEWV). In an IEWV problem, a platform needs to provide incentives to motivate high-quality solutions and truthful reporting of the solutions from the workers. A common approach to solve the IEWV problem is majority voting, where each worker is rewarded according to whether his solution matches the majority’s solution. However, previous work has not considered workers with heterogeneous solution accuracy. This is unrealistic in many domains, where one would expect workers to differ in judgment, expertise, and reliability. Moreover, prior work has not considered how this heterogeneity affects a platform’s tradeoff between the quality of the workers’ solutions and the platform’s cost of achieving this. We address these gaps by studying the interactions between the mobile crowdsourcing platform and workers as a two-stage Stackelberg game. In Stage I, the platform chooses the reward level for majority voting. In Stage II, the workers decide their effort levels and reporting strategies. We show that as a worker’s solution accuracy increases, he is more likely, in equilibrium, to exert effort and truthfully report his solution. However, given a fixed total worker population, surprisingly, the platform’s payoff may decrease in the number of high-accuracy workers. We further characterize the value of knowing the workers’ solution accuracy in terms of improving the platform’s optimal reward design and maximizing its payoff. Knowing such information enables a more effective aggregation of the workers’ solutions. We further design a discriminatory reward policy to incentivize heterogeneous workers. Surprisingly, such a discriminatory policy can improve both the platform’s and the workers’ payoffs, and hence improve the social welfare.
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
IEEE Trans. Mob. Comput.1
2021 Strategic Information Revelation in Crowdsourcing Systems Without Verification
abstract
We study a crowdsourcing problem where the platform aims to incentivize distributed workers to provide high-quality and truthful solutions without the ability to verify the solutions. While most prior work assumes that the platform and workers have symmetric information, we study an asymmetric information scenario where the platform has informational advantages. Specifically, the platform knows more information regarding workers' average solution accuracy, and can strategically reveal such information to workers. Workers will utilize the announced information to determine the likelihood that they obtain a reward if exerting effort on the task. We study two types of workers: (1) naive workers who fully trust the announcement, and (2) strategic workers who update prior belief based on the announcement. For naive workers, we show that the platform should always announce a high average accuracy to maximize its payoff. However, this is not always optimal for strategic workers, as it may reduce the credibility of the platform's announcement and hence reduce the platform's payoff. Interestingly, the platform may have an incentive to even announce an average accuracy lower than the actual value when facing strategic workers. Another counter-intuitive result is that the platform's payoff may decrease in the number of high-accuracy workers.
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
INFOCOM1
2020 Online Crowd Learning with Heterogeneous Workers via Majority Voting
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
WiOpt1
2019 Crowdsourcing with Heterogeneous Workers in Social Networks
abstract
Many online social networking platforms are leveraging crowdsourcing to enhance the user experience. These platforms seek to incentivize heterogeneous workers to exert efforts to complete tasks (e.g., moderation of posts and articles) and truthfully report their solutions. Output agreement mechanism (e.g., majority voting) is a common approach to this end. In an output agreement mechanism, a worker is rewarded according to whether his solution matches those of his peers. However, prior related work has not studied the workers' heterogeneous solution accuracy and how this heterogeneity affects the platform's payoff. We fill this void by modeling and analyzing the interactions between the platform and workers as a two-stage Stackelberg game. In Stage I, the platform chooses the reward level for the majority voting to maximize its payoff. In Stage II, the workers decide their effort levels and reporting strategies to maximize their payoffs. We show that as a worker's solution accuracy increases, he is more likely to exert effort and truthfully report his solution under the equilibrium reward mechanism. However, given a fixed total worker population, it is surprising that the platform's overall payoff does not monotonically increase in the number of high-accuracy workers. This is because a larger number of high-accuracy workers brings marginally decreasing benefit to the platform, but the rewards required to incentivize them may significantly grow. Moreover, we show that as the solutions of the high-accuracy workers become more accurate, the platform needs a smaller number of such workers to achieve the maximum payoff.
Chao Huang 0028, Haoran Yu 0001, Jianwei Huang 0001, Randall Berry
GLOBECOM1