Kaixuan Huang

dblp:244/2561 · DBLP profile ↗
← Back
29ranked-venue papers
6as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 14 since 2021Computer networks · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An Alternating Directional Dual-RBF Approach for Joint Multi-BSs and Multi-RISs Deployment
Tao Yu 0008, Shunqing Zhang, Jihong Li, Kaixuan Huang, Wen Chen 0001, Qingqing Wu 0001
ICC5
2026 Few-Shot Cross-Domain Indoor Localization via Multimodal Feature Refinement
abstract
Fingerprint-based indoor localization is a critical enabling technology for Internet of Things (IoT) applications, where the primary challenges stem from complex environmental variability and prohibitive costs of data collection and labeling. This paper introduces a cross-domain multi-modal indoor localization framework that effectively combines visual and WiFi signals using few-shot learning techniques, achieving improved localization performance with minimal training data. We derive an upper bound on the generalized transfer localization error. Based on this bound, our learning-based approach applies feature-level knowledge distillation from pre-trained localization models. This process systematically calibrates discrepancies in feature distributions between the source and target environments. As a result, the proposed method significantly reduces the dependence on large labeled datasets. Experimental results demonstrate that our proposed method achieves substantial improvements over state-of-the-art localization models, with a mean localization error of 0.247 meters across diverse indoor environments, while requiring substantially fewer labeled samples in the target domain.
Kaixuan Huang, Jian (Andrew) Zhang, Guangjin Pan, Shunqing Zhang
IEEE Internet Things J.1
2026 Embodied LLM Agents Learn to Cooperate in Organized Teams
abstract
Large language models (LLMs) have emerged as integral tools for reasoning, planning, and decision-making, drawing upon their extensive world knowledge and proficiency in language-related tasks. LLMs thus hold tremendous potential for natural language interaction within multiagent systems to foster cooperation. However, LLM agents tend to over-report and comply with any instruction, which may result in information redundancy and confusion in multiagent cooperation. Inspired by human organizations, this article introduces a framework that imposes prompt-based organization structures on LLM agents to mitigate these problems. Through a series of experiments with embodied LLM agents and human–agent collaboration, our results highlight the impact of designated leadership on team efficiency, shedding light on the leadership qualities displayed by LLM agents and their spontaneous cooperative behaviors. Further, we harness the potential of LLMs to propose enhanced organizational prompts, via acriticize-reflectprocess, resulting in novel organization structures that reduce communication costs and enhance team efficiency.
Kaixuan Huang, Natalia Vélez, Qingyun Wu, Huazheng Wang, Thomas L. Griffiths 0001, Mengdi Wang 0001
IEEE Trans. Comput. Soc. Syst.2
2026 TFSCL: A Novel Time-Frequency Similarity Contrastive Learning Method With Hybrid Augmentation for Robust and Accurate Specific Emitter Identification
abstract
As the Internet of Things (IoT) and Sixth Generation (6G) technologies advance rapidly, the cryptographic identification of electronic devices has become a critical issue in information security. Radio frequency fingerprint (RFF)-based specific emitter identification (SEI) has emerged as a prominent physical-layer authentication technique. To enhance the stability and accuracy of multi-target recognition in complex electromagnetic environments, a novel technique for individual specific emitter identification based on time-frequency similarity contrastive learning (TFSCL) is proposed. In this study, we present a novel pre-training method utilizing a deep complex-valued pyramid network (DCPN) to enhance the extraction and reconstruction of time series and frequency domain sequences. The DCPN enables contrastive learning of signal features in both the temporal and frequency domains, significantly reducing computational complexity and improving pretraining performance. Additionally, we first introduce the Time-Frequency Synchronization Data Added (TFS-DA), a Time-Frequency Hybrid Data Added Technique that employs Gray code to generate random sequences, effectively improving feature representation in both domains. Empirical results demonstrate that the proposed method achieves an accuracy rate of 97.12% on an automatic-dependent surveillance-broadcast (ADS-B) dataset that contains 10 categories with only data labeled 10%. On a LoRa dataset containing 30 categories with only data labeled 10%, the accuracy rate reaches 77.06%.
Kaixuan Huang, Yongtao Ma, Jialu Zhu, Yuxiang Han
IEEE Trans. Inf. Forensics Secur.1
2026 Large Wireless Localization Model (LWLM): A Foundation Model for Positioning in 6G Networks
abstract
Accurate and robust localization is a critical enabler for emerging 5G and 6G applications, including autonomous driving, extended reality (XR), and smart manufacturing. While data-driven approaches have shown promise, most existing models require large amounts of labeled data and struggle to generalize across deployment scenarios and wireless configurations. To address these limitations, we propose a foundation-model-based solution tailored for wireless localization. We first analyze how different self-supervised learning (SSL) tasks acquire general-purpose and task-specific semantic features based on information bottleneck (IB) theory. Building on this foundation, we design a pretraining methodology for the proposed Large Wireless Localization Model (LWLM). Specifically, we propose an SSL framework that jointly optimizes three complementary objectives: (i) spatial-frequency masked channel modeling (SF-MCM), (ii) domain-transformation invariance (DTI), and (iii) position-invariant contrastive learning (PICL). These objectives jointly capture the underlying semantics of wireless channel from multiple perspectives. We further design lightweight decoders for key downstream tasks, including time-of-arrival (ToA) estimation, angle-of-arrival (AoA) estimation, single base station (BS) localization, and multiple BS localization. Comprehensive experimental results confirm that LWLM consistently surpasses both model-based and supervised learning baselines across all localization tasks. In particular, LWLM achieves 26.0%--87.5% improvement over transformer models without pretraining, and exhibits strong generalization under label-limited fine-tuning and unseen BS configurations, confirming its potential as a foundation model for wireless localization.
Guangjin Pan, Kaixuan Huang, Hui Chen 0014, Shunqing Zhang, Christian Häger, Henk Wymeersch
IEEE Trans. Wirel. Commun.2
2025 SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
abstract
Evaluating aligned large language models' (LLMs) ability to recognize and reject unsafe user requests is crucial for safe, policy-compliant deployments. Existing evaluation efforts, however, face three limitations that we address with **SORRY-Bench**, our proposed benchmark. **First**, existing methods often use coarse-grained taxonomies of unsafe topics, and are over-representing some fine-grained topics. For example, among the ten existing datasets that we evaluated, tests for refusals of self-harm instructions are over 3x less represented than tests for fraudulent activities. SORRY-Bench improves on this by using a fine-grained taxonomy of 44 potentially unsafe topics, and 440 class-balanced unsafe instructions, compiled through human-in-the-loop methods. **Second**, evaluations often overlook the linguistic formatting of prompts, like different languages, dialects, and more --- which are only implicitly considered in many evaluations. We supplement SORRY-bench with 20 diverse linguistic augmentations to systematically examine these effects. **Third**, existing evaluations rely on large LLMs (e.g., GPT-4) for evaluation, which can be computationally expensive. We investigate design choices for creating a fast, accurate automated safety evaluator. By collecting 7K+ human annotations and conducting a meta-evaluation of diverse LLM-as-a-judge designs, we show that fine-tuned 7B LLMs can achieve accuracy comparable to GPT-4 scale LLMs, with lower computational cost. Putting these together, we evaluate over 50 proprietary and open-weight LLMs on SORRY-Bench, analyzing their distinctive safety refusal behaviors. We hope our effort provides a building block for systematic evaluations of LLMs' safety refusal capabilities, in a balanced, granular, and efficient manner. Benchmark demo, data, code, and models are available through [https://sorry-bench.github.io](https://sorry-bench.github.io).
Tinghao Xie, Xiangyu Qi, Yi Zeng 0005, Yangsibo Huang, T. W. U. Madhushani, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng 0007, Ruoxi Jia 0001, Bo Li 0026, Kai Li 0001, Danqi Chen 0001, Peter Henderson 0002, Prateek Mittal
ICLR6
2025 MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
abstract
Large language models have demonstrated impressive performance on challenging mathematical reasoning tasks, which has triggered the discussion of whether the performance is achieved by true reasoning capability or memorization. To investigate this question, prior work has constructed mathematical benchmarks when questions undergo simple perturbations – modifications that still preserve the underlying reasoning patterns of the solutions. However, no work has explored hard perturbations, which fundamentally change the nature of the problem so that the original solution steps do not apply. To bridge the gap, we construct MATH-P-Simple and MATH-P-Hard via simple perturbation and hard perturbation, respectively. Each consists of 279 perturbed math problems derived from level-5 (hardest) problems in the MATH dataset (Hendrycks et al., 2021). We observe significant performance drops on MATH-P-Hard across various models, including o1-mini (-16.49%) and gemini-2.0-flash-thinking (-12.9%). We also raise concerns about a novel form of memorization where models blindly apply learned problem-solving skills without assessing their applicability to modified contexts. This issue is amplified when using original problems for in-context learning. We call for research efforts to address this challenge, which is critical for developing more robust and reliable reasoning models. The project is available at https://math-perturb.github.io/.
Kaixuan Huang, Jiacheng Guo, Jiawei Ge 0003, Tianle Cai, Hui Yuan 0002, Runzhe Wang, Ming Yin 0003, Shange Tang, Yangsibo Huang, Chi Jin 0001, Chiyuan Zhang, Mengdi Wang 0001
ICML1
2025 Emergent Symbolic Mechanisms Support Abstract Reasoning in Large Language Models
abstract
Many recent studies have found evidence for emergent reasoning capabilities in large language models (LLMs), but debate persists concerning the robustness of these capabilities, and the extent to which they depend on structured reasoning mechanisms. To shed light on these issues, we study the internal mechanisms that support abstract reasoning in LLMs. We identify an emergent symbolic architecture that implements abstract reasoning via a series of three computations. In early layers, symbol abstraction heads convert input tokens to abstract variables based on the relations between those tokens. In intermediate layers, symbolic induction heads perform sequence induction over these abstract variables. Finally, in later layers, retrieval heads predict the next token by retrieving the value associated with the predicted abstract variable. These results point toward a resolution of the longstanding debate between symbolic and neural network approaches, suggesting that emergent reasoning in neural networks depends on the emergence of symbolic mechanisms.
Yukang Yang, Declan Campbell, Kaixuan Huang, Mengdi Wang 0001, Jonathan D. Cohen 0003, Taylor W. Webb
ICML3
2025 Deep Reinforcement Learning for Efficient and Fair Allocation of Healthcare Resources
abstract
The scarcity of health care resources, such as ventilators, often leads to the unavoidable consequence of rationing, particularly during public health emergencies or in resource-constrained settings like pandemics. The absence of a universally accepted standard for resource allocation protocols results in governments relying on varying criteria and heuristic-based approaches, often yielding suboptimal and inequitable outcomes. This study addresses the societal challenge of fair and effective critical care resource allocation by leveraging deep reinforcement learning to optimize policy decisions. We propose a transformer-based deep Q-network that integrates individual patient disease progression and interaction effects among patients to enhance allocation decisions. Our method aims to improve both fairness and overall patient outcomes. Experiments using metrics such as normalized survival rates and interracial allocation rate differences demonstrate that our approach significantly reduces excess deaths and achieves more equitable resource allocation compared to severity- and comorbidity-based protocols currently in use. Our findings highlight the potential of deep reinforcement learning to address critical health care challenges.
Yikuan Li, Chengsheng Mao, Kaixuan Huang, Hanyin Wang, Mengdi Wang 0001, Yuan Luo 0001
IJCAI3
2025 Transformer-Based Multi-Agent Reinforcement Learning Method With Credit-Oriented Strategy Differentiation
abstract
The problem of Multi-Agent Reinforcement Learning (MARL) shows a high level of both complexity in the environment and coordination between agents. In order to scale the algorithm to large-scale agent scenarios, neural networks designed for MARL are typically implemented with parameter sharing. These characteristics result in the challenges of partial observability, credit assignment and strategy homogenization. In this paper, a Transformer-Based Multi-Agent Reinforcement Learning Method With Credit-Oriented Strategy Differentiation (TMRC) is presented to address each of these challenges. First, we design a Temporal-Spatial Encoding module and an Attention-Based Value Decomposition module based on the Transformer architecture. The former leverages both temporal and spatial observation information, compensating for the missing environmental perspectives due to partial observability. The latter is designed to identify each agent’s individual contribution in complex interactions, effectively optimizing the credit assignment process. Then, we propose a Credit-Oriented Strategy Differentiation module that differentiates the entity representations of each agent based on their current task differences, allowing agents to have distinct real-time strategies, effectively mitigating the issue of strategy homogenization. We evaluate the proposed method on the SMAC benchmark. It demonstrates better final performance, faster convergence, and greater stability compared to other comparative methods. Additionally, a series of experiments are conducted to validate the effectiveness of the proposed modules. Our code is available at https://github.com/Hkxuan/TMRC.git.
Kaixuan Huang, Bo Jin 0001, Haiyin Piao, Ziqi Wei 0001
IROS1
2025 A Unified QoS-Aware Multiplexing Framework for Next-Generation Immersive Communication With Legacy Wireless Applications
abstract
Immersive communication, including emerging augmented reality, virtual reality, and holographic telepresence, has been identified as a key service for enabling next-generation wireless applications. To align with legacy wireless applications, such as enhanced mobile broadband or ultra-reliable low-latency communication, network slicing has been widely adopted. However, attempting to statistically isolate the above types of wireless applications through different network slices may lead to throughput degradation and increased queue backlog. To address these challenges, we establish a unified QoS-aware framework that supports immersive communication and legacy wireless applications simultaneously. Based on the Lyapunov drift theorem, we transform the original long-term throughput maximization problem into an equivalent short-term throughput maximization weighted by virtual queue length. Moreover, to cope with the challenges introduced by the interaction between large-timescale network slicing and short-timescale resource allocation, we propose an adaptive adversarial slicing (Ad2S) scheme for networks with invarying channel statistics. To track the network channel variations, we also propose a measurement extrapolation-Kalman filter (ME-KF)-based method and refine our scheme into Ad2S-non-stationary refinement (Ad2S-NR). Through extended numerical examples, we demonstrate that our proposed schemes achieve 3.86 Mbps throughput improvement and 63.96% latency reduction with 24.36% convergence time reduction. Within our framework, the trade-off between total throughput and user service experience can be achieved by tuning systematic parameters.
Jihong Li, Shunqing Zhang, Tao Yu 0008, Guangjin Pan, Kaixuan Huang, Xiaojing Chen 0001, Yanzan Sun, Junyu Liu, Jiandong Li 0001, Derrick Wing Kwan Ng
IEEE Internet Things J.5
2025 FedMWAD: Module-wise weighted aggregation federated learning combined with Ditto for patient-independent seizure prediction
Yulan Ding, Wenshan Zhao, Kaixuan Huang
Inf. Process. Manag.3
2024 Visual Adversarial Examples Jailbreak Aligned Large Language Models
abstract
Warning: this paper contains data, prompts, and model outputs that are offensive in nature. Recently, there has been a surge of interest in integrating vision into Large Language Models (LLMs), exemplified by Visual Language Models (VLMs) such as Flamingo and GPT-4. This paper sheds light on the security and safety implications of this trend. First, we underscore that the continuous and high-dimensional nature of the visual input makes it a weak link against adversarial attacks, representing an expanded attack surface of vision-integrated LLMs. Second, we highlight that the versatility of LLMs also presents visual attackers with a wider array of achievable adversarial objectives, extending the implications of security failures beyond mere misclassification. As an illustration, we present a case study in which we exploit visual adversarial examples to circumvent the safety guardrail of aligned LLMs with integrated vision. Intriguingly, we discover that a single visual adversarial example can universally jailbreak an aligned LLM, compelling it to heed a wide range of harmful instructions (that it otherwise would not) and generate harmful content that transcends the narrow scope of a `few-shot' derogatory corpus initially employed to optimize the adversarial example. Our study underscores the escalating adversarial risks associated with the pursuit of multimodality. Our findings also connect the long-studied adversarial vulnerabilities of neural networks to the nascent field of AI alignment. The presented attack suggests a fundamental adversarial challenge for AI alignment, especially in light of the emerging trend toward multimodality in frontier foundation models.
Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson 0002, Mengdi Wang 0001, Prateek Mittal
AAAI2
2024 Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
abstract
Large language models (LLMs) show inherent brittleness in their safety mechanisms, as evidenced by their susceptibility to jailbreaking and even non-malicious fine-tuning. This study explores this brittleness of safety alignment by leveraging pruning and low-rank modifications. We develop methods to identify critical regions that are vital for safety guardrails, and that are disentangled from utility-relevant regions at both the neuron and rank levels. Surprisingly, the isolated regions we find are sparse, comprising about $3$ % at the parameter level and $2.5$ % at the rank level. Removing these regions compromises safety without significantly impacting utility, corroborating the inherent brittleness of the model's safety mechanisms. Moreover, we show that LLMs remain vulnerable to low-cost fine-tuning attacks even when modifications to the safety-critical regions are restricted. These findings underscore the urgent need for more robust safety strategies in LLMs.
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang 0001, Peter Henderson 0002
ICML2
2024 A Theoretical Perspective for Speculative Decoding Algorithm
abstract
Transformer-based autoregressive sampling has been the major bottleneck for slowing down large language model inferences. One effective way to accelerate inference is Speculative Decoding, which employs a small model to sample a sequence of draft tokens and a large model to validate. Given its empirical effectiveness, the theoretical understanding of Speculative Decoding is falling behind. This paper tackles this gap by conceptualizing the decoding problem via markov chain abstraction and studying the key properties, output quality and inference acceleration, from a theoretical perspective. Our analysis covers the theoretical limits of speculative decoding, batch algorithms, and output quality-inference acceleration tradeoffs. Our results reveal the fundamental connections between different components of LLMs via total variation distances and show how they jointly affect the efficiency of decoding algorithms.
Ming Yin 0003, Minshuo Chen, Kaixuan Huang, Mengdi Wang 0001
NeurIPS3
2024 A Novel Cross-band CSI Prediction Scheme for Multi-band Fingerprint based Localization
abstract
Because of the advantages of computation complexity compared with traditional localization algorithms, fingerprint based localization is getting increasing demand. Expanding the fingerprint database from the frequency domain by channel reconstruction can improve localization accuracy. However, in a mobility environment, the channel reconstruction accuracy is limited by the time-varying parameters. In this paper, we proposed a system to extract the time-varying parameters based on space-alternating generalized expectation maximization (SAGE) algorithm, then used variational auto-encoder (VAE) to reconstruct the channel state information on another channel. The proposed scheme is tested on the data generated by the deep-MIMO channel model. Mathematical analysis for the viability of our system is also shown in this paper.
Ruihao Yuan, Kaixuan Huang, Yuru Duan, Shunqing Zhang
WCNC2
2024 Heterogeneous Feature Fusion Approach for Multi-Modal Indoor Localization
abstract
The demand for high-precision localization con-tinues to grow rapidly with the development of information technology. Localization techniques based on wireless signals and visible light images have become the mainstream approach for achieving accurate and precise localization. However, directly utilizing multi-modal data for localization often overlooks the complex relationships between different modalities, particularly in terms of spatial and temporal features at varying scales. In this paper, we present a novel high-precision indoor localization method that effectively aligns the spatiotemporal dimensions of different modalities and extracts features efficiently using a shared neural network. To further enhance the extraction of relevant features from the multi-modal data, we propose a fusion network based on multi-modal channels that effectively minimize disparities, thereby significantly improving localization accuracy. Through extensive verification using a prototype system, our pro-posed solution demonstrates outstanding localization accuracy, achieving an average localization error of only 0.22m.
Kaixuan Huang, Shunqing Zhang
WCNC2
2023 A Variational Auto-Encoder Enabled Multi-Band Channel Prediction Scheme for Indoor Localization
abstract
Indoor localization is getting increasing demands for various cutting-edged technologies, like Virtual/Augmented reality and smart home. Traditional model-based localization suffers from significant computational overhead, so fingerprint localization is getting increasing attention, which needs lower computation cost after the fingerprint database is built. However, the accuracy of indoor localization is limited by the complicated indoor environment which brings the multipath signal refraction. In this paper, we provided a scheme to improve the accuracy of indoor fingerprint localization from the frequency domain by predicting the channel state information (CSI) values from another transmitting channel and spliced the multi-band information together to get more precise localization results. We tested our proposed scheme on COST 2100 simulation data and real time orthogonal frequency division multiplexing (OFDM) WiFi data collected from an office scenario.
Ruihao Yuan, Kaixuan Huang, Pan Yang 0026, Shunqing Zhang
ICC2
2023 Deep Reinforcement Learning for Cost-Effective Medical Diagnosis
Yikuan Li, Joseph C. Kim, Kaixuan Huang, Yuan Luo 0001, Mengdi Wang 0001
ICLR4
2023 Score Approximation, Estimation and Distribution Recovery of Diffusion Models on Low-Dimensional Data
abstract
Diffusion models achieve state-of-the-art performance in various generation tasks. However, their theoretical foundations fall far behind. This paper studies score approximation, estimation, and distribution recovery of diffusion models, when data are supported on an unknown low-dimensional linear subspace. Our result provides sample complexity bounds for distribution estimation using diffusion models. We show that with a properly chosen neural network architecture, the score function can be both accurately approximated and efficiently estimated. Further, the generated distribution based on the estimated score function captures the data geometric structures and converges to a close vicinity of the data distribution. The convergence rate depends on subspace dimension, implying that diffusion models can circumvent the curse of data ambient dimensionality.
Minshuo Chen, Kaixuan Huang, Tuo Zhao, Mengdi Wang 0001
ICML2
2023 Reward-Directed Conditional Diffusion: Provable Distribution Estimation and Reward Improvement
abstract
We explore the methodology and theory of reward-directed generation via conditional diffusion models. Directed generation aims to generate samples with desired properties as measured by a reward function, which has broad applications in generative AI, reinforcement learning, and computational biology. We consider the common learning scenario where the dataset consists of majorly unlabeled data and a small set of data with noisy reward labels. Our approach leverages a learned reward function on the smaller data set as a pseudolabeler to label the unlabelled data. After pseudo-labelling, a conditional diffusion model (CDM) is trained on the data and samples are generated by setting a target value $a$ as the condition in CDM. From a theoretical standpoint, we show that this directed generator can effectively learn and sample from the reward-conditioned data distribution: 1. our model is capable of recovering the data's latent subspace representation. 2. the model generates samples moving closer to the user-specified target. The improvement in rewards of samples is influenced by a interplay between the strength of the reward signal, the distribution shift, and the cost of off-support extrapolation. We provide empirical results to validate our theory and highlight the relationship between the strength of extrapolation and the quality of generated samples.
Hui Yuan 0002, Kaixuan Huang, Chengzhuo Ni, Minshuo Chen, Mengdi Wang 0001
NeurIPS2
2023 Hybrid Cascaded and Feature-Level Fusion Scheme for Multi-Modal Indoor Localization
abstract
Smartphone based indoor localization has been widely explored for meeting the demand of high-precision low-cost indoor localization. Previous methods focus mainly on improving the localization accuracy of single sensor based localization, which may hinder their applications. In this article, we propose a novel encoder-decoder architecture for high-precision low-cost indoor localization. We first leverage two modal-specific encoders for feature extraction. Then, we propose two feature-level fusion strategies for feature fusion. Finally, we leverage two task-specific decoders for both position and orientation prediction. During training, we adopt WiFi-aided learning to provide a more reliable label. We test the proposed method in the corridor environment of a typical building. Experiments results show that our method can achieve less than a half meter localization accuracy, and meanwhile enjoys the run-time efficiency.
Kaixuan Huang, Shunqing Zhang
VTC2023-Spring2
2021 High Precision Indoor Localization with Dummy Antennas - An Experimental Study
abstract
With the rising demand for indoor localization, high precision technique-based fingerprints became increasingly important nowadays. The newest advanced localization system makes effort to improve localization accuracy in the time or frequency domain, for example, the UWB localization technique can achieve centimeter-level accuracy but have a high cost. Therefore, we present a spatial domain extension-based scheme with low cost and verify the effectiveness of antennas extension in localization accuracy. In this paper, we achieve sub-meter level localization accuracy using a single AP by extending three radio links of the modified laptops to more antennas. Moreover, the experimental results show that the localization performance is superior as the number of antennas increases with the help of spatial domain extension and angular domain assisted.
Kaixuan Huang, Chenlu Xiang, Shunqing Zhang, Shugong Xu, Xianfeng Ma, Qinglong Xian
GLOBECOM1
2021 Fast Federated Learning in the Presence of Arbitrary Device Unavailability
abstract
Federated learning (FL) coordinates with numerous heterogeneous devices to collaboratively train a shared model while preserving user privacy. Despite its multiple advantages, FL faces new challenges. One challenge arises when devices drop out of the training process. In this case, the convergence of popular FL algorithms such as FedAvg is severely influenced by the straggling devices. To tackle this challenge, we study federated learning algorithms in the presence of arbitrary device unavailability and propose an algorithm named Memory-augmented Impatient Federated Averaging (MIFA). Our algorithm efficiently avoids excessive latency induced by inactive devices, and corrects the gradient bias using the memorized latest updates from them. We prove that MIFA achieves minimax optimal convergence rates on non-i.i.d. data for both strongly convex and non-convex smooth functions. We also provide an explicit characterization of the improvement over baseline algorithms through a case study, and validate the results by numerical experiments on real-world datasets.
Xinran Gu, Kaixuan Huang, Jingzhao Zhang, Longbo Huang
NeurIPS2
2021 Going Beyond Linear RL: Sample Efficient Neural Function Approximation
abstract
Deep Reinforcement Learning (RL) powered by neural net approximation of the Q function has had enormous empirical success. While the theory of RL has traditionally focused on linear function approximation (or eluder dimension) approaches, little is known about nonlinear RL with neural net approximations of the Q functions. This is the focus of this work, where we study function approximation with two-layer neural networks (considering both ReLU and polynomial activation functions). Our first result is a computationally and statistically efficient algorithm in the generative model setting under completeness for two-layer neural networks. Our second result considers this setting but under only realizability of the neural net function class. Here, assuming deterministic dynamics, the sample complexity scales linearly in the algebraic dimension. In all cases, our results significantly improve upon what can be attained with linear (or eluder dimension) methods.
Baihe Huang, Kaixuan Huang, Sham M. Kakade, Jason D. Lee, Runzhe Wang, Jiaqi Yang 0001
NeurIPS2
2021 Optimal Gradient-based Algorithms for Non-concave Bandit Optimization
abstract
Bandit problems with linear or concave reward have been extensively studied, but relatively few works have studied bandits with non-concave reward. This work considers a large family of bandit problems where the unknown underlying reward function is non-concave, including the low-rank generalized linear bandit problems and two-layer neural network with polynomial activation bandit problem.For the low-rank generalized linear bandit problem, we provide a minimax-optimal algorithm in the dimension, refuting both conjectures in \cite{lu2021low,jun2019bilinear}. Our algorithms are based on a unified zeroth-order optimization paradigm that applies in great generality and attains optimal rates in several structured polynomial settings (in the dimension). We further demonstrate the applicability of our algorithms in RL in the generative model setting, resulting in improved sample complexity over prior approaches.Finally, we show that the standard optimistic algorithms (e.g., UCB) are sub-optimal by dimension factors. In the neural net setting (with polynomial activation functions) with noiseless reward, we provide a bandit algorithm with sample complexity equal to the intrinsic algebraic dimension. Again, we show that optimistic approaches have worse sample complexity, polynomial in the extrinsic dimension (which could be exponentially worse in the polynomial degree).
Baihe Huang, Kaixuan Huang, Sham M. Kakade, Jason D. Lee, Runzhe Wang, Jiaqi Yang 0001
NeurIPS2
2020 On the Convergence of FedAvg on Non-IID Data
Xiang Li 0050, Kaixuan Huang, Shusen Wang, Zhihua Zhang 0004
ICLR2
2020 Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel Perspective
abstract
Deep residual networks (ResNets) have demonstrated better generalization performance than deep feedforward networks (FFNets). However, the theory behind such a phenomenon is still largely unknown. This paper studies this fundamental problem in deep learning from a so-called ``neural tangent kernel'' perspective. Specifically, we first show that under proper conditions, as the width goes to infinity, training deep ResNets can be viewed as learning reproducing kernel functions with some kernel function. We then compare the kernel of deep ResNets with that of deep FFNets and discover that the class of functions induced by the kernel of FFNets is asymptotically not learnable, as the depth goes to infinity. In contrast, the class of functions induced by the kernel of ResNets does not exhibit such degeneracy. Our discovery partially justifies the advantages of deep ResNets over deep FFNets in generalization abilities. Numerical results are provided to support our claim.
Kaixuan Huang, Yuqing Wang 0005, Molei Tao, Tuo Zhao
NeurIPS1
2019 Push-Based Network-efficient Hadoop YARN Scheduling Mechanism for In-Memory Computing
abstract
In the big data era, data-intensive cluster computing systems like Hadoop, have gained much popularity, and YARN, the second generation of Hadoop becomes the general resource manager in the Hadoop ecosystem. In the distributed computing scenarios, data locality (scheduling tasks on where the data resides) is essential to the performance since higher data locality brings lower network transmission cost and higher throughput. However, we find that the native YARN scheduling mechanism has little data locality and the delay scheduling strategy leads to the long-tail effect while achieving data locality for in-memory computing scenarios. Therefore, in this paper we propose the push-based YARN scheduling mechanism for the in-memory computing environment. First, we classify the Resource Requests into various categories. Then, we prune the non-local Resource Requests to achieve fast datalocality in-memory computation. Finally, we push the left longtail Resource Requests to the data-locality nodes to avoid the long-tail effect. The experimental results demonstrate that the proposed scheduling mechanism achieves nearly 100% datalocality percentage comparing to the native YARN scheduling mechanism that only achieves 10% 20% data-locality percentage. Under the identical data-locality percentage, the proposed push based scheduling mechanism promotes nearly 20% throughput and reduces nearly 10% application running time comparing to the existing delay scheduling mechanism used in YARN.
Rong Gu 0001, Kaixuan Huang, Chunfeng Yuan, Yihua Huang 0001
ICPADS2