Minxue Tang

dblp:250/9350 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 68% Language models and text generation · 16% Reinforcement learning · 9%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 100%
Network and information security
3 papers
Security and privacy of machine learning · 84% Privacy and data protection · 16%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
federated learning
1.222023
Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction · ICML 2023
FedCor: Correlation-Based Active Client Selection Strategy for Heterogeneous Federated Learning · CVPR 2022
Natural language and speech › Language models and text generation › trustworthy language model
large language model privacy
0.912025
Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility · ICLR 2025
Security and privacy of machine learning
model stealing
0.812024
ModelGuard: Information-Theoretic Defense Against Model Extraction Attacks · USENIX Security Symposium 2024
Machine learning › Efficient and distributed learning › federated learning › client selection
client sampling
0.712023
Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction · ICML 2023
Machine learning › Efficient and distributed learning › federated learning
data heterogeneity
0.712023
Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction · ICML 2023
Machine learning › Efficient and distributed learning › federated learning
client selection
0.612022
FedCor: Correlation-Based Active Client Selection Strategy for Heterogeneous Federated Learning · CVPR 2022
Machine learning › Efficient and distributed learning › federated learning
heterogeneous federated learning
0.612022
FedCor: Correlation-Based Active Client Selection Strategy for Heterogeneous Federated Learning · CVPR 2022
Electronic design automation › machine learning for EDA
design quality prediction
0.612022
Towards collaborative intelligence: routability estimation based on decentralized private data · DAC 2022
Electronic design automation
machine learning for EDA
0.612022
Towards collaborative intelligence: routability estimation based on decentralized private data · DAC 2022
Electronic design automation › physical design › routing
routability
0.612022
Towards collaborative intelligence: routability estimation based on decentralized private data · DAC 2022
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.412019
Hierarchical Reinforcement Learning with Advantage-Based Auxiliary Rewards · NeurIPS 2019
Machine learning › Trustworthy machine learning
privacy
0.212024
ModelGuard: Information-Theoretic Defense Against Model Extraction Attacks · USENIX Security Symposium 2024
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.212022
FedCor: Correlation-Based Active Client Selection Strategy for Heterogeneous Federated Learning · CVPR 2022
Security and privacy of machine learning
federated learning
0.212022
Towards collaborative intelligence: routability estimation based on decentralized private data · DAC 2022
Privacy and data protection
privacy-preserving machine learning
0.212022
Towards collaborative intelligence: routability estimation based on decentralized private data · DAC 2022
Machine learning › Reinforcement learning › sparse reward reinforcement learning
long-horizon sparse-reward tasks
0.112019
Hierarchical Reinforcement Learning with Advantage-Based Auxiliary Rewards · NeurIPS 2019

Methods — techniques the papers use, named apart from their topics

memory editing · 1.7machine unlearning · 1.7information theory · 1.5neural network · 1.1federated learning · 1.1class-imbalance reduction · 0.7class-balanced sampling · 0.7gaussian process · 0.6correlation modeling · 0.6auxiliary reward · 0.4advantage function · 0.4
YearPublicationVenuePosition
2025 Proactive Privacy Amnesia for Large Language Models: Safeguarding PII with Negligible Impact on Model Utility
abstract
With the rise of large language models (LLMs), increasing research has recognized their risk of leaking personally identifiable information (PII) under malicious attacks. Although efforts have been made to protect PII in LLMs, existing methods struggle to balance privacy protection with maintaining model utility. In this paper, inspired by studies of amnesia in cognitive science, we propose a novel approach, Proactive Privacy Amnesia (PPA), to safeguard PII in LLMs while preserving their utility. This mechanism works by actively identifying and forgetting key memories most closely associated with PII in sequences, followed by a memory implanting using suitable substitute memories to maintain the LLM’s functionality. We conduct evaluations across multiple models to protect common PII, such as phone numbers and physical addresses, against prevalent PII-targeted attacks, demonstrating the superiority of our method compared with other existing defensive techniques. The results show that our PPA method completely eliminates the risk of phone number exposure by 100% and significantly reduces the risk of physical address exposure by 9.8% – 87.6%, all while maintaining comparable model utility performance.
Martin Kuo, Jingyang Zhang, Minxue Tang, Louis DiValentin, Aolin Ding, Jingwei Sun 0002, Amin Hass, Tianlong Chen 0001, Yiran Chen 0001, Hai Li 0001
ICLR4
2024 ModelGuard: Information-Theoretic Defense Against Model Extraction Attacks
Minxue Tang, Anna Dai, Louis DiValentin, Aolin Ding, Amin Hass, Neil Zhenqiang Gong, Yiran Chen 0001, Hai Li 0001
USENIX Security Symposium1
2023 Fed-CBS: A Heterogeneity-Aware Client Sampling Mechanism for Federated Learning via Class-Imbalance Reduction
abstract
Due to the often limited communication bandwidth of edge devices, most existing federated learning (FL) methods randomly select only a subset of devices to participate in training at each communication round. Compared with engaging all the available clients, such a random-selection mechanism could lead to significant performance degradation on non-IID (independent and identically distributed) data. In this paper, we present our key observation that the essential reason resulting in such performance degradation is the class-imbalance of the grouped data from randomly selected clients. Based on this observation, we design an efficient heterogeneity-aware client sampling mechanism, namely, Federated Class-balanced Sampling (Fed-CBS), which can effectively reduce class-imbalance of the grouped dataset from the intentionally selected clients. We first propose a measure of class-imbalance which can be derived in a privacy-preserving way. Based on this measure, we design a computation-efficient client sampling strategy such that the actively selected clients will generate a more class-balanced grouped dataset with theoretical guarantees. Experimental results show that Fed-CBS outperforms the status quo approaches in terms of test accuracy and the rate of convergence while achieving comparable or even better performance than the ideal setting where all the available clients participate in the FL training.
Ang Li 0005, Minxue Tang, Jingwei Sun 0002, Xiang Chen 0010, Fan Zhang 0069, Changyou Chen, Yiran Chen 0001, Hai Li 0001
ICML3
2022 FedCor: Correlation-Based Active Client Selection Strategy for Heterogeneous Federated Learning
abstract
Client-wise data heterogeneity is one of the major issues that hinder effective training in federated learning (FL). Since the data distribution on each client may vary dramatically, the client selection strategy can significantly influence the convergence rate of the FL process. Active client selection strategies are popularly proposed in recent studies. However, they neglect the loss correlations between the clients and achieve only marginal improvement compared to the uniform selection strategy. In this work, we propose FedCoran FLframework built on a correlation-based client selection strategy, to boost the convergence rate of FL. Specifically, we first model the loss correlations between the clients with a Gaussian Process (GP). Based on the GP model, we derive a client selection strategy with a significant reduction of expected global loss in each round. Besides, we develop an efficient GP training method with a low communication overhead in the FL scenario by utilizing the covariance stationarity. Our experimental results show that compared to the state-of-the-art method, FedCorr can improve the convergence rates by 34% ~ 99% and 26% ~ 51% on FMNIST and CIFAR-10, respectively.
Minxue Tang, Xuefei Ning, Yitu Wang, Jingwei Sun 0002, Yu Wang 0002, Hai Li 0001, Yiran Chen 0001
CVPR1
2022 Towards collaborative intelligence: routability estimation based on decentralized private data
abstract
Applying machine learning (ML) in design flow is a popular trend in Electronic Design Automation (EDA) with various applications from design quality predictions to optimizations. Despite its promise, which has been demonstrated in both academic researches and industrial tools, its effectiveness largely hinges on the availability of a large amount of high-quality training data. In reality, EDA developers have very limited access to the latest design data, which is owned by design companies and mostly confidential. Although one can commission ML model training to a design company, the data of a single company might be still inadequate or biased, especially for small companies. Such data availability problem is becoming the limiting constraint on future growth of ML for chip design. In this work, we propose an Federated-Learning based approach for well-studied ML applications in EDA. Our approach allows an ML model to be collaboratively trained with data from multiple clients but without explicit access to the data for respecting their data privacy. To further strengthen the results, we co-design a customized ML model FLNet and its personalization under the decentralized training scenario. Experiments on a comprehensive dataset show that collaborative training improves accuracy by 11% compared with individual local models, and our customized model FLNet significantly outperforms the best of previous routability estimators in this collaborative training flow.
Jingyu Pan, Chen-Chia Chang, Zhiyao Xie, Ang Li 0005, Minxue Tang, Tunhou Zhang, Jiang Hu 0001, Yiran Chen 0001
DAC5
2019 Hierarchical Reinforcement Learning with Advantage-Based Auxiliary Rewards
abstract
Hierarchical Reinforcement Learning (HRL) is a promising approach to solving long-horizon problems with sparse and delayed rewards. Many existing HRL algorithms either use pre-trained low-level skills that are unadaptable, or require domain-specific information to define low-level rewards. In this paper, we aim to adapt low-level skills to downstream tasks while maintaining the generality of reward design. We propose an HRL framework which sets auxiliary rewards for low-level skill training based on the advantage function of the high-level policy. This auxiliary reward enables efficient, simultaneous learning of the high-level policy and low-level skills without using task-specific knowledge. In addition, we also theoretically prove that optimizing low-level skills with this auxiliary reward will increase the task return for the joint policy. Experimental results show that our algorithm dramatically outperforms other state-of-the-art HRL methods in Mujoco domains. We also find both low-level and high-level policies trained by our algorithm transferable.
Siyuan Li 0003, Minxue Tang, Chongjie Zhang
NeurIPS3