Yinchuan Li

dblp:236/4930 · DBLP profile ↗
← Back
40ranked-venue papers
8as first author
34since 2021 · last 2026
0000-0002-4263-5130ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 11 since 2021Computer networks · 6 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Personalized Federated Learning with Bidirectional Communication Compression via One-Bit Random Sketching
abstract
Federated Learning (FL) enables collaborative training across decentralized data, but faces key challenges of bidirectional communication overhead and client-side data heterogeneity. To address communication costs while embracing data heterogeneity, we propose pFed1BS, a novel personalized federated learning framework that achieves extreme communication compression through one-bit random sketching. In personalized FL, the goal shifts from training a single global model to creating tailored models for each client. In our framework, clients transmit highly compressed one-bit sketches, and the server aggregates and broadcasts a global one-bit consensus. To enable effective personalization, we introduce a sign-based regularizer that guides local models to align with the global consensus while preserving local data characteristics. To mitigate the computational burden of random sketching, we employ the Fast Hadamard Transform for efficient projection. Theoretical analysis guarantees that our algorithm converges to a stationary neighborhood of the global potential function. Numerical simulations demonstrate that pFed1BS substantially reduces communication costs while achieving competitive performance compared to advanced communication-efficient FL algorithms.
Xu Zhang 0011, Guanghui Qiu, Yinchuan Li, Kaiyuan Feng
AAAI5
2026 Boosting Cross-problem Generalization in Diffusion-Based Neural Combinatorial Solver via Inference Time Adaptation
abstract
Diffusion-based Neural Combinatorial Optimization (NCO) has demonstrated effectiveness in solving NP-complete (NPC) problems by learning discrete diffusion models for solution generation, eliminating hand-crafted domain knowledge. Despite their success, existing NCO methods face significant challenges in both cross-scale and cross-problem generalization, and high training costs compared to traditional solvers. While recent studies on diffusion models have introduced training-free guidance approaches that leverage pre-defined guidance functions for conditional generation, such methodologies have not been extensively explored in combinatorial optimization. To bridge this gap, we propose a training-free inference time adaptation framework (DIFU-Ada) that enables both the zero-shot cross-problem transfer and cross-scale generalization capabilities of diffusion-based NCO solvers without requiring additional training. We provide theoretical analysis that helps understanding the cross-problem transfer capability. Our experimental results demonstrate that a diffusion solver, trained exclusively on the Traveling Salesman Problem (TSP), can achieve competitive zero-shot transfer performance across different problem scales on TSP variants, such as Prize Collecting TSP (PCTSP) and the Orienteering Problem (OP), through inference time adaptation.
Haoyu Lei, Kaiwen Zhou 0001, Yinchuan Li, Zhitang Chen, Farzan Farnia
AAAI3
2025 GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
abstract
Bin Xie, Rui Shao, Gongwei Chen, Kaiwen Zhou, Yinchuan Li, Jie Liu, Min Zhang, Liqiang Nie. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Rui Shao 0001, Gongwei Chen, Kaiwen Zhou 0001, Yinchuan Li, Jie Liu 0001, Min Zhang 0005, Liqiang Nie
ACL (1)5
2025 Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
abstract
Despite the significant success of imitation learning in robotic manipulation, its application to bimanual tasks remains highly challenging. Existing approaches mainly learn a policy to predict a distant next-best end-effector pose (NBP) and then compute the corresponding joint rotation angles for motion using inverse kinematics. However, they suffer from two important issues: (1) rarely considering the physical robotic structure, which may cause self-collisions or interferences, and (2) overlooking the kinematics constraint, which may result in the predicted poses not conforming to the actual limitations of the robot joints. In this paper, we propose Kinematics enhanced Spatial-TemporAl gRaph Diffuser (KStar Diffuser). Specifically, (1) to incorporate the physical robot structure information into action prediction, KStar Diffuser maintains a dynamic spatial-temporal graph according to the physical bimanual joint motions at continuous timesteps. This dynamic graph serves as the robot-structure condition for denoising the actions; (2) to make the NBP learning objective consistent with kinematics, we introduce the differentiable kinematics to provide the reference for optimizing KStar Diffuser. This module regularizes the policy to predict more reliable and kinematics-aware next end-effector poses. Experimental results show that our method effectively leverages the physical structural information and generates kinematics-aware actions in both simulation and real-world.
Qi Lv 0001, Xiang Deng 0002, Rui Shao 0001, Yinchuan Li, Jianye Hao, Longxiang Gao, Michael Yu Wang, Liqiang Nie
CVPR5
2025 Less is More: Empowering GUI Agent with Context-Aware Simplification
abstract
The research focus of GUI agents is shifting from text-dependent to pure-vision-based approaches, which, though promising, prioritize comprehensive pre-training data collection while neglecting contextual modeling challenges. We probe the characteristics of element and history contextual modeling in GUI agent and summarize: 1) the high-density and loose-relation of element context highlight the existence of many unrelated elements and their negative influence; 2) the high redundancy of history context reveals the inefficient history modeling in current GUI agents. In this work, we propose a context-aware simplification framework for building an efficient and effective GUI Agent, termed SimpAgent. To mitigate potential interference from numerous unrelated elements, we introduce a masking-based element pruning method that circumvents the intractable relation modeling through an efficient masking mechanism. To reduce the redundancy in historical information, we devise a consistency-guided history compression module, which enhances implicit LLM-based compression through innovative explicit guidance, achieving an optimal balance between performance and efficiency. With the above components, SimpAgent reduces 27% FLOPs and achieves superior GUI navigation performances. Comprehensive navigation experiments across diverse web and mobile environments demonstrate the effectiveness and potential of our agent.
Gongwei Chen, Xurui Zhou, Rui Shao 0001, Yibo Lyu, Kaiwen Zhou 0001, Shuai Wang 0020, Yinchuan Li, Zhongang Qi, Liqiang Nie
ICCV8
2025 3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
abstract
3D Affordance detection is a challenging problem with broad applications on various robotic tasks. Existing methods typically formulate the detection paradigm as a label-based semantic segmentation task. This paradigm relies on predefined labels and lacks the ability to comprehend complex natural language, resulting in limited generalization in open-world scene. To address these limitations, we reformulate the traditional affordance detection paradigm into \textit{Instruction Reasoning Affordance Segmentation} (IRAS) task. This task is designed to output a affordance mask region given a query reasoning text, which avoids fixed categories of input labels. We accordingly propose the \textit{3D-AffordanceLLM} (3D-ADLLM), a framework designed for reasoning affordance detection in 3D open-scene. Specifically, 3D-ADLLM introduces large language models (LLMs) to 3D affordance perception with a custom-designed decoder for generating affordance masks, thus achieving open-world reasoning affordance detection. In addition, given the scarcity of 3D affordance datasets for training large models, we seek to extract knowledge from general segmentation data and transfer it to affordance detection. Thus, we propose a multi-stage training strategy that begins with a novel pre-training task, i.e., \textit{Referring Object Part Segmentation}~(ROPS). This stage is designed to equip the model with general recognition and segmentation capabilities at the object-part level. Then followed by fine-tuning with the IRAS task, 3D-ADLLM obtains the reasoning ability for affordance detection. In summary, 3D-ADLLM leverages the rich world knowledge and human-object interaction reasoning ability of LLMs, achieving approximately an 8\% improvement in mIoU on open-vocabulary affordance detection tasks.
Hengshuo Chu, Xiang Deng 0002, Qi Lv 0001, Yinchuan Li, Jianye Hao, Liqiang Nie
ICLR5
2025 Ergodic Generative Flows
abstract
Generative Flow Networks (GFNs) were initially introduced on directed non-acyclic graphs to sample from an unnormalized distribution density. Recent works have extended the theoretical framework for generative methods allowing more flexibility and enhancing application range. However, many challenges remain in training GFNs in continuous settings and for imitation learning (IL), including intractability of flow-matching loss, limited tests of non-acyclic training, and the need for a separate reward model in imitation learning. The present work proposes a family of generative flows called Ergodic Generative Flows (EGFs) which are used to address the aforementioned issues. First, we leverage ergodicity to build simple generative flows with finitely many globally defined transformations (diffeomorphisms) with universality guarantees and tractable flow-matching loss (FM loss). Second, we introduce a new loss involving cross-entropy coupled to weak flow-matching control, coined KL-weakFM loss. It is designed for IL training without a separate reward model. We evaluate IL-EGFs on toy 2D tasks and real-world datasets from NASA on the sphere, using the KL-weakFM loss. Additionally, we conduct toy 2D reinforcement learning experiments with a target reward, using the FM loss.
Leo Maxime Brunswic, Mateo Clémente, Rui Heng Yang, Adam Sigal, Amir Rasouli, Yinchuan Li
ICML6
2025 STAR: Learning Diverse Robot Skill Abstractions through Rotation-Augmented Vector Quantization
abstract
Transforming complex actions into discrete skill abstractions has demonstrated strong potential for robotic manipulation.Existing approaches mainly leverage latent variable models, e.g., VQ-VAE, to learn skill abstractions through learned vectors (codebooks), while they suffer from codebook collapse and modeling the causal relationship between learned skills. To address these limitations, we present **S**kill **T**raining with **A**ugmented **R**otation (**STAR**), a framework that advances both skill learning and composition to complete complex behaviors. Specifically, to prevent codebook collapse, we devise rotation-augmented residual skill quantization (RaRSQ).It encodes relative angles between encoder outputs into the gradient flow by rotation-based gradient mechanism. Points within the same skill code are forced to be either pushed apart or pulled closer together depending on gradient directions.Further, to capture the casual relationship between skills, we present causal skill transformer (CST) which explicitly models dependencies between skill representations through an autoregressive mechanism for coherent action generation.Extensive experiments demonstrate the superiority of STAR on both LIBERO benchmark and realworld tasks, with around 12% improvement over the baselines.
Qi Lv 0001, Rui Shao 0001, Xiang Deng 0002, Yinchuan Li, Jianye Hao, Liqiang Nie
ICML5
2025 RA-DP: Rapid Adaptive Diffusion Policy for Training-Free High-frequency Robotics Replanning
abstract
Diffusion models exhibit impressive scalability in robotic task learning, yet they struggle to adapt to novel, highly dynamic environments. This limitation primarily stems from their constrained replanning ability: they either operate at a low frequency due to a time-consuming iterative sampling process, or are unable to adapt to unforeseen feedback in case of rapid replanning. To address these challenges, we propose RA-DP, a novel diffusion policy framework with training-free high-frequency replanning ability that solves the above limitations by adapting to unforeseen dynamic environments. Specifically, our method integrates guidance signals, which are often easily obtained in the new environment during the diffusion sampling process, and utilizes a novel action queue mechanism to generate replanned actions at every denoising step without retraining, thus forming a complete training-free framework for robot motion adaptation in unseen environments. We conduct extensive evaluations in both common simulation benchmarks and real-world environments. Our results indicate that RA-DP outperforms the state-of-the-art diffusion-based methods in terms of replanning frequency and success rate. At the end, we show that our framework is theoretically compatible with any training-free guidance signal, hence increasing its applicability to a wide range of robotics tasks.
Rui Heng Yang, Yinchuan Li, Amir Rasouli
IROS4
2025 UltraVSR: Achieving Ultra-Realistic Video Super-Resolution with Efficient One-Step Diffusion Space
Yong Liu 0031, Jinshan Pan, Yinchuan Li, Qingji Dong, Chao Zhu 0007, Yu Guo 0006, Fei Wang 0008
ACM Multimedia3
2025 Two-Steps Diffusion Policy for Robotic Manipulation via Genetic Denoising
abstract
Diffusion models, such as diffusion policy, have achieved state-of-the-art results in robotic manipulation by imitating expert demonstrations. While diffusion models were originally developed for vision tasks like image and video generation, many of their inference strategies have been directly transferred to control domains without adaptation. In this work, we show that by tailoring the denoising process to the specific characteristics of embodied AI tasks—particularly the structured, low-dimensional nature of action distributions---diffusion policies can operate effectively with as few as 5 neural function evaluations (NFE). Building on this insight, we propose a population-based sampling strategy, genetic denoising, which enhances both performance and stability by selecting denoising trajectories with low out-of-distribution risk. Our method solves challenging tasks with only 2 NFE while improving or matching performance. We evaluate our approach across 14 robotic manipulation tasks from D4RL and Robomimic, spanning multiple action horizons and inference budgets. In over 2 million evaluations, our method consistently outperforms standard diffusion-based policies, achieving up to 20\% performance gains with significantly fewer inference steps.
Mateo Clémente, Leo Maxime Brunswic, Rui Heng Yang, Yasser H. Khalil, Haoyu Lei, Amir Rasouli, Yinchuan Li
NeurIPS8
2025 Conditioning Matters: Training Diffusion Policies is Faster Than You Think
abstract
Diffusion policies have emerged as a mainstream paradigm for building vision-language-action (VLA) models. Although they demonstrate strong robot control capabilities, their training efficiency remains suboptimal. In this work, we identify a fundamental challenge in conditional diffusion policy training: when generative conditions are hard to distinguish, the training objective degenerates into modeling the marginal action distribution, a phenomenon we term loss collapse. To overcome this, we propose Cocos, a simple yet general solution that modifies the source distribution in the conditional flow matching to be condition-dependent. By anchoring the source distribution around semantics extracted from condition inputs, Cocos encourages stronger condition integration and prevents the loss collapse. We provide theoretical justification and extensive empirical results across simulation and real-world benchmarks. Our method achieves faster convergence and higher success rates than existing approaches, matching the performance of large-scale pre-trained VLAs using significantly fewer gradient steps and parameters. Cocos is lightweight, easy to implement, and compatible with diverse policy architectures, offering a general-purpose improvement to diffusion policy training.
Zibin Dong, Yinchuan Li, Hang Zhao 0021, Jianye Hao
NeurIPS3
2025 Proximalized Preference Optimization for Diverse Feedback Types: A Decomposed Perspective on DPO
abstract
Direct alignment methods typically train large language models (LLMs) by contrasting the likelihoods of preferred and dispreferred responses. While effective at capturing relative preferences, these methods are widely observed to suppress the absolute likelihoods of example responses. As a result, aligned models can deviate from expected patterns, exhibiting reward‑hacking effect even without an explicit reward model. This fundamental limitation of contrastive alignment, termed likelihood underdetermination, motivates us to revisit direct preference optimization (DPO)—the seminal direct alignment method. Interestingly, we show that the DPO loss admits a principled decomposition. The reformulated loss not only extends naturally to a broader range of feedback types, but also unveils the root cause of likelihood underdetermination. Specifically, we identify that standard DPO implicitly oversimplifies a regularizer in the reformulated loss; restoring this full term effectively resolves the underdetermination. Building on these insights, we introduce PRoximalized PReference Optimization (PRO), a unified alignment method that accommodates diverse feedback types while eliminating likelihood underdetermination through an efficient approximation of the full regularizer. Empirical evaluations demonstrate the consistent superiority of PRO over existing methods across pairwise, binary and scalar feedback.
Kaiyang Guo, Yinchuan Li, Zhitang Chen
NeurIPS2
2024 A Theory of Non-acyclic Generative Flow Networks
abstract
GFlowNets is a novel flow-based method for learning a stochastic policy to generate objects via a sequence of actions and with probability proportional to a given positive reward. We contribute to relaxing hypotheses limiting the application range of GFlowNets, in particular: acyclicity (or lack thereof). To this end, we extend the theory of GFlowNets on measurable spaces which includes continuous state spaces without cycle restrictions, and provide a generalization of cycles in this generalized context. We show that losses used so far push flows to get stuck into cycles and we define a family of losses solving this issue. Experiments on graphs and continuous tasks validate those principles.
Leo Maxime Brunswic, Yinchuan Li, Yushun Xu, Shangling Jui, Lizhuang Ma
AAAI2
2024 Sparse Federated Learning With Hierarchical Personalization Models
abstract
Federated learning (FL) can achieve privacy-safe and reliable collaborative training without collecting users’ private data. Its excellent privacy security potential promotes a wide range of federated learning (FL) applications in Internet of Things (IoT), wireless networks, mobile devices, autonomous vehicles, and cloud medical treatment. However, the FL method suffers from poor model performance on non-independent and identically distributed (non-i.i.d.) data and excessive traffic volume. To this end, we propose a personalized FL algorithm using a hierarchical proximal mapping based on the moreau envelop, named sparse federated learning with hierarchical personalized models (sFedHP), which significantly improves the acrlong GM performance facing diverse data. A continuously differentiable approximated$\ell _{1}$-norm is also used as the sparse constraint to reduce the communication cost. Convergence analysis shows that sFedHP’s convergence rate is state-of-the-art with linear speedup and the sparse constraint only reduces the convergence rate to a small extent while significantly reducing the communication cost. Experimentally, we demonstrate the benefits of sFedHP compared with the federated averaging (FedAvg), hierarchical fedavg (HierFAVG), and personalized FL methods based on local customization, including FedAMP, FedProx, per- FedAvg, pFedMe, and pFedGP.
Xiaofeng Liu 0009, Qing Wang 0015, Yunfeng Shao 0001, Yinchuan Li
IEEE Internet Things J.4
2024 Meta Generative Flow Networks with personalization for task-specific adaptation
Xinyuan Ji, Xu Zhang 0011, Wei Xi 0003, Haozhi Wang, Olga Gadyatskaya, Yinchuan Li
Inf. Sci.6
2024 Multi-agent Continuous Control with Generative Flow Networks
Yinchuan Li, Shunyu Liu 0001, Xu Zhang 0011, Yunfeng Shao 0001, Chao Wu 0001
Neural Networks2
2024 Towards Effective Clustered Federated Learning: A Peer-to-Peer Framework With Adaptive Neighbor Matching
abstract
In federated learning (FL), clients may have diverse objectives, and merging all clients' knowledge into one global model will cause negative transfer to local performance. Thus, clustered FL is proposed to group similar clients into clusters and maintain several global models. In the literature, centralized clustered FL algorithms require the assumption of the number of clusters and hence are not effective enough to explore the latent relationships among clients. In this paper, without assuming the number of clusters, we propose a peer-to-peer (P2P) FL algorithm namedPANM. InPANM, clients communicate with peers to adaptively form an effective clustered topology. Specifically, we present two novel metrics for measuring client similarity and a two-stage neighbor matching algorithm based Monte Carlo method and Expectation Maximization under the Gaussian Mixture Model assumption. We have conducted theoretical analyses ofPANMon the probability of neighbor estimation and the error gap to the clustered optimum. We have also implemented extensive experiments under both synthetic and real-world clustered heterogeneity. Theoretical analysis and empirical experiments show that the proposed algorithm is superior to the P2P FL counterparts, and it achieves better performance than the centralized cluster FL method.PANMis effective even under extremely low communication budgets.
Zexi Li 0001, Jiaxun Lu, Didi Zhu, Yunfeng Shao 0001, Yinchuan Li, Yongheng Wang, Chao Wu 0001
IEEE Trans. Big Data6
2024 Device Activity Detection and Channel Estimation for Millimeter-Wave Massive MIMO
abstract
Millimeter-Wave Massive MIMO is important for beyond 5G or 6G wireless communication networks. The goal of this paper is to establish successful communication between the cellular base stations and devices, focusing on the problem of joint user activity detection and channel estimation. Different from traditional compressed sensing (CS) methods that only use the sparsity of user activities, we develop several Approximate Message Passing (AMP) based CS algorithms by exploiting the sparsity of user activities and mmWave channels. First, a group soft-thresholding AMP is presented to utilize only the user activity sparsity. Second, a hard-thresholding AMP is proposed based on the on-grid CS approach. Third, a super-resolution AMP algorithm is proposed based on atomic norm, in which a greedy method is proposed as a super-resolution denoiser. And we smooth the denoiser based on Monte Carlo sampling to have Lipschitz continuity and present state evolution results. Extensive simulation results show that the proposed method outperforms the previous state-of-the-art methods.
Yinchuan Li, Yuancheng Zhan, Le Zheng, Xiaodong Wang 0001
IEEE Trans. Commun.1
2024 MAP: Model Aggregation and Personalization in Federated Learning With Incomplete Classes
abstract
In some real-world applications, data samples are usually distributed on local devices, where federated learning (FL) techniques are proposed to coordinate decentralized clients without directly sharing users’ private data. FL commonly follows the parameter server architecture and contains multiple personalization and aggregation procedures. The natural data heterogeneity across clients, i.e., Non-I.I.D. data, challenges both the aggregation and personalization goals in FL. In this paper, we focus on a special kind of Non-I.I.D. scene where clients own incomplete classes, i.e., each client can only access a partial set of the whole class set. The server aims to aggregate a complete classification model that could generalize to all classes, while the clients are inclined to improve the performance of distinguishing their observed classes. For better model aggregation, we point out that the standard softmax will encounter several problems caused by missing classes and propose “restricted softmax” as an alternative. For better model personalization, we point out that the hard-won personalized models are not well exploited and propose “inherited private model” to store the personalization experience. Our proposed algorithm named MAP could simultaneously achieve the aggregation and personalization goals in FL. Abundant experimental studies verify the superiorities of our algorithm.
Xin-Chun Li, Shaoming Song, Yinchuan Li, Bingshuai Li, Yunfeng Shao 0001, Yang Yang 0074, De-Chuan Zhan
IEEE Trans. Knowl. Data Eng.3
2024 Sparse Personalized Federated Learning
abstract
Federated learning (FL) is a collaborative machine learning technique to train a global model (GM) without obtaining clients' private data. The main challenges in FL are statistical diversity among clients, limited computing capability among clients' equipment, and the excessive communication overhead between the server and clients. To address these challenges, we propose a novel sparse personalized FL scheme via maximizing correlation (FedMac). By incorporating an approximated $\ell _{1}$ -norm and the correlation between client models and GM into standard FL loss function, the performance on statistical diversity data is improved and the communicational and computational loads required in the network are reduced compared with nonsparse FL. Convergence analysis shows that the sparse constraints in FedMac do not affect the convergence rate of the GM, and theoretical results show that FedMac can achieve good sparse personalization, which is better than the personalized methods based on the $\ell _{2}$ -norm. Experimentally, we demonstrate the benefits of this sparse personalization architecture compared with the state-of-the-art personalization methods (e.g., FedMac, respectively, achieves 98.95%, 99.37%, 90.90%, 89.06%, and 73.52% accuracy on the MNIST, FMNIST, CIFAR-100, Synthetic, and CINIC-10 datasets under non-independent and identically distributed (i.i.d.) variants).
Xiaofeng Liu 0009, Yinchuan Li, Qing Wang 0015, Xu Zhang 0011, Yunfeng Shao 0001, Yanhui Geng
IEEE Trans. Neural Networks Learn. Syst.2
2023 Universal Domain Adaptation via Compressive Attention Matching
abstract
Universal domain adaptation (UniDA) aims to transfer knowledge from the source domain to the target domain without any prior knowledge about the label set. The challenge lies in how to determine whether the target samples belong to common categories. The mainstream methods make judgments based on the sample features, which overemphasizes global information while ignoring the most crucial local objects in the image, resulting in limited accuracy. To address this issue, we propose a Universal Attention Matching (UniAM) framework by exploiting the self-attention mechanism in vision transformer to capture the crucial object information. The proposed framework introduces a novel Compressive Attention Matching (CAM) approach to explore the core information by compressively representing attentions. Furthermore, CAM incorporates a residual-based measurement to determine the sample commonness. By utilizing the measurement, UniAM achieves domain-wise and category-wise Common Feature Alignment (CFA) and Target Class Separation (TCS). Notably, UniAM is the first method utilizing the attention in vision transformer directly to perform classification tasks. Extensive experiments show that UniAM outperforms the current state-of-the-art methods on various benchmark datasets.
Didi Zhu, Yinchuan Li, Junkun Yuan, Zexi Li 0001, Kun Kuang 0001, Chao Wu 0001
ICCV2
2023 DAG Matters! GFlowNets Enhanced Explainer for Graph Neural Networks
Yinchuan Li, Jianye Hao
ICLR2
2023 CFlowNets: Continuous Control with Generative Flow Networks
Yinchuan Li, Haozhi Wang, Jianye Hao
ICLR1
2023 Generative Flow Networks for Precise Reward-Oriented Active Learning on Graphs
abstract
Many score-based active learning methods have been successfully applied to graph-structured data, aiming to reduce the number of labels and achieve better performance of graph neural networks based on predefined score functions. However, these algorithms struggle to learn policy distributions that are proportional to rewards and have limited exploration capabilities. In this paper, we innovatively formulate the graph active learning problem as a generative process, named GFlowGNN, which generates various samples through sequential actions with probabilities precisely proportional to a predefined reward function. Furthermore, we propose the concept of flow nodes and flow features to efficiently model graphs as flows based on generative flow networks, where the policy network is trained with specially designed rewards. Extensive experiments on real datasets show that the proposed approach has good exploration capability and transferability, outperforming various state-of-the-art methods.
Yinchuan Li, Yunfeng Shao 0001, Yan Zheng 0002, Jianye Hao
IJCAI1
2023 Generalized Universal Domain Adaptation with Generative Flow Networks
abstract
We introduce a new problem in unsupervised domain adaptation, termed as Generalized Universal Domain Adaptation (GUDA), which aims to achieve precise prediction of all target labels including unknown categories. GUDA bridges the gap between label distribution shift-based and label space mismatch-based variants, essentially categorizing them as a unified problem, guiding to a comprehensive framework for thoroughly solving all the variants. The key challenge of GUDA is developing and identifying novel target categories while estimating the target label distribution. To address this problem, we take advantage of the powerful exploration capability of generative flow networks and propose an active domain adaptation algorithm named GFlowDA, which selects diverse samples with probabilities proportional to a reward function. To enhance the exploration capability and effectively perceive the target label distribution, we tailor the states and rewards, and introduce an efficient solution for parent exploration and state transition. We also propose a training paradigm for GUDA called Generalized Universal Adversarial Network (GUAN), which involves collaborative optimization between GUAN and GFlowNet. Theoretical analysis highlights the importance of exploration, and extensive experiments on benchmark datasets demonstrate the superiority of GFlowDA.
Didi Zhu, Yinchuan Li, Yunfeng Shao 0001, Jianye Hao, Fei Wu 0001, Kun Kuang 0001, Jun Xiao 0001, Chao Wu 0001
ACM Multimedia2
2022 Federated Learning with Position-Aware Neurons
abstract
Federated Learning (FL) fuses collaborative models from local nodes without centralizing users' data. The permutation invariance property of neural networks and the non-i.i.d. data across clients make the locally updated parameters imprecisely aligned, disabling the coordinate-based parameter averaging. Traditional neurons do not explicitly consider position information. Hence, we propose Position-Aware Neurons (PANs) as an alternative, fusing position-related values (i.e., position encodings) into neuron outputs. PANs couple themselves to their positions and minimize the possibility of dislocation, even updating on heterogeneous data. We turn on/off PANs to disable/enable the permutation invariance property of neural networks. PANs are tightly coupled with positions when applied to FL, making parameters across clients pre-aligned and facilitating coordinate-based parameter averaging. PANs are algorithm-agnostic and could universally improve existing FL algorithms. Furthermore, “FL with PANs” is simple to implement and computationally friendly.
Xin-Chun Li, Yichu Xu, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, De-Chuan Zhan
CVPR5
2022 Personalized Federated Learning via Variational Bayesian Inference
abstract
Federated learning faces huge challenges from model overfitting due to the lack of data and statistical diversity among clients. To address these challenges, this paper proposes a novel personalized federated learning method via Bayesian variational inference named pFedBayes. To alleviate the overfitting, weight uncertainty is introduced to neural networks for clients and the server. To achieve personalization, each client updates its local distribution parameters by balancing its construction error over private data and its KL divergence with global distribution from the server. Theoretical analysis gives an upper bound of averaged generalization error and illustrates that the convergence rate of the generalization error is minimax optimal up to a logarithmic factor. Experiments show that the proposed method outperforms other advanced personalized methods on personalized models, e.g., pFedBayes respectively outperforms other SOTA algorithms by 1.25%, 0.42% and 11.71% on MNIST, FMNIST and CIFAR-10 under non-i.i.d. limited data.
Xu Zhang 0011, Yinchuan Li, Wenpeng Li, Kaiyang Guo, Yunfeng Shao 0001
ICML2
2022 Avoid Overfitting User Specific Information in Federated Keyword Spotting
abstract
Keyword spotting (KWS) aims to discriminate a specific wakeup word from other signals precisely and efficiently for different users.Recent works utilize various deep networks to train KWS models with all users' speech data centralized without considering data privacy.Federated KWS (FedKWS) could serve as a solution without directly sharing users' data.However, the small amount of data, different user habits, and various accents could lead to fatal problems, e.g., overfitting or weight divergence.Hence, we propose several strategies to encourage the model not to overfit user-specific information in FedKWS.Specifically, we first propose an adversarial learning strategy, which updates the downloaded global model against an overfitted local model and explicitly encourages the global model to capture user-invariant information.Furthermore, we propose an adaptive local training strategy, letting clients with more training data and more uniform class distributions undertake more local update steps.Equivalently, this strategy could weaken the negative impacts of those users whose data is less qualified.Our proposed FedKWS-UI could explicitly and implicitly learn user-invariant information in FedKWS.Abundant experimental results on federated Google Speech Commands verify the effectiveness of FedKWS-UI.
Xin-Chun Li, Jin-Lin Tang, Shaoming Song, Bingshuai Li, Yinchuan Li, Yunfeng Shao 0001, Le Gan, De-Chuan Zhan
INTERSPEECH5
2022 S2RL: Do We Really Need to Perceive All States in Deep Multi-Agent Reinforcement Learning?
abstract
Collaborative multi-agent reinforcement learning (MARL) has been widely used in many practical applications, where each agent makes a decision based on its own observation. Most mainstream methods treat each local observation as an entirety when modeling the decentralized local utility functions. However, they ignore the fact that local observation information can be further divided into several entities, and only part of the entities is helpful to model inference. Moreover, the importance of different entities may change over time. To improve the performance of decentralized policies, the attention mechanism is used to capture features of local information. Nevertheless, existing attention models rely on dense fully connected graphs and cannot better perceive important states. To this end, we propose a sparse state based MARL (S2RL) framework, which utilizes a sparse attention mechanism to discard irrelevant information in local observations. The local utility functions are estimated through the self-attention and sparse attention mechanisms separately, then are combined into a standard joint value function and auxiliary joint value function in the central critic. We design the S2RL framework as a plug-and-play module, making it general enough to be applied to various methods. Extensive experiments on StarCraft II show that S2RL can significantly improve the performance of many state-of-the-art methods.
Yinchuan Li, Jiahui Li 0003, Kun Kuang 0001, Furui Liu, Yunfeng Shao 0001, Chao Wu 0001
KDD2
2022 Asymmetric Temperature Scaling Makes Larger Networks Teach Well Again
abstract
Knowledge Distillation (KD) aims at transferring the knowledge of a well-performed neural network (the {\it teacher}) to a weaker one (the {\it student}). A peculiar phenomenon is that a more accurate model doesn't necessarily teach better, and temperature adjustment can neither alleviate the mismatched capacity. To explain this, we decompose the efficacy of KD into three parts: {\it correct guidance}, {\it smooth regularization}, and {\it class discriminability}. The last term describes the distinctness of {\it wrong class probabilities} that the teacher provides in KD. Complex teachers tend to be over-confident and traditional temperature scaling limits the efficacy of {\it class discriminability}, resulting in less discriminative wrong class probabilities. Therefore, we propose {\it Asymmetric Temperature Scaling (ATS)}, which separately applies a higher/lower temperature to the correct/wrong class. ATS enlarges the variance of wrong class probabilities in the teacher's label and makes the students grasp the absolute affinities of wrong classes to the target class as discriminative as possible. Both theoretical analysis and extensive experimental results demonstrate the effectiveness of ATS. The demo developed in Mindspore is available at \url{https://gitee.com/lxcnju/ats-mindspore} and will be available at \url{https://gitee.com/mindspore/models/tree/master/research/cv/ats}.
Xin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li, Bingshuai Li, Yunfeng Shao 0001, De-Chuan Zhan
NeurIPS4
2022 Parametric Translational Compensation for ISAR Imaging Based on Cascaded Subaperture Integration With Application to Asteroid Imaging
abstract
Translational compensation is a crucial step in inverse synthetic aperture radar (ISAR) imaging. However, nonparametric compensation methods may collapse under the condition of a low signal-to-noise ratio (SNR), and high-order parametric compensation methods usually have a large computational load. In this article, two cascaded integration methods are proposed to solve the problems above based on the generalized Radon-Fourier transform (GRFT). The first method models the translational motion as polynomial and utilizes the proposed subaperture GRFT (SAGRFT), which is a fast implementation of the GRFT, to estimate the translational parameters. The SAGRFT divides the full aperture into several subapertures and realizes coherent integration by implementing moving target detection (MTD) within subapertures and the GRFT among subapertures. In addition, the distribution property of the blind speed side lobes (BSSLs) generated by the SAGRFT is analyzed. On this basis, the second method based on metaheuristic algorithms is proposed to further accelerate the parameter estimation process and solve the BSSLs problem. The proposed methods can not only play a role in stealth target imaging but also be utilized in near-Earth asteroid (NEA) imaging in radar astronomy. Finally, the numerical simulations of asteroid imaging and the experimental results of plane imaging are demonstrated to verify the performance advantages of the proposed methods.
Zegang Ding, Yinchuan Li, Peng-Jie You
IEEE Trans. Geosci. Remote. Sens.3
2021 Unfolded Deep Neural Network (UDNN) for High Mobility Channel Estimation
abstract
High mobility channel estimation is crucial for beyond 5G(B5G) or 6G wireless communication networks. This paper is concerned with channel estimation of high mobility OFDM communication systems. First, a two-dimensional compressed sensing problem is formulated by approximately linearizing the channel as a product of an overcomplete dictionary with a sparse vector, in which the Doppler effect caused by the high mobility channel is considered. To solve the problem that the traditional compressed sensing algorithms have too many iterations and are time consuming, we propose an unfolded deep neural network (UDNN) as the fast solver, which is inspired by the structure of iterative shrinkage-thresholding algorithm (ISTA). All the parameters in UDNN (e.g. nonlinear transforms, shrinkage thresholds, measurement matrices, etc.) are learned end-to-end, rather than being hand-crafted. Experiments demonstrate that the proposed UDNN performs better than ISTA for OFDM high mobility channel estimation, while maintaining extremely fast computational speed.
Yinchuan Li, Xiaodong Wang 0001, Robert L. Olesen
WCNC1
2021 SAR Parametric Super-Resolution Image Reconstruction Methods Based on ADMM and Deep Neural Network
abstract
The compressed sensing (CS)-based synthetic aperture radar (SAR) imaging methods have emerged as the standard approach to obtain super-resolution (SR) SAR images and achieve extraordinary performances. However, they face three challenges. First, this kind of method is mainly based on the point scattering model and not suitable for characterizing the line-segment-scattering and surface-scattering features of distributed targets. Second, the hyperparameters in these methods are hard to tune to optimal values. Third, due to a large amount of calculation, these methods are difficult to apply in practice. In this article, to solve these problems, we introduce the line-segment-scatterers (LSSs) and rectangular-plate-scatterers (RPSs) in SAR echo model to develop the SAR hybrid echo model and propose two SAR parametric SR image reconstruction methods based on solving a CS problem, where three penalties are utilized to exploit the sparsity of the point scatterers, LSSs, and RPSs, respectively. At the core of the first method is a direct solver called multicomponent alternating direction method of multipliers (MC-ADMM) solver that solves the CS problem quickly and iteratively based on closed derivative expressions. In contrast, the second method maps the MC-ADMM solver into a deep unfolded neural network, i.e., the parametric SR imaging network (PSRI-Net), which is faster, and the parameters can be automatically set to the optimum. Since all the parameters of the MC-ADMM solver are learned discriminatively through end-to-end training in PSRI-Net. Extensive simulation and practical experiments are carried out to demonstrate the effectiveness of the proposed methods.
Yangkai Wei, Yinchuan Li, Zegang Ding, Yan Wang 0011, Tao Zeng 0001, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.2
2020 Multidimensional Spectral Super-Resolution With Prior Knowledge With Application to High Mobility Channel Estimation
abstract
The problem of estimating high-mobility channels is a special case of the general problem of recovering multi-dimensional (MD) complex sinusoids. This paper is concerned with estimation of multiple frequencies with prior knowledge from incomplete and/or noisy samples. Suppose that it is known a priori that the frequencies lie in some given intervals, we develop efficient super-resolution estimators by exploiting such prior knowledge based on frequency-selective (FS) atomic norm minimization. We study the MD Vandermonde decomposition of block Toeplitz matrices in which the frequencies are restricted to lie in given intervals. We then propose to solve the FS atomic norm minimization problems for the low-rank spectral tensor recovery by converting them into semidefinite programs based on the MD Vandermonde decomposition. We also develop fast solvers for solving these semidefinite programs via the alternating direction method of multipliers (ADMM), where each iteration involves a number of refinement steps to utilize the prior knowledge. Extensive simulation results are presented to illustrate the high performance of the proposed methods.
Yinchuan Li, Xiaodong Wang 0001, Zegang Ding
IEEE J. Sel. Areas Commun.1
2020 Multi-Angle SAR Sparse Image Reconstruction With Improved Attributed Scattering Model
abstract
The traditional synthetic aperture radar (SAR) sparse imaging methods are based on the point scattering model. However, this model is not suitable for many distributed targets with large variations in scattering characteristics at different angles, i.e., many distributed targets can no longer be considered as a combination of a series of ideal point scatterers under multi-angle observations. To solve this problem, by introducing the improved attributed scattering model into the traditional SAR echo model, we propose our multi-angle sparse image reconstruction method (MASIRM). Through modeling the illuminated scene with point scatterers and line-segment-scatterers, a multi-angle echo model is first presented. By generating an adaptive mixed dictionary and applying the pattern-coupled sparse Bayesian learning, the MASIRM obtains more geometric information of the distributed target with higher quality sparse SAR images. Real data experiments demonstrate that MASIRM performs favorably against traditional imaging methods.
Yangkai Wei, Yinchuan Li, Xinliang Chen, Zegang Ding
IEEE Geosci. Remote. Sens. Lett.2
2020 Multi-Target Position and Velocity Estimation Using OFDM Communication Signals
abstract
In this paper, we consider a passive radar system that estimates the positions and velocities of multiple moving targets by using OFDM signals transmitted by a totally un-coordinated and un-synchronizated illuminator and multiple receivers. It is assumed that data demodulation is performed separately based on the direct-path signal, and the error-prone estimated data symbols are made available to the passive radar receivers, which estimate the positions and velocities of the targets in two stages. First, we formulate a problem of joint estimation of the delay-Doppler of reflectors and the demodulation errors, by exploiting two types of sparsities of the system, namely, the numbers of reflectors (i.e., targets and clutters) and demodulation errors are both small. This problem is non-convex and a conjugate gradient descent method is proposed to solve it. Then in the second stage we determine the positions and velocities of targets based on the estimated delay-Doppler in the first stage. For the second stage, two methods are proposed: the first is based on numerically solving a set of nonlinear equations, while the second is based on the neural network, which is more efficient. The performance of the proposed algorithms is evaluated through extensive simulations.
Yinchuan Li, Xiaodong Wang 0001, Zegang Ding
IEEE Trans. Commun.1
2020 Near-Field Phase Cross Correlation Focusing Imaging and Parameter Estimation for Penetrating Radar
abstract
Penetrating radar systems are widely used to image the objects that are buried inside mediums (such as walls, ground, and so on). However, due to the phase error caused by the refraction of mediums, the object images obtained by directly applying the imaging methods that ignore the existence of the mediums are defocused, which affects the recognition of small objects. Conventional focus imaging algorithms typically obtain a focused image by iteratively searching for optimal compensation, which is very time-consuming and can result in overfocusing. To solve this problem, a near-field phase cross correlation (PCC) focusing imaging algorithm is proposed in this article. First, a free-space imaging model is established to replace the unknown half-space (air-to-medium) imaging model. Then, by analyzing the relationship between the free-space and the unknown half-space imaging models, the phase error and a reference depth in free-space imaging model are determined. The focused image can then be obtained by compensating for the phase error, which is directly estimated by the proposed PCC method. The permittivity and object depth can then be estimated based on the slope of the phase error and the reference depth. Furthermore, the proposed algorithm can be extended to the multi-layered model in a straightforward manner. Extensive simulations and experiments are presented to validate the proposed methods. The results show that the proposed PCC algorithm can effectively compensate the phase error and obtain high-quality images, and the permittivity and object depth can be accurately estimated in single-layer medium cases.
Zegang Ding, Yinchuan Li, Yin Xiang, Tao Zeng 0001, Teng Long 0001
IEEE Trans. Geosci. Remote. Sens.2
2019 Spectrum Recovery for Clutter Removal in Penetrating Radar Imaging
abstract
Penetrating radar systems are widely employed to scan the objects that are placed behind or buried inside mediums (such as walls, ground, and so on). As the clutter is much stronger than the target echo, clutter removal must be performed before imaging. The moving average subtraction, spatial notch filtering, and singular value decomposition methods are commonly used to remove clutter. However, the drawback is that these methods eliminate some of the target spectrum information, which causes target energy losses and generates side lobes. To solve this problem, two spectrum recovery methods are proposed in this paper. The first method recovers the spectrum magnitude and phase via sinc interpolation and linear fitting, respectively, which is fast and suitable for real-time processing. Although the second method recovers the spectrum based on matrix completion with prior information, which is more accurate and more computational expensive. Extensive simulations and experiments are presented to validate the proposed methods. The results show that the proposed methods can improve various traditional clutter removal methods, the side lobes are clearly suppressed, and the signal-to-clutter ratio is significantly improved.
Yinchuan Li, Xiaodong Wang 0001, Zegang Ding, Xu Zhang 0011, Yin Xiang, Xiaopeng Yang 0002
IEEE Trans. Geosci. Remote. Sens.1
2019 Interference Removal for Radar/Communication Co-Existence: The Random Scattering Case
abstract
In this paper, we consider an un-cooperative spectrum sharing scenario, where a radar system is to be overlaid to a pre-existing wireless communication system. Given the order of magnitude of the transmitted powers in play, we focus on the issue of interference mitigation at the communication receiver. We explicitly account for the reverberation produced by the (typically high-power) radar transmitter whose signal hits scattering centers (whether targets or clutter) producing interference onto the communication receiver, which is assumed to operate in an un-synchronized and un-coordinated scenario. We first show that the receiver design amounts to solve a joint (non-convex) interference removal and data demodulation problem. Next, we introduce two algorithms exploiting sparsity of a proper representation of the interference and the vector containing demodulation errors of the data block. The first algorithm is basically a relaxed constrained atomic norm minimization, while the latter relies on a two-stage processing structure and is based on alternating minimization. The merits of these algorithms are demonstrated through extensive simulations; interestingly, the two-stage alternating minimization algorithm turns out to achieve satisfactory performance with moderate computational complexity.
Yinchuan Li, Le Zheng, Marco Lops, Xiaodong Wang 0001
IEEE Trans. Wirel. Commun.1