Quanjun Yin

dblp:147/8406 · DBLP profile ↗
← Back
50ranked-venue papers
1as first author
42since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Computer networks · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Autonomous UAV Visual Object Search in City Space: Benchmark and Agentic Methodology
abstract
Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects based on visual inputs without external guidance. Existing approaches struggle in complex urban environments due to redundant semantic processing, similar object ambiguity, and the exploration-exploitation dilemma. To advance research and support the AVOS task, we introduce CityAVOS, the first benchmark dataset for autonomous search of static urban objects. It features 2,420 tasks of varying difficulty across six object categories, designed to rigorously evaluate UAV search strategies. To solve the AVOS task, we also propose PRPSearcher (Perception-Reasoning-Planning Searcher), a novel agentic method powered by multi-modal large language models (MLLMs) that enables a UAV agent to think and reason like humans on visual cues when searching for objects. Specifically, PRPSearcher constructs three specialized maps: an object-centric dynamic semantic map enhancing spatial perception, a 3D cognitive map based on semantic "attraction" values for target reasoning, and a 3D uncertainty map for balanced exploration-exploitation search. Moreover, we propose a denoising mechanism to mitigate interference from similar objects and design an Inspiration Promote Thought prompting mechanism for adaptive action planning. Experimental results on CityAVOS demonstrate that PRPSearcher surpasses existing baselines in both success rate and search efficiency (on average: +37.69% SR, +28.96% SPL, -30.69% MSS, and -46.40% NE). Our work paves the way for future advances in embodied visual target search.
Yatai Ji, Zhengqiu Zhu, Beidan Liu, Chen Gao 0001, Sihang Qiu, Yue Hu 0016, Quanjun Yin
AAAI9
2026 Learn to Relax with Large Language Models: Solving Constraint Optimization Problems via Bidirectional Coevolution
abstract
Large Language Model (LLM)-based optimization has recently shown promise for autonomous problem solving, yet most approaches still cast LLMs as passive constraint checkers rather than proactive strategy designers, limiting their effectiveness on complex Constraint Optimization Problems (COPs).To address this, we present AutoCO, an end-to-end Automated Constraint Optimization method that tightly couples operations-research principles of constraint relaxation with LLM reasoning.A core innovation is a unified triplerepresentation that binds relaxation strategies, algorithmic principles, and executable codes.This design enables the LLM to synthesize, justify, and instantiate relaxation strategies that are both principled and executable.To navigate fragmented solution spaces, AutoCO employs a bidirectional global-local coevolution mechanism, synergistically coupling Monte Carlo Tree Search (MCTS) for global relaxationtrajectory exploration with Evolutionary Algorithms (EAs) for local solution intensification.This continuous exchange of priors and feedback explicitly balances diversification and intensification, thus preventing premature convergence.Extensive experiments on three challenging COP benchmarks validate AutoCO's consistent effectiveness and superior performance, especially in hard regimes where current methods degrade.Results highlight Au-toCO as a principled and effective path toward proactive, verifiable LLM-driven optimization.
Beidan Liu, Zhengqiu Zhu, Chen Gao 0001, Tianle Pu, Quanjun Yin
ACL (1)7
2026 CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
abstract
Haotian Xu, Yue Hu, Zhengqiu Zhu, Chen Gao, Ziyou Wang, Junreng Rao, Wenhao Lu, Weishi Li, Quanjun Yin, Yong Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yue Hu 0016, Zhengqiu Zhu, Chen Gao 0001, Ziyou Wang, Junreng Rao, Wenhao Lu, Quanjun Yin, Yong Li 0008
ACL (1)9
2026 DSPIGCN: Dual-stream Physics-informed Graph Convolutional Network for Reliable Pedestrian Trajectory Prediction
Runkang Guo, Bin Chen 0003, Zhengqiu Zhu, Chen Gao 0001, Quanjun Yin
KDD (1)6
2026 Natural Language-Guided Autonomous Agents for Counterterrorism Simulation via Deep Reinforcement Learning
Xinmeng Li, Kai Xu 0014, Yue Hu 0016, Miao Zhang 0037, Quanjun Yin
SIMULTECH5
2026 Trajectory-level reward relabelling and progressive sub-task flight control for engineering-grade simulation of adversarial aerial engagements
Hesong Huang, Xinmeng Li, Quanjun Yin
Eng. Appl. Artif. Intell.6
2026 Indoor cooperative exploration using robot swarms enhanced by human understanding
Zhengqiu Zhu, Yaqiong Zhou, Sihang Qiu, Quanjun Yin
Expert Syst. Appl.6
2026 Boosting the Performance of Decentralized Federated Learning via Catalyst Acceleration
abstract
Decentralized Federated Learning has emerged as an alternative to centralized architectures due to its faster training, privacy preservation, and reduced communication overhead. In decentralized communication, the server aggregation phase in Centralized Federated Learning shifts to the client side, which means that clients connect with each other in a peer-to-peer manner. However, compared to the centralized mode, data heterogeneity in Decentralized Federated Learning will cause larger variances between aggregated models, which leads to slow convergence in training and poor generalization performance in tests. To address these issues, we introduce Catalyst Acceleration and propose an acceleration Decentralized Federated Learning algorithm called DFedCata. It consists of two main components: the Moreau envelope function, which primarily addresses parameter inconsistencies among clients caused by data heterogeneity, and Nesterov's extrapolation step, which accelerates the aggregation phase. Theoretically, we prove the optimization error bound and generalization error bound of the algorithm, providing a further understanding of the nature of the algorithm and the theoretical perspectives on the hyperparameter choice. Empirically, we demonstrate the advantages of the proposed algorithm in both convergence speed, computational cost, and generalization performance on CIFAR10/100 and Tiny-ImageNet with various non-iid data distributions. Moreover, extensive experiments are conducted to validate the theoretical properties of DFedCata, showing strong consistency between theory and empirical observations.
Qinglun Li, Miao Zhang 0037, Yingqi Liu, Quanjun Yin, Li Shen 0008, Xiaochun Cao
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 GeoNav: Empowering MLLMs with dual-scale geospatial reasoning for language-goal aerial navigation
Yue Hu 0016, Chen Gao 0001, Zhengqiu Zhu, Quanjun Yin
Pattern Recognit.6
2026 Geometric Contrastive Ensemble Distillation with calibration margin induction
Yang Yang 0080, Junyao Hou, Xiang Li 0067, Zhenghua Chen, Yingxue Gao, Chao Wang 0003, Min Wu 0008, Quanjun Yin
Pattern Recognit.10
2026 Safe and Energy-Efficient Trajectory Planning for Heterogeneous Multi-UAV Enabled Mobile Edge Computing
abstract
Mobile edge computing (MEC) has recently gained significant attention as a promising solution for processing delay sensitive and resource-intensive computational jobs. Existing system schedulers in MEC networks typically assume homogeneous service providers, uniformly distributed user equipment (UE), and identical service requirements, making them unsuitable for practical MEC scenarios where jobs are randomly generated with varying service and completion time requirements. Thus, in this work, we jointly optimize job scheduling and resource allocation in a heterogeneous multi-unmanned aerial vehicle (UAV) enabled MEC network, considering practical factors such as diverse service requirements of jobs, unknown distribution of UEs, and spatial-temporal job arrivals. We aim to reduce the overall job miss rate and the average energy consumption of both UAVs and UEs by jointly planning safe UAV trajectories and onboard resource allocation. To learn uncertain and dynamic UE-side states (e.g., job arrivals and mobility patterns) and ensure the UAV's safety during the flight, we propose a multi-agent safe reinforcement learning algorithm that combines a Shared Soft Actor-Critic architecture for extracting features of heterogeneous UAVs and a two-agent Markov Game of Intervention mechanism for collision avoidance, named SSAC-MGI. In particular, SSAC MGI further incorporates a fine-grained resource allocation scheme to improve onboard resource utilization and reduce job miss rate. Extensive real trace-driven simulations based on Alibaba cluster data validate the effectiveness and superiority of SSAC-MGI, compared with several state-of-the-art algorithms.
Riheng Jia, Quanjun Yin, Zhonglong Zheng, Minglu Li 0001
IEEE Trans. Mob. Comput.3
2026 Asynchronous Task Scheduling and Resource Allocation for UAV-Enabled Mobile Edge Computing Networks
abstract
Mobile edge computing (MEC) is promising in handling delay-sensitive or resource-intensive tasks in mobile internet. Existing system schedulers in MEC networks usually schedule all service providers in a synchronous manner, which may not suit the practical scenario where tasks are randomly generated and require different computational resources and service times. In this work, we jointly optimize the task scheduling and resource allocation in an unmanned aerial vehicle (UAV)-enabled MEC network, where multiple UAVs asynchronously and cooperatively deliver task offloading and computing services to edge devices (EDs), for maximizing the average per-UAV energy utility and minimizing the overall task missing ratio. To enhance scheduling efficiency and jointly optimize task scheduling and resource allocation, we develop an asynchronous layered multi-agent proximal policy optimization (AL-MAPPO) algorithm, by incorporating the multi-UAV asynchronous action execution mechanism and a discrete-continuous layered action space into the general MAPPO framework. AL-MAPPO enables each UAV to perform flexible task scheduling and fine-grained resource allocation asynchronously. Extensive trace-driven simulations based on Alibaba Cluster Data V2017 validate the effectiveness of AL-MAPPO, compared with several baseline algorithms.
Riheng Jia, Quanjun Yin, Zhonglong Zheng, Minglu Li 0001
IEEE Trans. Serv. Comput.3
2025 PychoAgent: Psychology-driven LLM Agents for Explainable Panic Prediction on Social Media during Sudden Disaster Events
abstract
Mengzhu Liu, Zhengqiu Zhu, Chuan Ai, Chen Gao, Xinghong Li, Lingnan He, Kaisheng Lai, Yingfeng Chen, Xin Lu, Yong Li, Quanjun Yin. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Mengzhu Liu, Zhengqiu Zhu, Chuan Ai, Chen Gao 0001, Xinghong Li, Lingnan He, Kaisheng Lai, Xin Lu 0002, Yong Li 0008, Quanjun Yin
EMNLP11
2025 Unveiling the Power of Multiple Gossip Steps: A Stability-Based Generalization Analysis in Decentralized Training
abstract
Decentralized training removes the centralized server, making it a communication-efficient approach that can significantly improve training efficiency, but it often suffers from degraded performance compared to centralized training. Multi-Gossip Steps (MGS) serve as a simple yet effective bridge between decentralized and centralized training, significantly reducing experiment performance gaps. However, the theoretical reasons for its effectiveness and whether this gap can be fully eliminated by MGS remain open questions. In this paper, we derive upper bounds on the generalization error and excess error of MGS using stability analysis, systematically answering these two key questions. 1). Optimization Error Reduction: MGS reduces the optimization error bound at an exponential rate, thereby exponentially tightening the generalization error bound and enabling convergence to better solutions. 2). Gap to Centralization: Even as MGS approaches infinity, a non-negligible gap in generalization error remains compared to centralized mini-batch SGD ($\mathcal{O}(T^{\frac{c\beta}{c\beta +1}}/{n m})$ in centralized and $\mathcal{O}(T^{\frac{2c\beta}{2c\beta +2}}/{n m^{\frac{1}{2c\beta +2}}})$ in decentralized). Furthermore, we provide the first unified analysis of how factors like learning rate, data heterogeneity, node count, per-node sample size, and communication topology impact the generalization of MGS under non-convex settings without the bounded gradients assumption, filling a critical theoretical gap in decentralized training. Finally, promising experiments on CIFAR datasets support our theoretical findings.
Qinglun Li, Yingqi Liu, Miao Zhang 0037, Xiaochun Cao, Quanjun Yin, Li Shen 0008
NeurIPS5
2025 Multi-agent reinforcement learning for task offloading with hybrid decision space in multi-access edge computing
Miao Zhang 0037, Quanjun Yin, Lujia Yin, Yong Peng 0006
Ad Hoc Networks3
2025 DFedGFM: Pursuing global consistency for Decentralized Federated Learning via global flatness and global momentum
Qinglun Li, Miao Zhang 0037, Tao Sun 0005, Quanjun Yin, Li Shen 0008
Neural Networks4
2025 DFedADMM: Dual Constraint Controlled Model Inconsistency for Decentralize Federated Learning
abstract
To address the communication burden issues associated with Federated Learning (FL), Decentralized Federated Learning (DFL) discards the central server and establishes a decentralized communication network, where each client communicates only with neighboring clients. However, existing DFL methods still suffer from two major challenges: local inconsistency and local heterogeneous overfitting, which existing DFL methods have not fundamentally addressed. To tackle these issues, we propose novel DFL algorithms, DFedADMM and its enhanced version DFedADMM-SAM, to improve the performance for DFL. The DFedADMM algorithm employs primal-dual optimization (ADMM) by utilizing dual variables to control the model inconsistency raised from the decentralized heterogeneous data distributions. The DFedADMM-SAM algorithm further improves on DFedADMM by employing a Sharpness-Aware Minimization (SAM) optimizer, which uses gradient perturbations to generate locally flat models and searches for models with uniformly low loss values to mitigate local heterogeneous overfitting. Theoretically, we derive convergence rates of $\mathcal {O}(\frac{1}{\sqrt{KT}}+\frac{1}{KT(1-\psi )^{2}})$O(1KT+1KT(1-ψ)2) and $ \mathcal {O}(\frac{1}{\sqrt{KT}}+\frac{1}{KT(1-\psi )^{2}}+ \frac{1}{T^{3/2}K^{1/2}})$O(1KT+1KT(1-ψ)2+1T3/2K1/2) in the non-convex setting for DFedADMM and DFedADMM-SAM, respectively, where $1 - \psi$1-ψ represents the spectral gap of the gossip matrix. Empirically, extensive experiments on MNIST, CIFAR10, and CIFAR100 datasets demonstrate that our algorithms exhibit superior performance in terms of generalization, convergence speed, and communication overhead compared to existing state-of-the-art (SOTA) optimizers in DFL.
Qinglun Li, Li Shen 0008, Guanghao Li 0002, Quanjun Yin, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Asymmetrically Decentralized Federated Learning
abstract
To address the communication burden and privacy concerns associated with the centralized server in Federated Learning (FL), Decentralized Federated Learning (DFL) has emerged, which discards the server with a peer-to-peer (P2P) communication framework, significantly expanding the application scenarios of FL. However, most existing DFL algorithms are based on symmetric topologies, such as ring and grid topology, which can easily lead to deadlocks and are susceptible to the impact of network link quality in practice. To address these issues, we propose DFedSGPSM, a transitional framework that converts symmetric DFL optimizers into asymmetric variants. By adopting the Push-Sum protocol in asymmetric network topologies, our framework successfully circumvents the deadlock and link-quality issues prevalent in symmetric configurations. To further validate the effectiveness of our algorithm framework, we integrate the local momentum (in DFedAvgM) and SAM (in DFedSAM) from existing symmetric DFL optimizer into DFedSGPSM to accelerate training and pursue smooth local minimum, which enables existing symmetric DFL optimizers to be seamlessly integrated into asymmetric DFL. Theoretical analysis proves that DFedSGPSM achieves a linear speedup rate of$\mathcal{O}(\frac{1}{\sqrt{nT}})$in the non-convex setting. This analysis also reveals crucial issues such as tighter upper bounds achieved with improved topological connectivity. Empirically, extensive experiments conducted on the MNIST, CIFAR10&100 datasets demonstrate the superior performance of our proposed algorithm compared to several existing SOTA optimizers in terms of generalization.
Qinglun Li, Miao Zhang 0037, Quanjun Yin, Li Shen 0008, Xiaochun Cao
IEEE Trans. Computers4
2025 SwimVG: Step-Wise Multimodal Fusion and Adaption for Visual Grounding
abstract
Visual grounding aims to ground an image region through natural language, which heavily relies on cross-modal alignment. Most existing methods transfer visual/linguistic knowledge separately by fully fine-tuning uni-modal pre-trained models, followed by a simple stack of visual-language transformers for multimodal fusion. However, these approaches not only limit adequate interaction between visual and linguistic contexts, but also incur significant computational costs. Therefore, to address these issues, we explore a step-wise multimodal fusion and adaption framework, namely SwimVG. Specifically, SwimVG proposes step-wise multimodal prompts (Swip) and cross-modal interactive adapters (CIA) for visual grounding, replacing the cumbersome transformer stacks for multimodal fusion. Swip can improve the alignment between the vision and language representations step by step, in a token-level fusion manner. In addition, weight-level CIA further promotes multimodal fusion by cross-modal interaction. Swip and CIA are both parameter-efficient paradigms, and they fuse the cross-modal features from shallow to deep layers gradually. Experimental results on four widely-used benchmarks demonstrate that SwimVG achieves remarkable abilities and considerable benefits in terms of efficiency.
Liangtao Shi, Ting Liu 0018, Xiantao Hu, Yue Hu 0016, Quanjun Yin, Richang Hong
IEEE Trans. Multim.5
2025 The Risk of Federated Learning to Skew Fine-Tuning Features and Underperform Robustness
abstract
To tackle the scarcity and privacy issues associated with domain-specific datasets, the integration of federated learning in conjunction with fine-tuning (FT) has emerged as a practical solution. However, our findings reveal that federated learning has the risk of skewing FT features and compromising the out-of-distribution (OOD) robustness of pretrained models. By introducing three robustness indicators and conducting experiments across diverse robust datasets, we elucidate these phenomena by scrutinizing the ability of data representations, transferability, and deviations within the model. To mitigate the negative impact of practical federated learning on model robustness, we introduce a general noisy projection (GNP)-based robust algorithm, ensuring no deterioration of accuracy on the target distribution. Specifically, the key strategy for enhancing model robustness entails the transfer of robustness from the pretrained model to the fine-tuned model, coupled with adding a small amount of Gaussian noise to augment the representative capacity of the model. The comprehensive experimental results demonstrate that our approach markedly enhances the robustness across diverse scenarios, encompassing various parameter-efficient FT (PEFT) methods and confronting different levels of label distribution skew and quantity distribution skew.
Mengyao Du, Miao Zhang 0037, Yuwen Pu, Qingming Li, Shouling Ji, Quanjun Yin
IEEE Trans. Neural Networks Learn. Syst.6
2024 MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression Comprehension
abstract
Referring Expression Comprehension (REC), which aims to ground a local visual region via natural language, is a task that heavily relies on multimodal alignment.Most existing methods utilize powerful pre-trained models to transfer visual/linguistic knowledge by full fine-tuning.However, full fine-tuning the entire backbone not only breaks the rich prior knowledge embedded in the pre-training, but also incurs significant computational costs.Motivated by the recent emergence of Parameter-Efficient Transfer Learning (PETL) methods, we aim to solve the REC task in an effective and efficient manner.Directly applying these PETL methods to the REC task is inappropriate, as they lack the specific-domain abilities for precise local visual perception and visual-language alignment.Therefore, we propose a novel framework of Multimodal Prior-guided Parameter Efficient Tuning, namely MaPPER.Specifically, MaPPER comprises Dynamic Prior Adapters guided by an aligned prior, and Local Convolution Adapters to extract precise local semantics for better visual perception.Moreover, the Prior-Guided Text module is proposed to further utilize the prior for facilitating the cross-modal alignment.Experimental results on three widely-used benchmarks demonstrate that MaPPER achieves the best accuracy compared to the full fine-tuning and other PETL methods with only 1.41% tunable backbone parameters.Our code is available at https://github.com/liuting20/MaPPER.
Zunnan Xu, Liangtao Shi, Quanjun Yin
EMNLP6
2024 DAP: Domain-Aware Prompt Learning for Vision-and-Language Navigation
abstract
Following language instructions to navigate in unseen environments is a challenging task for autonomous embodied agents. With strong representation capabilities, pretrained vision-and-language models are widely used in VLN. However, most of them are trained on web-crawled generalpurpose datasets, which incurs a considerable domain gap when used for VLN tasks. To address the problem, we propose a novel and model-agnostic Domain-Aware Prompt learning (DAP) framework. For equipping the pretrained models with specific object-level and scene-level cross-modal alignment in VLN tasks, DAP applies a low-cost prompt tuning paradigm to learn soft visual prompts for extracting in-domain image semantics. Specifically, we first generate a set of in-domain image-text pairs with the help of the CLIP model. Then we introduce soft visual prompts in the input space of the visual encoder in a pretrained model. DAP injects in-domain visual knowledge into the visual encoder of the pretrained model in an efficient way. Experimental results on both R2R and REVERIE show the superiority of DAP compared to existing state-of-the-art methods.
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin
ICASSP6
2024 DARA: Domain- and Relation-Aware Adapters Make Parameter-Efficient Tuning for Visual Grounding
abstract
Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved performance, but also introduced a significant burden on computational costs during fine-tuning. In this paper, we explore applying parameter-efficient transfer learning (PETL) to efficiently transfer the pre-trained vision-language knowledge to VG. Specifically, we propose DARA, a novel PETL method comprising Domain-aware Adapters (DA Adapters) and Relation-aware Adapters (RA Adapters) for VG. DA Adapters first transfer intra-modality representations to be more fine-grained for the VG domain. Then RA Adapters share weights to bridge the relation between two modalities, improving spatial reasoning. Empirical results on widely-used benchmarks demonstrate that DARA achieves the best accuracy while saving numerous updated parameters compared to the full fine-tuning and other PETL methods. Notably, with only 2.13% tunable backbone parameters, DARA improves average accuracy by 0.81% across the three benchmarks compared to the baseline model. Our code is available at https://github.com/liuting20/DARA.
Ting Liu 0018, Xuyang Liu 0002, Siteng Huang, Honggang Chen, Quanjun Yin, Long Qin 0004, Yue Hu 0016
ICME5
2024 Enhancing Multimodal Sentiment Analysis via Learning from Large Language Model
abstract
Multimodal sentiment analysis (MSA) detects human sentiments by understanding data from multiple modalities, such as text and images. Existing research primarily strives for an effective multimodal fusion framework to derive informative representations. However, these methods neglect the necessity of exploiting external knowledge to aid in analyzing sentiments. As a result, the lack of external commonsense embarrasses these models when the opinion cues come in an implicit and obscure manner. To address the limitation, in this paper, we propose an Auxiliary Rationale Knowledge enhanced framework, namely ARK, which improves MSA models via learning from a multimodal large language model (MLLM). Specifically, based on text-image pairs, we employ Chain-of-Thought prompting to generate image descriptions and rationales from the MLLM as auxiliary knowledge, thus enriching the original samples with commonsense knowledge encoded within the MLLM. By combining the source text with image descriptions, we are able to effectively handle MSA through a Text+Text paradigm. In this paradigm, smaller pre-trained language models (LMs) can be tasked for sentiment classification via prompt-tuning. Besides, rationales are leveraged as additional supervision to facilitate the learning of reasoning abilities by LMs. Experimental results demonstrate that our proposed method outperforms current state-of-the-art approaches across four datasets. Our data and code are available at https://github.com/ningpang/ArkMSA.
Ning Pang, Wansen Wu, Yue Hu 0016, Kai Xu 0014, Quanjun Yin, Long Qin 0004
ICME5
2024 PANDA: Prompt-Based Context- and Indoor-Aware Pretraining for Vision and Language Navigation
Ting Liu 0018, Yue Hu 0016, Wansen Wu, Youkai Wang, Kai Xu 0014, Quanjun Yin
MMM (1)6
2024 Reinforcement learning from suboptimal demonstrations based on Reward Relabeling
Yong Peng 0006, Junjie Zeng 0004, Yue Hu 0016, Quanjun Yin
Expert Syst. Appl.5
2024 Vision-language navigation: a survey and taxonomy
Wansen Wu, Tao Chang, Xinmeng Li, Quanjun Yin, Yue Hu 0016
Neural Comput. Appl.4
2024 Conversational Crowdsensing in the Age of Industry 5.0: A Parallel Intelligence and Large Models Powered Novel Sensing Approach
abstract
The transition from cyber-physical-system-based (CPS-based) Industry 4.0 to cyber-physical-social-system-based (CPSS-based) Industry 5.0 brings new requirements and opportunities to current sensing approaches, especially in light of recent progress in large language models (LLMs) and retrieval augmented generation (RAG). Therefore, the advancement of parallel intelligence powered crowdsensing intelligence (CSI) is witnessed, which is currently advancing toward linguistic intelligence. In this article, we propose a novel sensing paradigm, namely conversational crowdsensing, for Industry 5.0 (especially for social manufacturing). It can alleviate workload and professional requirements of individuals and promote the organization and operation of diverse workforce, thereby facilitating faster response and wider popularization of crowdsensing systems. Specifically, we design the architecture of conversational crowdsensing to effectively organize three types of participants (biological, robotic, and digital) from diverse communities. Through three levels of effective conversation (i.e., interhuman, human–AI, and inter-AI), complex interactions and service functionalities of different workers can be achieved to accomplish various tasks across three sensing phases (i.e., requesting, scheduling, and executing). Moreover, we explore the foundational technologies for realizing conversational crowdsensing, encompassing LLM-based multiagent systems, scenarios engineering and conversational human–AI cooperation. Finally, we present potential applications of conversational crowdsensing and discuss its implications. We envision that conversations in natural language will become the primary communication channel during crowdsensing process, enabling richer information exchange and cooperative problem-solving among humans, robots, and AI.
Zhengqiu Zhu, Sihang Qiu, Kai Xu 0014, Quanjun Yin, Jincai Huang 0001, Zhong Liu 0002, Fei-Yue Wang 0001
IEEE Trans. Comput. Soc. Syst.5
2024 Visual Grounding With Dual Knowledge Distillation
abstract
Visual grounding is a task that seeks to predict the specific location of an object or region described by a linguistic expression within an image. Despite the recent success, existing methods still suffer from two problems. First, most methods use independently pre-trained unimodal feature encoders for extracting expressive feature embeddings, thus resulting in a significant semantic gap between unimodal embeddings and limiting the effective interaction of visual-linguistic contexts. Second, existing attention-based approaches equipped with the global receptive field have a tendency to neglect the local information present in the images. This limitation restricts the semantic understanding required to distinguish between referred objects and the background, consequently leading to inadequate localization performance. Inspired by the recent advance in knowledge distillation, in this paper, we propose a DUal knowlEdge disTillation (DUET) method for visual grounding models to bridge the cross-modal semantic gap and improve localization performance simultaneously. Specifically, we utilize the CLIP model as the teacher model to transfer the semantic knowledge to a student model, in which the vision and language modalities are linked into a unified embedding space. Besides, we design a self-distillation method for the student model to acquire localization knowledge by performing the region-level contrastive learning to make the predicted region close to the positive samples. To this end, this work further proposes a Semantics-Location Aware sampling mechanism to generate high-quality self-distillation samples. Extensive experiments on five datasets and ablation studies demonstrate the state-of-the-art performance of DUET and its orthogonality with different student models, thereby making DUET adaptable to a wide range of visual grounding architectures. Our code are available on DUET.
Wansen Wu, Meng Cao 0002, Yue Hu 0016, Yong Peng 0006, Long Qin 0004, Quanjun Yin
IEEE Trans. Circuits Syst. Video Technol.6
2024 Intelligent Trajectory Design and Charging Scheduling in Wireless Rechargeable Sensor Networks With Obstacles
abstract
Wireless rechargeable sensor networks (WRSNs) are promising in maintaining sustainable large-area monitoring tasks. Mobile chargers (MCs) are commonly used in WRSNs to replenish energy to nodes due to its flexibility and easy maintenance. Most existing works on WRSNs focus on designing offline or model-based online charging methods, which need the exact system information to conduct the optimization. However, in practical WRSNs, the exact system information such as the nodes' locations and energy consumption rates may not be easily accessible to the optimizer due to their unpredictability and high dynamics. Thus, in this work, we jointly optimize the MC's trajectory design and charging scheduling in a general and practical WRSN with inaccessibility to the exact system information, such that the charging utility of the MC is maximized. To address this problem, we introduce the model-free reinforcement learning (RL) technique, which enables the MC to learn to jointly optimize its moving trajectory and charging scheduling by interacting with the environment and tracking feedback signals from nodes and obstacles in real time. Specifically, we develop a soft actor-critic based mobile security policy intervened algorithm (SAC-MSPI) based on a novel safe RL framework, which maximizes the MC's charging utility while maintaining the safe movement (not hitting obstacles) for the MC during the entire charging period. Extensive evaluation results show that the proposed SAC-MSPI algorithm outperforms existing main RL solutions and traditional algorithms with respect to the charging utility maximization as well as the collision avoidance.
Riheng Jia, Quanjun Yin, Zhonglong Zheng, Minglu Li 0001
IEEE Trans. Mob. Comput.3
2024 Low-Cost, High-Reliability Deployment for Cloud Applications With Low-Frequency Periodic Requests
abstract
Low-frequency periodic requests are common in cloud-based enterprise applications. These infrequent requests often leave microservices idle for extended periods, leading to low resource utilization. Furthermore, the randomness of response times may decrease the reliability of the cloud platform. Intuitively, the periodic nature of requests allows for the agile deployment of microservices to promptly free up occupied computing resources. Thus, the key lies in designing low-cost, high-reliability microservice deployment schemes. Traditional approaches relying on specialized expertise are impractical because of intricate interdependencies within microservice frameworks. To address this, the Microservice Deployment Problem for Low-frequency Periodic Requests (MDP-LPR) is formulated, and a Mixed Integer Programming (MIP) model is developed. A deployment framework leveraging statistical analysis and Monte Carlo simulation is proposed to ensure high reliability. Furthermore, a two-stage heuristic algorithm named Relaxation and Precision Mixed Algorithm (RPMA) is introduced to generate low-cost deployment schemes. Finally, experiments are conducted on real-world workflows. The results show that the RPMA outperforms its counterparts in generating low-cost deployment schemes, and the proposed deployment framework enables the automatic acquisition of low-cost, high-reliability deployment schemes.
Zhu Xiang, Lujia Yin, Miao Zhang 0037, Quanjun Yin
IEEE Trans. Serv. Comput.5
2023 Credit-based Differential Privacy Stochastic Model Aggregation Algorithm for Robust Federated Learning via Blockchain
abstract
By encapsulating model parameters in blocks when training machine learning models collaboratively, blockchain is recognized as a promising enabling technology to facilitate reliable federated learning under a distributed and untrusted environment. However, storing updated models of each worker in blockchain induces potential privacy risks, such as membership inference attacks. Besides, the volatile network conditions in the distributed environment may cause the deterioration of system robustness. This paper aims at addressing the privacy and robustness issues mentioned above. Specifically, a Credit-based Differential Privacy stochastic model aggregation algorithm combined with SIGN operation (Cre-DPSIGN) is adopted in our peer-to-peer network, which can realize the trade-off between privacy and accuracy. Furthermore, leveraging the transparency and tamper-proofing of blockchain, we design practical and reliable smart contracts for unbiased sampling based on the credit of workers to improve system robustness against Byzantine workers. In addition, we have demonstrated that the use of biased differential privacy mechanisms can lead to performance degradation. Therefore, we have introduced two unbiased differential privacy mechanisms and have proven their convergence and privacy guarantee. Extensive experiments conducted on MNIST datasets show that our algorithm can achieve byzantine fault tolerance rate with a private loss ϵ = 0.4. Compared with the state-of-the-art, aka DP-RSA (IJCAI-22), Cre-DPSIGN shows lower privacy loss consumption and better system robustness.
Mengyao Du, Miao Zhang 0037, Lin Liu 0018, Kai Xu 0014, Quanjun Yin
ICPP5
2023 Dynamic Multi-modal Prompting for Efficient Visual Grounding
Wansen Wu, Ting Liu 0018, Youkai Wang, Kai Xu 0014, Quanjun Yin, Yue Hu 0016
PRCV (7)5
2022 FedHiSyn: A Hierarchical Synchronous Federated Learning Framework for Resource and Data Heterogeneity
abstract
Federated Learning (FL) enables training a global model without sharing the decentralized raw data stored on multiple devices to protect data privacy. Due to the diverse capacity of the devices, FL frameworks struggle to tackle the problems of straggler effects and outdated models. In addition, the data heterogeneity incurs severe accuracy degradation of the global model in the FL training process. To address aforementioned issues, we propose a hierarchical synchronous FL framework, i.e., FedHiSyn. FedHiSyn first clusters all available devices into a small number of categories based on their computing capacity. After a certain interval of local training, the models trained in different categories are simultaneously uploaded to a central server. Within a single category, the devices communicate the local updated model weights to each other based on a ring topology. As the efficiency of training in the ring topology prefers devices with homogeneous resources, the classification based on the computing capacity mitigates the impact of straggler effects. Besides, the combination of the synchronous update of multiple categories and the device communication within a single category help address the data heterogeneity issue while achieving high accuracy. We evaluate the proposed framework based on MNIST, EMNIST, CIFAR10 and CIFAR100 datasets and diverse heterogeneous settings of devices. Experimental results show that FedHiSyn outperforms six baseline methods, e.g., FedAvg, SCAFFOLD, and FedAT, in terms of training accuracy and efficiency.
Yue Hu 0016, Miao Zhang 0037, Ji Liu 0003, Quanjun Yin, Yong Peng 0006, Dejing Dou
ICPP5
2022 FedGosp: A Novel Framework of Gossip Federated Learning for Data Heterogeneity
abstract
Federated learning (FL) provides the possibility to solve the problem of data privacy, but it suffers much from the data heterogeneity among different participants. Currently, some promising FL algorithms improve the effectiveness of learning under the non independent-and-identically-distributed (Non-IID) data settings. However, they require a large number of communication rounds between the server and clients for an acceptable accuracy. Inspired by the training paradigm of gossip learning, this paper proposes a new FL framework, named FedGosp. It first classifies the clients into different categories based on the model weights trained by the locally stored data. Then FedGosp utilizes the communication not only between clients and the server, but also between different classes of clients themselves. This training process enables instilling knowledge about various data distributions in the passed models. We evaluate the performance of FedGosp in multiple Non-IID settings on CIFAR10 and MNIST datasets, and compare it with the recently popular algorithms such as SCAFFOLD, FedAvg and FedProx. The experimental results show that FedGosp can improve the model accuracy by 6.53% and save 5.6 × communication costs at best compared to the second-ranked baseline.
Yue Hu 0016, Miao Zhang 0037, Li Li 0064, Tao Chang, Quanjun Yin
SMC6
2022 Who are the 'silent spreaders'?: contact tracing in spatio-temporal memory models
Yue Hu 0016, Budhitama Subagdja, Ah-Hwee Tan, Hiok Chai Quek, Quanjun Yin
Neural Comput. Appl.5
2022 Vision-Based Topological Mapping and Navigation With Self-Organizing Neural Networks
abstract
Spatial mapping and navigation are critical cognitive functions of autonomous agents, enabling one to learn an internal representation of an environment and move through space with real-time sensory inputs, such as visual observations. Existing models for vision-based mapping and navigation, however, suffer from memory requirements that increase linearly with exploration duration and indirect path following behaviors. This article presents e -TM, a self-organizing neural network-based framework for incremental topological mapping and navigation. e -TM models the exploration trajectories explicitly as episodic memory, wherein salient landmarks are sequentially extracted as "events" from streaming observations. A memory consolidation procedure then performs a playback mechanism and transfers the embedded knowledge of the environmental layout into spatial memory, encoding topological relations between landmarks. Fusion adaptive resonance theory (ART) networks, as the building block of the two memory modules, can generalize multiple input patterns into memory templates and, therefore, provide a compact spatial representation and support the discovery of novel shortcuts through inferences. For navigation, e -TM applies a transfer learning paradigm to integrate human demonstrations into a pretrained locomotion network for smoother movements. Experimental results based on VizDoom, a simulated 3-D environment, have shown that, compared to semiparametric topological memory (SPTM), a state-of-the-art model, e -TM reduces the time costs of navigation significantly while learning much sparser topological graphs.
Yue Hu 0016, Budhitama Subagdja, Ah-Hwee Tan, Quanjun Yin
IEEE Trans. Neural Networks Learn. Syst.4
2022 Efficient Flow-Based Scheduling for Geo-Distributed Simulation Tasks in Collaborative Edge and Cloud Environments
abstract
Edge computing is a good complement to cloud computing for deploying large-scale geo-distributed simulation applications, which are very sensitive to the communication delay among different simulation components (also called tasks in this paper) and users. We mainly focus on the efficient scheduling of simulation components in collaborative edge and cloud environments. As components should be deployed jointly with the consideration of capacity constraints of hosts, it is actually an NP-complete multi-dimensional bin packing problem. Meanwhile, dynamic changes of component and host states require the low deployment latency of scheduling algorithms. Unfortunately, most of the existing schedulers for modern clusters are queue-based, in which tasks are scheduled sequentially, thus lacking the ability to process tightly coupled tasks jointly. Other batching-based placement algorithms are usually time-consuming. This paper describes Pond, a novel flow-based scheduler with the awareness of interactions among tasks and users as well as heterogeneous multi-dimensional resources. First, characteristics of distributed simulation tasks are analysed and the scheduling problem is formulated as a min-cost max-flow (MCMF) problem over the flow network by mapping the communication overhead among tasks and users to the costs of arcs in the network. Considering the inherent defects of existing flow-based schedulers in dealing with multi-dimensional resources, a new method based on dominant resource is proposed and some problem specific heuristics are also designed. Extensive simulation experiments based on Alibaba production trace and some random synthetic parameters are conducted. Results show that Pond can reduce the average communication cost for each task significantly in a quite low deployment latency compared with some baselines.
Miao Zhang 0037, Yong Peng 0006, Jiancheng Zhu, Quanjun Yin
IEEE Trans. Parallel Distributed Syst.4
2021 Generation and Extraction Combined Dialogue State Tracking with Hierarchical Ontology Integration
abstract
Recently, the focus of dialogue state tracking has expanded from single domain to multiple domains.The task is characterized by the shared slots between domains.As the scenario gets more complex, the out-of-vocabulary problem also becomes more severe.Current models are not satisfactory for addressing the challenges of ontology integration between domains and out-of-vocabulary problems.To address the problem, we explore the hierarchical semantics of the ontology and enhance the interrelation between slots with masked hierarchical attention.In state value decoding stage, we address the out-of-vocabulary problem by combining generation method and extraction method together.We evaluate the performance of our model on two representative datasets, MultiWOZ in English and CrossWOZ in Chinese.The results show that our model yields a significant performance gain over current state-of-the-art state tracking model and it is more robust to out-of-vocabulary problem compared with other methods.
Xinmeng Li, Wansen Wu, Quanjun Yin
EMNLP (1)4
2021 Interpretable Goal Recognition for Path Planning with ART Networks
abstract
Goal recognition for path planning is an important task of intention identification and situation awareness, requiring an observer to predict the goal of an evader given observations of its movements. While existing models based on planning or Markov Decision Process (MDP) show superior performance over traditional library based methods, they require much effort in model design and can hardly provide legible decision rules for their users. To make the system more user-friendly while preserving accuracy of goal inference, this paper proposes a novel self-organizing neural network based inference model, which learns compact rule sets through generalizing the streaming observations of an evader. More critically, the system manifests a high level of interpretability with the linguistic if-then rule base, making it easily comprehensible for human decision makers. We conducted extensive experiments on a large-scale real-world road network. Results show that the proposed model produces accuracy comparable to those of two state-of-the-art methods while uniquely providing legible inference rules and strong robustness against multiple goals with missing data.
Yue Hu 0016, Kai Xu 0014, Budhitama Subagdja, Ah-Hwee Tan, Quanjun Yin
IJCNN5
2021 A discrete PSO-based static load balancing algorithm for distributed simulations in a cloud environment
Miao Zhang 0037, Yong Peng 0006, Quanjun Yin, Xu Xie 0005
Future Gener. Comput. Syst.4
2021 Regarding Goal Bounding and Jump Point Search
abstract
Jump Point Search (JPS) is a well known symmetry-breaking algorithm that can substantially improve performance for grid-based optimal pathfinding. When the input grid is static further speedups can be obtained by combining JPS with goal bounding techniques such as Geometric Containers (instantiated as Bounding Boxes) and Compressed Path Databases. Two such methods, JPS+BB and Two-Oracle Path PlannING (Topping), are currently among the fastest known approaches for computing shortest paths on grids. The principal drawback for these algorithms is the overhead costs: each one requires an all-pairs precomputation step, the running time and subsequent storage costs of which can be prohibitive. In this work we consider an alternative approach where we precompute and store goal bounding data only for grid cells which are also jump points. Since the number of jump points is usually much smaller than the total number of grid cells, we can save up to orders of magnitude in preprocessing time and space. Considerable precomputation savings do not necessarily mean performance degradation. For a second contribution we show how canonical orderings, partial expansion strategies and enhanced intermediate pruning can be leveraged to improve online query performance despite a reduction in preprocessed data. The combination of faster preprocessing and stronger online reasoning leads to three new and highly performant algorithms: JPS+BB+ and Two-Oracle Pathfinding Search (TOPS) based on search, and Topping+ based on path extraction. We give a theoretical analysis showing that each method is complete and optimal. We also report convincing gains in a comprehensive empirical evaluation that includes almost all current and cutting-edge algorithms for grid-based pathfinding.
Yue Hu 0016, Daniel Harabor, Long Qin 0004, Quanjun Yin
J. Artif. Intell. Res.4
2019 Online probabilistic goal recognition and its application in dynamic shortest-path local network interdiction
Kai Xu 0014, Yunxiu Zeng, Qi Zhang 0017, Quanjun Yin, Lin Sun 0008, Kaiming Xiao
Eng. Appl. Artif. Intell.4
2017 Bridging the Gap between Observation and Decision Making: Goal Recognition and Flexible Resource Allocation in Dynamic Network Interdiction
abstract
Goal recognition, which is the task of inferring an agent’s goals given some or all of the agent’s observed actions, is one of the important approaches in bridging the gap between the observation and decision making within an observe-orient-decide-act cycle. Unfortunately, few researches focus on how to improve the utilization of knowledge produced by a goal recognition system. In this work, we propose a Markov Decision Process-based goal recognition approach tailored to a dynamic shortest-path local network interdiction (DSPLNI) problem. We first introduce a novel DSPLNI model and its solvable dual form so as to incorporate real-time knowledge acquired from goal recognition system. Then a Markov Decision Process-based goal recognition model along with its dynamic Bayesian network representation and the applied goal inference method is proposed to identify the evader’s real goal within the DSPLNI context. Based on that, we further propose an efficient scalable technique in maintaining action utility map used in fast goal inference, and develop a flexible resource assignment mechanism in DSPLNI using knowledge from goal recognition system. Experimental results show the effectiveness and accuracy of our methods both in goal recognition and dynamic network interdiction.
Kai Xu 0014, Kaiming Xiao, Quanjun Yin, Yabing Zha, Cheng Zhu 0002
IJCAI3
2017 Learning real-time search on c-space GVDs
Quanjun Yin, Long Qin 0004, Yong Peng 0006
Frontiers Comput. Sci.1
2015 Aggregating Opinions to Optimize Multi-objective Urban Tactical Position Selection
abstract
In this research-in-progress paper we present a new real-world domain for studying the aggregation of different opinions: optimal urban tactical position selection (TPS). This is an important and foreseeable real world application, not only because cities have been viewed as centers of gravity by military planners throughout history, but also because the military significance of cities has increased proportionally as the global urbanization does. We first present a mapping between the domain of engineering research and that of the agent models present in the literature and use genetic multi-objective optimization method to generate Pareto TPS plans. Further we study the importance of forming diverse teams when aggregating opinions of different problem solvers for tactical position selection, and also the relationships of the number of problem solvers with time ratio and the solution efficiency. We show that a diverse team of problem solvers is able to provide better force deployment plans for early-stage decision makers to choose from. We also find that opinion aggregation methods, like approval voting, help to allocate a difficult problem solving among several computing resources, and at the same time ensuring the efficiency of solutions. Finally, we present next steps for a deeper exploration of our questions.
Kai Xu 0014, Lin Sun 0008, Quanjun Yin
DS-RT3
2015 Scheduling parallel jobs with tentative runs and consolidation in the cloud
Xiaocheng Liu, Yabing Zha, Quanjun Yin, Yong Peng 0006, Long Qin 0004
J. Syst. Softw.3
2014 A Path Planning Algorithm Based on Parallel Particle Swarm Optimization
Weitao Dang, Kai Xu 0014, Quanjun Yin
ICIC (1)3
2014 Multi-agent intention recognition using logical hidden semi-Markov models
Shi-guang Yue, Yabing Zha, Quanjun Yin
SIMULTECH3
2013 Dynamic Obstacle-Avoiding Path Planning for Robots Based on Modified Potential Field Method
Qi Zhang 0017, Shi-guang Yue, Quanjun Yin, Yabing Zha
ICIC (2)3