VLDB 2026 Research / reviewers in the wild / expert
Bo Ding 0001
dblp:37/516-1
· DBLP profile ↗
76ranked-venue papers
2as first author
52since 2021 · last 2026
0000-0002-1236-8318ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 since 2021Systems, architecture and hardware · 14 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 since 2021Computer networks · 5 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EvoNarrator: Modeling Scientific Evolution for Feasible Hypothesis GenerationabstractXiaoying Le, Pengfei Qian, Yuanzhao Zhai, Xu Zhang, Qian Liu, Feng Dawei, Bo Ding. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiaoying Le, Pengfei Qian, Yuanzhao Zhai, Bo Ding 0001 |
ACL (1) | 7 |
| 2026 | Divergence or Convergence? A Deep Insight into the Crowd Collaboration and its Productivity in Open Source Software based on EntropyabstractThe Fork and Pull-Request model is widely used in collaborative development of open source software (OSS), fostering innovation through independent repository copies, but it can also lead to inefficiencies and fragmentation. A key underexplored aspect is the integration effectiveness—the degree to which distributed original commits across forks are effectively integrated back into the main repository. It plays a critical role in OSS project productivity but remains poorly understood. In response, we introduce convergence entropy, a novel metric that quantifies the integration effectiveness by measuring the similarity between distributions of original and merged commits across forks, adjusted for integration ratio. This metric highlights not only the volume of contributions but also their diversity and coordination, offering a unique lens to understand forking practices. Moreover, we explore the relationship between convergence entropy and three dimensions of OSS project productivity, showing significant correlations. We also observe that other factors can alter this dynamic. Tao Wang 0158, Xunhui Zhang, Yang Zhang 0026, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001 |
CHI | 6 |
| 2026 | Diagnosing LLM Benchmark: A Psychometric Analysis of Difficulty and DiscriminationabstractBenchmarks have established themselves as the standard for evaluating and tracking the progress of Large Language Models. However, the community increasingly faces inconsistent model rankings across nominally similar tasks, suggesting structural misalignments in evaluation instruments. Current paradigms, which rely heavily on aggregate scalar scores, often treat benchmarks as black boxes and fail to diagnose the root causes of these discrepancies. In this work, we propose a structural diagnostic framework grounded in Item Response Theory to evaluate the benchmarks themselves. By mapping items into a latent skill space, we assess benchmark quality along two critical dimensions: calibration, which measures the alignment of difficulty with model capabilities, and efficiency, which evaluates the discriminative power of items. Our analysis of two math benchmarks yields two findings. First, distinct difficulty topologies quantitatively explain why models exhibit conflicting rankings. Second, an observed negative exponential saturation pattern shows that larger item counts can yield diminishing reliability gains. These findings show that quantity does not equal quality, providing a principled roadmap for constructing leaner, more rigorous benchmarks by resolving difficulty mismatches and pruning redundant dimensions. Jiacheng Qin, Bo Ding 0001, Yuanzhao Zhai |
ICPR (6) | 4 |
| 2026 | Model compression-driven instruction data mining framework with integrated adversarial attack strategies
Qiang Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
Expert Syst. Appl. | 6 |
| 2026 | Uncertainty-penalized reinforcement learning from human feedback with diversified reward LoRA ensembles
Yuanzhao Zhai, Han Zhang 0025, Yue Yu 0001, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
Inf. Process. Manag. | 7 |
| 2026 | Pay more attention to the robustness of LLMs on adversarial prompt for instruction data mining
Qiang Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
Neural Networks | 6 |
| 2026 | Rethinking Obscured Sub-Optimality in Analytic Learning for Exemplar-Free Class-Incremental LearningabstractExemplar-free Class-Incremental Learning (EFCIL) poses a significant challenge in mitigating catastrophic forgetting, due to the absence of exemplars. Recently, analytic learning-based methods propose a recursive alignment procedure to execute EFCIL in a phase-invariant manner and show state-of-the-art performance. However, they heavily rely on a frozen feature extractor trained with the initial dataset to avoid the misalignment between feature and label spaces, ignoring the importance of acquiring generalizable features across incremental tasks for performance improvement. To tackle this, we rethink the obscured sub-optimality of analytic learning-based methods, particularly through empirical reevaluation, and then introduce the Multi-head analytic learning (Muheal) approach. Muheal forms the multi-head model with a delicate feature extractor, thereby introducing a feature optimization procedure and a forgetting compensation module to balance the learning and forgetting. Specifically, within the feature optimization procedure, the feature extractor seeks to learn more generalizable features in a self-supervised manner using the fully-connected classification head. An analytic learning-based classification head follows to align the feature-label space. Additionally, we employ the compensation module to generate and align pseudo-features with a replicated analytic head, thus preventing overfitting and testing. Comprehensive experiments on several benchmark datasets have demonstrated that Muheal significantly outperforms existing state-of-the-art EFCIL methods and is comparable, if not superior, to methods that use replay techniques. Zijian Gao, Kele Xu, Xingxing Zhang 0001, Huiping Zhuang, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Dynamic Confidence Variance for Generalized Coreset in Active LearningabstractActive Learning (AL) aims to reduce data annotation costs by selecting the most informative samples from an unlabeled data pool. Traditional AL methods often rely on a single snapshot to identify uncertain or representative samples, often overlooking the poor generalization of a single model. Recent AL studies have attempted to address this issue by tracking a broader range of training dynamics for data selection, typically using averaging or accumulating manner. However, both our theoretical and experimental analyses reveal that these methods obscure the variability inherent in the training process, potentially prioritizing hard-to-learn samples that result in poor generalization. In this paper, we propose a novel AL method termed as Dynamic Confidence Variance (DCoV), that seamlessly integrates variability with the training dynamic to effectively identify a well-generalized Coreset. DCoV leverages the variance of the model’s prediction confidence throughout the training process for active sampling and model training. Our theoretical analysis demonstrates that DCoV provides a lower bound on the population risk of the model learned from selected labeled subset, spanning the entire training process. Extensive experiments demonstrate that our approach significantly outperforms existing state-of-the-art AL methods on various balanced and imbalanced benchmark datasets across various modalities. Tianjiao Wan, Zijian Gao, Xudong Gong, Bo Ding 0001, Huaimin Wang 0001, Kele Xu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Let the Blocks Fly (Flying Blocks): A Highly Efficient and Practical Consensus Protocol for Authoritative BlockchainsabstractIn recent years, blockchain has been increasingly applied to authoritative institutions (i.e., authoritative blockchains) to strengthen their authority and reputation by providing reliable and secure data to increase transparency, reducing fraud, and enhancing efficiency for distributed applications (DApps) like electronics certificate, land registration, and e-voting, etc. Blockchain systems in these scenarios often have the features of small node-size, high node-reputation and high node-performance, e.g., government blockchains. However, as one of the core technologies of blockchain, distributed consensus protocols are often designed for large-scale business DApps; their efficiency can be further improved when applied to authoritative blockchains. In this paper, taking into account the essential features of blockchain applications for institutions like government departments, we propose a consensus protocol known as Flying Blocks (FB) to further enhances the efficiency and practicality. FB combines the advantages of the delayed confirmation from Nakamoto consensus with the traditional voting-based BFT consensus to simplify the consensus process and reduce the network resources consumption. To evaluate the protocol's performance, we conduct comparison experiments with Hotstuff and RAFT. The results demonstrate that FB outperforms Hotstuff. Furthermore, to the best of our knowledge, FB is the first BFT protocol that outperforms RAFT, a CFT consensus protocol, deployed in Blockchain systems in terms of transaction processing efficiency. Xiang Fu 0002, Liaoliao Feng, Huaimin Wang 0001, Bo Ding 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | A Non-Intrusive Multi-Objective Task Scheduling Method for JointCloud EnvironmentabstractThe advent of advanced technologies, such as large models, has precipitated a surging demand for computational resources, thereby driving the evolution of cloud computing services from single-cloud architectures to multi-cloud paradigms. The accompanying question is how to enable resources from different cloud service providers to collaborate efficiently. Although existing works have explored multi-cloud scheduling methods, most of these works are centralized scheduling, where the decision-maker can schedule computing resources across all clouds. However, this is nearly impossible given the current situation in which computing resources in different clouds come from various cloud service providers. In response to this challenge, JointCloud, a multi-cloud cooperation architecture, has been proposed, which aims at enhancing the cooperation among multiple clouds to provide efficient multi-cluster services. Following the idea of JointCloud, proposes a multi-objective evolutionary algorithm (MOEA) based method for task scheduling in multi-cluster cloud computing environments, without intervening in the intra-cluster scheduling scheme. In the proposed method, we construct a mathematical model with the optimization objectives of minimizing overall waiting time and load imbalance between clusters based on the actual operation data in China Computing NET (C$^{2}$NET). In addition, we also develop an MOEA specifically tailored to address this problem. The performance of the proposed MOEA and existing state-of-the-art MOEAs is examined on the proposed problems. Comparison results highlight the promising performance of the proposed MOEA, the specifically tailored algorithm in effectively addressing the multi-cluster task scheduling problem. In addition, we also compared the results of the MOEAs with the results of three classical scheduling methods, the results proved the effectiveness of the MOEA-based method on this problem. Lianghao Li, Haibo Mi, Bo Ding 0001, Huaimin Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2026 | DockerFill: Automatically Completing Dockerfile Code With Syntax-Aware Multi-Task LearningabstractAs a kind of infrastructure-as-code, Dockerfile specifies the structure and functionality of a built Docker image and thus plays an important role in the containerized software development process. Nowadays developers need to spend extra time and effort configuring their Dockerfiles in addition to their regular coding work, which requires knowledge and skills orthogonal to those entailed in other software-related experiences. Poorly written Dockerfile code often introduces errors and maintenance costs. However, little automated support is available for assisting developers in configuring Dockerfiles. In this study, we first conduct an online survey to investigate Docker developers’ perceptions of Dockerfile writing, highlighting the needs and potential benefits of Dockerfile auto-completion techniques. Then, we introduceDOCKERFILL, a pre-trained model based approach that provides completion suggestions for Dockerfile-specific code.DOCKERFILLleverages multi-layer Transformer architecture with syntax-aware multi-task learning, which includes contextual file information and three pre-training tasks, i.e., masked language modeling, syntax type identification, and masked identifier prediction. To evaluateDOCKERFILL’s effectiveness, we collect a dataset of 6,350 high-quality real-world Dockerfiles. Our empirical results show that DOCKERFILL provides up to 52.38% accuracy for token-level completion and 19.69% exact match for line-level completion, outperforming the baselines by 7.32%-37.67% and 1.97%-19.69%, respectively. Also,DOCKERFILLobtains significantly higher human evaluation scores compared to the baselines. Yiwen Wu 0001, Yang Zhang 0026, Tao Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2025 | Enhancing Decision-Making for LLM Agents via Step-Level Q-Value ModelsabstractAgents significantly enhance the capabilities of standalone Large Language Models (LLMs) by perceiving environments, making decisions, and executing actions. However, LLM agents still face challenges in tasks that require multiple decision-making steps. Estimating the value of actions in specific tasks is difficult when intermediate actions are neither appropriately rewarded nor penalized. In this paper, we propose leveraging a task-relevant Q-value model to guide action selection. Specifically, we first collect decision-making trajectories annotated with step-level Q values via Monte Carlo Tree Search (MCTS) and construct preference data. We then use another LLM to fit these preferences through step-level Direct Policy Optimization (DPO), which serves as the Q-value model. During inference, at each decision-making step, LLM agents select the action with the highest Q value before interacting with the environment. We apply our method to various open-source and API-based LLM agents, demonstrating that Q-value models significantly improve their performance. Notably, the performance of the agent built with Phi-3-mini-4k-instruct improved by 103% on WebShop and 75% on HotPotQA when enhanced with Q-value models, even surpassing GPT-4o-mini. Additionally, Q-value models offer several advantages, such as generalization to different LLM agents and seamless integration with existing prompting strategies. Yuanzhao Zhai, Tingkai Yang, Kele Xu, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001 |
AAAI | 6 |
| 2025 | ALMP: Automatic Layer-By-Layer Mixed-Precision Quantization for Large Language Models
Huanxi Liu, Bo Ding 0001 |
ICIC (23) | 6 |
| 2025 | VVC-Gym: A Fixed-Wing UAV Reinforcement Learning Environment for Multi-Goal Long-Horizon ProblemsabstractMulti-goal long-horizon problems are prevalent in real-world applications. The additional goal space introduced by multi-goal problems intensifies the spatial complexity of exploration; meanwhile, the long interaction sequences in long-horizon problems exacerbate the temporal complexity of exploration. Addressing the great exploration challenge posed by multi-goal long-horizon problems depends not only on the design of algorithms but also on the design of environments and the availability of demonstrations to assist in training. To facilitate the above research, we propose a multi-goal long-horizon Reinforcement Learning (RL) environment based on realistic fixed-wing UAV's velocity vector control, named VVC-Gym, and generate multiple demonstration sets of various quality. Through experimentation, we analyze the impact of different environment designs on training, assess the quantity and quality of demonstrations and their influence on training, and assess the effectiveness of various RL algorithms, providing baselines on VVC-Gym and its corresponding demonstrations. The results suggest that VVC-Gym is suitable for studying: (1) the influence of environment designs on addressing multi-goal long-horizon problems with RL. (2) the assistance that demonstrations can provide in overcoming the exploration challenges of multi-goal long-horizon problems. (3) the RL algorithm designs with the least possible impact from environment designs on the efficiency and effectiveness of training. Xudong Gong, Kele Xu, Zhangjun Sun, Xing Zhou 0004, Bo Ding 0001, Huaimin Wang 0001 |
ICLR | 8 |
| 2025 | Improving the Continuity of Goal-Achievement Ability via Policy Self-Regularization for Goal-Conditioned Reinforcement LearningabstractThis paper addresses the challenge of discontinuity in goal-achievement capabilities observed in Goal-conditioned Reinforcement Learning (GCRL) algorithms. Through a theoretical analysis, we identify that the reuse of successful trajectories or policies during training can aid in achieving adjacent goals of achievable goals. However, the policy discrepancy between achievable and adjacent goals must be carefully managed to avoid both overly trivial and excessively large differences, which can respectively hinder policy performance. To tackle this issue, we propose a margin-based policy self-regularization approach that optimizes the policy discrepancies between adjacent desired goals to a minimal acceptable threshold. This method can be integrated into popular GCRL algorithms, such as GC-SAC, HER, and GC-PPO. Systematic evaluations across two robotic arm control tasks and a complex fixed-wing aircraft control task demonstrate that our approach significantly improves the continuity of goal-achievement abilities of GCRL algorithms, thereby enhancing their overall performance. Xudong Gong, Sen Yang 0003, Kele Xu, Bo Ding 0001, Huaimin Wang 0001, Yong Dou |
ICML | 5 |
| 2025 | GRACE: Graph-Adapted Case-Augmented Execution for Tool Use in Large Language ModelsabstractLarge Language Models (LLMs) have shown remarkable capability. However, their static parametric memory inherently impedes alignment with rapidly evolving real-world tools and APIs. Existing tool-augmented LLMs, although effective, frequently fail to identify accurate and executable tool combinations within an ever-expanding and dynamically changing toolbox, leading to cascading failures and hallucinations. These limitations are further exacerbated by semantic misalignment between natural-language queries and symbolic tool specifications. In this paper, we present Graph-Adapted Case-Augmented Execution (GRACE), a lightweight yet effective framework that equips LLMs with precise, scalable, and continually adaptive tooluse capabilities. GRACE introduces a demand-driven retrieval mechanism in which the LLM first composes a structured demand vector that explicitly aligns natural-language intent with the symbolic vocabulary of tool APIs, thereby narrowing the semantic gap and enabling markedly more precise retrieval. The demand vector is scored against every node of the tool graph; the LLM then selects the highest-ranked tool together with its dependency-closed subgraph, ensuring that all prerequisite constraints are satisfied and every retrieved tool can be executed immediately. Additionally, GRACE maintains a case library that stores execution traces and enables long-term memory without model retraining. Whenever tools are added, deprecated, or modified, both the graph and case library are updated, preserving long-term memory alignment with the evolving external ecosystem. Extensive experiments on ToolSandbox and ToolQA-D demonstrate that GRACE consistently surpasses both proprietary prompting baselines and fine-tuned models, achieving 76.1 % exact-match accuracy on ToolSandbox and 70.7 % under the dynamic tool environment on ToolQA-D, thereby establishing new state-of-the-art results while maintaining parameter efficiency and plug-and-play deployment. Yuanzhao Zhai, Bo Ding 0001 |
ICPADS | 4 |
| 2025 | V-Pilot: A Velocity Vector Control Agent for Fixed-Wing UAVs from Imperfect DemonstrationsabstractThis paper addresses the challenge of Velocity Vector Control (VVC) for fixed-wing UAVs using Reinforcement Learning (RL) in the presence of imperfect demonstrations. The multi-objective and long-horizon nature of VVC introduces significant spatial and temporal complexities, complicating RL's exploration. While demonstration-based RL methods can help mitigate exploration challenges, their effectiveness is often limited by the quality of the provided demonstrations. To tackle this, we propose V-Pilot, a novel approach that integrates: (1) a controller equipped with a control law model to reduce action oscillation, thus alleviating temporal exploration issues, and (2) a VVC-specific training workflow for iterative policy refinement and demonstration quality improvement. This framework is designed to enhance the performance of demonstration-based RL under imperfect demonstrations. We evaluate V-Pilot on the fixed-wing UAV RL environment, VVCGym. Experimental results demonstrate that V-Pilot outperforms PID and Behavioral Cloning across multiple performance metrics. Xudong Gong, Kele Xu, Xing Zhou 0004, Bo Ding 0001, Huaimin Wang 0001 |
ICRA | 6 |
| 2025 | A Review of Multi-Objective Optimization for Cloud Environment Storage OptimizationabstractThe advent of cloud computing offers a novel mode of managing computing resources, enabling users to flexibly utilize required computing and storage resources through the network. Simultaneously, it allows large-scale computing centers and data centers to more effectively utilize their computing resources. With the rapid development of cloud computing technology, the exponential growth of massive data storage needs from numerous users has brought challenges in storage cost, performance, reliability, and security. Conventional single-objective optimization approaches, which concentrate exclusively on a singular performance metric, are increasingly inadequate to address the intricate requirements of cloud storage systems. In contrast, multi-objective optimization methods can simultaneously optimize multiple aspects of concern to users or operators, providing more comprehensive solutions for cloud computing services. In this paper, we first introduce the basic concepts of multi-objective optimization problems. Subsequently, we present a comprehensive review of multi-objective optimization applications in cloud storage systems, including the setting of optimization objectives, algorithm selection, and comparison methods of experimental results. Finally, this paper summarizes and discusses the current implementation of multi-objective evolutionary algorithms in cloud storage optimization, and provides an outlook on future research directions. Lianghao Li, Haibo Mi, Kun Wang 0059, Bo Ding 0001, Huaimin Wang 0001 |
JCC | 4 |
| 2025 | Max-Informative Unlabeled Sample Replay for Semi-Supervised Class-Incremental Learning in Audio ClassificationabstractAudio classification constitutes a critical task aimed at assigning meaningful labels to audio recordings. Despite commendable efforts in this field, prior endeavors have encountered notable challenges. Firstly, the majority of existing solutions operate under the assumption of a fixed vocabulary for classification, neglecting the crucial need for systems to adapt continuously to dynamic data streams. Secondly, these approaches often assume that input data is fully annotated, a presumption that frequently diverges from practical scenarios. In response to these challenges, we present a novel semi-supervised class-incremental learning framework that utilizes memory replay with unlabeled data and introduces temporal consistency regularization to better alleviate catastrophic forgetting. Additional, we employ Iterative Projection and Matching to select the most informative samples for storage, addressing the issues of reservoir sampling while maintaining a balanced buffer. We showcase the framework's efficacy through a series of comprehensive experiments conducted on both the ESC-50 and Google Speech Commands datasets. The results demonstrate the remarkable performance of our proposed approach, particularly in few-shot scenarios. Qiang Wang 0006, Bo Ding 0001, Huaimin Wang 0001 |
JCC | 5 |
| 2025 | Preference-Strength-Aware Self-Improving Alignment with Generative Preference ModelsabstractSelf-improving alignment leveraging large language models (LLMs) to automatically generate synthetic preference data has garnered significant attention as a means of reducing reliance on human labelers. These methods typically employ the LLM-as-a-judge mechanism, where the LLM generates responses and then employs itself to judge which response best aligns with the given prompt for curating the binary self-preferred dataset. However, these methods encounter two major challenges: (1) LLM-as-a-judge often produces error-prone evaluations, resulting in low-quality preference annotation, and (2) their optimization strategies often overlook the strength of preferences within binary pairs, leading to overfitting. This paper proposes a novel method, Preference-Strength-aware Optimization (PSO), to address these issues. Specifically, PSO frames the preference annotation process as a judgment token prediction task using the generative preference model to produce reliable judgments. The predicted judgment token indicates the preferred response and its corresponding probability reflects the disparity between responses, referred to as preference strength. Based on this strength, we introduce a new preference-strength-aware loss to adaptively reweight the impact of different response pairs on optimization, concentrating the model's learning on high-quality response pairs. Our experiments demonstrate that PSO significantly improves performance in preference benchmarks, achieving stronger alignment with human preferences, reducing verbose responses, and mitigating overfitting. Furthermore, PSO exhibits robust generalization and sample efficiency, offering a scalable and promising solution for LLM alignment without relying on human-annotated preferences. Yuanzhao Zhai, Zhuo Zhang 0007, Cheng Yang 0004, Kele Xu, Yue Yu 0001, Wei Li 0022, Hui Wang 0030, Zenglin Xu, Bo Ding 0001, Huaimin Wang 0001 |
SIGIR | 10 |
| 2025 | Empowering Large Language Model Agent through Step-Level Self-Critique and Self-TrainingabstractLarge Language Model (LLM) agents frequently produce sub-optimal actions when tackling complex, multi-step decision-making tasks. Employing self-critique to identify flaws and suggest enhancements is an effective strategy for refining actions. Although trajectory-level critique is commonly employed, it often fails to identify flawed steps accurately. In this paper, we introduce SLSC-MCTS, a method that integrates Monte Carlo Tree Search with Step-Level Self-Critique to enhance LLM agents during both testing and self-training phases. During decision tree expansion with SLSC-MCTS, the LLM agent initially generates an action, receives environmental feedback, and subsequently generates further actions via self-critique and refinement. Through multiple episodes of SLSC-MCTS, LLM agents can effectively utilize step-level critiques while disregarding ineffective ones based on node values, thereby incorporating the critiques more robustly. Additionally, our method further empowers LLM agents in a self-training manner, collecting training data from the constructed decision tree to iteratively fine-tune the LLM agents. The self-training data gathered via SLSC-MCTS is diverse and high-quality, which further enhances the reasoning, critiquing, and refining abilities of LLM agents. Experimental results demonstrate that SLSC-MCTS significantly improves LLM agents during testing, surpassing state-of-the-art baselines and achieving shorter task completion trajectories across information retrieval benchmarks such as WebShop and HotPotQA. After three iterations of self-training, LLM agents established by Llama-3.1-8B-Instruct show substantial improvement, even surpassing human experts in WebShop. Yuanzhao Zhai, Huanxi Liu, Zhuo Zhang 0007, Kele Xu, Cheng Yang 0004, Bo Ding 0001, Huaimin Wang 0001 |
SIGIR | 8 |
| 2025 | Memory replay with unlabeled data for semi-supervised class-incremental learning via temporal consistency
Qiang Wang 0006, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
Frontiers Comput. Sci. | 4 |
| 2025 | Dual Temporal Masked Modeling for KPI Anomaly Detection via Similarity AggregationabstractWith the expanding scale of current industries, monitoring systems centered around Key Performance Indicators (KPIs) play an increasingly crucial role. KPI anomaly detection can monitor the potential risks according to KPI data and has garnered widespread attention due to its rapid responsiveness and adaptability to dynamic changes. Considering the absence of labels and the high cost of manual annotation of KPI data, the self-supervised approaches are proposed. Among them, mask modeling methods draw great attention and can learn the intrinsic distribution of data without relying on prior assumptions. However, conventional mask modeling often overlooks the examination of relationships between unsynchronized variables, treating them with equal importance, and inducing inaccurate detection results. To address this, this paper proposes a Dual Masked modeling Approach combined with Similarity Aggregation, named DMASA. Starting from a self-supervised approach based on mask modeling, DMASA incorporates spectral residual techniques to explore inter-variable dependencies and aggregates information from similar data to eliminate interference from irrelevant variables in anomaly detection. Extensive experiments on eight datasets and state-of-the-art results demonstrate the effectiveness of our approach. Our code is available athttps://github.com/colaudiolab/GT-DMASA. Zijian Gao, Kele Xu, Xu Wang 0064, Peichang Shi, Bo Ding 0001 |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2024 | Optimistic Model Rollouts for Pessimistic Offline Policy OptimizationabstractModel-based offline reinforcement learning (RL) has made remarkable progress, offering a promising avenue for improving generalization with synthetic model rollouts. Existing works primarily focus on incorporating pessimism for policy optimization, usually via constructing a Pessimistic Markov Decision Process (P-MDP). However, the P-MDP discourages the policies from learning in out-of-distribution (OOD) regions beyond the support of offline datasets, which can under-utilize the generalization ability of dynamics models. In contrast, we propose constructing an Optimistic MDP (O-MDP). We initially observed the potential benefits of optimism brought by encouraging more OOD rollouts. Motivated by this observation, we present ORPO, a simple yet effective model-based offline RL framework. ORPO generates Optimistic model Rollouts for Pessimistic offline policy Optimization. Specifically, we train an optimistic rollout policy in the O-MDP to sample more OOD model rollouts. Then we relabel the sampled state-action pairs with penalized rewards, and optimize the output policy in the P-MDP. Theoretically, we demonstrate that the performance of policies trained with ORPO can be lower-bounded in linear MDPs. Experimental results show that our framework significantly outperforms P-MDP baselines by a margin of 30%, achieving state-of-the-art performance on the widely-used benchmark. Moreover, ORPO exhibits notable advantages in problems that require generalization. Yuanzhao Zhai, Yiying Li, Zijian Gao, Xudong Gong, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
AAAI | 7 |
| 2024 | Temporal Inconsistency-Based Active LearningabstractDeep supervised learning has demonstrated strong capabilities; however, such progress relies on massive and expensive data annotation. Active Learning (AL) has been introduced to selectively annotate samples, thus reducing the human labeling effort. Previous AL research has focused on employing recently trained models to design sampling strategies, based on uncertainty or representativeness. Drawing inspiration from the issue of model forgetting, we propose a novel AL framework called Temporal Inconsistency-Based Active Learning (TIR-AL). In this framework, multiple snapshots of the models across consecutive cycles are jointly utilized to select samples with higher temporal inconsistency, by computing the proposed self-weighted nuclear norm metric. Furthermore, we introduce a consistency regularization term to mitigate the issue of forgetting. Together, these components make full use of the potential of data and facilitate effective interaction within the AL loop. To demonstrate the efficacy of TIR-AL, we conducted a set of experiments illustrating how our approach outperforms state-of-the-art methods without incurring any additional training costs. Tianjiao Wan, Yutao Dou, Kele Xu, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001 |
ICASSP | 5 |
| 2024 | Transformer-Inspired Lightweight Model for Efficient Time Series ForecastingabstractAccuracy and efficiency are pivotal considerations in the field of time series forecasting. Through the integration of meticulously designed temporal components, the Transformer-based models have significantly enhanced the accuracy of time series prediction. However, due to the utilization of attention mechanism, these models suffer heightened complexity and limited practicality. To address the complexity challenge and develop a more practical and lightweight model, we propose the integration of the core designs of Transformer-based models into the framework of CNNs. This integration entails the assimilation of low-level features from shallower layers into the prediction process, effectively addressing the challenge of inadequate resolution within deep feature maps. Furthermore, we leverage depthwise convolution to capture finer temporal details and adopt a channel-sharing prediction strategy to reduce the parameter count of the model. Empirical results, derived from experiments conducted on seven distinct datasets, substantiate the superior performance of our stream-lined model. Notably, it outperforms competing models in terms of both accuracy and model complexity. Xu Wang 0064, Kele Xu, Bo Ding 0001 |
ICASSP | 4 |
| 2024 | Iterative Regularized Policy Optimization with Imperfect DemonstrationsabstractImitation learning heavily relies on the quality of provided demonstrations. In scenarios where demonstrations are imperfect and rare, a prevalent approach for refining policies is through online fine-tuning with reinforcement learning, in which a Kullback–Leibler (KL) regularization is often employed to stabilize the learning process. However, our investigation reveals that on the one hand, imperfect demonstrations can bias the online learning process, the KL regularization will further constrain the improvement of online policy exploration. To address the above issues, we propose Iterative Regularized Policy Optimization (IRPO), a framework that involves iterative offline imitation learning and online reinforcement exploration. Specifically, the policy learned online is used to serve as the demonstrator for successive learning iterations, with a demonstration boosting to consistently enhance the quality of demonstrations. Experimental validations conducted across widely used benchmarks and a novel fixed-wing UAV control task consistently demonstrate the effectiveness of IRPO in improving both the demonstration quality and the policy performance. Our code is available at https://github.com/GongXudong/IRPO. Xudong Gong, Kele Xu, Yuanzhao Zhai, Chengkang Yao, Bo Ding 0001, Huaimin Wang 0001 |
ICML | 7 |
| 2024 | M3ixTS: Mixing of Multi-patch and Multi-view For Time Series Forecasting
Lianghao Li, Haibo Mi, Bo Ding 0001 |
ICONIP (3) | 4 |
| 2024 | Selective Learning for Sample-Efficient Training in Multi-Agent Sparse Reward Tasks (Extended Abstract)
Xinning Chen, Xuan Liu 0001, Yanwen Ba, Shigeng Zhang, Bo Ding 0001, Kenli Li 0001 |
IJCAI | 5 |
| 2024 | FP4-Quantization: Lossless 4bit Quantization for Large Language ModelsabstractLarge language models(LLMs) have demonstrated exceptional performance across a wide range of tasks. However, their extensive computational and storage requirements hinder their widespread deployment. To address this, low-bit quantization has emerged as a highly effective approach to reducing the inference cost of LLMs. Nevertheless, the existing repertoire of 4-bit quantization techniques is plagued by a substantial decline in model precision. In this paper, we introduce a novel 4-bit weight quantization method, FP4-Quantization, which leverages a 4-bit floating-point(FP4) representation that aligns better with the weight distribution characteristics of LLMs. Furthermore, it incorporates a Low-Rank Quantization Error Correction(LREC), involving progressive fine-tuning of low-rank parameters to rectify quantization errors, thereby enabling the achievement of precision-preserving 4-bit weight-only quantization. Our Experimental results on multiple zero-shot tasks demonstrate that FP4-Quantization achieves 4-bit weight quantization with an accuracy degradation of less than 0.5%. Huanxi Liu, Bo Ding 0001 |
JCC | 5 |
| 2024 | Tracing Training Progress: Dynamic Influence Based Selection for Active LearningabstractActive learning (AL) aims to select highly informative data points from an unlabeled dataset for annotation, mitigating the need for extensive human labeling effort. However, classical AL methods heavily rely on human expertise to design the sampling strategy, inducing limited scalability and generalizability. Many efforts have sought to address this limitation by directly connecting sample selection with model performance improvement, typically through influence function. Nevertheless, these approaches often ignore the dynamic nature of model behavior during training optimization, despite empirical evidence highlights the importance of dynamic influence to track the sample contribution. This oversight can lead to suboptimal selection, hindering the generalizability of model. In this study, we explore the dynamic influence based data selection strategy by tracing the impact of unlabeled instances on model performance throughout the training process. Our theoretical analyses suggest that selecting samples with higher projected gradients along the accumulated optimization direction at each checkpoint leads to improved performance. Furthermore, to capture a wider range of training dynamics without incurring excessive computational or memory costs, we introduce an additional dynamic loss term designed to encapsulate more generalized training progress information. These insights are integrated into a universal and task-agnostic AL framework termed Dynamic Influence Scoring for Active Learning (DISAL). Comprehensive experiments across various tasks have demonstrated that DISAL significantly surpasses existing state-of-the-art AL methods, demonstrating its ability to facilitate more efficient and effective learning in different domains. Tianjiao Wan, Kele Xu, Long Lan, Zijian Gao, Bo Ding 0001, Huaimin Wang 0001 |
ACM Multimedia | 6 |
| 2024 | Goal-Conditioned On-Policy Reinforcement LearningabstractExisting Goal-Conditioned Reinforcement Learning (GCRL) algorithms are built upon Hindsight Experience Replay (HER), which densifies rewards through hindsight replay and leverages historical goal-achieving information to construct a learning curriculum. However, when the task is characterized by a non-Markovian reward (NMR), whose computation depends on multiple steps of states and actions, HER can no longer densify rewards by treating a single encountered state as the hindsight goal. The lack of informative rewards hinders policy learning, resulting in rolling out failed trajectories. Consequently, the replay buffer is overwhelmed with failed trajectories, impeding the establishment of an applicable curriculum. To circumvent these limitations, we deviate from existing HER-based methods and propose an on-policy GCRL framework, GCPO, which is applicable to both multi-goal Markovian reward (MR) and NMR problems.
GCPO consists of (1) Pre-training from Demonstrations, which pre-trains the policy to possess an initial goal-achieving capability, thereby diminishing the difficulty of subsequent online learning. (2) Online Self-Curriculum Learning, which first estimates the policy's goal-achieving capability based on historical evaluation information and then selects progressively challenging goals for learning based on its current capability. We evaluate GCPO on a challenging multi-goal long-horizon task: fixed-wing UAV velocity vector control. Experimental results demonstrate that GCPO is capable of effectively addressing both multi-goal MR and NMR problems. Xudong Gong, Kele Xu, Bo Ding 0001, Huaimin Wang 0001 |
NeurIPS | 4 |
| 2024 | Subtraction of Hyperledger Fabric: A blockchain-based lightweight storage mechanism for digital evidences
Xiang Fu 0002, Haoliang Ma, Bo Ding 0001, Huaimin Wang 0001, Peichang Shi |
J. Syst. Archit. | 3 |
| 2024 | Less confidence, less forgetting: Learning with a humbler teacher in exemplar-free Class-Incremental learning
Zijian Gao, Kele Xu, Huiping Zhuang, Li Liu 0036, Xinjun Mao, Bo Ding 0001, Huaimin Wang 0001 |
Neural Networks | 6 |
| 2024 | Automated Data Augmentation for Audio ClassificationabstractAudio classification is a challenging task that requires categorizing audio data based on its content or characteristics. Existing approaches for audio classification rely either on supervised learning or fine-tuning based on self-supervised learning, both of which require manually labeled data. However, manually labeling audio datasets is a time-consuming and expensive process that limits the dataset's size. Moreover, the diversity of sound categories and class imbalances can further impede classification performance. To overcome these challenges, researchers have proposed various audio data augmentation methods. However, most of these methods focus less on augmentations combination and design and rely solely on waveform-based or spectrogram-based approaches. This paper presents an Automated Audio Augmentation (AAA) method for audio classification, which generates learnable and composable augmentation policies suitable for the audio classification task and can be employed in a plug-and-play manner. This method leverages both waveform-level and spectrogram-level augmentation, and a Bayesian optimization algorithm is proposed to search for composed augmentation policies. To the best of our knowledge, this is the first attempt to propose an automatic data augmentation method for audio classification tasks. Through large-scale empirical studies, we demonstrate that the proposed method outperforms previous competitive methods by a significant margin. We improve the average performance of multiple datasets by 6.421% and by 7.330% on few-shot scenarios, respectively. Yanjie Sun, Kele Xu, Chaorun Liu, Yong Dou, Huaimin Wang 0001, Bo Ding 0001, Qinghua Pan |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Selective Learning for Sample-Efficient Training in Multi-Agent Sparse Reward TasksabstractLearning effective strategies in sparse reward tasks is one of the fundamental challenges in reinforcement learning. This becomes extremely difficult in multi-agent environments, as the concurrent learning of multiple agents induces the non-stationarity problem and sharply increased joint state space. Existing works have attempted to promote multi-agent cooperation through experience sharing. However, learning from a large collection of shared experiences is inefficient as there are only a few high-value states in sparse reward tasks, which may instead lead to the curse of dimensionality in large-scale multi-agent systems. This paper focuses on sparse-reward multi-agent cooperative tasks and proposes an effective experience-sharing method, Multi-Agent Selective Learning (MASL), to boost sample-efficient training by reusing valuable experiences from other agents. MASL adopts a retrogression-based selection method to identify high-value traces of agents from the team rewards, based on which some recall traces are generated and shared among agents to motivate effective exploration. Moreover, MASL selectively considers information from other agents to cope with the non-stationarity issue while enabling efficient training for large-scale agents. Experimental results show that MASL significantly improves sample efficiency compared with state-of-the-art MARL algorithms in cooperative tasks with sparse rewards. Xinning Chen, Xuan Liu 0001, Yanwen Ba, Shigeng Zhang, Bo Ding 0001, Kenli Li 0001 |
ECAI | 5 |
| 2023 | Complementary Learning System Based Intrinsic Reward in Reinforcement LearningabstractDeep reinforcement learning has achieved encouraging performance in many realms. However, one of its primary challenges is the sparsity of extrinsic rewards, which is still far from solved. Complementary learning system theory suggests that effective human learning relies on two complementary learning systems utilizing short-term and long-term memories. Inspired by the fact that humans evaluate curiosity by comparing current observations with historical information, we propose a novel intrinsic reward, namely CLS-IR, which aims to address the problems caused by sparse extrinsic rewards. Specifically, we train a self-supervised predictive model with short-term and long-term memories via exponential moving averages. We employ the information gain between the two memories as the intrinsic reward, which does not incur additional training costs but leads to better exploration. To investigate the effectiveness of CLS-IR, we conduct extensive experimental evaluations; the results demonstrate that CLS-IR can achieve state-of-the-art performance on Atari games and DeepMind Control Suite. Zijian Gao, Kele Xu, Hongda Jia, Tianjiao Wan, Bo Ding 0001, Xinjun Mao, Huaimin Wang 0001 |
ICASSP | 5 |
| 2023 | Progressive Diversifying Policy for Multi-Agent Reinforcement LearningabstractMulti-Agent Reinforcement Learning (MARL) has recently achieved promising performance in many collaborative decision making tasks. However, one of the main bottleneck challenges for MARL is the sparsity of the team reward, which can lead to the homogenization of agents’ behaviors. To address these issues, we propose a Progressive Diversifying Policy (PDP) algorithm in this paper. Specifically, we propose to actively amplify the diversity between agents’ policies during learning and exploit diversity as an additional intrinsic reward for MARL. Furthermore, we propose a progressive diversity boosting policy to find a better team policy. Leveraging the aforementioned improvements, our method can handle sparse team rewards and alleviate the homogeneous behaviors of agents. We conduct experiments on widely-used MARL environments and the results show that PDP can provide state-of-the-art performance while maintaining a competitive convergence speed. Shaoqi Sun, Yuanzhao Zhai, Kele Xu, Bo Ding 0001 |
ICASSP | 5 |
| 2023 | Diversifying Message Aggregation in Multi-Agent Communication Via Normalized Tensor Nuclear Norm RegularizationabstractThe use of graph attention networks (GAT) in communication-enhanced multi-agent reinforcement learning (Comm-MARL) has become prevalent. While successful, GAT can lead to homogeneity in the strategies of message aggregation, which can severely limit multi-agent coordination. To address this challenge, we study the adjacency tensor of the communication graph. Then we define a new nuclear tensor rank and its convex surrogate, the normalized tensor nuclear norm to measure the homogeneity of message aggregation. Leveraging the norm, we further propose a plug-and-play regularizer on the adjacency tensor, named Normalized Tensor Nuclear Norm Regularization (NTNNR), to actively enrich the diversity of message aggregation during the training stage. NTNNR is agnostic to specific Comm-MARL algorithms and can be flexibly integrated with different graph-attention methods. Empirical results demonstrate that aggregating messages using NTNNR-enhanced GAT can improve the efficiency of the training and achieve higher asymptotic performance than existing message aggregation methods. Yuanzhao Zhai, Kele Xu, Bo Ding 0001, Zijian Gao, Huaimin Wang 0001 |
ICASSP | 3 |
| 2023 | Action Prediction for Cooperative Exploration in Multi-agent Reinforcement Learning
Yanqiang Zhang, Bo Ding 0001 |
ICONIP (2) | 3 |
| 2023 | Bi-level Multi-Agent Actor-Critic Methods with ransformersabstractRecently, deep multi-agent reinforcement learning methods have witnessed great progress, including multi-agent actor-critic methods. However, it’s worth noticing there is a performance gap between multi-agent actor-critic methods and state-of-the-art value-based methods. In this paper, we investigate the causes and attribute inferior performance to issues of contribution-mismatch and indiscriminate guidance. To overcome these problems, we introduce a novel bi-level multi-agent actorcritic reinforcement learning approach with transformers, called BMT. Specifically, we propose a simple but efficient bi-level optimization mechanism to learn both global critic and agentspecific critic, thus jointly guiding the policy update. In addition, we adopt the transformer-based model as the policy network to decouple complicated relationships and generate flexible policy. BMT is also general enough to be plugged into any actor-critic multi-agent reinforcement learning approach, such as MAPPO, and equips it with strong expression. On multiple benchmarks including multi-agent particle environments and a challenging set of StarCraft II micromanagement tasks, large-scale empirical experiments demonstrate that BMT-based multi-agent reinforcement learning methods achieve superior performance over both state-of-the-art actor-critic and value-based approaches. Tianjiao Wan, Haibo Mi, Zijian Gao, Yuanzhao Zhai, Bo Ding 0001 |
JCC | 5 |
| 2022 | Exploring Policy Diversity in Parallel Actor-Critic LearningabstractExploration is a critical challenge for deep reinforcement learning methods. Although existing works such as actor-critic algorithms have made much progress, most still suffer from the sample inefficiency problem in complex environments where rewards are sparse. Parallel sampling, which uses multiple actors with the same policy interacting with the environment, is an effective approach to improve sample efficiency. However, parallel parameter-sharing actors collect similar samples, which generally hinders the improvement of the overall exploration process. In this paper, we propose a Policy Diversity enhanced approach for parallel Actor-Critic (PDAC). Specifically, we extend the parallel actor-critic architecture to the PDAC framework composed of a shared critic and parallel distinct actors. Then we introduce the KL-divergence of the action probability distribution between parallel actors as the intrinsic reward to encourage actors to explore diverse strategies. We evaluate our approach in multiple challenging procedurally-generated tasks and compare it with state-of-the-art algorithms. Experiments show that PDAC makes significant progress in the comparison, in terms of cumulative rewards and sample efficiency. Yanqiang Zhang, Yuanzhao Zhai, Gongqian Zhou, Bo Ding 0001, Songwang Liu |
ICTAI | 4 |
| 2022 | Goal Consistency: An Effective Multi-Agent Cooperative Method for Multistage TasksabstractAlthough multistage tasks involving multiple sequential goals are common in real-world applications, they are not fully studied in multi-agent reinforcement learning (MARL). To accomplish a multi-stage task, agents have to achieve cooperation on different subtasks. Exploring the collaborative patterns of different subtasks and the sequence of completing the subtasks leads to an explosion in the search space, which poses great challenges to policy learning. Existing works designed for single-stage tasks where agents learn to cooperate only once usually suffer from low sample efficiency in multi-stage tasks as agents explore aimlessly. Inspired by human’s improving cooperation through goal consistency, we propose Multi-Agent Goal Consistency (MAGIC) framework to improve sample efficiency for learning in multi-stage tasks. MAGIC adopts a goal-oriented actor-critic model to learn both local and global views of goal cognition, which helps agents understand the task at the goal level so that they can conduct targeted exploration accordingly. Moreover, to improve exploration efficiency, MAGIC employs two-level goal consistency training to drive agents to formulate a consistent goal cognition. Experimental results show that MAGIC significantly improves sample efficiency and facilitates cooperation among agents compared with state-of-art MARL algorithms in several challenging multistage tasks. Xinning Chen, Xuan Liu 0001, Shigeng Zhang, Bo Ding 0001, Kenli Li 0001 |
IJCAI | 4 |
| 2022 | Uncertainty Estimation based Intrinsic Reward For Efficient Reinforcement LearningabstractFor reinforcement learning, the extrinsic reward is a core factor for the learning process which however can be very sparse or completely missing. In response, researchers have proposed the idea of intrinsic reward, such as encouraging the agent to visit novel states through prediction error. However, the deep prediction model can provide over-confident and miscalibrated predictions. To mitigate the impact of inaccurate prediction, previous research applied deep ensembles and achieved superior results, despite the increased computation and storage space. In this paper, inspired by the uncertainty estimation, we leverage Monte Carlo Dropout to generate intrinsic reward from the perspective of uncertainty estimation with the goal to decrease the demands for computing resources while retaining superior performance. Utilizing the simple yet effective approach, we conduct extensive experiments across a variety of benchmark environments. The experimental results suggest that our method provides a competitive performance in final score and is faster in running speed, while requiring much fewer computing resources and storage space. Tianjiao Wan, Peichang Shi, Bo Ding 0001, Zijian Gao |
JCC | 4 |
| 2022 | Improving scalability of multi-agent reinforcement learning with parameters sharingabstractImproving the scalability of a multi-agent system is one of the key challenges for applying reinforcement learning to learn an effective policy. Parameter sharing is a common approach used to improve the efficiency of learning by reducing the volume of policy network parameters that need to be updated. However, sharing parameters also reduces the variance between agents’ policies, which further restricts the diversity of their behaviors. In this paper, we introduce a policy parameter sharing approach, it maintains a policy network for each agent, and only updates one of them. The differentiated behavior of agents is maintained by the policy, while sharing parameters are updated through a soft way. Experiments in foraging scenarios demonstrate that our method can effectively improve the performance and also the scalability of the multi-agent systems. Bo Ding 0001, Peichang Shi |
JCC | 2 |
| 2022 | Masked Modeling-based Audio Representation for ACM Multimedia 2022 Computational Paralinguistics ChallengEabstractIn this paper, we present our solution for ACM Multimedia 2022 Computational Paralinguistics Challenge. Our method employs the self-supervised learning paradigm, as it achieves promising results in computer vision and audio signal processing. Specifically, we firstly explore modifying the Swin Transformer architecture to learn general representation for the audio signals, accompanied with random masking on the log-mel spectrogram. The main goal of the pretext task is to predict the masked parts, by combining the advantages of the Swin-Transformer and masked modeling. For the downstream tasks, we utilize the labelled datasets to fine-tune the pre-trained model. Compared with the competitive baselines, our approach can provide significant performance improvements without ensembling. Kang You, Kele Xu, Boqing Zhu, Ming Feng, Bo Liu 0014, Bo Ding 0001 |
ACM Multimedia | 8 |
| 2021 | A Hashgraph-Based Knowledge Sharing Approach for Mobile Robot Swarm
Xiao Shu, Bo Ding 0001, Xiang Fu 0002, Zhen Li 0011 |
CollaborateCom (2) | 2 |
| 2021 | Improving Ultrasound Tongue Contour Extraction Using U-Net and Shape Consistency-Based RegularizerabstractB-mode ultrasound tongue imaging is widely used to visualize the tongue motion, due to its appearing properties. Extracting the tongue surface contour in the B-mode ultrasound image is still a challenge, while it is a prerequisite for further quantitative analysis. Recently, deep learning-based approach has been adopted in this task. However, the standard deep models fail to address faint contour when the ultrasound wave goes parallel to the tongue surface. To address the faint or missing contours in the sequence, we explore the shape consistency-based regularizer, which can take sequential information into account. By incorporating the regularizer, the deep model not only can extract frame-specific contours, but also can enforce the similarity between the contours extracted from adjacent frames. Extensive experiments are conducted both on the synthetic and real ultrasound tongue imaging dataset and the results demonstrate the effectiveness of proposed method. To better promote the research in this field, we have released our code at1. Ming Feng, Kele Xu, Huaimin Wang 0001, Bo Ding 0001 |
ICASSP | 5 |
| 2021 | EAD: An Efficient Anomaly Detection Algorithm for Multivariate Time SeriesabstractAnomaly detection based on deep learning has been widely used in IT infrastructure management. Driven by large amount of data, deep learning (DL) based algorithms can achieve higher accuracy compared to rule-based algorithms. However, the computational complexity of these algorithms is much higher than traditional rule-based ones, which will cost lots of time and computing resources. This limits the application of such algorithms in some actual systems. Therefore, it is necessary to improve the detection efficiency (e.g., execution time, resource occupation, etc.) of existing DL-based algorithms to make them more practical. In this work, we propose an efficient detection approach to solve this problem by combining DL-based algorithms and rule-based algorithms. Specificly, we apply rule-based algorithms to filter the original data roughly and then exploit the DL-based algorithms to make the final decision. We evaluate our approach using two state-of-the-art deep learning anomaly detection approaches with three real-world datasets. The results show that the proposed approach can significantly improve detection efficiency, saving about 80%~95% of execution time under the premise that the accuracy is nearly unchanged. Dehong Ma, Bo Ding 0001, Hui Liu 0052 |
ICTAI | 2 |
| 2021 | Multi-Actor-Attention-Critic Reinforcement Learning for Central Place Foraging SwarmsabstractMultiple agents with relatively low cost, decentralized control, and robustness have the advantages of completing a foraging task more efficiently than a single advanced robot. Despite many foraging algorithms are efficient in multiple robot systems, most are pre-designed or not very adaptive to different environments since they have to evolve the parameters of the foraging algorithm in each different environment. Besides, designing an efficient collision avoidance strategy for multiple agents is a challenge. Addressing these issues, we introduce the multi-actor-attention-critic(MAAC) reinforcement learning method into the multiple foraging agents. We train the foraging strategy for multiple simulated agents. We compare our approach with existing foraging algorithms for multiple robots, the Central Place Foraging Algorithm (CPFA) and the Distributed Deterministic Spiral Algorithm (DDSA). Experimental results demonstrate that our approach outperforms the two algorithms. Also, we illustrate that our approach has a better performance in avoiding obstacles and adapting to different environments. Kele Xu, Bo Ding 0001, Zijian Gao |
IJCNN | 4 |
| 2021 | Accurate Respiration Monitoring for Mobile Users With Commercial RFID DevicesabstractVital signs (e.g., respiration rate or heartbeat rate) sensing is of great importance to implement pervasive in-home healthcare. Traditional vital signs monitoring approaches usually require users to wear some dedicated sensors. These approaches are intrusive and inconvenient to use, especially for elderly people. Some non-intrusive vital signs monitoring approaches based on wireless sensing have been proposed in recent years. However, these approaches require the target user to be in situ during the monitoring process, which greatly limits their utilization in practical scenarios where the target users usually move around. In this paper, we propose RF-RMM, an RFID-based approach to accurate and continuous respiration monitoring for mobile users. The major challenge in respiration monitoring for moving people is that the tiny body displacement caused by the user's respiration is overwhelmed by the user's entire body movement. To address this issue, we propose a novel approach that uses a pair of tags to eliminate the effect of the user's body movement. We fuse the data from the paired tags to cancel the effect of the user's entire body movement and retain only the displacement caused by the user's respiration. Another challenging issue in implementing RF-RMM is how to resolve the phase ambiguity problem when the target user moves around, which becomes more serious than in the static case. We propose a distance tracking algorithm to track the phase transition during the user's movement, according to which the phase ambiguity problem can be well handled. We implement RF-RMM on commercial RFID devices and conduct extensive real-world experiments to evaluate its performance. The results show that RF-RMM achieves accurate respiration rate monitoring with an average error of 0.54 BPM in estimating different users' respiration rate and an average relative error of less than 13% in estimating the user's individual breath length. Shigeng Zhang, Xuan Liu 0001, Bo Ding 0001, Song Guo 0001, Jianxin Wang 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2021 | Cloudroid Swarm: A QoS-Aware Framework for Multirobot Cooperation OffloadingabstractComputation offloading has been widely recognized as an effective way to promote the capabilities of resource‐constrained mobile devices. Recent years have seen a renewal of the importance of this technology in the emerging field of mobile robots, supporting resource‐intensive robot applications. However, cooperating to solve complex tasks in the physical world, which is a significant feature of a robot swarm compared to traditional mobile computing devices, has not received in‐depth attention in research concerned with traditional computation offloading. In this study, we propose an approach named cooperation offloading, which offloads the intensive communication among robots as well as the computation for compute‐intensive and data‐intensive tasks. We analyze the performance gain of cooperation offloading by formalizing multirobot cooperative models; in addition, we study offloading decisions. Based on this approach, we design a cloud robotic framework named Cloudroid Swarm and develop several QoS‐aware mechanisms to provide a general solution to cooperation offloading with QoS assurance in multirobot cooperative scenes. We implement Cloudroid Swarm to transparently migrate multirobot applications to cloud servers without any code modification. We evaluate our framework using three different multirobot cooperative applications. Our results show that Cloudroid Swarm can be applied to various robotic applications and real‐world environments and bring significant benefits in terms of both network optimization and task performance. Besides, our framework has good scalability and can do support as many as 256 robot entities simultaneously. Yuanzhao Zhai, Bo Ding 0001, Pengfei Zhang 0006 |
Wirel. Commun. Mob. Comput. | 2 |
| 2020 | Using Configuration Semantic Features and Machine Learning Algorithms to Predict Build Result in Cloud-Based Container EnvironmentabstractContainer technologies are being widely used in large scale production cloud environments, of which Docker has become the de-facto industry standard. In practice, Docker builds often break, and a large amount of efforts are put into troubleshooting broken builds. Prior studies have evaluated the rate at which builds in large organizations fail. However, there is still a lack of early warning methods for predicting the Docker build result before the build starts. This paper provides a first attempt to propose an automatic method named PDBR. It aims to use the configuration semantic features extracted by AST and the machine learning algorithms to predict build result in the cloud-based container environment. The evaluation experiments based on more than 36,000 collected Docker builds show that PDBR achieves 73.45%-91.92% in F1 and 29.72%-72.16% in AUC. We also demonstrate that different ML classifiers have significant and large effects on the PDBR AUC performance. Yiwen Wu 0001, Yang Zhang 0026, Junsheng Chang, Bo Ding 0001, Tao Wang 0006, Huaimin Wang 0001 |
ICPADS | 4 |
| 2020 | Cooperative Offloading for Multiple Robot ApplicationsabstractComputation offloading has been widely recognized as an effective way to promote the capabilities of resource-constrained mobile devices. The past years have seen a renewed importance of this technology in the emerging field of mobile robots. However, a significant feature of robots compared to traditional mobile computing devices (e.g. smartphones) is that they must collaborate to solve complex tasks in the physical world in many cases, which implies intensive data exchange among the robot peers. This characteristic has not been dealt with in-depth in traditional computation offloading research. In this paper, we propose an approach called Cooperative Offloading, which takes into account the cooperation among mobile devices as well as the communication it brings in computation offloading. Firstly, we propose the offloading decision approach that deals with offloading from a set of robots that cooperate to finish a specific task. Secondly, we present a set of mechanisms to optimize the data transfer path on network topology in multi-robot applications, which can significantly reduce the bandwidth consumption of the robot wireless network. Based on the cooperative offloading approach, we realized a cloud robotics framework called Cloudroid Swarm. Evaluations based on real-life applications have shown that Cloudroid Swarm brings more than five times performance promotion compared to the setup without offloading or with individual computation offloading. Yuanzhao Zhai, Bo Ding 0001, Pengfei Zhang 0006, Qingtong Wu, Peichang Shi, Huaimin Wang 0001 |
JCC | 2 |
| 2020 | Improving Policy Generalization for Teacher-Student Reinforcement Learning
Xudong Gong, Hongda Jia, Xing Zhou 0004, Bo Ding 0001, Jie Xu 0007 |
KSEM (2) | 5 |
| 2019 | Solving multi-scenario cardinality constrained optimization problems via multi-objective evolutionary algorithms
Xing Zhou 0004, Huaimin Wang 0001, Wei Peng 0005, Bo Ding 0001, Rui Wang 0017 |
Sci. China Inf. Sci. | 4 |
| 2019 | Balanced connected task allocations for multi-robot systems: An exact flow-based integer program and an approximate tree-based genetic algorithm
Xing Zhou 0004, Huaimin Wang 0001, Bo Ding 0001, Tianjiang Hu, SuNing Shang |
Expert Syst. Appl. | 3 |
| 2019 | An Energy-Aware Offloading Framework for Edge-Augmented Mobile RFID SystemsabstractInternet of Things (IoT) have been widely used in many fields including smart city, industry Internet and automatic driving. Because IoT end devices usually have only limited capability in computation and power supply, they are not suitable to execute energy-consuming computational tasks. In many cases, we need to offload computational tasks from IoT end devices to edge servers in order to save energy consumption on the end devices. This process is usually termed as computing offloading. In this paper, we study computing offloading in radio frequency identification (RFID) systems built with mobile readers. We analyze the energy consumption characteristics of different components in mobile RFID systems, based on which we propose a framework to perform energy-aware offloading for such systems. By using tag searching as an example, we illustrate how our framework can help offload computational intensive tasks to edge servers to save energy consumption on mobile readers while satisfying the constraint on total execution time. Simulation results shown that the energy consumption of mobile readers can be greatly reduced by using our offloading framework. Xuan Liu 0001, Quan Yang, Juan Luo, Bo Ding 0001, Shigeng Zhang |
IEEE Internet Things J. | 4 |
| 2019 | Range-Based Localization for Sparse 3-D Sensor NetworksabstractLocalization plays a pivotal role in wireless sensor networks. Many range-based localization algorithms have been proposed for 2-D sensor networks or densely deployed 3-D sensor networks. However, range-based localization in sparse 3-D sensor networks is still a challenging problem, because the sparseness of the network makes it difficult to obtain a proper order of nodes to be sequentially localized. The patch-and-stitching localization strategy can conquer the sparseness problem in 2-D networks, but for 3-D networks it is still unknown how to uniquely merge two patches when there are not enough common nodes. In this paper, we solve this challenging problem by deriving the conditions under which two subnetworks can be uniquely merged. In the proposed approach, we treat the translation parameters as unknowns and form a set of equations with which the unknowns can be uniquely solved. The novelty of our algorithm also lies in that we exploit both common nodes and connecting edges among adjacent subnetworks to merge them, resulting in very high chances that two subnetworks can be merged. We conduct extensive simulation experiments to evaluate the performance of the proposed algorithm. The results show that the proposed algorithm could localize more than 90% of nodes in sparse 3-D networks with average node degree of 11 and anchor ratio of 5%, while the best existing solution can localize only 52% of nodes in the same situation. Xuan Liu 0001, Jiangjin Yin, Shigeng Zhang, Bo Ding 0001, Song Guo 0001, Kun Wang 0005 |
IEEE Internet Things J. | 4 |
| 2019 | Multi-objective evolutionary computation for topology coverage assessment problem
Xing Zhou 0004, Huaimin Wang 0001, Bo Ding 0001, Wei Peng 0005, Rui Wang 0017 |
Knowl. Based Syst. | 3 |
| 2018 | Learning to Cooperate in Decentralized Multi-robot Exploration of Dynamic Environments
Mingyang Geng, Xing Zhou 0004, Bo Ding 0001, Huaimin Wang 0001, Lei Zhang 0200 |
ICONIP (7) | 3 |
| 2018 | How Many Robots are Enough: A Multi-Objective Genetic Algorithm for the Single-Objective Time-Limited Complete Coverage ProblemabstractComplete coverage, which is the foundation of many robotic applications, aims to cover an area as quickly as possible. This study investigates the time-limited version of multi-robot complete coverage problem, that is, to find the least number of robots and allocate tasks properly to them such that they can finish a known mission within the time limit. This version of problem can be tackled straightforwardly based on optimizing the task-allocation to a fixed number of robots and enumerating the number. However, the number-fixed problem is NP-hard and the existing algorithm for the number-fixed problem allows intersecting tasks (possibly causing robots' interference) and endures high approximation factor. In this study, the time-limited complete coverage problem is tackled with a multi-objective approach, instead of enumerating robots' number and optimizing each number-fixed problem one by one. The multi-objective GA, Mofint, at first estimates the lower and upper bounds of the number of robots. It abstracts each task as a weighted node of a graph. Then, Mofint evolves individuals, each individual being a forest containing a certain number (within the bounds) of non-intersecting trees. Mofint can finally obtain higher precision than existing work with less time: the approximation factor for Mofint is 1.1 to 1.5 times the ideal allocation when robots' number is fixed, while for existing work is 1.5 to 2. Due to its higher precision, the least number of robots obtained in the experiments by Mofint is 0.6 times of existing work. Xing Zhou 0004, Huaimin Wang 0001, Bo Ding 0001 |
ICRA | 3 |
| 2018 | Cloud-Based Framework for Scalable and Real-Time Multi-Robot SLAMabstractIn the past decade, multi-robot simultaneous localization and mapping (SLAM) has been widely studied. However, the problem of collaborative SLAM with a large number of robots, such as dozens of robots, is far from being well solved. The challenges stem from not only the computation complexity in large-scale map merging but also the inefficiency to enable the parallel computing in this process, which is indispensable for us to make avail of the frontier of computing technology such as powerful cloud infrastructure. To effectively address these challenges, especially the latter one, we propose a scalable and real-time multi-robot visual SLAM framework based on the cloud robotic paradigm. The prominent feature of our framework is that it can distribute the SLAM process to multiple computing hosts in a cluster, which enables map building in parallel. To eliminate the bottleneck from data sharing between different sub-tasks, we also introduce diversified messaging pattern for various messaging scenarios, as well as the consistency policies for map data. The evaluations on the prototype of our framework, have shown that our method can do support as many as 256 robot entities simultaneously, without any compromising on the precision of poses estimation and map building. Pengfei Zhang 0006, Huaimin Wang 0001, Bo Ding 0001, SuNing Shang |
ICWS | 3 |
| 2018 | Unsupervised Learning of Depth and Pose Estimation based on Continuous Frame WindowabstractWe present an unsupervised learning framework for the task of monocular depth and camera motion estimation from video sequences. In common with recent work, we use an unsupervised end-to-end learning method, requiring monocular video sequences for training. What makes the difference is, our approach not only uses image reconstruction as the supervisory signal but also exploits the pose estimation method which was used in traditional SLAM approach to enhance the supervisory signal and add training constraints. In pose estimation, a continuous frame window is set to construct the pose graph. Our method uses single-view depth and multi-view pose networks, with a loss based on reconstructing nearby images to the target using the predicted depth and pose. During training, the networks are thus coupled by the loss but can be applied independently at test time. Our evaluation of experiments on the KITTI dataset proves the effectiveness of our method: 1) monocular depth performs superior to the supervised methods that use ground-truth depth data for training and the existing unsupervised learning method. Our method performs comparably with the supervised methods that use ground-truth pose data for training. 2) pose estimation performs almost the same compared to established SLAM systems under comparable input settings. SuNing Shang, Huaimin Wang 0001, Pengfei Zhang 0006, Bo Ding 0001 |
IJCNN | 4 |
| 2018 | RoboCloud: augmenting robotic visions for open environment modeling using Internet knowledge
Yiying Li, Huaimin Wang 0001, Bo Ding 0001, Wei Zhou 0107 |
Sci. China Inf. Sci. | 3 |
| 2017 | Cloudroid: A Cloud Framework for Transparent and QoS-Aware Robotic Computation OutsourcingabstractMany robotic tasks require heavy computation, which can easily exceed the robot's onboard computer capability. A promising solution to address this challenge is outsourcing thecomputation to the cloud. However, exploiting the potential ofcloud resources in robotic software is difficult, because it in-volves complex code modification and extensive (re)configurationprocedures. Moreover, quality of service (QoS) such as timeliness, which is critical to robot's behavior, have to be considered. Inthis paper, we propose a transparent and QoS-aware softwareframework called Cloudroid for cloud robotic applications. Thisframework supports direct deployment of existing robotic soft-ware packages to the cloud, transparently transforming theminto Internet-accessible cloud services. And with the automati-cally generated service stubs, robotic applications can outsourcetheir computation to the cloud without any code modification. Furthermore, the robot and the cloud can cooperate to maintainthe specific QoS property such as request response time, evenin a highly dynamic and resource-competitive environment. Weevaluated Cloudroid based on a group of typical robotic scenariosand a set of software packages widely adopted in real-worldrobot practices. Results show that robots capability can beenhanced significantly without code modification and specific QoSobjectives can be guaranteed. In certain tasks, the "cloud + robot" setup shows improved performance in orders of magnitudecompared with the robot native setup. Ben Hu, Huaimin Wang 0001, Pengfei Zhang 0006, Bo Ding 0001, Huimin Che |
CLOUD | 4 |
| 2017 | Cloud-Based Knowledge Sharing in Cooperative Robot Tracking of Multiple Targets with Deep Neural Network
Hui Bao, Huaimin Wang 0001, Bo Ding 0001, SuNing Shang |
ICONIP (6) | 3 |
| 2017 | Enabling Imagination: Generative Adversarial Network-Based Object Finding in Robotic Tasks
Huimin Che, Ben Hu, Bo Ding 0001, Huaimin Wang 0001 |
ICONIP (6) | 3 |
| 2017 | Learning from Internet: Handling Uncertainty in Robotic Environment ModelingabstractUncertainty is a great challenge for environment perception of autonomous robots. For instance, while building semantic maps (i.e., maps with semantic labels such as object names), the robot may encounter unexpected objects of which it has no knowledge. It will lead to inevitable failures in traditional environment modeling software. The abundant knowledge being accumulated on the Internet has the potential to assist robots to handle such kind of uncertainly. However, existing researches have not touched this issue yet. This paper proposes a cloud-based semantic mapping engine named SemaCloud, which can not only augment robot's environment modeling capability by the rich cloud resources but also cope with uncertainty by utilizing the Internet knowledge on necessary. It adopts a state-of-art Deep Neural Network (DNN) for real-time and accurate recognition of pre-trained objects. If an object is beyond the knowledge of this DNN, a special mechanism named QoS-aware cloud phase transition is triggered to seek help from existing recognition services on the Internet. By a set of carefully-designed algorithms, it can maximize benefits and minimize the negative impacts on the Quality of Service (QoS) properties of robotic applications, which is essential to many robot scenarios. The experiments on both open datasets and real robots show that our work can handle uncertainly successfully in robotic semantic mapping without sacrificing critical real-time constraints. Yiying Li, Huaimin Wang 0001, Bo Ding 0001, Huimin Che |
Internetware | 3 |
| 2015 | Auxo: an architecture-centric framework supporting the online tuning of software adaptivity
Huaimin Wang 0001, Bo Ding 0001, Dian-xi Shi, Jiannong Cao 0001, Alvin Chan Toong Shoon |
Sci. China Inf. Sci. | 2 |
| 2014 | MABP: an optimal resource allocation approach in data center networks
Xiaoling Li 0002, Huaimin Wang 0001, Bo Ding 0001, Xiaoyong Li 0002 |
Sci. China Inf. Sci. | 3 |
| 2014 | Resource allocation with multi-factor node ranking in data center networks
Xiaoling Li 0002, Huaimin Wang 0001, Bo Ding 0001, Xiaoyong Li 0002 |
Future Gener. Comput. Syst. | 3 |
| 2012 | Topology awareness algorithm for virtual network mappingabstractNetwork virtualization is recognized as an effective way to overcome the ossification of the Internet. However, the virtual network mapping problem (VNMP) is a critical challenge, focusing on how to map the virtual networks to the substrate network with efficient utilization of infrastructure resources. The problem can be divided into two phases: node mapping phase and link mapping phase. In the node mapping phase, the existing algorithms usually map those virtual nodes with a complete greedy strategy, without considering the topology among these virtual nodes, resulting in too long substrate paths (with multiple hops). Addressing this problem, we propose a topology awareness mapping algorithm, which considers the topology among these virtual nodes. In the link mapping phase, the new algorithm adopts the k -shortest path algorithm. Simulation results show that the new algorithm greatly increases the long-term average revenue, the acceptance ratio, and the long-term revenue-to-cost ratio ( R/C ). Xiaoling Li 0002, Huaimin Wang 0001, Changguo Guo, Bo Ding 0001, Xiaoyong Li 0002, Wen-qi Bi, Shuang Tan |
J. Zhejiang Univ. Sci. C | 4 |
| 2010 | Taming software adaptability with architecture-centric frameworkabstractIn many cases, we would like to enhance the predefined adaptability of a running application, for example, to enable it to cope with a strange environment. To make such kind of runtime modifications is a challenging task. In existing engineering practices, the online policy upgrade approach just focuses on the modification of adaptation decision logic and lacks system-level means to assess the validity of an upgrade. This paper proposes a framework for adaptive software that supports the online reconfiguration of each concern in the “sensing-decision-execution” adaptation loop. To achieve this goal, our framework supports an architecture style which encapsulates adaptation concerns as software architecture elements. And then, it maintains a runtime architecture model to enable the dynamic reconfiguration of those elements as well as help to ensure the validity of a change. A third party can selectively add, remove or replace part of this model to enhance the running application's adaptability. We validated this framework by two cases extracted from real life. Bo Ding 0001, Huaimin Wang 0001, Dian-xi Shi, Jiannong Cao 0001 |
PerCom | 1 |
| 2009 | Towards Unanticipated Adaptation: An Architecture-Based ApproachabstractOver its lifetime, adaptive software may have to deal with the environment not anticipated during the original development. In such cases, we should introduce new adaptive code, for example, to detect the strange contexts or update the out-of-date adaptation decision logic. This paper proposes an engineering approach facilitates this kind of post-delivery modifications based on software architecture techniques. Our approach introduces a component model separates different adaptation concerns (sensing, decision and execution) as different types of software architecture elements. The clear separation lays the foundation for the independent maintenance of each concern. And then, with the aid of a container supports the instantiation and run-time modification of the software architecture model, those concerns can be bound together without recompiling the whole software, even while it is running. Our approach enables the fine-grained, low-cost modifications of delivered adaptive software in the case that an unanticipated environment emerges. Bo Ding 0001, Huaimin Wang 0001, Dian-xi Shi, Xiang Rao |
SERA | 1 |
| 2008 | Component Based Context ModelabstractContext awareness is one of the most fundamental issues in pervasive computing. In this paper, component based context model based on middleware architecture is proposed. Moreover, OWL-based context ontology for modelling context information to easily share and reuse context knowledge is presented. By giving fire alarm scenario for our prototype, the proposed component based context architecture can be deployed in different context-aware application and can provide a middleware support for context representation and knowledge sharing. Bo Ding 0001, Huaimin Wang 0001, Dian-xi Shi |
WAIM | 2 |