VLDB 2026 Research / reviewers in the wild / expert
Wei Song 0008
dblp:62/1539-8
· DBLP profile ↗
41ranked-venue papers
12as first author
38since 2021 · last 2026
0000-0002-0828-7486ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 8 first-author · 23 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Neural Network-Aided Differential Evolution With Double Q-Learning for Dynamic OptimizationabstractAs the search space and hence the optimum vary through time, dynamic optimization problems (DOPs) bring tremendous difficulties. Changes in DOPs often manifest as diverse dynamics. Consequently, regulating individuals’ search to adapt to diverse dynamics is crucial to tackle DOPs. Besides, due to the inherent population nature in dynamic optimization algorithms (DOAs), loss of global and local diversities is a critical issue deteriorating the performance of DOAs. Faced with these difficulties, this article proposes a neural network-aided differential evolution with double Q-learning (NNDE-DQ), in which an evolutionary regulation network (ERN) is designed to maintain high global and local diversities over time and regulate individuals’ search that can adapt to diverse dynamics. NNDE-DQ first partitions the search space into multiple subspaces and in each subspace distant individuals are selected as the centers of ERN’s hidden nodes activated by radial basis function. Every input individual is mutated with two randomly selected hidden node centers from different subspaces as differential individuals, facilitating the maintenance of a high global diversity due to very distinct differential terms of the population. Moreover, each mutated individual selects a hidden node center from the subspace located by the mutated individual to undergo crossover. Due to distant hidden node centers in each subspace, a high local diversity can be maintained by individuals’ crossover. Besides, DQ is introduced to acquire ERN’s desired output by interactively estimating individuals’ state-action information, enabling ERN to learn the regulation of individuals’ search in dynamic environments and hence adapt to diverse dynamics. The experimental results demonstrate that NNDE-DQ significantly improves the performance in solving various DOPs comparing to seven state-of-the-art DOAs. Wei Song 0008, Yaochu Jin, Yinan Guo 0001, Shengxiang Yang |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2026 | Towards Privacy-Preserving Top-$k$k Location-Based Dominating Queries Over Encrypted DataabstractWith the growth of cloud computing infrastructure, its cost-efficient paradigm is driving a growing wave of small and medium-sized enterprises to migrate data and services to cloud based platforms. However, due to privacy concerns, data encryption prior to outsourcing is an important means of protection, which in turn requires performing queries over the encrypted data. While several approaches do offer support for privacy preserving skyline or top-k queries, they typically struggle with efficiency when extended to secure top-k dominating queries, due to the inherent nature of combining the advantages of both top k and skyline queries. Nevertheless, this nature renders them as a more practical and promising alternative for location-based services. To address this, we introduce STLD, a secure top-k location-based dominating query scheme. Specifically, we develop an innovative index structure called Secure Aggregate R-tree (SAR-tree) by utilizing the Paillier cryptosystem and introducing meticulously crafted noise, while also incorporating principles from aggregate R-trees and semi-blind R-trees. Leveraging this structure, we propose a series of secure sub-protocols to facilitate top-k dominating queries, accompanied by optimization techniques to mitigate latency associated with computationally intensive dominating operations. Given an encrypted query, STLD not only efficiently answers the query but also guarantees the privacy of data(sets), results, queries and access patterns. Finally, STLD undergoes rigorous theoretical security and complexity analysis, complemented by empirical evaluations that demonstrate its performance and feasibility, achieving a reduction in query cost by 40%-60% compared to multiple competing methods. Zuan Wang, Xiaofeng Ding 0001, Wei Song 0008, Pan Zhou 0001, Lin Chen 0033, Youliang Tian, Kim-Kwang Raymond Choo |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Multipattern Learning and Collaboration-Based Evolutionary Optimizer for Large-Scale Multiobjective OptimizationabstractRecently, machine learning-embedded large-scale multiobjective evolutionary algorithms (LMOEAs) have shown great promise in solving large-scale multiobjective optimization problems (LMOPs). However, the fast convergence of the population to the true Pareto-optimal front (POF) and even distribution of the obtained Pareto-optimal solutions (POSs) on the POF are not adequately considered when tackling an LMOP. Besides, existing LMOEAs typically pair solutions with a matching rule and employ a network to learn the evolution pattern among the obtained solution pairs. It is difficult to learn various evolution patterns through a simple network, which hinders the collaboration of different patterns for enhancing the search capability. Facing such difficulties, this article proposes an LMOEA with multipattern learning and collaboration (LMOEA-MLC), where a single-hidden-layer multioutput network (SMN) is established to learn inductive and hybrid evolution patterns. Specifically, two inductive ones can be learned with the solution pairs built by two matching rules toward fast convergence and even distribution, respectively. Moreover, the solution pairs considering the fusion of the two inductive ones are collected, enabling SMN to learn a hybrid one and thus making a tradeoff between fast convergence and even distribution. Besides, the learned evolution patterns collaborate to enhance the search capability due to the distinct patterns. To enhance learning speed, SMN’s parameters are updated by an incremental random vector functional link (IRVFL). In our experiments, comprehensive comparisons with eight state-of-the-art LMOEAs demonstrate the significant performance improvement of LMOEA-MLC in handling LMOPs. Wei Song 0008, Mingshuo Song, Haojie Zhou, Xiaoyan Sun 0002, Yaochu Jin, Songbai Liu, Qiuzhen Lin, Shengxiang Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2025 | Discriminator-Guided Embodied Planning for LLM AgentabstractLarge Language Models (LLMs) have showcased remarkable reasoning capabilities in various domains, yet face challenges in complex embodied tasks due to the need for a coherent long-term policy and context-sensitive environmental understanding. Previous work performed LLM refinement relying on outcome-supervised feedback, which can be costly and ineffective. In this work, we introduce a novel framework, Discriminator-Guided Action Optimization (DGAP), for facilitating the optimization of LLM action plans via step-wise signals. Specifically, we employ a limited set of demonstrations to enable the discriminator to learn a score function, which assesses the alignment between LLM-generated actions and the underlying optimal ones at every step. Based on the discriminator, LLMs are prompted to generate actions that maximize the score, utilizing historical action-score pair trajectories as guidance. Under mild conditions, DGAP resembles critic-regularized optimization and has been demonstrated to achieve a stronger policy than the LLM planner. In experiments across different LLMs (GPT-4, Llama3-70B) in ScienceWorld and VirtualHome, our method achieves superior performance and better efficiency than previous methods. Haofu Qian, Chenjia Bai, Jiatao Zhang, Fei Wu 0001, Wei Song 0008, Xuelong Li 0001 |
ICLR | 5 |
| 2025 | FCRF: Flexible Constructivism Reflection for Long-Horizon Robotic Task Planning with Large Language ModelsabstractAutonomous error correction is critical for domestic robots to achieve reliable execution of complex long-horizon tasks. Prior work has explored self-reflection in Large Language Models (LLMs) for task planning error correction; however, existing methods are constrained by inflexible self-reflection mechanisms that limit their effectiveness. Motivated by these limitations and inspired by human cognitive adaptation, we propose the Flexible Constructivism Reflection Framework (FCRF), a novel Mentor-Actor architecture that enables LLMs to perform flexible self-reflection based on task difficulty, while constructively integrating historical valuable experience with failure lessons. We evaluated FCRF on diverse domestic tasks through simulation in AlfWorld and physical deployment in the real-world environment. Experimental results demonstrate that FCRF significantly improves overall performance and self-reflection flexibility in complex long-horizon robotic tasks. Website at https://mongoosesyf.github.io/FCRF.github.io/ Jiatao Zhang, Zeng Gu, Qingmiao Liang, Tuocheng Hu, Wei Song 0008, Shiqiang Zhu |
IROS | 6 |
| 2025 | Towards Privacy-Preserving Range Queries with Secure Learned Spatial Index over Encrypted DataabstractWith the growing reliance on cloud services for large-scale data management, preserving the security and privacy of outsourced datasets has become increasingly critical. While encrypting data and queries can prevent direct content exposure, recent research reveals that adversaries can still infer sensitive information via access pattern and search path analysis. However, existing solutions that offer strong access pattern privacy often incur substantial performance overhead. In this paper, we propose a novel privacy-preserving range query scheme over encrypted datasets, offering strong security guarantees while maintaining high efficiency. To achieve this, we develop secure l earned spatial index (SLS-INDEX), a secure learned index that integrates the Paillier cryptosystem with a hierarchical prediction architecture and noise-injected buckets, enabling data-aware query acceleration in the encrypted domain. To further obfuscate query execution paths, SLS-INDEX- based Range Queries (SLRQ) employs a permutation-based secure bucket prediction protocol. Additionally, we introduce a secure point extraction protocol that generates candidate results to reduce the overhead of secure computation. We provide formal security analysis under realistic leakage functions and implement a prototype to evaluate its practical performance. Extensive experiments on both real-world and synthetic datasets demonstrate that SLRQ significantly outperforms existing solutions in query efficiency while ensuring dataset, query, result, and access pattern privacy. Zuan Wang, Juntao Lu, Jiazhuang Wu, Youliang Tian, Wei Song 0008, Qiuxian Li |
TrustCom | 5 |
| 2025 | Referring Expression Comprehension in semi-structured human-robot interaction
Tianlei Jin, Qiwei Meng, Qiulan Huang, Fangtai Guo, Shu Kong, Wei Song 0008, Jiakai Zhu, Jason Gu |
Expert Syst. Appl. | 7 |
| 2025 | Feature Bank-Guided Reconstruction for Anomaly DetectionabstractVisual surface anomaly detection targets the location of anomalies, with numerous methods available to address the challenge. Reconstruction-based methods are popular for their adaptability and interpretability. However, reconstruction-based methods currently struggle with the challenge of achieving low image fidelity and a tendency to reconstruct anomalies. To overcome these challenges, we introduces the Feature Bank-guided Reconstruction method (FBR), incorporating three innovative modules: anomaly simulation, feature bank module, and a cross-fused Discrete Cosine Transform channel attention module. Guided by these modules, our method is capable of reconstructing images with enhanced robustness. The experimental results validate the effectiveness of the proposed approach, which not only achieves outstanding performance on the BeanTech AD dataset with an 96.4% image-AUROC and a 97.3% pixel-AUROC, but also demonstrates competitive performance on the MVTec AD dataset with a 99.5% image-AUROC and a 98.3% pixel-AUROC. Tao Zhang 0010, Wei Song 0008 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Multiform Differential Evolution With Elite-Guided Knowledge Transfer for Coal Mine Integrated Energy Systems Constrained DispatchabstractThe dispatch optimization of coal mine integrated energy system is challenging due to high dimensionality, strong coupling constraints, and multiobjective. Existing constrained multiobjective evolutionary algorithms struggle with locating multiple small and irregular feasible regions when solving the dispatch problem. To address this issue, we here develop a multiform EA framework that incorporates the dispatch-correlated domain knowledge to effectively deal with strong constraints and multiobjective optimization. Possible evolutionary multiform construction strategy based on complex constraint relationship analysis and handling, i.e., constraint-coupled spatial decomposition, constraint strength classification, and constraint handling technique, is first explored. Within the multiform evolutionary optimization framework, two strategies, i.e., an elite-guided knowledge transfer by designing a special crowding distance mechanism to select dominant individuals from each task and a neighborhood-driven dual mutation to effectively balance the diversity and convergence of each optimized task for the differential evolution algorithm, are further developed. The performance of the proposed algorithm in feasibility, convergence, and diversity is demonstrated in a case study of a coal mine integrated energy system (IES) by comparing with CPLEX solver and eight state-of-the-art constrained multiobjective EAs. Canyun Dai, Xiaoyan Sun 0002, Hejuan Hu, Wei Song 0008, Yong Zhang 0016, Dun-Wei Gong |
IEEE Trans. Evol. Comput. | 4 |
| 2025 | Multisource and Hidden Source-Based Knowledge Transfer for Solving Dynamic Multiobjective Optimization ProblemsabstractRecently, transfer-learning-based dynamic multiobjective optimization algorithms (TL-DMOAs) have been shown to be very promising in solving dynamic multiobjective optimization problems (DMOPs). However, it is difficult for them to model knowledge capable of delineating the Pareto optimal solutions (POSs) found in each historical environment, because the POSs’ distribution cannot be adequately reflected. Besides, existing TL-DMOAs normally focus on acquiring knowledge from historical environments, but neglect correlations behind them for excavating potential knowledge, restricting the performance in generating high-quality initial populations (HIPs). To address these issues, herein a DMOA with multisource and hidden source-based knowledge transfer (DMOA-MHKT) is proposed. First, we design a knowledge extraction strategy by introducing mean shift, a nonparametric clustering method, to cluster the historical POSs. As clusters’ representatives, the cluster centers are considered to represent environmental knowledge, because they can adequately reflect the POSs’ distribution. Second, the most similar historical environment through environmental match and the last one are selected as two explicit sources. In the former source the POSs’ cluster centers are treated as its knowledge. By contrast, based on the POSs’ cluster centers and knee points in the latter source, a scoring method is designed to generate environmental knowledge by depicting the dynamics between two continuous environments. Third, after aligning knowledge of the explicit sources, a hidden source is learned by excavating correlations and potential knowledge behind them, facilitating the generalization enhancement in generating HIPs. The experimental results especially performance comparisons with seven state-of-the-art DMOAs demonstrate that DMOA-MHKT brings significant improvements in solving DMOPs. Wei Song 0008, Xiaoyan Sun 0002, Yaochu Jin, Khin Wee Lai |
IEEE Trans. Evol. Comput. | 1 |
| 2024 | Neural Network-Assisted Particle Swarm Dynamic OptimizationabstractIn the field of optimization, many problems change over time and they are referred to dynamic optimization prob-lems (DOPs). faced with DOPs, how to adapt to environmental changes and find the optima are challenging. In this paper, we propose a neural network-assisted particle swarm optimization (NN-PSO) algorithm and the search-guided neural networks (SGNN) are designed. Each particle selects its local and global learning targets based on the hidden nodes of SGNN. Besides, in the output layer, SGNN adjusts the local acceleration coefficient of each particle and hence the global one considering the relation between the two coefficients. A reinforcement learning manner is employed to obtain the desired output of SGNN. We define the significance and crowding degree metrics of hidden nodes, which aims to obtain a compact structure of SGNN. Incremental learning is leveraged to ensure the network approximation ability. In our experiments, we compare the proposed NN-PSO with five state-of-the-art dynamic optimization algorithms on moving peaks benchmark (MPB) benchmark test suite. The experimental results demonstrate that our algorithm achieves significant performance improvement in solving DOPs. Wei Song 0008, Mingshuo Song |
CEC | 2 |
| 2024 | Dynamic Multi-Task Interactive Evolutionary Optimization Algorithm with Search Space AlignmentabstractThe interactive evolutionary multi-task optimization approach assisted by surrogate models has proven successful in enhancing individualized recommendation performance. However, in light of the dynamically changing user preferences, it becomes imperative to further develop more powerful knowledge sharing strategy to improve the interactive multi-task optimization efficiency. A probability model-assisted search space alignment among multiple tasks method is proposed for knowledge transfer across multi-task environment. Additionally, a diversity maintaining mechanism of the transferred population is proposed to improve the quality and diversity of the initial population in a preference varied multi-task environment. The research demonstrates the effectiveness of the proposed approach in achieving effective knowledge transfer across multi-task environmen. Furthermore, the utilization of a high-quality initial population not only enhances the evolutionary search efficiency of the algorithm but also significantly improves the accuracy, diversity., and novelty of personalized recommendation. Weidong Wu, Xiaoyan Sun 0002, Yong Zhang 0016, Wei Song 0008 |
CEC | 4 |
| 2024 | M2ConceptBase: A Fine-Grained Aligned Concept-Centric Multimodal Knowledge BaseabstractMultimodal knowledge bases (MMKBs) provide cross-modal aligned knowledge crucial for multimodal tasks. However, the images in existing MMKBs are generally collected for entities in encyclopedia knowledge graphs. Therefore, detailed groundings of visual semantics with linguistic concepts are lacking, which are essential for the visual concept cognition ability of multimodal models. Addressing this gap, we introduce M2 ConceptBase, the first concept-centric MMKB. M2 ConceptBase models concepts as nodes with associated images and detailed textual descriptions. We propose a context-aware multimodal symbol grounding approach to align concept-image and concept-description pairs using context information from image-text datasets. Comprising 951K images and 152K concepts, M2 ConceptBase links each concept to an average of 6.27 images and a single description, ensuring comprehensive visual and textual semantics. Human studies confirm more than 95% alignment accuracy, underscoring its quality. Additionally, our experiments demonstrate that M2 ConceptBase significantly enhances VQA model performance on the OK-VQA task. M2 ConceptBase also substantially improves the fine-grained concept understanding capabilities of multimodal large language models through retrieval augmentation in two concept-related tasks, highlighting its value. Zhiwei Zha, Jiaan Wang, Zhixu Li, Xiangru Zhu, Wei Song 0008, Yanghua Xiao |
CIKM | 5 |
| 2024 | Flickr30K-CFQ: A Compact and Fragmented Query Dataset for Text-image Retrieval
Yaoxian Song, Xuwu Wang, Xiangru Zhu, Zhixu Li, Wei Song 0008, Tiefeng Li |
DASFAA (3) | 6 |
| 2024 | MLDT: Multi-Level Decomposition for Complex Long-Horizon Robotic Task Planning with Open-Source Large Language Model
Jiatao Zhang, Lanling Tang, Guilin Qi, Wei Song 0008 |
DASFAA (5) | 8 |
| 2024 | Leveraging the efficiency of multi-task robot manipulation via task-evoked planner and reinforcement learningabstractMulti-task learning has expanded the boundaries of robotic manipulation, enabling the execution of increasingly complex tasks. However, policies learned through reinforcement learning exhibit limited generalization and narrow distributions, which restrict their effectiveness in multi-task training. Addressing the challenge of obtaining policies with generalization and stability represents a non-trivial problem. To tackle this issue, we propose a planning-guided reinforcement learning method. It leverages a task-evoked planner(TEP) and a reinforcement learning approach with planner’s guidance. TEP utilizes reusable samples as the source, with the aim of learning reachability information across different task scenarios. Then in reinforcement learning, TEP assesses and guides the Actor towards better outputs and smoothly enhances the performance in multi-task benchmarks. We evaluate this approach within the Meta-World framework and compare it with prior works in terms of learning efficiency and effectiveness. Depending on experimental results, our method has more efficiency, higher success rates, and demonstrates more realistic behavior. Haofu Qian, Jiatao Zhang, Jason Gu, Wei Song 0008, Shiqiang Zhu |
ICRA | 6 |
| 2024 | Aligning Knowledge Graph with Visual Perception for Object-goal NavigationabstractObject-goal navigation is a challenging task that requires guiding an agent to specific objects based on first-person visual observations. The ability of agent to comprehend its surroundings plays a crucial role in achieving successful object finding. However, existing knowledge-graph-based navigators often rely on discrete categorical one-hot vectors and vote counting strategy to construct graph representation of the scenes, which results in misalignment with visual images. To provide more accurate and coherent scene descriptions and address this misalignment issue, we propose the Aligning Knowledge Graph with Visual Perception (AKGVP) method for object-goal navigation. Technically, our approach introduces continuous modeling of the hierarchical scene architecture and leverages visual-language pre-training to align natural language description with visual perception. The integration of a continuous knowledge graph architecture and multimodal feature alignment empowers the navigator with a remarkable zero-shot navigation capability. We extensively evaluate our method using the AI2-THOR simulator and conduct a series of experiments to demonstrate the effectiveness and efficiency of our navigator. Nuo Xu 0006, Wen Wang 0017, Zheyuan Lin, Wei Song 0008, Chunlong Zhang, Jason Gu, Chao Li 0028 |
ICRA | 6 |
| 2024 | FLTRNN: Faithful Long-Horizon Task Planning for Robotics with Large Language ModelsabstractRecent planning methods based on Large Language Models typically employ the In-Context Learning paradigm. Complex long-horizon planning tasks require more context(including instructions and demonstrations) to guarantee that the generated plan can be executed correctly. However, in such conditions, LLMs may overlook(unfaithful) the rules in the given context, resulting in the generated plans being invalid or even leading to dangerous actions. In this paper, we investigate the faithfulness of LLMs for complex long-horizon tasks. Inspired by human intelligence, we introduce a novel framework named FLTRNN. FLTRNN employs a language-based RNN structure to integrate task decomposition and memory management into LLM planning inference, which could effectively improve the faithfulness of LLMs and make the planner more reliable. We conducted experiments in VirtualHome household tasks. Results show that our model significantly improves faithfulness and success rates for complex long-horizon tasks. Website at https://tannl.github.io/FLTRNN.github.io/ Jiatao Zhang, Lanling Tang, Qiwei Meng, Haofu Qian, Wei Song 0008, Shiqiang Zhu, Jason Gu |
ICRA | 7 |
| 2024 | Learning to Guide Particle Search for Dynamic Multiobjective OptimizationabstractDynamic multiobjective optimization problems (DMOPs) are characterized by multiple objectives that change over time in varying environments. More specifically, environmental changes can be described as various dynamics. However, it is difficult for existing dynamic multiobjective algorithms (DMOAs) to handle DMOPs due to their inability to learn in different environments to guide the search. Besides, solving DMOPs is typically an online task, requiring low computational cost of a DMOA. To address the above challenges, we propose a particle search guidance network (PSGN), capable of directing individuals' search actions, including learning target selection and acceleration coefficient control. PSGN can learn the actions that should be taken in each environment through rewarding or punishing the network by reinforcement learning. Thus, PSGN is capable of tackling DMOPs of various dynamics. Additionally, we efficiently adjust PSGN hidden nodes and update the output weights in an incremental learning way, enabling PSGN to direct particle search at a low computational cost. We compare the proposed PSGN with seven state-of-the-art algorithms, and the excellent performance of PSGN verifies that it can handle DMOPs of various dynamics in a computationally very efficient way. Wei Song 0008, Shaocong Liu, Xinjie Wang 0002, Yinan Guo 0001, Shengxiang Yang, Yaochu Jin |
IEEE Trans. Cybern. | 1 |
| 2024 | Whole-Body Inverse Kinematics and Operation-Oriented Motion Planning for Robot Mobile ManipulationabstractHigh DoF mobile manipulation of robots is a nonlinear, nonchain redundant problem. In this article, we focus on two subissues of robot mobile manipulation: whole-body inverse kinematics (whole-body IK) and operation-oriented motion planning (OOMP). Whole-body IK solves the robot arm joint configuration and the mobile base position configuration according to the target pose. OOMP generates a feasible trajectory from the current pose to the target pose. The trajectory can avoid obstacles and touch operated objects. We introduce neural network optimization (NNO) methods with two variations to solve whole-body IK and OOMP, respectively. For whole-body IK, we design a fully connected network (FCN) to predict ten DoF of position and joint configurations based on the target pose. We use these ten DoF configurations to derive the predicted pose for online optimization. For OOMP, we design a GRU-based network to generate trajectories based on the initial and goal states. We mainly adopt sphere masks to modify the point cloud properties of the target object dynamically. During optimization, the trajectory keeps away from point clouds but approaches sphere masks. Finally, we conduct extensive experiments both on a Franka Panda robot and a mobile dual-arm robot. The results demonstrate the superior performance of our NNO method on whole body IK and OOMP, and implement mobile manipulation in different environments successfully. Tianlei Jin, Jiakai Zhu, Shiqiang Zhu, Zaixing He, Shuyou Zhang 0001, Wei Song 0008, Jason Gu |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | Scene-Driven Multimodal Knowledge Graph Construction for Embodied AIabstractEmbodied AI is one of the most popular studies in artificial intelligence and robotics, which can effectively improve the intelligence of real-world agents (i.e. robots) serving human beings. Scene knowledge is important for an agent to understand the surroundings and make correct decisions in the varied open world. Currently, knowledge base for embodied tasks is missing and most existing work use general knowledge base or pre-trained models to enhance the intelligence of an agent. For conventional knowledge base, it is sparse, insufficient in capacity and cost in data collection. For pre-trained models, they face the uncertainty of knowledge and hard maintenance. To overcome the challenges of scene knowledge, we propose a scene-driven multimodal knowledge graph (Scene-MMKG) construction method combining conventional knowledge engineering and large language models. A unified scene knowledge injection framework is introduced for knowledge representation. To evaluate the advantages of our proposed method, we instantiate Scene-MMKG considering typical indoor robotic functionalities (Manipulation andMobility), namedManipMob-MMKG. Comparisons in characteristics indicate our instantiated ManipMob-MMKG has broad superiority on data-collection efficiency and knowledge quality. Experimental results on typical embodied tasks show that knowledge-enhanced methods using our instantiated ManipMob-MMKG can improve the performance obviously without re-designing model structures complexly. Yaoxian Song, Penglei Sun, Zhixu Li, Wei Song 0008, Yanghua Xiao, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Interaction-and-Response Network for Distantly Supervised Relation ExtractionabstractDistantly supervised relation extraction (DSRE) aims to identify semantic relations from massive plain texts. A broad range of the prior research has leveraged a series of selective attention mechanisms over sentences in a bag to extract relation features without considering dependencies among the relation features. As a result, potential discriminative information existed in the dependencies is ignored, causing a decline in the performance of extracting entity relations. In this article, we focus on going beyond the selective attention mechanisms and propose a new framework termed interaction-and-response network (IR-Net) that adaptively recalibrates the features of sentence, bag, and group levels by explicitly modeling interdependencies among the features on each level. The IR-Net consists of a series of interactive and responsive modules throughout feature hierarchy, seeking to strengthen its power of learning salient discriminative features for distinguishing entity relations. We conduct extensive experiments on three benchmark DSRE datasets, including NYT-10, NYT-16, and Wiki-20m. The experimental results demonstrate that the IR-Net brings obvious improvements in performance when comparing ten state-of-the-art DSRE methods for entity relation extraction. Wei Song 0008, Weishuai Gu, Fuxin Zhu, Soon Cheol Park |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Particle Search Control Network for Dynamic OptimizationabstractIn dynamic optimization problems (DOPs), environmental changes can be characterized as various dynamics. Faced with different dynamics, existing dynamic optimization algorithms (DOAs) are difficult to tackle, because they are incapable of learning in each environment to control the search. Besides, diversity loss is a critical issue in solving DOPs. Maintaining a high-diversity over dynamic environments is reasonable as it can address such an issue automatically. In this article, we propose a particle search control network (PSCN) to maintain a high-diversity over time and control two key search actions of each input individual, i.e., locating the local learning target and adjusting the local acceleration coefficient. Specifically, PSCN adequately considers the diversity to generate subpopulations located by hidden node centers, where each center is assessed by significance-based criteria and distance-based criteria. The former enable a small intrasubpopulation distance and a big search scope (subpopulation width) for each subpopulation, while the latter make each center distant from other existing centers. In each subpopulation, the best-found position is selected as the local learning target. In the output layer, PSCN determines the action of adjusting the local acceleration coefficient of each individual. Reinforcement learning is introduced to obtain the desired output of PSCN, enabling the network to control the search by learning in different iterations of each environment. The experimental results especially performance comparisons with eight state-of-the-art DOAs demonstrate that PSCN brings significant improvements in performance of solving DOPs. Wei Song 0008, Shaocong Liu, Xiaofeng Ding 0001, Yinan Guo 0001, Shengxiang Yang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2024 | HEU-Net: hybrid attention residual block-based network with external skip connections for metal corrosion semantic segmentation
Tiancheng Zhu, Shiqiang Zhu, Hongliang Ding, Wei Song 0008, Cunjun Li |
Vis. Comput. | 5 |
| 2023 | Scale-Aware Graph Convolutional Network for Fine-Grained Image ClassificationabstractFine-grained Image Classification (FGIC) is a hot research topic in computer vision. Currently, FGIC faces several challenges, such as similar appearances, cluttered backgrounds, and pose variations. To effectively address these challenges, we propose a framework called Scale-Aware Graph Convolutional Network (SAGCN) to capture subtle differences in images. Leveraging the characteristics of fine-grained images, we design two core modules, namely Scale-Aware Selection Module (SASM) and Spatial Semantic Correlation Module (SSCM). SASM aggregates multi-scale information of fine-grained images by fusing features from multiple layers. SSCM establishes semantic-spatial relationships by propagating information among different parts of the fine-grained image. Furthermore, we propose a Pairwise Appearance Similarity Loss (PAS-Loss) to distinguish easily confused categories. Extensive experiments demonstrate that our method achieves state-of-the-art results on benchmark datasets. Wei Song 0008 |
ICIS | 2 |
| 2023 | Fast Contextual Scene Graph Generation with Unbiased Context AugmentationabstractScene graph generation (SGG) methods have historically suffered from long-tail bias and slow inference speed. In this paper, we notice that humans can analyze relationships between objects relying solely on context descriptions, and this abstract cognitive process may be guided by experience. For example, given descriptions of cup and table with their spatial locations, humans can speculate possible relationshipsor. Even without visual appearance information, some impossible predicates like flying in and looking at can be empirically excluded. Accordingly, we propose a contextual scene graph generation (C-SGG) method without using visual information and introduce a context augmentation method. We propose that slight perturbations in the position and size of objects do not essentially affect the relationship between objects. Therefore, at the context level, we can produce diverse context descriptions by using a context augmentation method based on the original dataset. These diverse context descriptions can be used for unbiased training of C-SGG to alleviate long-tail bias. In addition, we also introduce a context guided visual scene graph generation (CV-SGG) method, which leverages the C-SGG experience to guide vision to focus on possible predicates. Through extensive experiments on the publicly available dataset, C-SGG alleviates long-tail bias and omits the huge computation of visual feature extraction to realize real-time SGG. CV-SGG achieves a great trade-off between common predicates and tail predicates. Tianlei Jin, Fangtai Guo, Qiwei Meng, Shiqiang Zhu, Xiangming Xi, Wen Wang 0017, Zonghao Mu, Wei Song 0008 |
CVPR | 8 |
| 2023 | FLASH: Low-Latency Serverless Model Inference with Multi-Core Parallelism in EdgeabstractLow response latency holds a pivotal role in the landscape of edge deep learning model inference, yet the constrained resources within edge computing environments often limit its full potential. Within the scope of this paper, we substantiate that even amid the resource limitations inherent to edge computing environments, it remains feasible to curtail model response latency by elevating multi-core parallel efficiency. Our research encompasses a comprehensive analysis of the parallel acceleration effects observed across models featuring diverse parameter magnitudes during the inference process. This analysis culminates in the development of FLASH, an online deep model inference system tailored for Serverless edge inference, strategically optimized through the utilization of multi-core parallelism. FLASH exhibits dynamic adaptability by modulating the number of CPU cores within computational instances in accordance with traffic request loads. It also employs a dynamic scaling mechanism to finely adjust model placement, ultimately facilitating inference acceleration and mitigating the concomitant cold start overhead. Empirical experimentation conducted across a spectrum of burst-level workloads serves to underscore FLASH’s capacity, resulting in an average reduction in response latency by 33% and a maximum reduction of 75%, while concurrently realizing a throughput enhancement of 2.94x. Yanying Lin, Yingfei Tang, Wei Song 0008, Kejiang Ye |
ICPADS | 6 |
| 2023 | KGNet: Knowledge-Guided Networks for Category-Level 6D Object Pose and Size EstimationabstractDespite the giant leap made in object 6D pose estimation and robotic grasping under structured scenarios, most approaches depend heavily on the exact CAD models of target objects beforehand, thereby limiting their wide applications. To address this, we propose a novel knowledge-guided network - KGNet to estimate the pose and size of category-level unseen objects. This network includes three primary innovations: knowledge-guided categorical model generation, pointwise deformation probability matrix and synergetic RGBD feature fusion, with the former two leveraging categorical object knowledge for unseen object reconstruction and the latter one facilitating pose-sensitive feature extraction. Exten-sive experiments on CAMERA25 and REAL275 verify their effectiveness, and KGNet achieves the SOTA performance on these two acknowledged benchmarks. Additionally, a real-world robotic grasping experiment is conducted, and its results further qualitatively prove the practicability and robustness of KGNet. Qiwei Meng, Jason Gu, Shiqiang Zhu, Jianfeng Liao, Tianlei Jin, Fangtai Guo, Wen Wang 0017, Wei Song 0008 |
ICRA | 8 |
| 2023 | RFFCE: Residual Feature Fusion and Confidence Evaluation Network for 6DoF Pose EstimationabstractIn this paper, we propose a novel RGBD-based object 6DoF pose estimation network - RFFCE. It is a two-stage method that firstly leverages deep neural networks for feature extraction and object points matching, and then the geometric principles are utilized for final pose computation. Our approach consists of three primary innovations: residual feature fusion for representative RGBD feature extraction; confidence evaluation and confidence-based paired points offsets regression for self-evaluation and self-optimization respectively. Their effectiveness is verified through an ablation study, and our RFFCE achieves the SOTA performance on LineMOD, Occlusion-LineMOD and YCB-Video datasets. Additionally, we also conduct a real-world object grasping experiment for visualization and qualitative evaluation of the RFFCE. Qiwei Meng, Shanshan Ji, Shiqiang Zhu, Tianlei Jin, Jason Gu, Wei Song 0008 |
ICRA | 7 |
| 2023 | Hierarchical Knowledge Transfer Network for Distantly Supervised Relation ExtractionabstractDistantly supervised relation extraction (DSRE) aims to identify the relation between the two entities (e.g. name and location). Most existing methods extract semantic features from each level separately, without taking into account the transfer of hierarchical knowledge obtained at various levels. As a result, a large amount of knowledge that can improve the quality of the feature representations is lost, resulting in decreased performance for predicting entity relations. In this paper, we propose a novel framework termed the Hierarchical Knowledge Transfer Network (HKTN) that is capable of transferring hierarchical knowledge learned from different levels to improve the performance of predicting entity relations. Specifically, the two representation refinement blocks with re-calibrators at the bag and group levels construct robust bag features and comprehensive group features, respectively. During the construction process, the high-level features are capable of guiding the learning of the bottom-level features using the two re-calibrators. As the construction of the high-level feature representations is based on the bottom-level feature representations, prediction-based contrastive learning fully excavates bottom-level features, which can improve the quality of the feature representation at each level. The experimental results demonstrate that our proposed HKTN achieves an obvious improvement on the two benchmark datasets, including NYT-10 and GDS. Wei Song 0008, Weishuai Gu |
IJCNN | 1 |
| 2023 | Towards Safe and Aggressive Motion Generation for Dynamic Targets Pick-and-PlaceabstractIn this paper, we present a framework to generate time-optimal trajectories for dynamic target pick-and-place tasks. We develop an optimization-based trajectory generation method for manipulators, which can conduct spatial-temporal deformation under user-defined requirements. We formulate the problem of dynamic target pick-and-place, in which the trajectory duration and jerk are optimized and terminal states are adjusted instead of being fixed. The motions are constrained within the mechanical limits and to avoid collisions. Constraints transcription is adopted to convert constraints to weighted penalties. Then the problem can be solved based on the trajectory generation method with a high-level optimizer. We integrate the proposed method with online perception into a robot arm platform, in which a conveyor belt is used to transport the objects. Simulations and real-world experiments are conducted under a range of object speeds. Results show that the proposed method achieves online grasping under the object velocity up to 0.5m/s with an average computing time of 190ms. Jianfeng Liao, Shiqiang Zhu, Wei Song 0008, Yinchun Huang |
IROS | 6 |
| 2023 | Predictive hierarchical reinforcement learning for path-efficient mapless navigation with moving target
Biao Luo 0001, Wei Song 0008, Chunhua Yang 0001 |
Neural Networks | 3 |
| 2023 | B2C-AFM: Bi-Directional Co-Temporal and Cross-Spatial Attention Fusion Model for Human Action RecognitionabstractHuman Action Recognition plays a driving engine of many human-computer interaction applications. Most current researches focus on improving the model generalization by integrating multiple homogeneous modalities, including RGB images, human poses, and optical flows. Furthermore, contextual interactions and out-of-context sign languages have been validated to depend on scene category and human per se. Those attempts to integrate appearance features and human poses have shown positive results. However, with human poses' spatial errors and temporal ambiguities, existing methods are subject to poor scalability, limited robustness, and sub-optimal models. In this paper, inspired by the assumption that different modalities may maintain temporal consistency and spatial complementarity, we present a novel Bi-directional Co-temporal and Cross-spatial Attention Fusion Model (B2C-AFM). Our model is characterized by the asynchronous fusion strategy of multi-modal features along temporal and spatial dimensions. Besides, the novel explicit motion-oriented pose representations called Limb Flow Fields (Lff) are explored to alleviate the temporal ambiguity regarding human poses. Experiments on publicly available datasets validate our contributions. Abundant ablation studies experimentally show that B2C-AFM achieves robust performance across seen and unseen human actions. The codes are available at https://github.com/gftww/B2C.git. Fangtai Guo, Tianlei Jin, Shiqiang Zhu, Xiangming Xi, Wen Wang 0017, Qiwei Meng, Wei Song 0008, Jiakai Zhu |
IEEE Trans. Image Process. | 7 |
| 2022 | Independent Relationship Detection for Real-Time Scene Graph Generation
Tianlei Jin, Wen Wang 0017, Shiqiang Zhu, Xiangming Xi, Qiwei Meng, Zonghao Mu, Wei Song 0008 |
ICONIP (4) | 7 |
| 2022 | Depth-aware gaze-following via auxiliary networks for roboticsabstractGaze-Following aims to predict the gaze target of a subject within an image, and information on orientation and depth greatly improves this task. However, previous methods require additional datasets to obtain depth or orientation information, leading to cumbersome training or inference processes. To this end, we propose an end-to-end depth-aware gaze-following approach that incorporates depth and orientation information without additional datasets. Our approach identifies a primary task, gaze-following, supervised by true labels from the gaze-following dataset and two auxiliary tasks, scene depth estimation and 3D orientation estimation, supervised by generated pseudo labels. Intermediate auxiliary features are integrated into the primary task network as implicit information. We propose a residual filter module for screening useful information that can enhance gaze-following prediction performance. Extensive experiments on GazeFollow and VideoAttentionTarget show that our approach achieves state-of-the-art results (0.120 Ave. Dist. achieved on GazeFollow and 0.104 L2 Dist. achieved on VideoAttentionTarget). Finally, we apply our approach to a real robot for understanding human attention and intention. Compared to the previous depth considered gaze-following method, our method saves half of the computation time. Tianlei Jin, Qizhi Yu, Shiqiang Zhu, Zheyuan Lin, Yuanhai Zhou, Wei Song 0008 |
Eng. Appl. Artif. Intell. | 7 |
| 2021 | A novel deep auto-encoder considering energy and label constraints for categorization
Wei Song 0008, Soon Cheol Park |
Expert Syst. Appl. | 1 |
| 2021 | A new deep auto-encoder using multiscale reconstruction errors and weight update correlation
Wei Song 0008, Wei Li 0121, Ziyu Hua, Fuxin Zhu |
Inf. Sci. | 1 |
| 2021 | Estimating the Optimal Number of Clusters Via Internal Validity Index
Shibing Zhou, Fei Liu 0001, Wei Song 0008 |
Neural Process. Lett. | 3 |
| 2018 | Taking advantage of multi-regions-based diagonal texture structure descriptor for image retrieval
Wei Song 0008, Yubing Zhang, Fei Liu 0001, ZhiLei Chai, Feng Ding 0001, Xuezhong Qian, Soon Cheol Park |
Expert Syst. Appl. | 1 |
| 2017 | Particle swarm optimization algorithm with environmental factors for clustering analysis
Wei Song 0008, Yingying Qiao |
Soft Comput. | 1 |
| 2011 | Fuzzy evolutionary optimization modeling and its applications to unsupervised categorization and extractive summarization
Wei Song 0008, Lim Cheon Choi, Soon Cheol Park, Xiaofeng Ding 0001 |
Expert Syst. Appl. | 1 |