VLDB 2026 Research / reviewers in the wild / expert
Fengyu Zhou 0002
dblp:13/7780-2
· DBLP profile ↗
32ranked-venue papers
1as first author
27since 2021 · last 2026
0000-0001-5140-7036ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 13 since 2021Databases, data management, data science and information retrieval · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fusing context and path with efficient negative sampling and retrieval for inductive link prediction
Guannan Si, Linnan Lu, Mingshen Li, Fengyu Zhou 0002 |
Expert Syst. Appl. | 5 |
| 2026 | STVRM : Spatio-temporal relational modeling with vision transformer for dynamic scene graph generation
Linnan Lu, Guannan Si, Mingshen Li, Fengyu Zhou 0002 |
Expert Syst. Appl. | 5 |
| 2025 | Generating fault signals for mobile robots based on multimodal knowledge and multi-channel correlation generative adversarial networkabstractThe imbalanced data limit the effectiveness of mobile robot fault diagnosis, while generating pseudo multi-sensor signals of mobile robot is an effective solution. However, existing generative methods often fail to balance the differences and correlations among channels across multi-sensor signals. To address these issues, a novel multimodal knowledge and multi-channel correlation generative adversarial network (MKMCGAN) is proposed to generate high-quality fault signals. Specifically, wavelet packet decomposition (WPD) are used to extract time-frequency features for each channel, then multi generator-discriminator pair strategy (MGDS) and a time-frequency analysis knowledge module (TFKM) are designed to bring higher similarity between the generated signal and the real signal. Subsequently, we construct a sensor data association graph and design a prior knowledge correlation module (PKM), which effectively consider the impact of inter-channel correlations on generated signals. Eventually, a novel multi-channel correlation generative adversarial network is proposed to extract time-frequency features and consider inter-channel correlations, which can generate high-quality fault signals. The effectiveness of MKMCGAN is thoroughly validated on datasets collected from a real robot fault diagnosis test bench. Experimental results indicate that MKMCGAN generates higher-quality signals compared to state-of-the-art methods. Xinyang Cui, Fengyu Zhou 0002, Longda Zhang, Xianfeng Yuan |
Adv. Eng. Informatics | 2 |
| 2025 | MGTN-DSI: A multi-sensor graph transfer network considering dual structural information for fault diagnosis under varying working conditions
Jianjie Liu, Xianfeng Yuan, Xilin Yang, Tianyi Ye, Xinxin Yao, Fengyu Zhou 0002 |
Adv. Eng. Informatics | 8 |
| 2025 | Towards dual-perspective alignment: A novel hierarchical selective adversarial network for transfer fault diagnosis
Xianfeng Yuan, Xilin Yang, Xinxin Yao, Jianjie Liu, Fengyu Zhou 0002, Peng Duan 0002 |
Adv. Eng. Informatics | 6 |
| 2025 | Protecting the interests of owners of intelligent fault diagnosis models: A style relationship-preserving privacy protection method
Xilin Yang, Xianfeng Yuan, Xinxin Yao, Jianjie Liu, Fengyu Zhou 0002 |
Expert Syst. Appl. | 6 |
| 2025 | EKCA-Cap: Improving dense captioning via external knowledge and context awareness
Zhenwei Zhu, Fengyu Zhou 0002, Saike Huang |
Neurocomputing | 2 |
| 2025 | Listen, Perceive, Grasp: CLIP-Driven Attribute-Aware Network for Language-Conditioned Visual Segmentation and GraspingabstractEndowing robots with the ability to understand natural language and execute grasping is a challenging task in a human-centric environment. Existing works on language-conditioned grasping achieve end-to-end grasping detection based on language. However, these works lack fine-grained visual grounding, resulting in cognitive deficits for robots. Moreover, they ignore the correlation between visual attributes of objects and grasping, leading to coarse grasp poses. To this end, we propose a CLIP-driven aTtribute-aware network (CTNet) for language-conditioned visual segmentation and grasping, enabling the robots to listen, perceive, and grasp the referred object in real-world applications. Specifically, we first employ Listen stage to understand basic linguistic and visual concepts. Subsequently, we introduce Perceive stage to mine multi-modal features and visual attribute cues (e.g., boundary and spatial location), then yield a language-conditioned segmentation mask. Further, we design Grasp stage to aggregate the perceived attribute information and refine the spatial location and grasping rectangle, generating a high-quality grasp pose. Lastly, we provide an extended large dataset Ref-OCID-Grasp to train and test our method, achieving a grasping accuracy of 97.76% and segmentation OIoU of 91.82%. The real-world robotic applications demonstrate the effectiveness of our proposed approach. The project, video, and dataset can be found athttps://ctnetgrasp.github.io. Note to Practitioners—Most of the existing grasping methods focus on clearing all objects in the workspace. However, as robots integrate into human society, robots should learn to grasp the desired object by understanding human language. Therefore, language-conditioned grasping is a significant skill for human-robot collaboration. The prior works directly complete the grasp detection through the language-grasp paradigm, but they ignore the discussion on whether the robot understands the concept of vision and language expression of the object. Therefore, this paper proposed the Listen-Perceive-Grasp paradigm, in which the Listen-Perceive stage is responsible for the conception alignment of the object in language expression and visual pixels, and the Perceive-Grasp stage achieves the constraining and refining the grasp detection by the perceived visual attributes such as boundary and shape. Experiments show that this method can obtain a refiner grasp pose in cluttered environments and perform language-conditioned grasping well in the real world. In future research, we will work on 6-DoF grasping and multi-object disambiguation conditioned on language. Jialong Xie, Jin Liu 0018, Saike Huang, Chaoqun Wang 0009, Fengyu Zhou 0002 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | HOGN-TVGN: Human-inspired Embodied Object Goal Navigation based on Time-varying Knowledge Graph Inference Networks for Robots
Baojiang Yang, Xianfeng Yuan, Zhongmou Ying, Boyi Song, Yong Song 0005, Fengyu Zhou 0002, Weihua Sheng |
Adv. Eng. Informatics | 7 |
| 2024 | Fault diagnosis of mobile robot based on dual-graph convolutional network with prior fault knowledge
Longda Zhang, Fengyu Zhou 0002, Peng Duan 0002, Xianfeng Yuan |
Adv. Eng. Informatics | 2 |
| 2024 | Complete feature learning and consistent relation modeling for few-shot knowledge graph completion
Jin Liu 0018, Chongfeng Fan, Fengyu Zhou 0002, Huijuan Xu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Joint multimodal entity-relation extraction based on temporal enhancement and similarity-gated attention
Guoxiang Wang, Jin Liu 0018, Jialong Xie, Zhenwei Zhu, Fengyu Zhou 0002 |
Knowl. Based Syst. | 5 |
| 2024 | SATR: Semantics-Aware Triadic Refinement network for referring image segmentation
Jialong Xie, Jin Liu 0018, Guoxiang Wang, Fengyu Zhou 0002 |
Knowl. Based Syst. | 4 |
| 2024 | Triadic temporal-semantic alignment for weakly-supervised video moment retrieval
Jin Liu 0018, Jialong Xie, Fengyu Zhou 0002, Shengfeng He |
Pattern Recognit. | 3 |
| 2024 | Video Question Answering With Semantic Disentanglement and ReasoningabstractVideo question answering aims to provide correct answers given complex videos and related questions, posting high requirements of the comprehension ability in both video and language processing. Existing works phrase this task as a multi-modal fusion process by aligning the video context with the whole question, ignoring the rich semantic details of nouns and verbs separately in the multi-modal reasoning process to derive the final answer. To fill this gap, in addition to the semantic alignment of the whole sentence, we propose to disentangle the semantic understanding of language, and reason over the corresponding frame-level and motion-level video features. We design an unified multi-granularity language module of residual structure to adapt the semantic understanding at different granularity with context exchange, e.g., word-level and sentence-level. To enhance the holistic question understanding for answer prediction, we also design a contrastive sampling approach by selecting irrelevant questions as negative samples to break the intrinsic correlations between questions and answers within the dataset. Notably, our model is competent for both multiple-choice and open-ended video question answering. We further employ a pre-trained language model to retrieve relevant knowledge as candidate answer context to facilitate open-ended VideoQA. Extensive quantitative and qualitative experiments on four public datasets (NextQA, MSVD, MSRVTT, and TGIF-QA-R) demonstrate the effective and superior performance of our proposed model. Our code will be released upon the paper’s acceptance. Jin Liu 0018, Guoxiang Wang, Jialong Xie, Fengyu Zhou 0002, Huijuan Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Question Type-Aware Debiasing for Test-Time Visual Question Answering Model AdaptationabstractIn Visual Question Answering (VQA), addressing language prior bias, where models excessively rely on superficial correlations between questions and answers, is crucial. This issue becomes more pronounced in real-world applications with diverse domains and varied question-answer distributions during testing. To tackle this challenge, Test-time Adaptation (TTA) has emerged, allowing pre-trained VQA models to adapt using unlabeled test samples. Current state-of-the-art models select reliable test samples based on fixed entropy thresholds and employ self-supervised debiasing techniques. However, these methods struggle with diverse answer spaces linked to different question types and may fail to identify biased samples that still leverage relevant visual context. In this paper, we propose Question type-guided Entropy Minimization and Debiasing (QED) as a solution for test-time VQA model adaptation. Our approach involves adaptive entropy minimization based on question types to improve the identification of fine-grained and unreliable samples. Additionally, we generate negative samples for each test sample and label them as biased if their answer entropy change rate significantly differs from positive test samples, subsequently removing them. We evaluate our approach on two public benchmarks, VQA-CP v2, and VQA-CP v1, and achieve new state-of-the-art results, with overall accuracy rates of 48.13% and 46.18%, respectively. Jin Liu 0018, Jialong Xie, Fengyu Zhou 0002, Shengfeng He |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Accurate and Efficient 3D Panoptic Mapping Using Diverse Information Modalities and Multidimensional Data Associationabstract3D Panoptic perception is essential for the understanding of real-world environment and plays an increasingly important role in the field of robotics. However, most existing methods heavily rely on image panoptic segmentation networks to acquire panoptic information of the environment, which is time-consuming and susceptible to interference. In this paper, we propose a novel and efficient panoptic mapping method based on multi-source information. Specifically, to improve the real-time performance of the system, we first apply lightweight object detection and semantic segmentation to extract 2D semantic and instance information from images. Second, a panoptic inference algorithm is designed that fully utilizes multi-source information, including geometry-based and learning-based information, to simultaneously reason about background and foreground objects in the environment. Finally, we take advantage of the scalability of the framework by introducing a multi-object tracking algorithm into the framework, thus providing the temporal information among consecutive frames to the data association module. Based on two popular datasets, extensive comparison experiments are conducted to illustrate the effectiveness of the proposed method. Experimental results show that compared with state-of-the-art panoptic mapping methods, the proposed method achieves superior performance in accuracy, real-timeness and stability. Furthermore, we also evaluate our method in real-world scenarios and CPU-only device to demonstrate the feasibility of its practical deployment. Zhongmou Ying, Xianfeng Yuan, Boyi Song, Yong Song 0005, Fengyu Zhou 0002, Weihua Sheng |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A Novel Data Augmentation Method Based on Denoising Diffusion Probabilistic Model for Fault Diagnosis Under Imbalanced DataabstractImbalanced data constitute a significant challenge in intelligent fault diagnosis cases because they can result in degraded diagnosis accuracy, which can in turn jeopardize the safety and reliability of industrial equipment. Generative adversarial networks (GANs) have been effectively used as common data augmentation methods to address this issue. However, their training process is difficult to perform and prone to mode collapse. Therefore, this article proposes a novel data augmentation method grounded in a diffusion model. The proposed method generates samples through physical simulation rather than adversarial training, which avoids the instability and mode collapse issues faced by GANs, leading to a more stable training process. Moreover, the proposed method utilizes the characteristics of gradual diffusion and random sampling to enhance the authenticity and diversity of sample generation. In addition, in terms of evaluating generation models, most existing works do not have a unified and thorough evaluation framework. Therefore, a comprehensive evaluation framework is proposed to effectively and comprehensively evaluate the performance of data augmentation models. Finally, the proposed method is evaluated using an open-source dataset and two actual testbeds to validate its effectiveness. The experimental results show that our method can generate higher quality and more diverse pseudosamples, and achieve superior fault diagnosis performance under imbalanced data. Specifically, our approach achieves diagnosis accuracies of 97.00%, 96.48%, and 98.30% on the three different datasets, all of which are superior to those of the compared state-of-the-art data augmentation algorithms. Xiongyan Yang, Tianyi Ye, Xianfeng Yuan, Xiaoxue Mei, Fengyu Zhou 0002 |
IEEE Trans. Ind. Informatics | 6 |
| 2023 | GranCATs: Cross-Lingual Enhancement through Granularity-Specific Contrastive AdaptersabstractMultilingual language models (MLLMs) have demonstrated remarkable success in various cross-lingual downstream tasks, facilitating the transfer of knowledge across numerous languages, whereas this transfer is not universally effective. Our study reveals that while existing MLLMs like mBERT can capturephrase-level alignments across the language families, they struggle to effectively capturesentence-level andparagraph-level alignments. To address this limitation, we propose GranCATs, Granularity-specific Contrastive AdapTers. We collect a new dataset that observes each sample at three distinct levels of granularity and employ contrastive learning as a pre-training task to train GranCATs on this dataset. Our objective is to enhance MLLMs' adaptation to a broader range of cross-lingual tasks by equipping them with improved capabilities to capture global information at different levels of granularity. Extensive experiments show that MLLMs with GranCATs yield significant performance advancements across various language tasks with different text granularities, including entity alignment, relation extraction, sentence classification and retrieval, and question-answering. These results validate the effectiveness of our proposed GranCATs in enhancing cross-lingual alignments across various text granularities and effectively transferring this knowledge to downstream tasks. Meizhen Liu 0001, Jiakai He, Xu Guo 0002, Jianye Chen, Siu Cheung Hui, Fengyu Zhou 0002 |
CIKM | 6 |
| 2023 | Be flexible! learn to debias by sampling and prompting for robust visual question answering
Jin Liu 0018, Chongfeng Fan, Fengyu Zhou 0002, Huijuan Xu 0001 |
Inf. Process. Manag. | 3 |
| 2023 | Question-conditioned debiasing with focal visual context fusion for visual question answering
Jin Liu 0018, Guoxiang Wang, Chongfeng Fan, Fengyu Zhou 0002, Huijuan Xu 0001 |
Knowl. Based Syst. | 4 |
| 2023 | Dynamic Deployment and Scheduling Strategy for Dual-Service Pooling-Based Hierarchical Cloud Service System in Intelligent BuildingsabstractDue to the excessive concentration of computing resources in the traditional centralized cloud service system, there will be three prominent problems of management confusion, construction cost and network delay. Therefore, we propose to virtualize regional edge computing resources in intelligent buildings as edge service pooling, then presents a hierarchical cloud platform with dual-service pooling structure and a dynamic strategy for the proposed model. The analytic hierarchy process (AHP) based quality of service (QoS) evaluation mechanism and the dynamic normal distribution selection method are adopted for service deployment. And the dynamic inertia particle swarm optimization (DI-PSO) algorithm is employed to realize task scheduling. Furthermore, the cloud platform and existing terminal server group are used to conduct platform structure comparison experiments, and the popular task scheduling algorithms are selected for simulation experiments. Experimental results of platform measurement show that the average service response time of different services can be improved by about 17.3 to 37.4 percent. The average occupancy ratio of computing resources can be reduced by about 5 percent. The simulation results show that the earliest completion time of single task list can be decreased by 11.3 to 20.9 percent, and the makespan of 100 task lists can be improved by 0.3 times. Hongchang Sun, Shengjun Wang, Fengyu Zhou 0002, Meizhen Liu 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2023 | Fault Diagnosis of Wheeled Robot Based on Prior Knowledge and Spatial-Temporal Difference Graph Convolutional NetworkabstractThe critical issue of wheeled robot fault diagnosis is to comprehensively evaluate its health condition using multisensor data, but traditional deep learning-based methods are hard to model the relationships among sensor measurements. Unlike these methods, the graph convolutional network (GCN), which uses the graph-structured data along with the association graph as input, is more efficient for relationship modeling. However, existing GCN-based fault diagnosis methods suffer from the following weaknesses: the association graphs are obtained according to the similarity of data samples or their features, which cannot guarantee accuracy; and these models are focused on spatial correlations and neglect temporal correlations. To address these problems, we propose to construct the association graph based on prior knowledge,i.e., a simplified mathematical model of the wheeled robot. Moreover, we develop a spatial-temporal difference graph convolutional network (STDGCN) for wheeled robot fault diagnosis. This network contains a difference layer that utilizes localized difference properties for feature enhancement, and the spatial-temporal graph convolutional modules are introduced to jointly capture the spatial-temporal correlations. To verify the effectiveness of the STDGCN for fault diagnosis, experiments are carried out, and the results show that the STDGCN achieves superior performance. Zhaoming Miao, Yingxiang Xia, Fengyu Zhou 0002, Xianfeng Yuan |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Knowledge graph attention mechanism for distant supervision neural relation extraction
Meizhen Liu 0001, Fengyu Zhou 0002, Jiakai He |
Knowl. Based Syst. | 2 |
| 2022 | Self-Attention Networks and Adaptive Support Vector Machine for aspect-level sentiment classification
Meizhen Liu 0001, Fengyu Zhou 0002, Jiakai He, Ke Chen 0022, Yang Zhao 0042, Hongchang Sun |
Soft Comput. | 2 |
| 2021 | Co-attention networks based on aspect and context for aspect-level sentiment analysis
Meizhen Liu 0001, Fengyu Zhou 0002, Ke Chen 0022, Yang Zhao 0042 |
Knowl. Based Syst. | 2 |
| 2021 | Heartbeats Classification Using Hybrid Time-Frequency Analysis and Transfer Learning Based on ResNetabstractThe classification of heartbeats is an important method for cardiac arrhythmia analysis. This study proposes a novel heartbeat classification method using hybrid time-frequency analysis and transfer learning based on ResNet-101. The proposed method has the following major advantages over the afore-mentioned methods: it avoids the need for manual features extraction in the traditional machine learning method, and it utilizes 2-D time-frequency diagrams which provide not only frequency and energy information but also preserve the morphological characteristic within the ECG recordings, and it owns enough deep to make better use of performance of CNN. The method deploys a hybrid time-frequency analysis of the Hilbert transform (HT) and the Wigner-Ville distribution (WVD) to transform 1-D ECG recordings into 2-D time-frequency diagrams which were then fed into a transfer learning classifier based on ResNet-101 for two classification tasks (i.e., 5 heartbeat categories assigned by the ANSI/AAMI standard (i.e., N, V, S, Q and F) and 14 original beat kinds of the MIT/BIH arrhythmia database). For 5 heartbeat categories classification, the results show the F1-score of N, V, S, Q and F categories areF$_{N}$0.9899,F$_{V}$0.9845,F$_{S}$0.9376,F$_{Q}$0.9968,F$_{F}$0.8889, respectively, and the overall F1-score is 0.9595 using the combination data balancing. The results show the average values for accuracy, sensitivity, specificity, predictive value and F1-score on test set for 14 beat kinds the MIT-BIH arrhythmia database are 99.75%, 91.36%, 99.85%, 90.81% and 0.9016, respectively. Compared with other methods, the proposed method can yield more accurate results. Yatao Zhang, Shoushui Wei, Fengyu Zhou 0002, Dong Li 0052 |
IEEE J. Biomed. Health Informatics | 4 |
| 2020 | Meta-Learning based prototype-relation network for few-shot classification
Fengyu Zhou 0002, Jin Liu 0018, Lianjie Jiang |
Neurocomputing | 2 |
| 2019 | Hybrid particle swarm optimization with spiral-shaped mechanism for feature selection
Ke Chen 0022, Fengyu Zhou 0002, Xianfeng Yuan |
Expert Syst. Appl. | 2 |
| 2016 | A high precision visual localization sensor and its working methodology for an indoor mobile robotabstractTo overcome the shortcomings of existing robot localization sensors, such as low accuracy and poor robustness, a high precision visual localization system based on infrared-reflective artificial markers is designed and illustrated in detail in this paper. First, the hardware system of the localization sensor is developed. Secondly, we design a novel kind of infrared-reflective artificial marker whose characteristics can be extracted by the acquisition and processing of the infrared image. In addition, a confidence calculation method for marker identification is proposed to obtain the probabilistic localization results. Finally, the autonomous localization of the robot is achieved by calculating the relative pose relation between the robot and the artificial marker based on the perspective-3-point (P3P) visual localization algorithm. Numerous experiments and practical applications show that the designed localization sensor system is immune to the interferences of the illumination and observation angle changes. The precision of the sensor is ±1.94 cm for position localization and ±1.64◦ for angle localization. Therefore, it satisfies perfectly the requirements of localization precision for an indoor mobile robot. Fengyu Zhou 0002, Xianfeng Yuan, Yang Yang 0023, Zhi-fei Jiang, Chen-lei Zhou |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2014 | Social Media as Sensor in Real World: Geolocate User with Microblog
Xueqin Sui, Zhumin Chen, Pengjie Ren, Jun Ma 0001, Fengyu Zhou 0002 |
NLPCC | 6 |
| 2014 | A method of abnormal habits recognition in intelligent space
Guohui Tian, Hao Wu 0065, Fengyu Zhou 0002 |
Eng. Appl. Artif. Intell. | 4 |