VLDB 2026 Research / reviewers in the wild / expert
Zhuo Zhou
dblp:25/10265
· DBLP profile ↗
18ranked-venue papers
5as first author
13since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Spatial pose estimation algorithm for orchard obscured apple for picking robots
Zhuo Zhou, Runying Zhang, Fuzeng Yang |
Expert Syst. Appl. | 2 |
| 2026 | Beyond the Horizon: Decoupling Multi-View UAV Action Recognition via Partial Order TransferabstractAction recognition using uncrewed aerial vehicles (UAVs) faces unique challenges due to substantial view variations along the vertical spatial axis. Unlike ground-based scenarios, UAVs capture actions from diverse altitudes, resulting in pronounced appearance discrepancies and reduced recognition robustness. To address this, we introduce a multi-view formulation tailored for UAV altitudes and empirically uncover a distinctive partial order among views, where recognition accuracy consistently declines as altitude increases. This key observation motivates the proposed Aero Partial Order Guided Network (Aerorder), which explicitly models and exploits the hierarchical structure of UAV views to enhance cross-altitude action recognition. Aerorder comprises three main components: (1) a View Partition (VP) module that groups views by altitude using the head-to-body ratio; (2) an Order-aware Feature Decoupling (OFD) module that disentangles action-relevant and view-specific representations under partial order guidance; and (3) an Action Partial Order Guide (APOG) that progressively transfers knowledge from easier (low-altitude) to harder (high-altitude) views. Extensive experiments on Drone-Action, MOD20, and UAV validate the superiority of Aerorder, achieving consistent improvements over state-of-the-art methods, up to 4.7% and 1.3% gains on Drone-Action and MOD20, respectively. Wenxuan Liu 0008, Zhuo Zhou, Xuemei Jia, Siyuan Yang 0001, Wenxin Huang, Xian Zhong, Chia-Wen Lin |
AAAI | 2 |
| 2026 | Real-time transaction scheduling in multiple-deep tier-to-tier shuttle-based storage and retrieval systems by using deep Q-learning
Zhun Xu, Liyun Xu, Huan Shao 0004, Gang Shang, Zhuo Zhou |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Hierarchical agent architecture-based large-scale AGV cluster real-time motion collaboration control in dynamic and complex production environments
Zhuo Zhou, Liyun Xu, Liqiang Liao, Minghai Yuan, Gang Shang, Zhun Xu |
Expert Syst. Appl. | 1 |
| 2026 | Attention-Emotion Assessment of ASD Children via Representation Learning Based on Cross-Modal Disentanglement and Attention AlignmentabstractComprehensive cognitive-affective profiling for precision ASD assessment is hindered by inadequate computational modeling of emotion-behavior dynamics. This gap persists due to limited methods for reconciling EEG, eye-tracking, and facial expression modalities-critical for capturing ASD heterogeneity drivers. This study proposes the Cognitive-Affective Ability Assessment (CA3) framework, focusing on attention control in conjunction with valence-arousal, to address these limitations through: (1)Ability-Specific Neural Disentanglementextracting attention- and emotion-specific representations from EEG using ET and facial expression recordings as anchors; (2)Symmetrical Cross-Ability Alignmentmodeling contextual dependencies via a symmetric cross-attention mechanism; and (3)Uncertainty-Aware Ability Integrationfusing per-ability predictions using Dirichlet modeling and Dempster-Shafer theory while quantifying confidence. Extensive experiments have been conducted to evaluate CA$^{3}$vs. state-of-the-arts counterparts on the multimodel datasets (EEG, ET and facial expression recordings) with CCNU (27 ASD vs. 30 typically developing (TD) children) and BNU (49 ASD vs. 48 TD), and the results demonstrate that: 1) In the ASD cognitive-affective abilities assessment task, accuracy is improved by 2.7% for attention control and 2.5% for emotion perception, and by 3.5% for cognitive-affective abilities assessment, compared with current state-of-the-art methods. 2) For attention control and valence-arousal, composite ability gaps are measured between ASD and typically developing children: 15.6% and 19.9% for mild ASD, 28.2% and 39.1% for moderate ASD, and 44.4% and 61.3% for severe ASD, respectively. Overall, the framework effectively bridges computational assessment with clinically actionable insights, enabling robust decision-making through personalized cognitive-affective ASD profiles. Zhuo Zhou, Dan Chen 0001, Zhiyi Yang, Tengfei Gao, Jingying Chen 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Prediction of seam tracking errors in the intelligent welding system: A rapid prediction method based on real-time monitoring data
Gang Shang, Liyun Xu, Zufa Li, Lizhen Xiao, Zhuo Zhou, Hanwu He |
Adv. Eng. Informatics | 5 |
| 2025 | Digital-twin-based AGV cluster dynamic scheduling for solar cell production workshop using deep reinforcement learning
Zhuo Zhou, Liyun Xu, Liqiang Liao, Zhun Xu |
Neurocomputing | 1 |
| 2024 | Bi-Causal: Group Activity Recognition via Bidirectional CausalityabstractCurrent approaches in Group Activity Recognition (GAR) predominantly emphasize Human Relations (HRs) while often neglecting the impact of Human-Object Inter-actions (HOIs). This study prioritizes the consideration of both HRs and HOIs, emphasizing their interdependence. Notably, employing Granger Causality Tests reveals the presence of bidirectional causality between HRs and HOIs. Leveraging this insight, we propose a Bidirectional-Causal GAR network. This network establishes a causality commu-nication channel while modeling relations and interactions, enabling reciprocal enhancement between human-object interactions and human relations, ensuring their mutual consistency. Additionally, an Interaction Module is devised to effectively capture the dynamic nature of human-object interactions. Comprehensive experiments conducted on two publicly available datasets showcase the superiority of our proposed method over state-of-the-art approaches. Our project page: https://angzong.github.io/bi-causal.github.io/ Youliang Zhang, Wenxuan Liu 0008, Danni Xu, Zhuo Zhou, Zheng Wang 0007 |
CVPR | 4 |
| 2023 | Joint Optimal Selection of Imaging Time Interval and Imaging Projection Plane Based on Short-Times Fraction Fourier Transform for Bistatic SAR Maritime Ship Target ImagingabstractDifferent sea states make maritime ship targets move three-dimensionally, causing imaging results severely defocused. Therefore, selection of proper imaging time has attracted global attention in the field of ISAR imaging. This paper proposes an imaging time and plane selection algorithm based on short-time fractional Fourier transform (STFrFT). By STFrFT, time-frequency curves of the target echo for different fractional orders are obtained, and probability density functions of frequency distributions are offered to estimate when Doppler frequency of the target changes most stably. Then, different imaging planes of different STFrFT orders are analyzed to select the optimal imaging time moment. Finally, by comparing image contrasts of various imaging lengths, the optimal imaging time is selected. In general, this method can select the optimal imaging moment, imaging length and imaging plane, and simulation verifies the effectiveness of this imaging method. Zhuo Zhou, Junao Li, Qing Yang 0032, Zhongyu Li 0001, Junjie Wu 0001, Jianyu Yang 0001 |
IGARSS | 3 |
| 2023 | Uncovering the Unseen: Discover Hidden Intentions by Micro-Behavior Graph ReasoningabstractThis paper introduces a new and challenging Hidden Intention Discovery (HID) task. Unlike existing intention recognition tasks, which are based on obvious visual representations to identify common intentions for normal behavior, HID focuses on discovering hidden intentions when humans try to hide their intentions for abnormal behavior. HID presents a unique challenge in that hidden intentions lack the obvious visual representations to distinguish them from normal intentions. Fortunately, from a sociological and psychological perspective, we find that the difference between hidden and normal intentions can be reasoned from multiple micro-behaviors, such as gaze, attention, and facial expressions. Therefore, we first discover the relationship between micro-behavior and hidden intentions and use graph structure to reason about hidden intentions. To facilitate research in the field of HID, we also constructed a seminal dataset containing a hidden intention annotation of a typical theft scenario for HID. Extensive experiments show that the proposed network improves performance on the HID task by 9.9% over the state-of-the-art method SBP. Zhuo Zhou, Wenxuan Liu 0008, Danni Xu, Zheng Wang 0007, Jian Zhao 0006 |
ACM Multimedia | 1 |
| 2023 | Dual-Recommendation Disentanglement Network for View Fuzz in Action RecognitionabstractMulti-view action recognition aims to identify action categories from given clues. Existing studies ignore the negative influences of fuzzy views between view and action in disentangling, commonly arising the mistaken recognition results. To this end, we regard the observed image as the composition of the view and action components, and give full play to the advantages of multiple views via the adaptive cooperative representation among these two components, forming a Dual-Recommendation Disentanglement Network (DRDN) for multi-view action recognition. Specifically, 1) For the action, we leverage a multi-level Specific Information Recommendation (SIR) to enhance the interaction among intricate activities and views. SIR offers a more comprehensive representation of activities, measuring the trade-off between global and local information. 2) For the view, we utilize a Pyramid Dynamic Recommendation (PDR) to learn a complete and detailed global representation by transferring features from different views. It is explicitly restricted to resist the fuzzy noise influence, focusing on positive knowledge from other views. Our DRDN aims for complete action and view representation, where PDR directly guides action to disentangle with view features and SIR considers mutual exclusivity of view and action clues. Extensive experiments have indicated that the multi-view action recognition method DRDN we proposed achieves state-of-the-art performance over powerful competitors on several standard benchmarks. The code will be available at https://github.com/51cloud/DRDN. Wenxuan Liu 0008, Xian Zhong, Zhuo Zhou, Kui Jiang, Zheng Wang 0007, Chia-Wen Lin |
IEEE Trans. Image Process. | 3 |
| 2022 | VCD: View-Constraint Disentanglement for Action RecognitionabstractAction recognition is a hot topic in computer vision due to its wide range of applications in urban surveillance. Although some methods are more advanced from an invariant view perspective, those approaches do not perform well for the viewpoint change. To address this issue, one possible solution is tantamount to track the view-invariant representation as it evolves with the performed action. However, the views’ and actions’ performance always complement each other, once simply looking for the view-invariant representation may cause some behavior information to be lost. In this paper, we propose the View-Constraint Disentanglement (VCD) framework for cross-view action recognition. Specifically, Constraint Disentanglement Module (CDM) is utilized to learn an action-invariant representation by discretizing view-specific representation and its normal distribution, which resolves the entangled relationship between view and action. Moreover, a novel Adaptive Distribution Module (ADM) is intended to befit enhance the high-correlation viewpoint variation information and refine the suitable weight. Extensive experiments are conducted on public benchmarks, indicating that our approach achieves better performance than other state-of-the-art approaches. Xian Zhong, Zhuo Zhou, Wenxuan Liu 0008, Kui Jiang, Xuemei Jia, Wenxin Huang, Zheng Wang 0007 |
ICASSP | 2 |
| 2021 | A Robust Method with DropBlock for Face Anti-SpoofingabstractFace anti-spoofing is an important component for reliable face recognition. Previous deep learning methods usually exploit a large network structure and are prone to be overfitting when the face anti-spoofing database is a small scale. And in some occasions, multi-frame information or other auxiliary information is not available. In order to address these issues, we propose an end-to-end deep learning approach with DropBlock layer to distinguish between a fake face and a genuine face, which makes use of frame-only RGB image. In this paper, we evaluate the proposed approach with different settings on three popular databases, i.e., REPLAY-ATTACK, CASIA-FASD, and CASIA-SURF. The proposed approach can achieve competitive results on these databases. The experimental results show that our model is more robust and can learn more generalized features for face anti-spoofing. Zhuo Zhou, Zhenhua Guo 0001 |
IJCNN | 2 |
| 2020 | Face Anti-Spoofing Using Spatial Pyramid PoolingabstractFace recognition system is vulnerable to many kinds of presentation attacks, so how to effectively detect whether the image is from the real face is particularly important. At present, many deep learning-based anti-spoofing methods have been proposed. But these approaches have some limitations, for example, global average pooling (GAP) easily loses local information of faces, single-scale features easily ignore information differences in different scales, while a complex network is prone to be overfitting. In this paper, we propose a face anti-spoofing approach using spatial pyramid pooling (SPP). Firstly, we use ResNet-18 with a small amount of parameter as the basic model to avoid overfitting. Further, we use spatial pyramid pooling module in the single model to enhance local features while fusing multi-scale information. The effectiveness of the proposed method is evaluated on three databases, CASIA-FASD, Replay-Attack and CASIA-SURF. The experimental results show that the proposed approach can achieve state-of-the-art performance. Zhuo Zhou, Zhenhua Guo 0001 |
ICPR | 2 |
| 2016 | An IoT-Based Online Monitoring System for Continuous Steel CastingabstractMonitoring solutions using the Internet of Things (IoT) techniques, can continuously gather sensory data, such as temperature and pressure, and provide abundant information for a monitoring center. Nevertheless, the heterogeneous and massive data bring significant challenges to real-time monitoring and decision making, particularly in time-sensitive industrial environments. This paper presents an online monitoring system based on an IoT system architecture which is composed of four layers: 1) sensing; 2) network; 3) service resource; and 4) application layers. It integrates various data processing techniques including protocol conversion, data filtering, and data conversion. The proposed system has been implemented and demonstrated through a real continuous steel casting production line, and integrated with the TeamCenter platform. Results indicate that the proposed solution well addresses the challenge of heterogeneous data and multiple communication protocols in real-world industrial environments. Feng Zhang 0013, Min Liu 0002, Zhuo Zhou, Weiming Shen 0001 |
IEEE Internet Things J. | 3 |
| 2013 | Quantum ant colony algorithm-based emergency evacuation path choice algorithmabstractThe evacuation path optimization in the disaster area plays an important role in reducing the human and social harm and saving aid time. In this paper, a novel algorithm for emergency evacuation path choice based on quantum ant colony algorithm (QACA) is proposed, and it avoids premature convergence and speeds up the convergence to the global optimal solution. In the proposed algorithm, Q-bit is used to represent the pheromone, and the rotation gate is used to update the pheromone. Simulation results show that the proposed algorithm is feasible and effective. Feng Zhang 0013, Min Liu 0002, Zhuo Zhou, Weiming Shen 0001 |
CSCWD | 3 |
| 2013 | A data processing framework for IoT based online monitoring systemabstractOnline monitoring system for continuous casting equipment is established based on IOT (Internet of things) sensing technology and communication technology. As the system contains a variety of sensor types and data transmission protocols, it will lead to a large amount of heterogeneous data and the data is difficult to integrate with applications in upper layer. A data processing framework is introduced into the system to deal with such problems. The framework focuses on protocol conversion, data processing methods and integration with applications in upper layer. Finally, application in online monitoring system proved the validity of the framework. Zhuo Zhou, Min Liu 0002, Feng Zhang 0013, Li Bai 0003, Weiming Shen 0001 |
CSCWD | 1 |
| 2011 | A Hybrid Quantum-Inspired Particle Swarm Evolution Algorithm and SQP Method for Large-Scale Economic Dispatch Problems
Qun Niu, Zhuo Zhou, Tingting Zeng |
ICIC (3) | 2 |