Wenming Zhe

dblp:305/2067 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0003-1753-5784ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Single-Frame Supervision for Spatio-Temporal Video Grounding
abstract
Spatio-Temporal Video Grounding (STVG) aims at localizing the spatio-temporal tube of a specific object in an untrimmed video given a free-form natural language query. As the annotation of tubes is labor intensive, researchers are motivated to explore weakly supervised approaches in recent works, which usually results in significant performance degradation. To achieve a less expensive STVG method with acceptable accuracy, this work investigates the "single-frame supervision" paradigm that requires a single frame labeled with a bounding box within the temporal boundary of the fully supervised counterpart as the supervisory signal. Based on the characteristics of the STVG problem, we propose a Two-Stage Multiple Instance Learning (T-SMILE) method, which creates pseudo labels by expanding the annotated frame to its contextual frames, thereby establishing a fully-supervised problem to facilitate further model training. The innovations of the proposed method are three-folded, including 1) utilizing multiple instance learning to dynamically select instances in positive bags for the recognition of starting and ending timestamps, 2) learning highly discriminative query features by incorporating spatial prior constraints in cross-attention, and 3) designing a curriculum learning-based strategy that iterative assigns dynamic weights to spatial and temporal branches, thereby gradually adapting to the learning branch with larger difficulty. To facilitate future research on this task, we also contribute a large-scale benchmark containing 12,469 videos on complex scenes with single-frame annotation. The extensive experiments on two benchmarks demonstrate that T-SMILE significantly outperforms all weakly-supervised methods. Remarkably, it also performs better than some fully-supervised methods associated with much more annotation labor costs.
Kun Liu 0016, Mengxue Qu, Yang Liu 0235, Yunchao Wei, Wenming Zhe, Yao Zhao 0001, Wu Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 ALNS Framework for Platform-Based Vehicle Scheduling in Logistics Industrial Park CPSs
abstract
The limitation of platform resources is usually the bottleneck of loading/unloading operations in logistics. Hence, the strategies for utilizing platforms significantly impact the efficiency of logistics, which further affects the throughput of logistics industrial parks. In order to improve the efficiency of logistics management and operation, platform-based vehicle scheduling is studied in this work, which aims to appropriately assign the vehicles for loading/unloading to specific platforms and time ranges, thereby optimizing the usage of platforms and avoiding unnecessary waiting. The recent trend of Industrial 4.0 has further promoted the applications of the cyber-physical system (CPS) infrastructure in modern logistics systems, allowing the scheduling based on the real-time states of entities and resources. However, this also poses new challenges such that the decisions to support large-scale scheduling tasks need to be made within limited time. In order to tackle these challenges and enable real-time decision making, this work proposes an innovative scheduling algorithm based on the adaptive large neighborhood search (ALNS) framework to optimize the operation time in the logistics industrial park with high efficiency. The proposed algorithm is compared against the Gurobi professional solver based on real-world data for vehicle scheduling in logistics indus-trial parks, where it is demonstrated to significantly outperform Gurobi on the efficiency and scalability in both day-ahead and real-time scenarios without any degradation of solution quality.
Botong Liu, Yuqiao Wang, Wenming Zhe
INDIN4
2022 Enhancing Vehicle State Recognition in Logistics Industrial Parks via Dynamic Hidden Markov Model
abstract
Platform-based vehicle recognition is a critical task in logistics scenarios that facilitates the efficient management of resources. Although recent advances in the computer vision domain can be conveniently adopted to recognize the identities of vehicles and the occupations of platforms, the efficacy is significantly compromised by the severe interference and noise at the platforms of logistics industrial parks. This work tackles these difficulties through concentrating on the sequential characteristics of vehicles during arrival and departure. An innovative dynamic hidden Markov model (DHMM) is proposed to estimate the real sequence of vehicle states from the noisy observations. A dynamic Viterbi algorithm is also developed to solve the proposed DHMM method with high efficiency. The proposed method is evaluated against multiple baselines through experiments, where it can recognize the vehicle states with high accuracy and is demonstrated to significantly outperform the baselines when the interference is strong.
Yang Liu 0064, Mingjie Guo, Shiyan Hu 0001, Wenming Zhe
ETFA4
2022 Three-Stage Root Cause Analysis for Logistics Time Efficiency via Explainable Machine Learning
abstract
The performance of logistics highly depends on the time efficiency, and hence, plenty of efforts have been devoted to ensuring the on-time delivery in modern logistics industry. However, the delay in logistics transportation and delivery can still happen due to various practical issues, which significantly impact the quality of logistics service. In order to address this issue, this work investigates the root causes impacting the time efficiency, thereby facilitating the operation of logistics systems such that resources can be appropriately allocated to improve the performance. The proposed solution comprises three stages, where statistical methods are employed in the first stage to analyze the pattern of on-time delivery rate and detect the abnormalities induced by non-ideality of operations. Subsequently, a machine learning model is trained to capture the underlying correlations between time efficiency and potential impacting factors. Finally, explainable machine learning techniques are utilized to quantify the contributions of the impacting factors to the time efficiency, thereby recognizing the root causes. The proposed method is comprehensively studied on the real JD Logistics data through experiments, where it can identify the root causes that impact the time efficiency of logistics delivery with high accuracy. Furthermore, it is also demonstrated to outperform the baselines including a recent state-of-the-art method.
Shiqi Hao, Yang Liu 0064, Yuan Wang 0058, Wenming Zhe
KDD5
2021 Computer Vision Based Conveyor Belt Congestion Recognition in Logistics Industrial Parks
abstract
Various automatic and intelligent technologies have been employed to facilitate the logistics operations in recent years promoted by the trend of industry 4.0. As the most frequently used automatic equipment in the logistics industrial parks, conveyor belt plays a critical role on the efficient sorting of packages. Due to the reasons like non-ideality of scheduling and inappropriate operations, conveyor belts can potentially be impacted by congestion, thereby inducing a series of consequences such as delay, lose and damage of packages. In order to tackle these issues, a computer vision-based method is proposed to recognize the congestion on conveyor belts. Other than the popular deep learning-based techniques, the proposed method comprehensively analyzes the characteristics of conveyor belt congestion using statistical approaches and extract informative features for decision making. Finally, the proposed method is evaluated on the data collected from real package sorting scenarios, where it outperforms the deep learning and conventional pattern recognition-based methods on both detection accuracy and capability of generalization.
Yingchun Niu, Yang Liu 0064, Li Zheng 0002, Wenming Zhe
ETFA6
2021 A Model Fusion Approach for Goods Information Inspection in Dual-Platform E-Commerce Systems
abstract
In nowadays, the large-scale e-commerce corporations tend to operate their own logistics networks to guarantee the speed, safety and economic efficiency of goods delivery. Thus, an e-commerce corporation usually needs to manage an online retailing platform and a logistics platform simultaneously. Despite the offered convenience, the cross-platform management faces grand challenges on security. The inappropriate philosophies for maintaining goods information and malicious behaviors of some merchants may induce mismatch between the goods and corresponding information, which consequently leads to incorrect delivery and the degradation of customer experience. In order to tackle this issue, an innovative model fusion method is proposed in this work for goods information inspection. It investigates the advantages of multiple natural language processing models as well as domain knowledge to extract informative text features, which are subsequently fed into a multi-layer perceptron for final decision on whether the goods information is accurate. Unlike the recent popular deep architectures, the proposed method leverages the complimentary effect of features from different sources utilizing a wide structure to achieve a superior inspection accuracy. Finally, the proposed method is validated using real JD Logistics data and is demonstrated to outperform the existing techniques. Furthermore, the experimental results also demonstrate that leveraging the complimentary effects can bring additional improvement compared to merely exploring deeper.
Yang Liu 0064, Zhuozhuo Zhao, Wenming Zhe
ETFA6