EDBT 2026 Demo / reviewers in the wild / expert
Jeremy S. Smith
dblp:81/8286
· DBLP profile ↗
36ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0002-0212-2365ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression ComprehensionabstractEmbodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20, 558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar. Runwei Guan, Ruixiao Zhang 0001, Ningwei Ouyang, Ka Lok Man, Xiaohao Cai, Ming Xu 0011, Jeremy S. Smith, Eng Gee Lim, Yutao Yue, Hui Xiong 0001 |
ICRA | 8 |
| 2025 | NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave RadarabstractRecently, visual grounding and multi-sensors setting have been incorporated into perception system for terrestrial autonomous driving systems and Unmanned Surface Vessels (USVs), yet the high complexity of modern learning-based visual grounding model using multi-sensors prevents such model to be deployed on USVs in the real-life. To this end, we design a low-power multi-task model named NanoMVG for waterway embodied perception, guiding both camera and 4D millimeter-wave radar to locate specific object(s) through natural language. NanoMVG can perform both box-level and mask-level visual grounding tasks simultaneously. Compared to other visual grounding models, NanoMVG achieves highly competitive performance on the WaterVG dataset, particularly in harsh environments. Moreover, the real-world experiments with deployment of NanoMVG on embedded edge device of USV demonstrates its fast inference speed for real-time perception and capability of boasting ultra-low power consumption for long endurance. Runwei Guan, Liye Jia, Haocheng Zhao, Shanliang Yao, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Yutao Yue |
IROS | 9 |
| 2025 | Referring flexible image restoration
Runwei Guan, Rongsheng Hu, Zhuhao Zhou, Tianlang Xue, Ka Lok Man, Jeremy S. Smith, Eng Gee Lim, Weiping Ding 0001, Yutao Yue |
Expert Syst. Appl. | 6 |
| 2025 | Exploring interaction concepts for human-object-interaction detection via global- and local-scale enhancing
Tianlun Luo, Qiao Yuan, Boxuan Zhu, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith, Eng Gee Lim |
Neurocomputing | 6 |
| 2025 | Simple yet effective: An explicit query-based relation learner for human-object-interaction detection
Tianlun Luo, Qiao Yuan, Boxuan Zhu, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith, Eng Gee Lim |
Neurocomputing | 6 |
| 2025 | WaterVG: Waterway Visual Grounding Based on Text-Guided Vision and mmWave RadarabstractWaterway perception is critical for the special operations and autonomous navigation of Unmanned Surface Vessels (USVs), but current perception schemes are sensor-based, neglecting the interaction between humans and USVs for embodied perception in various operations. Therefore, inspired by visual grounding, we present WaterVG, the inaugural visual grounding dataset tailored for USV-based waterway perception guided by human prompts. WaterVG contains a wealth of prompts describing multiple targets, with instance-level annotations, including bounding boxes and masks. Specifically, WaterVG comprises 11,568 samples and 34,987 referred targets, integrating both visual and radar characteristics. The text-guided two-sensor pattern provides a fine granularity of text prompts aligned with the visual and radar features of the referent targets, containing both qualitative and numeric descriptions. To enhance the endurance and maintain the normal operations of USVs in open waterways, we propose Potamoi, a low-power visual grounding model. Potamoi is a multi-task model employing a sophisticated Phased Heterogeneous Modality Fusion (PHMF) mechanism, which includes Adaptive Radar Weighting (ARW) and Multi-Head Slim Cross Attention (MHSCA). The ARW module utilizes a gating mechanism to adaptively extract essential radar features for fusion with visual inputs, ensuring prompt alignment. MHSCA, characterized by its low parameter count and computational efficiency (FLOPs), effectively integrates contextual information from both sensors with linguistic features, delivering outstanding performance in visual grounding tasks. Comprehensive experiments and evaluations on WaterVG demonstrate that Potamoi achieves state-of-the-art results compared to existing methods. The project is available athttps://github.com/GuanRunwei/WaterVG. Runwei Guan, Liye Jia, Shanliang Yao, Fengyufan Yang, Erick Purwanto, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Xuming Hu, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 10 |
| 2024 | ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave RadarabstractPanoptic Driving Perception (PDP) is critical for the autonomous navigation of Unmanned Surface Vehicles (USVs). A PDP model typically integrates multiple tasks, necessitating the simultaneous and robust execution of various perception tasks to facilitate downstream path planning. The fusion of visual and radar sensors is currently acknowledged as a robust and cost-effective approach. However, most existing research has primarily focused on fusing visual and radar features dedicated to object detection or utilizing a shared feature space for multiple tasks, neglecting the individual representation differences between various tasks. To address this gap, we propose a pair of Asymmetric Fair Fusion (AFF) modules with favorable explainability designed to efficiently interact with independent features from both visual and radar modalities, tailored to the specific requirements of object detection and semantic segmentation tasks. The AFF modules treat image and radar maps as irregular point sets and transform these features into a crossed-shared feature space for multitasking, ensuring equitable treatment of vision and radar point cloud features. Leveraging AFF modules, we propose a novel and efficient PDP model, ASY-VRNet, which processes image and radar features based on irregular super-pixel point sets. Additionally, we propose an effective multi-task learning method specifically designed for PDP models. Compared to other lightweight models, ASY-VRNet achieves state-of-the-art performance in object detection, semantic segmentation, and drivable-area segmentation on the WaterScenes benchmark. Our project is publicly available at https://github.com/GuanRunwei/ASY-VRNet. Runwei Guan, Shanliang Yao, Ka Lok Man, Yong Yue 0001, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
IROS | 6 |
| 2024 | Context-based local-global fusion network for 3D point cloud classification and segmentation
Junwei Wu 0001, Mingjie Sun, Chenru Jiang, Jiejie Liu, Jeremy S. Smith |
Expert Syst. Appl. | 5 |
| 2024 | A deep top-down framework towards generalisable multi-view pedestrian detection
Ming Xu 0011, Yuchen Ling, Jeremy S. Smith, Yuyao Yan, Xinheng Wang 0001 |
Neurocomputing | 4 |
| 2024 | FindVehicle and VehicleFinder: a NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval systemabstractAbstract Natural language (NL) based vehicle retrieval is a task aiming to retrieve a vehicle that is most consistent with a given NL query from among all candidate vehicles. Because NL query can be easily obtained, such a task has a promising prospect in building an interactive intelligent traffic system (ITS). Current solutions mainly focus on extracting both text and image features and mapping them to the same latent space to compare the similarity. However, existing methods usually use dependency analysis or semantic role-labelling techniques to find keywords related to vehicle attributes. These techniques may require a lot of pre-processing and post-processing work, and also suffer from extracting the wrong keyword when the NL query is complex. To tackle these problems and simplify, we borrow the idea from named entity recognition (NER) and construct FindVehicle, a NER dataset in the traffic domain. It has 42.3k labelled NL descriptions of vehicle tracks, containing information such as the location, orientation, type and colour of the vehicle. FindVehicle also adopts both overlapping entities and fine-grained entities to meet further requirements. To verify its effectiveness, we propose a baseline NL-based vehicle retrieval model called VehicleFinder. Our experiment shows that by using text encoders pre-trained by FindVehicle, VehicleFinder achieves 87.7% precision and 89.4% recall when retrieving a target vehicle by text command on our homemade dataset based on UA-DETRAC [1]. From loading the command into VehicleFinder to identifying whether the target vehicle is consistent with the command, the time cost is 279.35 ms on one ARM v8.2 CPU and 93.72 ms on one RTX A4000 GPU, which is much faster than the Transformer-based system. The dataset is open-source via the link https://github.com/GuanRunwei/FindVehicle , and the implementation can be found via the link https://github.com/GuanRunwei/VehicleFinder-CTIM . Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
Multim. Tools Appl. | 7 |
| 2024 | PPM: A boolean optimizer for data association in multi-view pedestrian detection
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Yuchen Ling |
Pattern Recognit. | 4 |
| 2023 | From detection to understanding: A survey on representation learning for human-object interaction
Tianlun Luo, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith |
Neurocomputing | 4 |
| 2023 | MAN and CAT: mix attention to nn and concatenate attention to YOLO
Runwei Guan, Ka Lok Man, Haocheng Zhao, Ruixiao Zhang 0001, Shanliang Yao, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
J. Supercomput. | 6 |
| 2022 | 3D Random Occlusion and Multi-layer Projection for Deep Multi-camera Pedestrian Localization
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Xi Yang 0008 |
ECCV (10) | 4 |
| 2022 | A novel two-stream structure for video anomaly detection in smart city management
Ka Lok Man, Jeremy S. Smith, Steven Guan 0001 |
J. Supercomput. | 3 |
| 2021 | Multicamera pedestrian detection using logic minimization
Yuyao Yan, Ming Xu 0011, Jeremy S. Smith, Mo Shen, Jin Xi |
Pattern Recognit. | 3 |
| 2020 | Attentive Prototype Few-Shot Learning with Capsule Network-Based Embedding
Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu, Chaoyi Pang |
ECCV (28) | 2 |
| 2020 | Pose-robust Face Recognition by Deep Meta Capsule network-based Equivariant EmbeddingabstractDespite the exceptional success in face recognition related technologies, handling large pose variations still remains a key challenge. Current techniques for pose-robust face recognition either, directly extract pose-invariant features, or first synthesize a face that matches the target pose before feature extraction. It is more desirable to learn face representations equivariant to pose variations. To this end, this paper proposes a deep meta Capsule network-based Equivariant Embedding Model (DM-CEEM) with three distinct novelties. First, the proposed RB-CapsNet allows DM-CEEM to learn an equivariant embedding for pose variations and achieve the desired transformation for input face images. Second, we introduce a new version of a Capsule network called RB-CapsNet to extend CapsNet to perform a profile-to-frontal face transformation in deep feature space. Third, we train the DM-CEEM in a meta way by treating a single overall classification target as multiple sub-tasks that satisfy certain unknown probabilities. In each sub-task, we sample the support and query sets randomly. The experimental results on both controlled and in-the-wild databases demonstrate the superiority of DM-CEEM over state-of-the-art. Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu |
ICPR | 2 |
| 2020 | Moving shadow detection via binocular vision and colour clusteringabstractA pedestrian segmentation algorithm in the presence of cast shadows is presented in this study. The novelty of this algorithm lies in the fusion of multi‐view and multi‐plane homographic projections of foregrounds and the use of the fused data to guide colour clustering. This brings about an advantage over the existing binocular algorithms in that it can remove cast shadows while keeping pedestrians’ body parts, which occlude shadows. Phantom detection, which is inherent with the binocular method, is also investigated. Experimental results with real‐world videos have demonstrated the efficiency of this algorithm. Ming Xu 0011, Jeremy S. Smith, Yuyao Yan |
IET Comput. Vis. | 3 |
| 2020 | Image captioning via hierarchical attention mechanism and policy gradient optimization
Shiyang Yan, Yuan Xie 0006, Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu |
Signal Process. | 4 |
| 2019 | Image-Image Translation to Enhance Near Infrared Face RecognitionabstractWith the rapid development of facial recognition, the research field of near infrared (NIR) face recognition, which is less sensitive to illumination levels, has attracted increased attention. Unfortunately, directly applying the face recognition model trained using visible light (VIS) data to NIR face data does not produce a satisfactory performance. This is due to the domain bias between the NIR images and the VIS images. To this end, we created the Outdoor NIR-VIS Face (ONVF) database and Indoor NIR Face (INF) database to increase the number of near infrared facial images for system training and evaluation. In this paper, we propose an efficient NIR face recognition method, which consists of face detection and alignment, NIR-VIS image translation and face embedding. The NIR-VIS image conversion model is capable of transforming near-infrared facial images into their corresponding VIS images whilst maintaining sufficient identity information to enable existing VIS facial recognition models to perform recognition. Extensive experiments using the INF dataset and the CSIST database have demonstrated that the proposed method yields a consistent and competitive performance for near infrared face recognition. Fangyu Wu 0001, Weihang You, Jeremy S. Smith, Wenjin Lu |
ICIP | 3 |
| 2019 | Vehicle re-identification in still images: Application of semi-supervised learning and re-ranking
Fangyu Wu 0001, Shiyang Yan, Jeremy S. Smith |
Signal Process. Image Commun. | 3 |
| 2018 | Joint Semi-supervised Learning and Re-ranking for Vehicle Re-identificationabstractVehicle re-identification (re-ID) remains an unproblematic problem due to the complicated variations in vehicle appearances from multiple camera views. Most existing algorithms for solving this problem are developed in the fully-supervised setting, requiring access to a large number of labeled training data. However, it is impractical to expect large quantities of labeled data because the high cost of data annotation. Besides, re-ranking is a significant way to improve its performance when considering vehicle re-ID as a retrieval process. Yet limited effort has been devoted to the research of re-ranking in the vehicle re-ID. To address these problems, in this paper, we propose a semi-supervised learning system based on the Convolutional Neural Network (CNN) and re-ranking strategy for Vehicle re-ID. Specifically, we adopt the structure of Generative Adversarial Network (GAN) to obtain more vehicle images and enrich the training set, then a uniform label distribution will be assigned to the unlabeled samples according to the Label Smoothing Regularization for Outliers (LSRO), which regularizes the supervised learning model and improves the performance of re-ID. To optimize the re-ID results, an improved re-ranking method is exploited to optimize the initial rank list. Experimental results on publically available datasets, VeRi-776 and VehicleID, demonstrate that the method significantly outperforms the state-of-the-art. Fangyu Wu 0001, Shiyang Yan, Jeremy S. Smith |
ICPR | 3 |
| 2018 | Image Captioning using Adversarial Networks and Reinforcement LearningabstractImage captioning is a significant task in artificial intelligence which connects computer vision and natural language processing. With the rapid development of deep learning, the sequence to sequence model with attention, has become one of the main approaches for the task of image captioning. Nevertheless, a significant issue exists in the current framework: the exposure bias problem of Maximum Likelihood Estimation (MLE) in the sequence model. To address this problem, we use generative adversarial networks (GANs) for image captioning, which compensates for the exposure bias problem of MLE and also can generate more realistic captions. GANs, however, cannot be directly applied to a discrete task, like language processing, due to the discontinuity of the data. Hence, we use a reinforcement learning (RL) technique to estimate the gradients for the network. Also, to obtain the intermediate rewards during the process of language generation, a Monte Carlo roll-out sampling method is utilized. Experimental results on the COCO dataset validate the improved effect from each ingredient of the proposed model. The overall effectiveness is also evaluated. Shiyang Yan, Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu |
ICPR | 3 |
| 2018 | Multi-view visual surveillance and phantom removal for effective pedestrian detection
Jie Ren 0014, Ming Xu 0011, Jeremy S. Smith, Huimin Zhao 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Hierarchical Multi-scale Attention Networks for action recognition
Shiyang Yan, Jeremy S. Smith, Wenjin Lu |
Signal Process. Image Commun. | 2 |
| 2017 | CHAM: Action recognition using convolutional hierarchical attention modelabstractRecently, the soft attention mechanism, which was originally proposed in language processing, has been applied in computer vision tasks like image captioning. This paper presents improvements to the soft attention model by combining a con-volutional Long Short-Term Memory (LSTM) with a hierarchical system architecture to recognize action categories in videos. We call this model the Convolutional Hierarchical Attention Model (CHAM). The model applies a convolution-al operation inside the LSTM cell and an attention map generation process to recognize actions. The hierarchical architecture of this model is able to explicitly reason on multi-granularities of action categories. The proposed architecture achieved improved results on three publicly available datasets: the UCF sports dataset, the Olympic sports dataset and the HMDB51 dataset. Shiyang Yan, Jeremy S. Smith, Wenjin Lu |
ICIP | 2 |
| 2017 | Multiview pedestrian localisation via a prime candidate chart based on occupancy likelihoodsabstractA sound way to localize occluded people is to project the foregrounds from multiple camera views to a reference view by homographies and find the foreground intersections. However, this may give rise to phantoms due to foreground intersections from different people. In this paper, each intersection region is warped back to the original camera view and is associated with a candidate box of the average size of pedestrians at that location. Then a joint occupancy likelihood is calculated for each intersection region. In the second step, essential candidate boxes are identified first, each of which covers at least a part of the foreground that is not covered by another candidate box. The non-essential candidate boxes are selected to cover the remaining foregrounds in the order of their joint occupancy likelihoods. Experiments on benchmark video datasets have demonstrated the good performance of our algorithm in comparison with other state-of-the-art methods. Yuyao Yan, Ming Xu 0011, Jeremy S. Smith |
ICIP | 3 |
| 2017 | Action Recognition from Still Images Based on Deep VLAD Spatial Pyramids
Shiyang Yan, Jeremy S. Smith |
Signal Process. Image Commun. | 2 |
| 2012 | Pruning phantom detections from multiview foreground intersectionabstractHomography mapping and fusion of foreground regions from multiple camera views is an effective technique for moving object detection. However, the intersections of non-corresponding foreground regions frequently cause phantom detections. In this paper, an algorithm using colour template matching is proposed to identify such phantoms from the multiview foreground intersection. Experiments on real-world video sequences have been carried out. Jie Ren 0014, Ming Xu 0011, Jeremy S. Smith |
ICIP | 3 |
| 2012 | A colour statistical approach to phantom pruning in multi-view detectionabstractTo increase the robustness of detection in intelligent video surveillance systems, homography has been widely used to fuse foreground regions projected from multiple camera views to a reference view. However, the intersections of non-corresponding foreground regions can cause phantoms. This paper proposes a colour statistical approach to cope with this problem. This method is based on the Mahalanobis distance between the colour patches which correspond to the same foreground region in the reference view. This method can overcome the problems in the pixelwise colour correlation approach. Jie Ren 0014, Ming Xu 0011, Jeremy S. Smith |
SMC | 3 |
| 2012 | Robust localisation of pedestrians with cast shadows using homology in a monocular viewabstractIn this paper an object detection algorithm is proposed, which is robust in the presence of cast shadows and is based on geometric projections. The novelty of the work lies in the use of homology mapping of the foreground regions between different parallel planes within a monocular view, unlike some existing algorithms which depend on the use of multiple cameras. The results on an open video dataset are provided. Ming Xu 0011, Tianyuan Jia, Jeremy S. Smith |
SMC | 4 |
| 2012 | Cast shadow removal in motion detection by exploiting multiview geometryabstractAn object detection algorithm which can remove moving cast shadows is presented. It is based on the homography mapping of foreground regions from multiple cameras. Not only the homography for the ground plane but also those for multiple parallel planes are employed. Unlike the existing geometric approaches, this algorithm removes cast shadows while keeping the feet of moving objects. Ming Xu 0011, Tianyuan Jia, Jie Ren 0014, Jeremy S. Smith |
SMC | 5 |
| 2011 | A Multiview Approach to Robust Detection in the Presence of Cast ShadowsabstractThis paper presents an object detection algorithm using multiple cameras, which is robust in the presence of cast shadows. The information fusion is based on homography mapping of the foreground regions to a top view image. The homography is based on multiple planar planes parallel to the ground plane. Two novel approaches to estimating such homography have been proposed. The results on an open video dataset are demonstrated. Ming Xu 0011, Jie Ren 0014, Dongyong Chen, Jeremy S. Smith, Zhechi Liu |
ICIG | 4 |
| 2011 | Real-time detection via homography mapping of foreground polygons from multiple camerasabstractA real-time object detection algorithm by using multiple cameras is proposed. The information fusion is based on homographic transformation of the foreground information from multiple cameras to a reference image. Unlike the most recent algorithms which transmit and project foreground bitmap images, we approximate the contour of each fore ground region with a polygon and only transmit and project the polygon vertices. These polygons are rebuilt and fused in the reference image. This greatly reduces the requirement on network bandwidth and avoids homographic transformations at image levels. The results on an open video dataset are demonstrated. Ming Xu 0011, Jie Ren 0014, Dongyong Chen, Jeremy S. Smith, Guifen Wang |
ICIP | 4 |
| 2006 | Web-Based Distributed Embedded Gateway System DesignabstractWeb-Based distributed control and monitoring systems has been developed. The system offers multi-network integration, real-time operations and cross-platform solution. The commanding message is received from a long distance connection and sent to the local controller nodes to perform and required operations. Status information for the local controller nodes is gathered from individual nodes and sent back via the long distance connection Jeremy S. Smith, Tuo Li 0002 |
Web Intelligence | 2 |