Jeremy S. Smith

dblp:81/8286 · DBLP profile ↗
← Back
36ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0002-0212-2365ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 2 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression Comprehension
abstract
Embodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20, 558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar.
Runwei Guan, Ruixiao Zhang 0001, Ningwei Ouyang, Ka Lok Man, Xiaohao Cai, Ming Xu 0011, Jeremy S. Smith, Eng Gee Lim, Yutao Yue, Hui Xiong 0001
ICRA8
2025 NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave Radar
abstract
Recently, visual grounding and multi-sensors setting have been incorporated into perception system for terrestrial autonomous driving systems and Unmanned Surface Vessels (USVs), yet the high complexity of modern learning-based visual grounding model using multi-sensors prevents such model to be deployed on USVs in the real-life. To this end, we design a low-power multi-task model named NanoMVG for waterway embodied perception, guiding both camera and 4D millimeter-wave radar to locate specific object(s) through natural language. NanoMVG can perform both box-level and mask-level visual grounding tasks simultaneously. Compared to other visual grounding models, NanoMVG achieves highly competitive performance on the WaterVG dataset, particularly in harsh environments. Moreover, the real-world experiments with deployment of NanoMVG on embedded edge device of USV demonstrates its fast inference speed for real-time perception and capability of boasting ultra-low power consumption for long endurance.
Runwei Guan, Liye Jia, Haocheng Zhao, Shanliang Yao, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Yutao Yue
IROS9
2025 Referring flexible image restoration
Runwei Guan, Rongsheng Hu, Zhuhao Zhou, Tianlang Xue, Ka Lok Man, Jeremy S. Smith, Eng Gee Lim, Weiping Ding 0001, Yutao Yue
Expert Syst. Appl.6
2025 Exploring interaction concepts for human-object-interaction detection via global- and local-scale enhancing
Tianlun Luo, Qiao Yuan, Boxuan Zhu, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith, Eng Gee Lim
Neurocomputing6
2025 Simple yet effective: An explicit query-based relation learner for human-object-interaction detection
Tianlun Luo, Qiao Yuan, Boxuan Zhu, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith, Eng Gee Lim
Neurocomputing6
2025 WaterVG: Waterway Visual Grounding Based on Text-Guided Vision and mmWave Radar
abstract
Waterway perception is critical for the special operations and autonomous navigation of Unmanned Surface Vessels (USVs), but current perception schemes are sensor-based, neglecting the interaction between humans and USVs for embodied perception in various operations. Therefore, inspired by visual grounding, we present WaterVG, the inaugural visual grounding dataset tailored for USV-based waterway perception guided by human prompts. WaterVG contains a wealth of prompts describing multiple targets, with instance-level annotations, including bounding boxes and masks. Specifically, WaterVG comprises 11,568 samples and 34,987 referred targets, integrating both visual and radar characteristics. The text-guided two-sensor pattern provides a fine granularity of text prompts aligned with the visual and radar features of the referent targets, containing both qualitative and numeric descriptions. To enhance the endurance and maintain the normal operations of USVs in open waterways, we propose Potamoi, a low-power visual grounding model. Potamoi is a multi-task model employing a sophisticated Phased Heterogeneous Modality Fusion (PHMF) mechanism, which includes Adaptive Radar Weighting (ARW) and Multi-Head Slim Cross Attention (MHSCA). The ARW module utilizes a gating mechanism to adaptively extract essential radar features for fusion with visual inputs, ensuring prompt alignment. MHSCA, characterized by its low parameter count and computational efficiency (FLOPs), effectively integrates contextual information from both sensors with linguistic features, delivering outstanding performance in visual grounding tasks. Comprehensive experiments and evaluations on WaterVG demonstrate that Potamoi achieves state-of-the-art results compared to existing methods. The project is available athttps://github.com/GuanRunwei/WaterVG.
Runwei Guan, Liye Jia, Shanliang Yao, Fengyufan Yang, Erick Purwanto, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Xuming Hu, Yutao Yue
IEEE Trans. Intell. Transp. Syst.10
2024 ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
abstract
Panoptic Driving Perception (PDP) is critical for the autonomous navigation of Unmanned Surface Vehicles (USVs). A PDP model typically integrates multiple tasks, necessitating the simultaneous and robust execution of various perception tasks to facilitate downstream path planning. The fusion of visual and radar sensors is currently acknowledged as a robust and cost-effective approach. However, most existing research has primarily focused on fusing visual and radar features dedicated to object detection or utilizing a shared feature space for multiple tasks, neglecting the individual representation differences between various tasks. To address this gap, we propose a pair of Asymmetric Fair Fusion (AFF) modules with favorable explainability designed to efficiently interact with independent features from both visual and radar modalities, tailored to the specific requirements of object detection and semantic segmentation tasks. The AFF modules treat image and radar maps as irregular point sets and transform these features into a crossed-shared feature space for multitasking, ensuring equitable treatment of vision and radar point cloud features. Leveraging AFF modules, we propose a novel and efficient PDP model, ASY-VRNet, which processes image and radar features based on irregular super-pixel point sets. Additionally, we propose an effective multi-task learning method specifically designed for PDP models. Compared to other lightweight models, ASY-VRNet achieves state-of-the-art performance in object detection, semantic segmentation, and drivable-area segmentation on the WaterScenes benchmark. Our project is publicly available at https://github.com/GuanRunwei/ASY-VRNet.
Runwei Guan, Shanliang Yao, Ka Lok Man, Yong Yue 0001, Jeremy S. Smith, Eng Gee Lim, Yutao Yue
IROS6
2024 Context-based local-global fusion network for 3D point cloud classification and segmentation
Junwei Wu 0001, Mingjie Sun, Chenru Jiang, Jiejie Liu, Jeremy S. Smith
Expert Syst. Appl.5
2024 A deep top-down framework towards generalisable multi-view pedestrian detection
Ming Xu 0011, Yuchen Ling, Jeremy S. Smith, Yuyao Yan, Xinheng Wang 0001
Neurocomputing4
2024 FindVehicle and VehicleFinder: a NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval system
abstract
Abstract Natural language (NL) based vehicle retrieval is a task aiming to retrieve a vehicle that is most consistent with a given NL query from among all candidate vehicles. Because NL query can be easily obtained, such a task has a promising prospect in building an interactive intelligent traffic system (ITS). Current solutions mainly focus on extracting both text and image features and mapping them to the same latent space to compare the similarity. However, existing methods usually use dependency analysis or semantic role-labelling techniques to find keywords related to vehicle attributes. These techniques may require a lot of pre-processing and post-processing work, and also suffer from extracting the wrong keyword when the NL query is complex. To tackle these problems and simplify, we borrow the idea from named entity recognition (NER) and construct FindVehicle, a NER dataset in the traffic domain. It has 42.3k labelled NL descriptions of vehicle tracks, containing information such as the location, orientation, type and colour of the vehicle. FindVehicle also adopts both overlapping entities and fine-grained entities to meet further requirements. To verify its effectiveness, we propose a baseline NL-based vehicle retrieval model called VehicleFinder. Our experiment shows that by using text encoders pre-trained by FindVehicle, VehicleFinder achieves 87.7% precision and 89.4% recall when retrieving a target vehicle by text command on our homemade dataset based on UA-DETRAC [1]. From loading the command into VehicleFinder to identifying whether the target vehicle is consistent with the command, the time cost is 279.35 ms on one ARM v8.2 CPU and 93.72 ms on one RTX A4000 GPU, which is much faster than the Transformer-based system. The dataset is open-source via the link https://github.com/GuanRunwei/FindVehicle , and the implementation can be found via the link https://github.com/GuanRunwei/VehicleFinder-CTIM .
Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Jeremy S. Smith, Eng Gee Lim, Yutao Yue
Multim. Tools Appl.7
2024 PPM: A boolean optimizer for data association in multi-view pedestrian detection
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Yuchen Ling
Pattern Recognit.4
2023 From detection to understanding: A survey on representation learning for human-object interaction
Tianlun Luo, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith
Neurocomputing4
2023 MAN and CAT: mix attention to nn and concatenate attention to YOLO
Runwei Guan, Ka Lok Man, Haocheng Zhao, Ruixiao Zhang 0001, Shanliang Yao, Jeremy S. Smith, Eng Gee Lim, Yutao Yue
J. Supercomput.6
2022 3D Random Occlusion and Multi-layer Projection for Deep Multi-camera Pedestrian Localization
Ming Xu 0011, Yuyao Yan, Jeremy S. Smith, Xi Yang 0008
ECCV (10)4
2022 A novel two-stream structure for video anomaly detection in smart city management
Ka Lok Man, Jeremy S. Smith, Steven Guan 0001
J. Supercomput.3
2021 Multicamera pedestrian detection using logic minimization
Yuyao Yan, Ming Xu 0011, Jeremy S. Smith, Mo Shen, Jin Xi
Pattern Recognit.3
2020 Attentive Prototype Few-Shot Learning with Capsule Network-Based Embedding
Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu, Chaoyi Pang
ECCV (28)2
2020 Pose-robust Face Recognition by Deep Meta Capsule network-based Equivariant Embedding
abstract
Despite the exceptional success in face recognition related technologies, handling large pose variations still remains a key challenge. Current techniques for pose-robust face recognition either, directly extract pose-invariant features, or first synthesize a face that matches the target pose before feature extraction. It is more desirable to learn face representations equivariant to pose variations. To this end, this paper proposes a deep meta Capsule network-based Equivariant Embedding Model (DM-CEEM) with three distinct novelties. First, the proposed RB-CapsNet allows DM-CEEM to learn an equivariant embedding for pose variations and achieve the desired transformation for input face images. Second, we introduce a new version of a Capsule network called RB-CapsNet to extend CapsNet to perform a profile-to-frontal face transformation in deep feature space. Third, we train the DM-CEEM in a meta way by treating a single overall classification target as multiple sub-tasks that satisfy certain unknown probabilities. In each sub-task, we sample the support and query sets randomly. The experimental results on both controlled and in-the-wild databases demonstrate the superiority of DM-CEEM over state-of-the-art.
Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu
ICPR2
2020 Moving shadow detection via binocular vision and colour clustering
abstract
A pedestrian segmentation algorithm in the presence of cast shadows is presented in this study. The novelty of this algorithm lies in the fusion of multi‐view and multi‐plane homographic projections of foregrounds and the use of the fused data to guide colour clustering. This brings about an advantage over the existing binocular algorithms in that it can remove cast shadows while keeping pedestrians’ body parts, which occlude shadows. Phantom detection, which is inherent with the binocular method, is also investigated. Experimental results with real‐world videos have demonstrated the efficiency of this algorithm.
Ming Xu 0011, Jeremy S. Smith, Yuyao Yan
IET Comput. Vis.3
2020 Image captioning via hierarchical attention mechanism and policy gradient optimization
Shiyang Yan, Yuan Xie 0006, Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu
Signal Process.4
2019 Image-Image Translation to Enhance Near Infrared Face Recognition
abstract
With the rapid development of facial recognition, the research field of near infrared (NIR) face recognition, which is less sensitive to illumination levels, has attracted increased attention. Unfortunately, directly applying the face recognition model trained using visible light (VIS) data to NIR face data does not produce a satisfactory performance. This is due to the domain bias between the NIR images and the VIS images. To this end, we created the Outdoor NIR-VIS Face (ONVF) database and Indoor NIR Face (INF) database to increase the number of near infrared facial images for system training and evaluation. In this paper, we propose an efficient NIR face recognition method, which consists of face detection and alignment, NIR-VIS image translation and face embedding. The NIR-VIS image conversion model is capable of transforming near-infrared facial images into their corresponding VIS images whilst maintaining sufficient identity information to enable existing VIS facial recognition models to perform recognition. Extensive experiments using the INF dataset and the CSIST database have demonstrated that the proposed method yields a consistent and competitive performance for near infrared face recognition.
Fangyu Wu 0001, Weihang You, Jeremy S. Smith, Wenjin Lu
ICIP3
2019 Vehicle re-identification in still images: Application of semi-supervised learning and re-ranking
Fangyu Wu 0001, Shiyang Yan, Jeremy S. Smith
Signal Process. Image Commun.3
2018 Joint Semi-supervised Learning and Re-ranking for Vehicle Re-identification
abstract
Vehicle re-identification (re-ID) remains an unproblematic problem due to the complicated variations in vehicle appearances from multiple camera views. Most existing algorithms for solving this problem are developed in the fully-supervised setting, requiring access to a large number of labeled training data. However, it is impractical to expect large quantities of labeled data because the high cost of data annotation. Besides, re-ranking is a significant way to improve its performance when considering vehicle re-ID as a retrieval process. Yet limited effort has been devoted to the research of re-ranking in the vehicle re-ID. To address these problems, in this paper, we propose a semi-supervised learning system based on the Convolutional Neural Network (CNN) and re-ranking strategy for Vehicle re-ID. Specifically, we adopt the structure of Generative Adversarial Network (GAN) to obtain more vehicle images and enrich the training set, then a uniform label distribution will be assigned to the unlabeled samples according to the Label Smoothing Regularization for Outliers (LSRO), which regularizes the supervised learning model and improves the performance of re-ID. To optimize the re-ID results, an improved re-ranking method is exploited to optimize the initial rank list. Experimental results on publically available datasets, VeRi-776 and VehicleID, demonstrate that the method significantly outperforms the state-of-the-art.
Fangyu Wu 0001, Shiyang Yan, Jeremy S. Smith
ICPR3
2018 Image Captioning using Adversarial Networks and Reinforcement Learning
abstract
Image captioning is a significant task in artificial intelligence which connects computer vision and natural language processing. With the rapid development of deep learning, the sequence to sequence model with attention, has become one of the main approaches for the task of image captioning. Nevertheless, a significant issue exists in the current framework: the exposure bias problem of Maximum Likelihood Estimation (MLE) in the sequence model. To address this problem, we use generative adversarial networks (GANs) for image captioning, which compensates for the exposure bias problem of MLE and also can generate more realistic captions. GANs, however, cannot be directly applied to a discrete task, like language processing, due to the discontinuity of the data. Hence, we use a reinforcement learning (RL) technique to estimate the gradients for the network. Also, to obtain the intermediate rewards during the process of language generation, a Monte Carlo roll-out sampling method is utilized. Experimental results on the COCO dataset validate the improved effect from each ingredient of the proposed model. The overall effectiveness is also evaluated.
Shiyang Yan, Fangyu Wu 0001, Jeremy S. Smith, Wenjin Lu
ICPR3
2018 Multi-view visual surveillance and phantom removal for effective pedestrian detection
Jie Ren 0014, Ming Xu 0011, Jeremy S. Smith, Huimin Zhao 0001
Multim. Tools Appl.3
2018 Hierarchical Multi-scale Attention Networks for action recognition
Shiyang Yan, Jeremy S. Smith, Wenjin Lu
Signal Process. Image Commun.2
2017 CHAM: Action recognition using convolutional hierarchical attention model
abstract
Recently, the soft attention mechanism, which was originally proposed in language processing, has been applied in computer vision tasks like image captioning. This paper presents improvements to the soft attention model by combining a con-volutional Long Short-Term Memory (LSTM) with a hierarchical system architecture to recognize action categories in videos. We call this model the Convolutional Hierarchical Attention Model (CHAM). The model applies a convolution-al operation inside the LSTM cell and an attention map generation process to recognize actions. The hierarchical architecture of this model is able to explicitly reason on multi-granularities of action categories. The proposed architecture achieved improved results on three publicly available datasets: the UCF sports dataset, the Olympic sports dataset and the HMDB51 dataset.
Shiyang Yan, Jeremy S. Smith, Wenjin Lu
ICIP2
2017 Multiview pedestrian localisation via a prime candidate chart based on occupancy likelihoods
abstract
A sound way to localize occluded people is to project the foregrounds from multiple camera views to a reference view by homographies and find the foreground intersections. However, this may give rise to phantoms due to foreground intersections from different people. In this paper, each intersection region is warped back to the original camera view and is associated with a candidate box of the average size of pedestrians at that location. Then a joint occupancy likelihood is calculated for each intersection region. In the second step, essential candidate boxes are identified first, each of which covers at least a part of the foreground that is not covered by another candidate box. The non-essential candidate boxes are selected to cover the remaining foregrounds in the order of their joint occupancy likelihoods. Experiments on benchmark video datasets have demonstrated the good performance of our algorithm in comparison with other state-of-the-art methods.
Yuyao Yan, Ming Xu 0011, Jeremy S. Smith
ICIP3
2017 Action Recognition from Still Images Based on Deep VLAD Spatial Pyramids
Shiyang Yan, Jeremy S. Smith
Signal Process. Image Commun.2
2012 Pruning phantom detections from multiview foreground intersection
abstract
Homography mapping and fusion of foreground regions from multiple camera views is an effective technique for moving object detection. However, the intersections of non-corresponding foreground regions frequently cause phantom detections. In this paper, an algorithm using colour template matching is proposed to identify such phantoms from the multiview foreground intersection. Experiments on real-world video sequences have been carried out.
Jie Ren 0014, Ming Xu 0011, Jeremy S. Smith
ICIP3
2012 A colour statistical approach to phantom pruning in multi-view detection
abstract
To increase the robustness of detection in intelligent video surveillance systems, homography has been widely used to fuse foreground regions projected from multiple camera views to a reference view. However, the intersections of non-corresponding foreground regions can cause phantoms. This paper proposes a colour statistical approach to cope with this problem. This method is based on the Mahalanobis distance between the colour patches which correspond to the same foreground region in the reference view. This method can overcome the problems in the pixelwise colour correlation approach.
Jie Ren 0014, Ming Xu 0011, Jeremy S. Smith
SMC3
2012 Robust localisation of pedestrians with cast shadows using homology in a monocular view
abstract
In this paper an object detection algorithm is proposed, which is robust in the presence of cast shadows and is based on geometric projections. The novelty of the work lies in the use of homology mapping of the foreground regions between different parallel planes within a monocular view, unlike some existing algorithms which depend on the use of multiple cameras. The results on an open video dataset are provided.
Ming Xu 0011, Tianyuan Jia, Jeremy S. Smith
SMC4
2012 Cast shadow removal in motion detection by exploiting multiview geometry
abstract
An object detection algorithm which can remove moving cast shadows is presented. It is based on the homography mapping of foreground regions from multiple cameras. Not only the homography for the ground plane but also those for multiple parallel planes are employed. Unlike the existing geometric approaches, this algorithm removes cast shadows while keeping the feet of moving objects.
Ming Xu 0011, Tianyuan Jia, Jie Ren 0014, Jeremy S. Smith
SMC5
2011 A Multiview Approach to Robust Detection in the Presence of Cast Shadows
abstract
This paper presents an object detection algorithm using multiple cameras, which is robust in the presence of cast shadows. The information fusion is based on homography mapping of the foreground regions to a top view image. The homography is based on multiple planar planes parallel to the ground plane. Two novel approaches to estimating such homography have been proposed. The results on an open video dataset are demonstrated.
Ming Xu 0011, Jie Ren 0014, Dongyong Chen, Jeremy S. Smith, Zhechi Liu
ICIG4
2011 Real-time detection via homography mapping of foreground polygons from multiple cameras
abstract
A real-time object detection algorithm by using multiple cameras is proposed. The information fusion is based on homographic transformation of the foreground information from multiple cameras to a reference image. Unlike the most recent algorithms which transmit and project foreground bitmap images, we approximate the contour of each fore ground region with a polygon and only transmit and project the polygon vertices. These polygons are rebuilt and fused in the reference image. This greatly reduces the requirement on network bandwidth and avoids homographic transformations at image levels. The results on an open video dataset are demonstrated.
Ming Xu 0011, Jie Ren 0014, Dongyong Chen, Jeremy S. Smith, Guifen Wang
ICIP4
2006 Web-Based Distributed Embedded Gateway System Design
abstract
Web-Based distributed control and monitoring systems has been developed. The system offers multi-network integration, real-time operations and cross-platform solution. The commanding message is received from a long distance connection and sent to the local controller nodes to perform and required operations. Status information for the local controller nodes is gathered from individual nodes and sent back via the long distance connection
Jeremy S. Smith, Tuo Li 0002
Web Intelligence2