VLDB 2026 Research / reviewers in the wild / expert
Eng Gee Lim
dblp:130/1634 · also Enggee Lim
· DBLP profile ↗
86ranked-venue papers
0as first author
65since 2021 · last 2026
0000-0003-0199-7386ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 26 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 23 since 2021Computer networks · 20 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 9 since 2021Systems, architecture and hardware · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RIS-LLM: Reasoning-informed semantic modeling of electricity market price dynamics
Jieming Ma, Ka Lok Man, Steven Guan 0001, Eng Gee Lim |
Adv. Eng. Informatics | 8 |
| 2026 | A joint topology-data fusion graph network for robust traffic speed prediction with data anomalism
Ruiyuan Jiang, Dongyao Jia, Eng Gee Lim, Shangbo Wang |
Inf. Sci. | 3 |
| 2026 | ML-BF: Responsive and Dynamic Intrusion Detection towards Intelligent Connected Vehicles
Jia Liu 0074, Wenjun Fan, Eng Gee Lim, Yifan Dai 0006, Alexei Lisitsa 0001 |
Peer Peer Netw. Appl. | 3 |
| 2026 | MvP-Diff: Multivariate yet precise diffusion for anomaly images synthesis and segmentation
Siyue Yao, Eng Gee Lim, Siyue Yu, Jimin Xiao, Mingjie Sun |
Pattern Recognit. | 2 |
| 2026 | EmoSENSE: Modeling Sentiment-Semantic Knowledge With Hierarchical Reinforcement Learning for Emotional Image GenerationabstractEmotional image generation aims to create images that effectively reflect target emotions. A fundamental challenge in this task is the affective gap, which refers to the discrepancy between visual content and emotional states perceived by users. Existing methods generally assume strong and explicit associations between target emotions and specific objects (e.g., “monster” and “fear”), which limits their generalization ability when encountering uncommon emotion-object pairs. This limitation stems from two main factors: 1) Most existing approaches primarily focus on semantic alignment without explicitly modeling how emotions influence visual attributes such as brightness and colorfulness; 2) diffusion-based image generation methods have limited capability in handling diverse sentiment-semantic pairs. To address these challenges, we propose EmoSENSE, a novel hierarchical fuzzy reinforcement learning framework for the emotional image generation task. EmoSENSE consists of a high-level module and a low-level module, working collaboratively in a hierarchical structure to inject sentiment-semantic knowledge into emotional images. The high-level module quantifies sentiment-semantic correlations within a unified emotional space, connecting emotions to visual attributes. The low-level module refines this connection by optimizing a fuzzy-logic-based mapping between emotions and visual attributes through reinforcement learning, enabling flexible adaptation to diverse emotion-object pairs. Extensive qualitative and quantitative experiments on public dataset demonstrate that EmoSENSE significantly enhances both the visual quality and emotional expression ability of the generated images, achieving a 12.21% higher EmoAccuracy-8 classes than the previous state-of-the-art methods.https://github.com/forever3600/EmoSENSE. Junyi Guo, Qiufeng Wang 0001, Yaran Chen, Fangyu Wu 0001, Eng Gee Lim |
IEEE Trans. Affect. Comput. | 7 |
| 2026 | WormNet: An Automated Deep Learning Platform for Robust and High-Throughput C. elegans Behavior AnalysisabstractCaenorhabditis elegans is a model organism widely used in genetics, neuroscience, aging, and toxicology, where quantitative behavior analysis is essential for high-throughput screening and phenotype assessment. Automated worm tracking and analysis software based on traditional image processing has been developed to reduce manual workload, but such methods still degrade markedly under low contrast, lighting drift, and dense multi-worm conditions, limiting real-world scalability and robustness. Based on microscope imaging and deep learning, we develop WormNet, an integrated hardware–software platform for automatedC. elegansbehavior analysis with enhanced accuracy and throughput. The system combines a temperature-controlled motorized stage, bright-field microscope, and industrial camera, enabling stable acquisition of multi-worm videos under controlled experimental conditions. On the algorithmic side, a multi-scale backbone with a simplified snake convolution (SimSnake Conv) module is used to accurately detect worms, while a multi-scale cross-modal gated fusion (MCGF) module fuses detection features to refine instance masks. A downstream tracking and skeletonization pipeline then automatically calculates morphology and motility parameters, including body length, width, speed, and bending angle. Through experiments on two representative datasets, WormDrop-Micro and WormDish-Real, the system achieves maximum to 1.90% improvement in AP50and 28.50% improvement in mAP50–95over state-of-the-art (SOTA) detectors, and in terms of segmentation accuracy, the Dice coefficient was improved by up to 1.71%, and mIoU by 3.00%. The WormNet pipeline further achieved a processing speed of 20.45 FPS and supported batch analysis of multi-worm videos containing up to approximately 30 worms per field of view. With this performance, WormNet enables robust, fully automated, and high-throughputC. elegansbehavior analysis under diverse experimental conditions, making large-scale drug and toxicology screening studies feasible. Shiyan Li, Jinxin Gu, Yixue Qiao, Ian Sandall, Eng Gee Lim |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2026 | Finite Blocklength Relaying Communication With Unitary Beamforming and Energy Harvesting: Fairness Oriented Design
Yuanchen Wang, T. Aaron Gulliver, Yiyuan Xie, Chaowei Wang, Ruihong Jiang, Tingnan Bao, Eng Gee Lim, Ramy Samy |
IEEE Trans. Ind. Informatics | 8 |
| 2026 | From High-SNR Radar Signal to ECG: A Transfer Learning Model With Cardio-Focusing Algorithm for Scenarios With Limited DataabstractElectrocardiogram (ECG), as a crucial fine-grained cardiac feature, has been successfully recovered from radar signals in the literature, but the performance heavily relies on the high-quality radar signal and numerous radar-ECG pairs for training, restricting the applications in new scenarios due to data scarcity. Therefore, this work focuses on radar-based ECG recovery in new scenarios with limited data and proposes a cardio-focusing and-tracking (CFT) algorithm to precisely track the cardiac location to ensure an efficient acquisition of high quality radar signals. Furthermore, a transfer learning model (RFcardi) is proposed to extract cardio-related information from the radar signal without ECG ground truth based on the intrinsic sparsity of cardiac features, and only a few synchronous radar ECG pairs are required to fine-tune the pre-trained model for ECG recovery. The experimental results reveal that the proposed CFT can dynamically identify the cardiac location, and the RFcardi model can effectively generate faithful ECG recoveries after using a small number of radar-ECG pairs for training. The code and dataset will be made available after publication. Haocheng Zhao, Sijie Xiong, Rui Yang 0007, Eng Gee Lim, Yutao Yue |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly DetectionabstractExisting unsupervised distillation-based methods rely on the differences between encoded and decoded features to locate abnormal regions in test images. However, the decoder trained only on normal samples still reconstructs abnormal patch features well, degrading performance. This issue is particularly pronounced in unsupervised multi-class anomaly detection tasks. We attribute this behavior to ‘over-generalization’ (OG) of decoder: the significantly increasing diversity of patch patterns in multi-class training enhances the model generalization on normal patches, but also inadvertently broadens its generalization to abnormal patches. To mitigate ‘OG’, we propose a novel approach that leverages class-agnostic learnable prompts to capture common textual normality across various visual patterns, and then apply them to guide the decoded features towards a ‘normal’ textual representation, suppressing ‘over-generalization’ of the decoder on abnormal patterns. To further improve performance, we also introduce a gated mixture-of-experts module to specialize in handling diverse patch patterns and reduce mutual interference between them in multi-class training. Our method achieves competitive performance on the MVTec AD and VisA datasets, demonstrating its effectiveness. Xiaoyang Wang 0007, Huihui Bai 0001, Eng Gee Lim, Jimin Xiao |
AAAI | 4 |
| 2025 | POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationabstractWeakly Supervised Semantic Segmentation (WSSS) leverages Class Activation Maps (CAMs) to extract spatial information from image-level labels. However, CAMs primarily highlight the most discriminative foreground regions, leading to incomplete results. Prototype-based methods attempt to address this limitation by employing prototype CAMs instead of classifier CAMs. Nevertheless, existing prototype-based methods typically use a single prototype for each class, which is insufficient to capture all attributes of the foreground features due to the significant intra-class variations across different images. Consequently, these methods still struggle with incomplete CAM predictions. In this paper, we propose a novel framework called Prototypical Optimal Transport (POT) for WSSS. POT enhances CAM predictions by dividing features into multiple clusters and activating each cluster using its prototype. In this process, a similarity-aware optimal transport is employed to assign features to the most probable clusters. This similarity-aware strategy ensures the prioritization of significant cluster prototypes, thereby improving the accuracy of feature assignment. Additionally, we introduce an adaptive OT-based consistency loss to refine feature representations. This framework effectively overcomes the limitations of single-prototype methods, providing more complete and accurate CAM predictions. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) demonstrate that our method significantly improves the quality of CAMs and achieves state-of-the-art performances. The source code will be released https://github.com/jianwang91/POT. Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
CVPR | 5 |
| 2025 | HiGarment: Cross-Modal Harmony Based Diffusion Model for Flat Sketch to Realistic Garment ImageabstractDiffusion-based garment synthesis tasks primarily focus on the design phase in the fashion domain, while the garment production process remains largely underexplored. To bridge this gap, we introduce a new task: Flat Sketch to Realistic Garment Image (FS2RG), which generates realistic garment images by integrating flat sketches and textual guidance. FS2RG presents two key challenges: 1) fabric characteristics are solely guided by textual prompts, providing insufficient visual supervision for diffusion-based models, which limits their ability to capture fine-grained fabric details; 2) flat sketches and textual guidance may provide conflicting information, requiring the model to selectively preserve or modify garment attributes while maintaining structural coherence. To tackle this task, we propose HiGarment, a novel framework that comprises two core components: i) a multi-modal semantic enhancement mechanism that enhances fabric representation across textual and visual modalities, and ii) a harmonized cross-attention mechanism that dynamically balances information from flat sketches and text prompts, allowing controllable synthesis by generating either sketch-aligned (image-biased) or text-guided (text-biased) outputs. Furthermore, we collect Multi-modal Detailed Garment, the largest open-source dataset for garment generation. Experimental results and user studies demonstrate the effectiveness of HiGarment in garment synthesis. The code and dataset are available at https://github.com/Maple498/HiGarment. Junyi Guo, Fangyu Wu 0001, Huanda Lu, Qiufeng Wang 0001, Wenmian Yang, Eng Gee Lim, Dongming Lu |
ICCV | 7 |
| 2025 | Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic Segmentation
Jian Wang 0122, Tianhong Dai, Bingfeng Zhang, Siyue Yu, Eng Gee Lim, Jimin Xiao |
ICCV | 5 |
| 2025 | DecAD: Decoupling Anomalies in Latent Space for Multi-Class Unsupervised Anomaly Detection
Xiaoyang Wang 0007, Huihui Bai 0001, Eng Gee Lim, Jimin Xiao |
ICCV | 4 |
| 2025 | Talk2Radar: Bridging Natural Language with 4D mmWave Radar for 3D Referring Expression ComprehensionabstractEmbodied perception is essential for intelligent vehicles and robots in interactive environmental understanding. However, these advancements primarily focus on vision, with limited attention given to using 3D modeling sensors, restricting a comprehensive understanding of objects in response to prompts containing qualitative and quantitative queries. Recently, as a promising automotive sensor with affordable cost, 4D millimeter-wave radars provide denser point clouds than conventional radars and perceive both semantic and physical characteristics of objects, thereby enhancing the reliability of perception systems. To foster the development of natural language-driven context understanding in radar scenes for 3D visual grounding, we construct the first dataset, Talk2Radar, which bridges these two modalities for 3D Referring Expression Comprehension (REC). Talk2Radar contains 8,682 referring prompt samples with 20, 558 referred objects. Moreover, we propose a novel model, T-RadarNet, for 3D REC on point clouds, achieving State-Of-The-Art (SOTA) performance on the Talk2Radar dataset compared to counterparts. Deformable-FPN and Gated Graph Fusion are meticulously designed for efficient point cloud feature modeling and cross-modal fusion between radar and text features, respectively. Comprehensive experiments provide deep insights into radar-based 3D REC. We release our project at https://github.com/GuanRunwei/Talk2Radar. Runwei Guan, Ruixiao Zhang 0001, Ningwei Ouyang, Ka Lok Man, Xiaohao Cai, Ming Xu 0011, Jeremy S. Smith, Eng Gee Lim, Yutao Yue, Hui Xiong 0001 |
ICRA | 9 |
| 2025 | DriftRemover: Hybrid Energy Optimizations for Anomaly Images Synthesis and SegmentationabstractThis paper tackles the challenge of anomaly image synthesis and segmentation to generate various anomaly images and their segmentation labels to mitigate the issue of data scarcity. Existing approaches employ the precise mask to guide the generation, relying on additional mask generators, leading to increased computational costs and limited anomaly diversity. Although a few works use coarse masks as the guidance to expand diversity, they lack effective generation of labels for synthetic images, thereby reducing their practicality. Therefore, our proposed method simultaneously generates anomaly images and their corresponding masks by utilizing coarse masks and anomaly categories. The framework utilizes attention maps from synthesis process as mask labels and employs two optimization modules to tackle drift challenges, which are mismatches between synthetic results and real situations. Our evaluation demonstrates that our method improves pixel-level AP by 1.3% and F1-MAX by 1.8% in anomaly detection tasks on the MVTec dataset. Additionally, its successful application in practical scenarios highlights its effectiveness, improving IoU by 37.2% and F-measure by 25.1% with the Floor Dirt dataset. The code is available at https://github.com/JJessicaYao/DriftRemover. Siyue Yao, Mingjie Sun, Siyue Yu, Jimin Xiao, Eng Gee Lim |
IJCAI | 6 |
| 2025 | NanoMVG: USV-Centric Low-Power Multi-Task Visual Grounding based on Prompt-Guided Camera and 4D mmWave RadarabstractRecently, visual grounding and multi-sensors setting have been incorporated into perception system for terrestrial autonomous driving systems and Unmanned Surface Vessels (USVs), yet the high complexity of modern learning-based visual grounding model using multi-sensors prevents such model to be deployed on USVs in the real-life. To this end, we design a low-power multi-task model named NanoMVG for waterway embodied perception, guiding both camera and 4D millimeter-wave radar to locate specific object(s) through natural language. NanoMVG can perform both box-level and mask-level visual grounding tasks simultaneously. Compared to other visual grounding models, NanoMVG achieves highly competitive performance on the WaterVG dataset, particularly in harsh environments. Moreover, the real-world experiments with deployment of NanoMVG on embedded edge device of USV demonstrates its fast inference speed for real-time perception and capability of boasting ultra-low power consumption for long endurance. Runwei Guan, Liye Jia, Haocheng Zhao, Shanliang Yao, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Yutao Yue |
IROS | 8 |
| 2025 | More or Less? Effects of Visual Information Modulation on Context PerceptionabstractThis study examines the effects of two visual guidance techniques, Visual Enhancement and Visual Suppression, on user perception of contextual information in video content. Visual Enhancement introduces explicit visual cues to highlight target content, whereas Visual Suppression attenuates non-target elements, for example, by reducing their brightness. Both approaches aim to isolate specific objects from the background, directing attention to critical information within complex, dynamic scenes. Despite their growing usage, the relative effectiveness of these approaches in guiding attention and their impact on peripheral context awareness remain underexplored. To address this gap, we conducted a controlled user study with 27 participants. The results indicate that Visual Enhancement, through the addition of salient cues, more effectively directs user attention to target information than Visual Suppression. Our findings advance understanding of visual attention in dynamic environments and offer implications for designing visual guidance strategies. Jifan Yang, Fuqi Xie 0002, Zhaolin Lu, Yu Liu 0077, Martijn ten Bhömer, Eng Gee Lim, Lingyun Yu 0001 |
VINCI | 10 |
| 2025 | Referring flexible image restoration
Runwei Guan, Rongsheng Hu, Zhuhao Zhou, Tianlang Xue, Ka Lok Man, Jeremy S. Smith, Eng Gee Lim, Weiping Ding 0001, Yutao Yue |
Expert Syst. Appl. | 7 |
| 2025 | Exploring interaction concepts for human-object-interaction detection via global- and local-scale enhancing
Tianlun Luo, Qiao Yuan, Boxuan Zhu, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith, Eng Gee Lim |
Neurocomputing | 7 |
| 2025 | Simple yet effective: An explicit query-based relation learner for human-object-interaction detection
Tianlun Luo, Qiao Yuan, Boxuan Zhu, Steven Guan 0001, Rui Yang 0007, Jeremy S. Smith, Eng Gee Lim |
Neurocomputing | 7 |
| 2025 | IMCGNN: Information Maximization based Continual Graph Neural Networks for inductive node classification
Qiao Yuan, Steven Guan 0001, Tianlun Luo, Ka Lok Man, Eng Gee Lim |
Neurocomputing | 5 |
| 2025 | Updatable Signature with public tokens
Haotian Yin, Jie Zhang 0030, Wanxin Li, Yuji Dong, Eng Gee Lim, Dominik Wojtczak |
J. Inf. Secur. Appl. | 5 |
| 2025 | Auxiliary captioning: Bridging image-text matching and image captioning
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001 |
Signal Process. Image Commun. | 4 |
| 2025 | Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-Scarce ScenariosabstractThe scarcity of data in various scenarios, such as medical, industry and autonomous driving, leads to model overfitting and dataset imbalance, thus hindering effective detection and segmentation performance. Existing studies employ the generative models to synthesize more training samples to mitigate data scarcity. However, these synthetic samples are repetitive or simplistic and fail to provide "crucial information" that targets the downstream model's weaknesses. Additionally, these methods typically require separate training for different objects, leading to computational inefficiencies. To address these issues, we propose Crucial-Diff, a domain-agnostic framework designed to synthesize crucial samples. Our method integrates two key modules. The Scene Agnostic Feature Extractor (SAFE) utilizes a unified feature extractor to capture target information. The Weakness Aware Sample Miner (WASM) generates hard-to-detect samples using feedback from the detection results of downstream model, which is then fused with the output of SAFE module. Together, our Crucial-Diff framework generates diverse, high-quality training data, achieving a pixel-level AP of 83.63% and an F1-MAX of 78.12% on MVTec. On polyp dataset, Crucial-Diff reaches an mIoU of 81.64% and an mDice of 87.69%. Code is publicly available at https://github.com/JJessicaYao/Crucial-diff. Siyue Yao, Mingjie Sun, Eng Gee Lim, Ran Yi 0002, Baojiang Zhong, Moncef Gabbouj |
IEEE Trans. Image Process. | 3 |
| 2025 | Communication Strategy on Macro-and-Micro Traffic State in Cooperative Deep Reinforcement Learning for Regional Traffic Signal ControlabstractAdaptive Traffic Signal Control (ATSC) has become a popular research topic in intelligent transportation systems. Regional Traffic Signal Control (RTSC) using the Multi-agent Deep Reinforcement Learning (MADRL) technique has become a promising approach for ATSC due to its ability to achieve the optimum trade-off between scalability and optimality. Most existing RTSC approaches partition a traffic network into several disjoint regions, followed by applying centralized reinforcement learning techniques to each region. However, the pursuit of cooperation among RTSC agents still remains an open issue and no communication strategy for RTSC agents has been investigated. In this paper, we propose communication strategies to capture the correlation of micro-traffic states among lanes and the correlation of macro-traffic states among intersections. We first justify that the evolution equation of the RTSC process is Markovian via a system of store-and-forward queues. Next, based on the evolution equation, we propose two GAT-Aggregated (GA2) communication modules—GA2-Naive and GA2-Aug to extract both intra-region and inter-region correlations between macro and micro traffic states. While GA2-Naive only considers the movements at each intersection, GA2-Aug also considers the lane-changing behavior of vehicles. Two proposed communication modules are then aggregated into two existing novel RTSC frameworks—RegionLight and Regional-DRL. Experimental results demonstrate that both GA2-Naive and GA2-Aug effectively improve the performance of existing RTSC frameworks under both real and synthetic scenarios. Hyperparameter testing also reveals the robustness and potential of our communication modules in large-scale traffic networks. Hankang Gu, Shangbo Wang, Dongyao Jia, Yanrong Luo, Guoqiang Mao, Jianping Wang 0001, Eng Gee Lim |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2025 | WaterVG: Waterway Visual Grounding Based on Text-Guided Vision and mmWave RadarabstractWaterway perception is critical for the special operations and autonomous navigation of Unmanned Surface Vessels (USVs), but current perception schemes are sensor-based, neglecting the interaction between humans and USVs for embodied perception in various operations. Therefore, inspired by visual grounding, we present WaterVG, the inaugural visual grounding dataset tailored for USV-based waterway perception guided by human prompts. WaterVG contains a wealth of prompts describing multiple targets, with instance-level annotations, including bounding boxes and masks. Specifically, WaterVG comprises 11,568 samples and 34,987 referred targets, integrating both visual and radar characteristics. The text-guided two-sensor pattern provides a fine granularity of text prompts aligned with the visual and radar features of the referent targets, containing both qualitative and numeric descriptions. To enhance the endurance and maintain the normal operations of USVs in open waterways, we propose Potamoi, a low-power visual grounding model. Potamoi is a multi-task model employing a sophisticated Phased Heterogeneous Modality Fusion (PHMF) mechanism, which includes Adaptive Radar Weighting (ARW) and Multi-Head Slim Cross Attention (MHSCA). The ARW module utilizes a gating mechanism to adaptively extract essential radar features for fusion with visual inputs, ensuring prompt alignment. MHSCA, characterized by its low parameter count and computational efficiency (FLOPs), effectively integrates contextual information from both sensors with linguistic features, delivering outstanding performance in visual grounding tasks. Comprehensive experiments and evaluations on WaterVG demonstrate that Potamoi achieves state-of-the-art results compared to existing methods. The project is available athttps://github.com/GuanRunwei/WaterVG. Runwei Guan, Liye Jia, Shanliang Yao, Fengyufan Yang, Erick Purwanto, Ka Lok Man, Eng Gee Lim, Jeremy S. Smith, Xuming Hu, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2025 | Exploring Radar Data Representations in Autonomous Driving: A Comprehensive ReviewabstractWith the rapid advancements of sensor technology and deep learning, autonomous driving systems are providing safe and efficient access to intelligent vehicles as well as intelligent transportation. Among these equipped sensors, the radar sensor plays a crucial role in providing robust perception information in diverse environmental conditions. This review focuses on exploring different radar data representations utilized in autonomous driving systems. Firstly, we introduce the capabilities and limitations of the radar sensor by examining the working principles of radar perception and signal processing of radar measurements. Then, we delve into the generation process of five radar representations, including the ADC signal, radar tensor, point cloud, grid map, and micro-Doppler signature. For each radar representation, we examine the related datasets, methods, advantages and limitations. Furthermore, we discuss the challenges faced in these data representations and propose potential research directions. Above all, this comprehensive review offers an in-depth insight into how these representations enhance autonomous system capabilities, providing guidance for radar perception researchers. To facilitate retrieval and comparison of different data representations, datasets and methods, we provide an interactive website at https://radar-camera-fusion.github.io/radar. Shanliang Yao, Runwei Guan, Zitian Peng, Chenhang Xu, Yilu Shi, Weiping Ding 0001, Eng Gee Lim, Yong Yue 0001, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2025 | radarODE: An ODE-Embedded Deep Learning Model for Contactless ECG Reconstruction From Millimeter-Wave RadarabstractRadar-based cardiac monitoring has become a popular research direction recently, but the fine-grained electrocardiogram (ECG) signal is still hard to reconstruct from millimeter-wave radar signal. The key obstacle is to decouple cardiac activities in the electrical domain (i.e., ECG) from that in the mechanical domain (i.e., heartbeat), and most existing research only uses purely data-driven methods to map such domain transformation as a black box. Therefore, this work first proposes a signal model that considers the fine-grained cardiac feature sensed by radar, and a novel deep learning framework called radarODE is designed to extract both temporal and morphological features for generating ECG. In addition, ordinary differential equations are embedded in radarODE as a decoder to provide morphological prior, helping the convergence of the model training and improving the robustness under body movements. After being validated on the dataset, the proposed radarODE achieves better performance compared with the benchmark in terms of missed detection rate, root mean square error, Pearson correlation coefficient with improvements of 9%, 16% and 19%, respectively. The validation results imply that radarODE is capable of recovering ECG signals from radar signals with high fidelity and can potentially be implemented in real-life scenarios Runwei Guan, Rui Yang 0007, Yutao Yue, Eng Gee Lim |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | SFBM: Shared Feature Bias Mitigating for Long-Tailed Image RecognitionabstractLong-tailed distribution exists in real-world scenario and compromises the performance of recognition models. In this article, we point out that a neural network classifier has a shared feature bias, which tends to regard the shared features among different classes as head-class discriminative features, leading to misclassifications on tail-class samples under long-tailed scenarios. To solve this issue, we propose a shared feature bias mitigating (SFBM) framework. Specifically, we create two parallel classifiers trained concurrently with the baseline classifier, using our special training loss. The parallel classifier weight sums are then used for estimating the shared feature components in baseline classifier weights. Finally, we rectify the baseline classifier by removing the estimated shared feature components from it while supplementing the parallel classifier weights class by class to the rectified classifier weights, mitigating shared feature bias. Our proposed SFBM demonstrates broad compatibility with nearly all recognition methods while maintaining high computational efficiency, as it introduces no additional computation during inference. Extensive experiments on CIFAR10/100-LT, ImageNet-LT, and iNaturalist 2018 demonstrate that simply incorporating SFBM during the training phase consistently boosts the performance of various state-of-the-art methods by significant margins. The complete source code will be made publicly available at https://github.com/bzbz-bot/SFBM. Xinqiao Zhao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001, Jimin Xiao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | AlignMalloc: Warp-Aware Memory Rearrangement Aligned With UVM Prefetching for Large-Scale GPU Dynamic AllocationsabstractAs parallel computing tasks rapidly expand in both complexity and scale, the need for efficient GPU dynamic memory allocation becomes increasingly important. While progress has been made in developing dynamic allocators for substantial applications, their real-world applicability is still limited due to inefficient memory access behaviors. This paper introduces AlignMalloc, a novel memory management system that aligns with the Unified Virtual Memory (UVM) prefetching strategy, significantly enhancing both memory allocation and access performance in large-scale dynamic allocation scenarios. We analyze the fundamental inefficiencies in UVM access and first reveal the mismatch between memory access and UVM prefetching methods. To resolve this issue, AlignMalloc implements a warp-aware memory rearrangement strategy that exploits the regularity of warps to align with the UVM's static prefetching setup. Additionally, AlignMalloc introduces an OR tree-based structure within a host-co-managed framework to further optimize dynamic allocation. Comprehensive experiments demonstrate that AlignMalloc substantially outperforms current state-of-the-art systems, achieving up to$2.7 \times$improvement in dynamic allocation and$2.3 \times$in memory access. Additionally, eight real-world applications with diverse memory access patterns exhibit consistent performance enhancements, with average speedups$1.5 \times$. Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Eng Gee Lim, Keqin Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | Adversarial Erasing Transformer for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation has attracted a lot of attention recently. Previous methods can be divided into two types, which are single-stage training and multi-stage training. In this paper, we focus on multi-stage training for image-level weakly supervised semantic segmentation. Many recent methods have tried to use transformer architecture as the backbone for CAM generation since it can capture global relationships to refine CAM accurately. However, we observe that such a backbone still fails to generate complete and smooth CAM. We argue that this is because the attention mechanism in the transformer can only pay attention to the most discriminative relationships. It is difficult to capture semantic-level long-range pair-wise relationships under image-level supervision. Thus, we propose an adversarial erasing transformer network called AETN, where an erasing attention mechanism is designed to establish more extensive pair-wise relationships. To cope with erasing, more target features will be forced to activate. Thus, better feature representation can be obtained for more accurate CAM generation. Besides, to further help our network learn better feature representation, we propose a self-consistent learning mechanism based on different augmentations. In this way, our AETN outperforms recent methods. Our AETN achieves 73.0 mIoU on the PASCAL VOC 2012 val set and 73.9 mIoU on the PASCAL VOC 2012 test set. Code is available a https://github.com/siyueyu/AETN. Bingfeng Zhang, Siyue Yu, Xuru Gao, Mingjie Sun, Eng Gee Lim, Jimin Xiao |
ECAI | 5 |
| 2024 | ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave RadarabstractPanoptic Driving Perception (PDP) is critical for the autonomous navigation of Unmanned Surface Vehicles (USVs). A PDP model typically integrates multiple tasks, necessitating the simultaneous and robust execution of various perception tasks to facilitate downstream path planning. The fusion of visual and radar sensors is currently acknowledged as a robust and cost-effective approach. However, most existing research has primarily focused on fusing visual and radar features dedicated to object detection or utilizing a shared feature space for multiple tasks, neglecting the individual representation differences between various tasks. To address this gap, we propose a pair of Asymmetric Fair Fusion (AFF) modules with favorable explainability designed to efficiently interact with independent features from both visual and radar modalities, tailored to the specific requirements of object detection and semantic segmentation tasks. The AFF modules treat image and radar maps as irregular point sets and transform these features into a crossed-shared feature space for multitasking, ensuring equitable treatment of vision and radar point cloud features. Leveraging AFF modules, we propose a novel and efficient PDP model, ASY-VRNet, which processes image and radar features based on irregular super-pixel point sets. Additionally, we propose an effective multi-task learning method specifically designed for PDP models. Compared to other lightweight models, ASY-VRNet achieves state-of-the-art performance in object detection, semantic segmentation, and drivable-area segmentation on the WaterScenes benchmark. Our project is publicly available at https://github.com/GuanRunwei/ASY-VRNet. Runwei Guan, Shanliang Yao, Ka Lok Man, Yong Yue 0001, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
IROS | 7 |
| 2024 | A Lightweight and Responsive On-Line IDS Towards Intelligent Connected Vehicles System
Jia Liu 0074, Wenjun Fan, Yifan Dai 0006, Eng Gee Lim, Alexei Lisitsa 0001 |
SAFECOMP | 4 |
| 2024 | Leveraging Semi-supervised Learning for Enhancing Anomaly-based IDS in Automotive Ethernet
Jia Liu 0074, Wenjun Fan, Yifan Dai 0006, Eng Gee Lim, Zhoujin Pan, Alexei Lisitsa 0001 |
TrustCom | 4 |
| 2024 | Subspace-Based Semi-Blind Channel Estimation for User-Centric Cell-Free Massive MIMO SystemsabstractRecently, the user-centric (UC) cell-free (CF) massive multiple-input multiple-output (MIMO) system has emerged as one of the most promising enablers for future wireless networks. However, with the increase of the number of user equipments (UEs) in the network, pilot reuse for channel estimation causes interference, referred to as pilot contamination. In this paper, a subspace-based semi-blind channel estimation scheme is proposed for UC CF massive MIMO systems. By exploiting the uplink data signal received at each access point (AP), eigenvalue-decomposition (EVD) is employed on the covariance matrix of received signal to obtain the signal subspace. Based on the channel gains and pilot assignment results for each UE, a subspace selection scheme is developed to separate target signal from interference in subspace, which is followed by subspace projection to mitigate the pilot contamination. Simulations results verify the effectiveness of the proposed scheme in improving the uplink channel estimation accuracy of UC CF systems even without the prior knowledge of channel covariance matrix. Bowen Zhong, Xu Zhu 0001, Eng Gee Lim |
VTC Spring | 3 |
| 2024 | Discriminative Feature Enhancement Network for few-shot classification and beyond
Fangyu Wu 0001, Qiufeng Wang 0001, Qi Chen 0026, Eng Gee Lim |
Expert Syst. Appl. | 7 |
| 2024 | FindVehicle and VehicleFinder: a NER dataset for natural language-based vehicle retrieval and a keyword-based cross-modal vehicle retrieval systemabstractAbstract Natural language (NL) based vehicle retrieval is a task aiming to retrieve a vehicle that is most consistent with a given NL query from among all candidate vehicles. Because NL query can be easily obtained, such a task has a promising prospect in building an interactive intelligent traffic system (ITS). Current solutions mainly focus on extracting both text and image features and mapping them to the same latent space to compare the similarity. However, existing methods usually use dependency analysis or semantic role-labelling techniques to find keywords related to vehicle attributes. These techniques may require a lot of pre-processing and post-processing work, and also suffer from extracting the wrong keyword when the NL query is complex. To tackle these problems and simplify, we borrow the idea from named entity recognition (NER) and construct FindVehicle, a NER dataset in the traffic domain. It has 42.3k labelled NL descriptions of vehicle tracks, containing information such as the location, orientation, type and colour of the vehicle. FindVehicle also adopts both overlapping entities and fine-grained entities to meet further requirements. To verify its effectiveness, we propose a baseline NL-based vehicle retrieval model called VehicleFinder. Our experiment shows that by using text encoders pre-trained by FindVehicle, VehicleFinder achieves 87.7% precision and 89.4% recall when retrieving a target vehicle by text command on our homemade dataset based on UA-DETRAC [1]. From loading the command into VehicleFinder to identifying whether the target vehicle is consistent with the command, the time cost is 279.35 ms on one ARM v8.2 CPU and 93.72 ms on one RTX A4000 GPU, which is much faster than the Transformer-based system. The dataset is open-source via the link https://github.com/GuanRunwei/FindVehicle , and the implementation can be found via the link https://github.com/GuanRunwei/VehicleFinder-CTIM . Runwei Guan, Ka Lok Man, Feifan Chen, Shanliang Yao, Rongsheng Hu, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
Multim. Tools Appl. | 8 |
| 2024 | Prototype Guided Pseudo Labeling and Perturbation-based Active Learning for domain adaptive semantic segmentation
Junkun Peng, Mingjie Sun, Eng Gee Lim, Qiufeng Wang 0001, Jimin Xiao |
Pattern Recognit. | 3 |
| 2024 | Cross-frame feature-saliency mutual reinforcing for weakly supervised video salient object detection
Jian Wang 0122, Siyue Yu, Bingfeng Zhang, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
Pattern Recognit. | 6 |
| 2024 | Unified Multi-Modality Video Object Segmentation Using Reinforcement LearningabstractThe main task we aim to tackle is the multi-modality video object segmentation (VOS), which can be divided into two sub-tasks: mask-referred and language-referred VOS, where the first-frame mask-level or language-level label is utilized to provide the target information, respectively. Due to the huge gap between different modalities, existing works never come up with a unified framework for these two sub-tasks. In this work, such a unified framework is designed, where the visual and linguistic inputs are first spilt into a number of image patches and words, and then mapped into same-size tokens, which are equally processed by a self-attention based segmentation model. Furthermore, to highlight the significant information and discard the non-target or ambiguous one, unified multi-modality filter networks are further designed, and reinforcement learning is adopted to optimize such networks. Experiments show that new state-of-the-art performances are achieved by the proposed method: 52.8% ofJ&Fon Ref-YoutubeVOS dataset and 83.2% ofJSon YoutubeVOS dataset, respectively. The code will be released. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Cairong Zhao, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Class Activation Map Calibration for Weakly Supervised Semantic SegmentationabstractImage-level weakly supervised semantic segmentation (WSSS) has received substantial attention due to its cost-effective annotation process. In WSSS, Class Activation Maps (CAMs) generated via classifier weights tend to focus on the most discriminative region, while the CAMs derived from class prototypes are significantly enhanced to cover more complete regions. However, the prototype CAMs still exhibit limitations such as incomplete localization maps on target objects and the presence of background noise. In this paper, we propose a novel WSSS framework called Classifier-Prototype Mutual Calibration (CPMC) that leverages the characteristics of both classifier and prototype CAMs to address the above issues. Specifically, an iterative refinement strategy based on context feature dependency is applied to refine the original classifier CAMs, which helps to generate improved prototype CAMs. Subsequently, local prototypes are constructed based on the false negative regions and false positive regions extracted from the previous two CAMs, which contribute to completing missing parts of the target object and suppressing background noise respectively. Therefore, CPMC can alleviate the aforementioned issues. Extensive experimental results on standard WSSS benchmarks (PASCAL VOC and MS COCO) show that our method significantly improves the quality of CAMs and achieves state-of-the-art performance. Our source code will be released. Jian Wang 0122, Tianhong Dai, Xinqiao Zhao, Ángel F. García-Fernández, Eng Gee Lim, Jimin Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Large-Scale Traffic Signal Control Using Constrained Network Partition and Adaptive Deep Reinforcement LearningabstractMulti-agent Deep Reinforcement Learning (MADRL) based traffic signal control lbecomes a popular research topic in recent years. To alleviate the scalability issue of completely centralized reinforcement learning (RL) techniques and the non-stationarity issue of completely decentralized RL techniques on large-scale traffic networks, some literature utilizes a regional control approach where the whole network is firstly partitioned into multiple disjoint regions, followed by applying the centralized RL approach to each region. However, the existing partitioning rules either have no constraints on the topology of regions or require the same topology for all regions. Meanwhile, no existing regional control approach explores the performance of optimal joint action in an exponentially growing regional action space when intersections are controlled by 4-phase traffic signals (EW, EWL, NS, NSL). In this paper, we propose a novel RL training framework named RegionLight to tackle the above limitations. Specifically, the topology of regions is firstly constrained to a star network which comprises one center and an arbitrary number of leaves. Next, the network partitioning problem is modeled as an optimization problem to minimize the number of regions. Then, an Adaptive Branching Dueling Q-Network (ABDQ) model is proposed to decompose the regional control task into several joint signal control sub-tasks corresponding to particular intersections. Subsequently, these sub-tasks maximize the regional benefits cooperatively. Finally, the global control strategy for the whole network is obtained by concatenating the optimal joint actions of all regions. Experimental results demonstrate the superiority of our proposed framework over all baselines under both real and synthetic scenarios in all evaluation metrics. Hankang Gu, Shangbo Wang, Xiaoguang Ma, Dongyao Jia, Guoqiang Mao, Eng Gee Lim, Cheuk Pong Ryan Wong |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water SurfacesabstractAutonomous driving on water surfaces plays an essential role in executing hazardous and time-consuming missions, such as maritime surveillance, survivor rescue, environmental monitoring, hydrography mapping and waste cleaning. This work presents WaterScenes, the first multi-task 4D radar-camera fusion dataset for autonomous driving on water surfaces. Equipped with a 4D radar and a monocular camera, our Unmanned Surface Vehicle (USV) proffers all-weather solutions for discerning object-related information, including color, shape, texture, range, velocity, azimuth, and elevation. Focusing on typical static and dynamic objects on water surfaces, we label the camera images and radar point clouds at pixel-level and point-level, respectively. In addition to basic perception tasks, such as object detection, instance segmentation and semantic segmentation, we also provide annotations for free-space segmentation and waterline segmentation. Leveraging the multi-task and multi-modal data, we conduct benchmark experiments on the uni-modality of radar and camera, as well as the fused modalities. Experimental results demonstrate that 4D radar-camera fusion can considerably improve the accuracy and robustness of perception on water surfaces, especially in adverse lighting and weather conditions. WaterScenes dataset is public onhttps://waterscenes.github.io. Shanliang Yao, Runwei Guan, Zhaodong Wu, Yi Ni, Zile Huang, Ryan Wen Liu, Yong Yue 0001, Weiping Ding 0001, Eng Gee Lim, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2024 | Correction: ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim |
Vis. Comput. | 7 |
| 2024 | ITContrast: contrastive learning with hard negative synthesis for image-text matching
Fangyu Wu 0001, Qiufeng Wang 0001, Zhao Wang 0001, Siyue Yu, Yushi Li, Eng Gee Lim |
Vis. Comput. | 7 |
| 2023 | Interactive Rehabilitation Carpet for Children with Cerebral PalsyabstractChildren with cerebral palsy (CP) can experience complex gait deviations and need to go through intensive lower extremity rehabilitation exercises to develop and enhance their motor control in daily living. However, most of them cannot persist in the regular repetitive exercise sessions using hospital-based equipment. To provide a playful and attractive rehabilitation environment, an interactive carpet with interchangeable covers and varied step lengths is introduced to motivate children for lower extremity training. The vibrant colours, engaging games, visual and audio feedback are designed to increase the carpet-human interaction. This carpet can support gait exercise with five types of step lengths, which improves its accessibility and usability for children with CP. Yijia An, Qinglei Bu, Jie Sun 0024, Eng Gee Lim, Lijun Kong, Roshan Devaraj |
TEI | 4 |
| 2023 | Design and Development of a Mixed Reality Acupuncture Training SystemabstractThis paper looks at how mixed reality can be used for the improvement and enhancement of Chinese acupuncture practice through the introduction of an acupuncture training simulator. A prototype system developed for our study allows practitioners to insert virtual needles using their bare hands into a full-scale 3D representation of the human body with labelled acupuncture points. This provides them with a safe and natural environment to develop their acupuncture skills simulating the actual physical process of acupuncture. It also helps them to develop their muscle memory for acupuncture and better develops their memory of acupuncture points through a more immersive learning experience. We describe some of the design decisions and technical challenges overcome in the development of our system. We also present the results of a comparative user evaluation with potential users aimed at assessing the viability of such a mixed reality system being used as part of their training and development. The results of our evaluation reveal the training system outperformed in the enhancement of spatial understanding as well as improved learning and dexterity in acupuncture practice. These results go some way to demonstrating the potential of mixed reality for improving practice in therapeutic medicine. Qilei Sun, Jiayou Huang, Paul Craig, Lingyun Yu 0001, Eng Gee Lim |
VR | 6 |
| 2023 | Fully and Weakly Supervised Referring Expression Segmentation With End-to-End LearningabstractReferring Expression Segmentation (RES), which is aimed at localizing and segmenting the target according to the given language expression, has drawn increasing attention. Existing methods jointly consider the localization and segmentation steps, which rely on the fused visual and linguistic features for both steps. We argue that the conflict between the purpose of identifying an object and generating a mask limits the RES performance. To solve this problem, we propose a parallel position-kernel-segmentation pipeline to better isolate and then interact the localization and segmentation steps. In our pipeline, linguistic information will not directly contaminate the visual feature for segmentation. Specifically, the localization step localizes the target object in the image based on the referring expression, and then the visual kernel obtained from the localization step guides the segmentation step. This pipeline also enables us to train RES in a weakly-supervised way, where the pixel-level segmentation labels are replaced by click annotations on center and corner points. The position head is fully-supervised and trained with the click annotations as supervision, and the segmentation head is trained with weakly-supervised segmentation losses. To validate our framework on a weakly-supervised setting, we annotated three RES benchmark datasets (RefCOCO, RefCOCO+ and RefCOCOg) with click annotations. Our method is simple but surprisingly effective, outperforming all previous state-of-the-art RES methods on fully- and weakly-supervised settings by a large margin. The code and dataset will be released onhttps://github.com/detectiveli/PKS.git. Hui Li 0085, Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | MAN and CAT: mix attention to nn and concatenate attention to YOLO
Runwei Guan, Ka Lok Man, Haocheng Zhao, Ruixiao Zhang 0001, Shanliang Yao, Jeremy S. Smith, Eng Gee Lim, Yutao Yue |
J. Supercomput. | 7 |
| 2023 | Cycle-Free Weakly Referring Expression Grounding With Self-Paced LearningabstractIn this paper, we are tackling the weakly referring expression grounding task to localize the target object in an image according to a given query sentence, where the mapping between the query sentence and image regions is blind during the training period. Previous methods all follow a cyclic forward-backward pipeline to handle this task, where the query sentence is firstly converted to the result region through the forward module, and then the result region is converted back to a sentence through the backward module, with the difference between the reconstructed sentence and original query used as the loss to optimize the entire network. These existing methods, however, suffer from the deviation issue when the result region, generated through the forward module, totally deviates from the target area, but the backward module still reconstructs a similar sentence. The aforementioned loss function cannot penalize this kind of deviation because of the consistent prediction of the sentence. To overcome this limitation, we propose a cycle-free pipeline, where a region describer network is designed to predict the textual description for each candidate region, and a result region is selected according to the similarity between the predicted description and the query sentence. Furthermore, a self-paced learning mechanism is designed to avoid the drift issue during the warm-up period of the optimization process. The proposed method achieves a higher average accuracy on RefCOCO and RefCOCO+ datasets, compared with all previous state-of-the-art methods. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Starting Point Selection and Multiple-Standard Matching for Video Object Segmentation With Language AnnotationabstractIn this study, we investigate language-level video object segmentation, where first-frame language annotation is used to describe the target object. Because a language label is typically compatible with all frames in a video, the proposed method can choose the most suitable starting frame to mitigate initialization failure. Apart from extracting the visual feature from a static video frame, a motion-language score based on optical flow is also proposed to describe moving objects more accurately. Scores of multiple standards are then aggregated using an attention-based mechanism to predict the final result. The proposed method is evaluated on four widely-used video object segmentation datasets, including the DAVIS 2017, DAVIS 2016, SegTrack V2 and YouTubeObject datasets, and a novel accuracy measured as mean region similarity is obtained on both the DAVIS 2017 (67.2%) and DAVIS 2016 (83.5%) datasets. The code will be published. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Democracy Does Matter: Comprehensive Feature Mining for Co-Salient Object DetectionabstractCo-salient object detection, with the target of detecting co-existed salient objects among a group of images, is gaining popularity. Recent works use the attention mechanism or extra information to aggregate common co-salient features, leading to incomplete even incorrect responses for target objects. In this paper, we aim to mine comprehensive co-salient features with democracy and reduce background interference without introducing any extra information. To achieve this, we design a democratic prototype generation module to generate democratic response maps, covering sufficient co-salient regions and thereby involving more shared attributes of co-salient objects. Then a comprehensive prototype based on the response maps can be generated as a guide for final prediction. To suppress the noisy background information in the prototype, we propose a self-contrastive learning module, where both positive and negative pairs are formed without relying on additional classification information. Besides, we also design a democratic feature enhancement module to further strengthen the co-salient features by readjusting attention values. Extensive experiments show that our model obtains better performance than previous state-of-the-art methods, especially on challenging real-world cases (e.g., for CoCA, we obtain a gain of 2.0% for MAE, 5.4% for maximum F-measure, 2.3% for maximum E-measure, and 3.7% for S-measure) under the same settings. Source code is available at https://github.com/siyueyu/DCFM. Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim |
CVPR | 4 |
| 2022 | Clustering-based Pilot Assignment for User-Centric Cell-Free mmWave Massive MIMO SystemsabstractThe cell-free (CF) massive multiple-input multiple-output (MIMO) system, combined with millimeter wave (mmWave), is considered as one of the most promising enablers for future wireless networks. However, with the increase of user equipments (UEs) in the network, the pilot sequence reuse will bring interference during the channel estimation period, which is called pilot contamination. In this paper, an efficient pilot assignment scheme called the k-means clustering-based pilot assignment (KCPA) is proposed to mitigate the severe pilot contamination in user-centric (UC) CF massive MIMO systems at mmWave frequencies. By exploiting the knowledge of UE locations, the pilot contamination could be reduced through structured pilot assignment based on the UE clustering result. In addition, a pilot assignment scheme based on the Tabu search, called the k-means clustering-based Tabu search pilot assignment (KCTSPA) is devised to further improve the performance. Numerical results verify that the proposed pilot assignment schemes can greatly reduce the uplink channel estimation error of mmWave UC systems with a low complexity compared to the conventional pilot assignment schemes. Bowen Zhong, Xu Zhu 0001, Eng Gee Lim |
VTC Fall | 3 |
| 2022 | Grant-Free Communications With Adaptive Period for IIoT: Sparsity and Correlation-Based Joint Channel Estimation and Signal DetectionabstractIn this article, we investigate the grant-free communications with adaptive period for Industrial Internet of Things, where only a fraction of devices is active at a time. To the best of our knowledge, this is the first work to exploit the noncontinuous temporal correlation of the received signal for joint user activity detection (UAD), channel estimation, and signal detection, while all the previous work requires continuous transmission. Two schemes are proposed toward this purpose, namely, periodic block orthogonal matching pursuit (PBOMP) and periodic block sparse Bayesian learning (PBSBL), which outperform the previous schemes in terms of the success rate of UAD, bit error rate, and accuracy in period estimation and channel estimation. The Cramér–Rao lower bounds (CRLBs) of channel estimation by PBOMP and PBSBL are derived. It is shown that the two proposed approaches have close CRLBs and normalized mean-square error at high SNR. Yuanchen Wang, Xu Zhu 0001, Eng Gee Lim, Zhongxiang Wei, Yufei Jiang |
IEEE Internet Things J. | 3 |
| 2022 | Transformer-Based Language-Person Search With Multiple Region SlicingabstractLanguage-person search is an essential technique for applications like criminal searching, where it is more feasible for a witness to provide language descriptions of a suspect than providing a photo. Most existing works treat the language-person pair as a black-box, neither considering the inner structure in a person picture, nor the correlations between image regions and referring words. In this work, we propose a transformer-based language-person search framework with matching conducted between words and image regions, where a person picture is vertically separated into multiple regions using two different ways, including the overlapped slicing and the key-point-based slicing. The co-attention between linguistic referring words and visual features are evaluated via transformer blocks. Besides the obtained outstanding searching performance, the proposed method enables to provide interpretability by visualizing the co-attention between image parts in the person picture and the corresponding referring words. Without bells and whistles, we achieve the state-of-the-art performance on the CUHK-PEDES dataset with Rank-1 score of 57.67% and the PA100K dataset with mAP of 22.88%, with simple yet elegant design. Code is available onhttps://github.com/detectiveli/T-MRS. Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Improved Camshift Algorithm in AGV Vision-based Tracking with Edge Computing
Tongpo Zhang, Xiaokai Nie, Xu Zhu 0001, Eng Gee Lim, Fei Ma 0002, Limin Yu |
J. Supercomput. | 4 |
| 2022 | A Semi-Blind Multiuser SIMO GFDM System in the Presence of CFOs and IQ ImbalancesabstractIn this paper, we investigate an open topic of a multiuser single-input-multiple-output (SIMO) generalized frequency division multiplexing (GFDM) system in the presence of carrier frequency offsets (CFOs) and in-phase/quadrature-phase (IQ) imbalances. A low-complexity semi-blind joint estimation scheme of multiple channels, CFOs and IQ imbalances is proposed. By utilizing the subspace approach, CFOs and channels corresponding to$U$users are first separated into$U$groups. For each individual user, CFO is extracted by minimizing the smallest eigenvalue whose corresponding eigenvector is utilized to estimate channel blindly. The IQ imbalance parameters are estimated jointly with channel ambiguities by very few subcarriers. The proposed scheme is feasible for a wider range of receive antennas number and has no constraints on the assignment scheme of subsymbols and subcarriers, modulation type, cyclic prefix length and the number of subsymbols per GFDM symbol. Simulation results show that the proposed scheme significantly outperforms the existing methods in terms of bit error rate, outage probability, mean-square-errors of CFO estimation, channel and IQ imbalance estimation, while at much higher spectral efficiency and lower computational complexity. The Cramér-Rao lower bound is derived to verify the effectiveness of the proposed scheme, which is shown to be close to simulation results. Yujie Liu 0001, Xu Zhu 0001, Eng Gee Lim, Yufei Jiang, Yi Huang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2021 | Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency CoherenceabstractSparse labels have been attracting much attention in recent years. However, the performance gap between weakly supervised and fully supervised salient object detection methods is huge, and most previous weakly supervised works adopt complex training methods with many bells and whistles. In this work, we propose a one-round end-to-end training approach for weakly supervised salient object detection via scribble annotations without pre/post-processing operations or extra supervision data. Since scribble labels fail to offer detailed salient regions, we propose a local coherence loss to propagate the labels to unlabeled regions based on image features and pixel distance, so as to predict integral salient regions with complete object structures. We design a saliency structure consistency loss as self-consistent mechanism to ensure consistent saliency maps are predicted with different scales of the same image as input, which could be viewed as a regularization technique to enhance the model generalization ability. Additionally, we design an aggregation module (AGGM) to better integrate high-level features, low-level features and global context information for the decoder to aggregate various information. Extensive experiments show that our method achieves a new state-of-the-art performance on six benchmarks (e.g. for the ECSSD dataset: Fβ = 0.8995, Eξ = 0.9079 and MAE = 0.0489), with an average gain of 4.60% for F-measure, 2.05% for E-measure and 1.88% for MAE over the previous best performing method on this task. Source code is available at http://github.com/siyueyu/SCWSSOD. Siyue Yu, Bingfeng Zhang, Jimin Xiao, Eng Gee Lim |
AAAI | 4 |
| 2021 | Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement LearningabstractIn this paper, we are tackling the proposal-free referring expression grounding task, aiming at localizing the target object according to a query sentence, without relying on off-the-shelf object proposals. Existing proposal-free methods employ a query-image matching branch to select the highest-score point in the image feature map as the target box center, with its width and height predicted by another branch. Such methods, however, fail to utilize the contextual relation between the target and reference objects, and lack interpretability on its reasoning procedure. To solve these problems, we propose an iterative shrinking mechanism to localize the target, where the shrinking direction is decided by a reinforcement learning agent, with all contents within the current image patch comprehensively considered. Besides, the sequential shrinking processes enable to demonstrate the reasoning about how to iteratively find the target. Experiments show that the proposed method boosts the accuracy by 4.32% against the previous state-of-the- art (SOTA) method on the RefCOCOg dataset, where query sentences are long and complex with many targets referred by other reference objects. Mingjie Sun, Jimin Xiao, Eng Gee Lim |
CVPR | 3 |
| 2021 | Monoscopic vs. Stereoscopic Views and Display Types in the Teleoperation of Unmanned Ground Vehicles for Object AvoidanceabstractVirtual reality (VR) head-mounted displays (HMD) have recently been used to provide an immersive, first-person vision/view in real-time for manipulating remotely-controlled unmanned ground vehicles (UGV). The teleoperation of UGV can be challenging for operators when it is done in real time. One big challenge is for operators to perceive quickly and rapidly the distance of objects that are around the UGV while it is moving. In this research, we explore the use of monoscopic and stereoscopic views and display types (immersive and non-immersive VR) for operating vehicles remotely. We conducted two user studies to explore their feasibility and advantages. Results show a significantly better performance when using an immersive display with stereoscopic view for dynamic, real-time navigation tasks that require avoiding both moving and static obstacles. The use of stereoscopic view in an immersive display in particular improved user performance and led to better usability. Jialin Wang 0002, Hai-Ning Liang, Shan Luo 0001, Eng Gee Lim |
RO-MAN | 5 |
| 2021 | Discriminative Triad Matching and Reconstruction for Weakly Referring Expression GroundingabstractIn this paper, we are tackling the weakly-supervised referring expression grounding task, for the localization of a referent object in an image according to a query sentence, where the mapping between image regions and queries are not available during the training stage. In traditional methods, an object region that best matches the referring expression is picked out, and then the query sentence is reconstructed from the selected region, where the reconstruction difference serves as the loss for back-propagation. The existing methods, however, conduct both the matching and the reconstruction approximately as they ignore the fact that the matching correctness is unknown. To overcome this limitation, a discriminative triad is designed here as the basis to the solution, through which a query can be converted into one or multiple discriminative triads in a very scalable way. Based on the discriminative triad, we further propose the triad-level matching and reconstruction modules which are lightweight yet effective for the weakly-supervised training, making it three times lighter and faster than the previous state-of-the-art methods. One important merit of our work is its superior performance despite the simple and neat design. Specifically, the proposed method achieves a new state-of-the-art accuracy when evaluated on RefCOCO (39.21 percent), RefCOCO+ (39.18 percent) and RefCOCOg (43.24 percent) datasets, that is 4.17, 4.08 and 7.8 percent higher than the previous one, respectively. The code is available at https://github.com/insomnia94/DTWREG. Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu 0001, John Yannis Goulermas |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Progressive sample mining and representation learning for one-shot person re-identification
Hui Li 0085, Jimin Xiao, Mingjie Sun, Eng Gee Lim, Yao Zhao 0001 |
Pattern Recognit. | 4 |
| 2021 | Fast pixel-matching for video object segmentation
Siyue Yu, Jimin Xiao, Bingfeng Zhang, Eng Gee Lim, Yao Zhao 0001 |
Signal Process. Image Commun. | 4 |
| 2021 | Improving Multi-Hop Time Synchronization Performance in Wireless Sensor Networks Based on Packet-Relaying Gateways With Per-Hop Delay CompensationabstractBased on the reverse asymmetric time synchronization framework, we have proposed several schemes with a major focus on the energy efficiency and computational complexity of a large number of battery-powered, low-cost sensor nodes in wireless sensor networks (WSNs). To address the cumulative end-to-end synchronization error, we have also introduced an idea of compensating for the processing delays at packet-relaying gateways as an energy-efficient way of multi-hop extension of WSN time synchronization schemes. In this paper, we present a comprehensive analysis of the multi-hop extension of WSN time synchronization schemes based on packet-relaying gateways with the per-hop delay compensation and the results of extensive experiments for the energy-efficient time synchronization schemes based on the reverse asymmetric time synchronization framework together with the flooding time synchronization protocol as a representative of existing schemes. Experimental results based on a real testbed demonstrate that the multi-hop extension based on packet-relaying gateways with the per-hop delay compensation greatly improves the performance of time synchronization of all the schemes considered compared to the multi-hop extension based on the conventional time-translating gateways. Xintao Huan, Kyeong Soo Kim, Sanghyuk Lee, Eng Gee Lim, Alan Marshall 0001 |
IEEE Trans. Commun. | 4 |
| 2021 | Deep Learning Based Multistep Solar Forecasting for PV Ramp-Rate Control Using Sky ImagesabstractSolar forecasting is one of the most promising approaches to address the intermittent photovoltaic (PV) power generation by providing predictions before upcoming ramp events. In this article, a novel multistep forecasting (MSF) scheme is proposed for PV power ramp-rate control (PRRC). This method utilizes an ensemble of deep ConvNets without additional time series models (e.g., recurrent neural network (RNN) or long short-term memory) and exogenous variables, thus more suitable for industrial applications. The MSF strategy can make multiple predictions in comparison with a single forecasting point produced by a conventional method while maintaining the same high temporal resolution. Besides, stacked sky images that integrate temporal-spatial information of cloud motions are used to further improve the forecasting performance. The results demonstrate a favorable forecasting accuracy in comparison to the existing forecasting models with the highest skill score of 17.7%. In the PRRC application, the MSF-based PRRC can detect more ramp-rates violations with a higher control rate of 98.9% compared with the conventional forecasting-based control. Thus, the PV generation can be effectively smoothed with less energy curtailment on both clear and cloudy days using the proposed approach. Yang Du 0005, Xiaoyang Chen 0006, Eng Gee Lim, Huiqing Wen, Lin Jiang 0001, Wei Xiang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2020 | Fast Template Matching and Update for Video Object Tracking and SegmentationabstractIn this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely adopted to handle this task, and the challenges lie in the selection of the matching method to predict the result as well as to decide whether to update the target template using the newly predicted result. The existing methods, however, make these selections in a rough and inflexible way, compromising their performance. To overcome this limitation, we propose a novel approach which utilizes reinforcement learning to make these two decisions at the same time. Specifically, the reinforcement learning agent learns to decide whether to update the target template according to the quality of the predicted result. The choice of the matching method will be determined at the same time, based on the action history of the reinforcement learning agent. Experiments show that our method is almost 10 times faster than the previous state-of-the-art method with even higher accuracy (region similarity of 69.1% on DAVIS 2017 dataset). Mingjie Sun, Jimin Xiao, Eng Gee Lim, Bingfeng Zhang, Yao Zhao 0001 |
CVPR | 3 |
| 2020 | Compressive Sensing based User Activity Detection and Channel Estimation in Uplink NOMA SystemsabstractConventional request-grant based non-orthogonal multiple access (NOMA) incurs tremendous overhead and high latency. To enable grant-free access in NOMA systems, user activity detection (UAD) is essential. In this paper, we investigate compressive sensing (CS) aided UAD, by utilizing the property of quasi-time-invariant channel tap delays as the prior information. This does not require any prior knowledge of the number of active users like the previous approaches, and therefore is more practical. Two UAD algorithms are proposed, which are referred to as gradient based and time-invariant channel tap delays assisted CS (g-TIDCS) and mean value based and TIDCS (m-TIDCS), respectively. They achieve much higher UAD accuracy than the previous work at low signal-to-noise ratio (SNR). Based on the UAD results, we also propose a low-complexity CS based channel estimation scheme, which achieves higher accuracy than the previous channel estimation approaches. Yuanchen Wang, Xu Zhu 0001, Eng Gee Lim, Zhongxiang Wei, Yujie Liu 0001, Yufei Jiang |
WCNC | 3 |
| 2020 | Adaptive ROI generation for video object segmentation using reinforcement learning
Mingjie Sun, Jimin Xiao, Eng Gee Lim, Yanchun Xie, Jiashi Feng |
Pattern Recognit. | 3 |
| 2020 | A Beaconless Asymmetric Energy-Efficient Time Synchronization Scheme for Resource-Constrained Multi-Hop Wireless Sensor NetworksabstractThe ever-increasing number of WSN deployments based on a large number of battery-powered, low-cost sensor nodes, which are limited in their computing and power resources, puts the focus of WSN time synchronization research on three major aspects of accuracy, energy consumption, and computational complexity. In the literature, the latter two aspects haven't received much attention compared to the accuracy of WSN time synchronization. Especially in multi-hop WSNs, intermediate gateway nodes are overloaded with tasks for not only relaying messages but also a variety of computations for their offspring nodes as well as themselves. Therefore, not only minimizing the energy consumption but also lowering the computational complexity while maintaining the synchronization accuracy is crucial to the design of time synchronization schemes for resource-constrained sensor nodes. In this paper, focusing on the three aspects of WSN time synchronization, we introduce a framework of reverse asymmetric time synchronization for resource-constrained multi-hop WSNs and propose a beaconless energy-efficient time synchronization scheme based on reverse one-way message dissemination. Experimental results with a WSN testbed based on TelosB motes running TinyOS demonstrate that the proposed scheme conserves up to 95% energy consumption compared to the flooding time synchronization protocol while achieving microsecond-level synchronization accuracy. Xintao Huan, Kyeong Soo Kim, Sanghyuk Lee, Eng Gee Lim, Alan Marshall 0001 |
IEEE Trans. Commun. | 4 |
| 2020 | Wearable EBG-Backed Belt Antenna for Smart On-Body ApplicationsabstractThis article presents an innovative belt antenna with an electromagnetic band-gap (EBG) ground plane made of textile materials. The antenna can be applied in a smart belt system to set up a communication link with other electronic devices and/or host a variety of sensors to track human motions. The proposed belt antenna works at 2.45 GHz in the industrial, scientific, and medical radio band for Bluetooth low energy communications. Considering the effect the human body would have on the performance of a belt antenna, a textile ground plane is designed to be integrated into the trouser fabric behind the belt to provide isolation from the body and simultaneously improve antenna radiation characteristics. Through the application of the ground plane, the belt antenna achieves a maximum realized gain of 7.94 dBi and a minimum specific absorption rate of 0.04 W/kg at 0.5 W input power. During the design process, characteristic mode analysis is used to explore the underlining principle and further optimize the antenna performance. Two typical EBG structures are analyzed in detail for this application scenario. The suspended transmission line method is used to evaluate EBG performance variations when the textile ground plane is bent. A prototype of such a system is fabricated and tested. Experimental results shows that the belt antenna, together with the textile EBG ground plane, is an excellent candidate for a smart belt system with desirable radiation pattern, efficiency, and safety limit. Rui Pei, Mark Leach, Eng Gee Lim, Zhao Wang 0001, Chaoyun Song, Jingchen Wang, Wenzhang Zhang, Zhenzhen Jiang, Yi Huang 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2019 | Semi-Blind Joint Multi-CFO and Multi-Channel Estimation for GFDMA with Arbitrary Carrier AssignmentabstractWe propose a low-complexity semi-blind joint multi- carrier frequency offset (CFO) and multi-channel estimation scheme for uplink generalized frequency division multiple access (GFDMA) systems. To the best of our knowledge, this is the first work to investigate the estimation of both CFOs and channels for a wide range of GFDMA systems, allowing arbitrary carrier assignment, modulation type and cyclic prefix length, and a wide range of the number of receive antennas. Thanks to the orthogonality between noise subspace and each signal subspace of U users, a complex U-CFO and U-channel estimation problem is decomposed into 2U one-dimensional problems, and solved in a semi-blind manner. Also, the multi-CFO compensation is performed at receiver rather than transmitter, avoiding spectral overhead due to feedback of multiple CFOs. Simulation results show that the proposed scheme significantly outperforms the existing methods in terms of bit error rate (BER) and root-mean-square-errors (RMSEs) of CFO and channel estimation, at much lower computational complexity than the existing methods. Yujie Liu 0001, Xu Zhu 0001, Eng Gee Lim, Yufei Jiang, Yi Huang 0001 |
GLOBECOM | 3 |
| 2019 | Robust Semi-Blind Estimation of Channel and CFO for GFDM SystemsabstractWe propose a robust semi-blind estimation scheme of channel and carrier frequency offset (CFO) for generalized frequency division multiplexing (GFDM) systems. This, to the best of our knowledge, is the first work to propose an integral solution to channel and full-range CFO for a wide range of GFDM systems. Based on the derived equivalent system model with CFO included implicitly, a subspace based method is proposed to perform initial channel estimation blindly, which requires only a small number of received symbols to achieve the second order statistics of the received signal. Then, CFO estimation and channel ambiguity elimination are undertaken in series by utilizing a small number of nulls and pilots in a single sub-symbol. Both channel and CFO estimations are more robust against inter-carrier interference (ICI) and inter-symbol interference (ISI) caused by the nonorthogonal filter of GFDM, compared to the existing methods. The proposed scheme achieves a bit error rate (BER) performance close to the ideal case with perfect CFO and channel estimations especially at medium and high signal-to-noise-ratios (SNRs). Yujie Liu 0001, Xu Zhu 0001, Eng Gee Lim, Yufei Jiang, Yi Huang 0001 |
ICC | 3 |
| 2019 | Electric Vehicles Assisted Multi-Household Cooperative Demand Response StrategyabstractThe recent ongoing development of electrical vehicles (EVs) offers vast benefits not only in environmental protection and economics, but also in demand response (DR) management on consumer side. Adopting EVs in DR enables householders to alleviate the load burden while reducing electric bill simultaneously. In this paper, we utilize EVs as temporary energy storage facilities to assist the power transaction, which ensures the flexibility and economic benefit. An innovative EVs assisted DR strategy including a neighbor energy sharing (NES) model is proposed, to jointly optimize the load distribution via vehicle to home (V2H) and vehicle to neighbor (V2N) connections, and economic cost for a residential network with multi-household. The effectiveness of the proposed DR strategy is verified by numerical results in terms of load balancing and cost reduction. It also significantly outperforms the previous DR approaches. Xu Zhu 0001, Eng Gee Lim, Wolfgang Kellerer |
VTC Spring | 3 |
| 2019 | Fast Iterative Semi-Blind Receiver for URLLC in Short-Frame Full-Duplex Systems With CFOabstractWe propose an iterative semi-blind (ISB) receiver structure to enable ultra-reliable low-latency communications in short-frame full-duplex (FD) systems with carrier frequency offset (CFO). To the best of our knowledge, this is the first paper to propose an integral solution to channel estimation and CFO estimation for short-frame FD systems by utilizing a single pilot. By deriving an equivalent system model with the CFO included implicitly, a subspace-based blind channel estimation is proposed at the initial stage, followed by CFO estimation and channel ambiguities elimination. Then, the refinement of channel and the CFO estimates is conducted iteratively. The integer and fractional parts of CFO in the full range are estimated as a whole and in closed-form at each iteration. The proposed ISB receiver significantly outperforms the previous methods in terms of frame error rate, mean square errors of channel estimation and CFO estimation and output signal-to-interference-and-noise ratio, while at a halved spectral overhead. Cramér-Rao lower bounds are derived to verify the effectiveness of the proposed ISB receiver structure. It also demonstrates high-computational efficiency as well as the fast convergence speed. Yujie Liu 0001, Xu Zhu 0001, Eng Gee Lim, Yufei Jiang, Yi Huang 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2018 | Iterative Semi-Blind CFO Estimation, SI Cancelation and Signal Detection for Full-Duplex SystemsabstractWe propose an iterative semi-blind carrier frequency offset (CFO) estimation, self-interference (SI) cancelation and signal detection scheme for full-duplex (FD) orthogonal frequency division multiplexing (OFDM) systems. To the best of our knowledge, this is the first work to consider signal detection of FD systems in the presence of both CFO and SI. The CFO estimation, SI cancelation and signal detection are performed initially by a subspace based semi-blind method, which are then enhanced significantly by performing iterations among them. Its CFO compensation is performed on the desired signal estimate, avoiding the introduction of CFO to the SI. The pilots for the desired signal and SI are carefully designed to enable simultaneous transmission of them to achieve FD training mode. Simulation results show that, the proposed iterative scheme, with much lower training overhead, demonstrates a significant performance enhancement over the existing methods. By utilizing the second order statistics of the received signal, a much superior bit error rate (BER) performance can be achieved compared to the case with perfect SI cancelation and CFO compensation. Its output signal-to- interference-and-noise-ratio (SINR) is close to that with perfect SI cancelation, and robust against the input signal-to-interference ratio (SIR). Yujie Liu 0001, Xu Zhu 0001, Eng Gee Lim, Yufei Jiang, Yi Huang 0001 |
GLOBECOM | 3 |
| 2018 | High-Accuracy Joint Multi-CFO and Multi-TOA Estimation for Multiuser SIMO OFDM SystemsabstractA joint multi-carrier frequency offset (CFO) and multi- time of arrival (TOA) estimation algorithm for multiuser single-input multiple-output (SIMO) orthogonal frequency division multiplexing (OFDM) systems is proposed. With carefully designed pilots, multiple CFOs and TOAs of K users are separated jointly, dividing a complex 2K-dimensional estimation problem into 2K low-complexity mono-dimensional estimation problems. Two CFO estimation approaches, including a low-complexity closed-form solution and a high-accuracy null-subcarrier assisted accurate estimation approach, are proposed, where the integer and fractional parts of each CFO are estimated as a whole rather separately. Each TOA is estimated regardless of CFO by exploring the features of the inter-carrier interference matrix. The Cramer-Rao lower bounds (CRLBs) of multi-CFO and mutli-TOA estimation are derived for the first time for SIMO OFDM systems. Simulation results show that the proposed CFO and TOA estimators provide higher estimation accuracy than the existing approaches. They also achieve performances close to the CRLBs especially at high signal-to-noise-ratios (SNRs). Yujie Liu 0001, Xu Zhu 0001, Eng Gee Lim, Yufei Jiang, Yi Huang 0001 |
ICC | 3 |
| 2017 | Exploring the effectiveness of student-generated video tutorials in electronic lab-based teachingabstractLab-based teaching in which hands-on experiments are to be conducted by students takes an important part for a wide range of engineering and science disciplines. In our current practice, the lab-based teaching involves live demonstration and tutorials after the off-line lab manual review. This has become particularly problematic when the number of students is large and insufficiency on the lab-supporting system becomes a common issue. In the meantime, even with a small number of students, it can be interesting to prepare the lab in a one-to-one tutorial. Well-designed video tutorials eliminate the time and space constraints on learning and provide comprehensive details to students to enable them focus on deepening the understanding of concepts, rather than spending majority of time on trouble shooting during the lab. With the full technical support from the Digital Learning Resources Hub at Xi'an Jiaotong-Liverpool University, we propose to involve student volunteers to generate a series of customized video tutorials for our electronics lab-based teaching practice. From the viewpoint of students themselves, these video tutorials are carefully designed based on students' learning needs to seamlessly integrate a wide range of theoretical and practical information. Under the supervision of staff, volunteers will be able to repeat their learning cycles with a different role and enhance their own understanding and knowledge structures, promoting the student-centered education model. Some scenery-based video tutorials will be used in the online quizzes questions to better prepare the students. The generated video tutorials will be shared across a number of electronic engineering modules to further investigate the effectiveness of these video tutorials. The effectiveness can be further explored using online questionnaires and online quizzes and evaluated by comparative studies. Shaofeng Lu, Xiaoyang Wang 0007, Yang Du 0005, Eng Gee Lim |
FIE | 5 |
| 2017 | Energy-Efficient Time Synchronization Based on Asynchronous Source Clock Frequency Recovery and Reverse Two-Way Message Exchanges in Wireless Sensor NetworksabstractWe consider energy-efficient time synchronization in a wireless sensor network where a head node is equipped with a powerful processor and supplied power from outlet, and sensor nodes are limited in processing and battery-powered. It is thisasymmetrythat our study focuses on; unlike most existing schemes to save the power of all network nodes, we concentrate on battery-powered sensor nodes in minimizing energy consumption for time synchronization. We present a time synchronization scheme based on asynchronous source clock frequency recovery and reverse two-way message exchanges combined with measurement data report messages, where we minimize the number of message transmissions from sensor nodes while achieving sub-microsecond time synchronization accuracy through propagation delay compensation. We carry out the performance analysis of the estimation of both measurement time and clock frequency with lower bounds for the latter. Simulation results verify that the proposed scheme outperforms the schemes based on conventional two-way message exchanges with and without clock frequency recovery in terms of the accuracy of measurement time estimation and the number of message transmissions and receptions at sensor nodes as an indirect measure of energy efficiency. Kyeong Soo Kim, Sanghyuk Lee, Eng Gee Lim |
IEEE Trans. Commun. | 3 |
| 2016 | Wearable antenna design for bioinformationabstractThis paper is a study of wearable antenna design for medical applications. A literature review of existing wearable systems is performed, with specific attention paid to the antenna element. Two antennas working at 2.4 GHz were simulated using a software tool; firstly a basic rectangular patch on FR4 substrate and the other on soft textile material. The bending performance of the soft textile antenna was investigated. Rui Pei, Jing Chen Wang, Mark Leach, Zhao Wang 0001, Sanghyuk Lee, Eng Gee Lim |
CIBCB | 6 |
| 2016 | Semi-blind precoding aided ML CFO estimation for ICA based MIMO OFDM systemsabstractWe propose a semi-blind precoding aided maximum likelihood (ML) carrier frequency offset (CFO) estimation method and a precoding aided equalization based on independent component analysis (ICA) receiver structure for multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) wireless communication systems. By carefully designing a constant in the precoding process, the power between reference data and source data can be balanced to enable ML CFO estimation and ambiguity elimination for the ICA output signals at the receiver. This proposed semi-blind non-redundant structure is much more bandwidth-and-energy efficient than the pilot aided ML CFO estimation method, as no real-time training or extra transmission power is required. Simulation results show that the proposed scheme provides a bit error rate (BER) performance the same as the performance of the pilot aided ML CFO estimation method, and close to the ideal case with perfect channel state information (CSI) and no CFO. Yufei Jiang, Xu Zhu 0001, Eng Gee Lim, Yi Huang 0001, Zhongxiang Wei, Hai Lin 0001 |
ICC | 3 |
| 2014 | ICA based joint semi-blind equalization and CFO estimation for OFDMA systemsabstractWe propose a joint independent component analysis (ICA) based equalization and carrier frequency offset (CFO) estimation scheme for orthogonal frequency division multiple access (OFDMA) systems, which requires only a single pilot, resulting in a very low spectral overhead. On the one hand, a low-complexity ambiguity elimination method is proposed in the ICA equalized signals. By designing and exploring a fixed order of correlation in the transmitted signals, the permutation ambiguity problem is solved, while the remaining quadrant ambiguity is eliminated by one pilot symbol. One the another hand, linear CFO estimation is performed with a closed-form solution based on the phase difference between adjacent rows in the CFO-corrupted channel structure. Simulation results show that the proposed semi-blind ICA based scheme not only outperforms some existing CFO estimation approaches, but also provides a bit error rate (BER) performance, comparable to the ideal case with perfect channel state information (CSI) and no CFO. Yufei Jiang, Xu Zhu 0001, Eng Gee Lim, Yi Huang 0001, Hai Lin 0001 |
GLOBECOM | 3 |
| 2014 | Low-complexity frequency synchronization for ICA based semi-blind CoMP systems with ICI and phase rotation caused by multiple CFOsabstractWe propose a low-complexity frequency synchronization approach for semi-blind independent component analysis (ICA) based coordinated multi-point (CoMP) orthogonal frequency division multiplexing (OFDM) systems, with multiple carrier frequency offsets (CFOs). The key idea is to introduce a short pilot for both multi-CFO estimation and ambiguity elimination in the ICA equalized signals. First, by minimizing the cross-correlation between original pilots and received pilots with implicit phase rotation correction, the one-dimensional search based CFO estimation can be performed without trial ICI compensation and channel state information (CSI). Second, by maximizing the real part of the cross-correlation between ICA equalized pilots and original pilots, the same pilots can be used again to eliminate the permutation and quadrant ambiguity in the ICA equalized signals. Simulation results show that, with a very low training overhead of 2%, the proposed multi-CFO estimation approach outperforms existing methods. Also, the proposed semi-blind CoMP system can achieve a bit error rate (BER) performance close to the ideal case with perfect CSI and no CFO. Yufei Jiang, Xu Zhu 0001, Eng Gee Lim, Yi Huang 0001, Hai Lin 0001 |
ICC | 3 |
| 2013 | Semi-blind MIMO OFDM systems with precoding aided CFO estimation and ICA based equalizationabstractWe propose a semi-blind multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) system, with a precoding aided carrier frequency offset (CFO) estimation approach, and an independent component analysis (ICA) based equalization structure. A number of reference data sequences are carefully designed offline and are superimposed to source data via a non-redundant linear precoding process, which can kill two birds with one stone, without introducing any extra total transmit power and spectral overhead. First, the reference data sequences are selected from a pool of carefully designed orthogonal sequences. The CFO estimation is to minimize the sum cross-correlation between the CFO compensated signals and the rest orthogonal sequences in the pool. Second, the same reference data enable elimination of the permutation and quadrant ambiguity in the ICA equalized signals by maximizing the cross-correlation between the ICA equalized signals and the reference data. Simulation results show that, without extra bandwidth and power needed, the proposed semi-blind system achieves a bit error rate (BER) performance close to the ideal case with perfect channel state information (CSI) and no CFO. Also, the precoding aided CFO estimation outperforms the constant amplitude zero autocorrelation (CAZAC) sequences based CFO estimation approach, with no spectral overhead. Yufei Jiang, Xu Zhu 0001, Eng Gee Lim, Hai Lin 0001, Yi Huang 0001 |
GLOBECOM | 3 |
| 2013 | Joint semi-blind channel equalization and ICI mitigation for carrier aggregation based CoMP OFDMA systems with multiple CFOsabstractWe propose a joint semi-blind equalization and inter-carrier interference (ICI) mitigation scheme for multiple carrier frequency offsets (CFOs) corrupted signals in carrier aggregation (CA) based multiple access coordinated multi-point (CoMP) systems with OFDMA. The CFO-induced ICI is mitigated implicitly via the independent component analysis (ICA) based semi-blind equalization, without requiring an explicit process of estimation of multiple CFOs. Only a small number of pilots are used to resolve the remaining indeterminacies in the ICA equalized signals, introducing a very low training overhead. Simulation results show that the proposed semi-blind ICA based equalization scheme provides a bit error rate (BER) performance closed to the ideal case with perfect CSI and no CFO, and also outperforms the approach with the constant amplitude zero autocorrelation (CAZAC) sequences based explicit CFO estimation. Yufei Jiang, Xu Zhu 0001, Eng Gee Lim, Yi Huang 0001 |
ICC | 3 |
| 2012 | Semi-blind CoMP system with multiple-CFO estimation and ICA based equalizationabstractWe propose an orthogonal frequency division multiplexing (OFDM) based semi-blind coordinated multi-point (CoMP) system, with a low complexity estimation approach for multiple carrier frequency offsets (CFOs), and an equalization structure based on the independent component analysis (ICA). Our work is different in that a small number of well-designed pilot symbols are employed to kill two birds with one stone. On the one hand, using the structure of pilots, a complex multi-dimensional search for multiple CFOs is divided into a series of low-complexity mono-dimensional searches. On the other hand, the cross-correlation between the transmitted and the received pilot symbols is explored to allow elimination of the remaining ambiguity in the ICA equalized signals. Simulation results show that with a low training overhead of 3.1%, the proposed semi-blind system can achieve a bit error rate (BER) performance close to the ideal case with perfect channel state information (CSI) and no CFO at the receiver. Yufei Jiang, Xu Zhu 0001, Eng Gee Lim, Hai Lin 0001, Yi Huang 0001 |
GLOBECOM | 3 |
| 2012 | Pricing Bermudan Interest Rate Swaptions via Parallel Simulation under the Extended Multi-factor LIBOR Market Model
Ka Lok Man, Eng Gee Lim |
NPC | 3 |