Yong Yue 0001

dblp:172/3742 · DBLP profile ↗
← Back
29ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0001-7695-4538ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 6 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021Human-computer interaction and ubiquitous computing · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 A Mixture-of-Expert Model for Cross-Subject Motor Imagery Decoding
abstract
Motor Imagery Brain-Computer Interface (MIBCI) is one of the most widely used BCI paradigms. However, due to large inter-individual variability in EEG signals, EEGbased pattern recognition models face significant challenges in cross-subject generalization. In this study, we posit that although cross-subject MI-BCI is fundamentally a cross-domain task, the population is not homogeneous; instead, latent subpopulations exist in which subjects share more consistent classification boundaries. To exploit this structure, we propose a Mixture-of-Experts (MoE) framework that automatically partitions training trials into latent groups and trains expert classifiers specialized for each group, while simultaneously learning a gating network that assigns test trials to the most suitable expert(s). This design enables the system to adapt to subpopulation structure, mitigate negative transfer from dissimilar subjects, and better model inter-subject heterogeneity. Evaluations on two public MI-EEG datasets (EEGMMIDB and OpenBMI) using k-fold cross-subject protocols demonstrate that our MoE approach significantly improves classification accuracy compared to most baseline models.
Jingzhou Xu, Haoyu Wu 0001, Yong Yue 0001, Jun Qi 0001
BIBM3
2025 MiTPose: Multi-Granularity Guided Vision Transformer for Human Pose Estimation
abstract
Two-dimensional human pose estimation (HPE) has been extensively applied across various domains, including behavioral analysis, identity verification, and automated industrial manufacturing. Compared to convolutional neural networks (CNNs), the Vision Transformer (ViT) has demonstrated impressive results in human pose estimation. However, two main challenges arise: (1) the complexity of image size and parameters grows quadratically, making traditional ViT unsuitable for deployment on edge devices, and (2) the attention mechanism in transformers lacks the ability to capture local fine-grained perception. To address these issues, we propose a novel method Multi-Granularity guided Vision Transformer for human pose estimation (MiTPOSE), which integrates both CNN and transformer for feature encoding. Specifically, we introduce an improved SCConv encoder with Global Response Normalization, which consists of a spatial reconstruction unit and channel reconstruction unit to reduce redundant computations and enhance representative feature learning. Furthermore, we incorporate a novel Multi-Granularity block to address the shortcomings of traditional self-attention mechanisms in capturing local fine-grained details, and a lightweight decoder for keypoint detection. Comprehensive evaluations and tests on the COCO benchmark datasets demonstrate that MiTPose achieves competitive performance in pose estimation compared to state-of-the-art methods.
Qizhong Gao, Yize Liu, Zhuozhi Li, Yuhao Jin, Yong Yue 0001
INDIN7
2025 USVTrack: USV-Based 4D Radar-Camera Tracking Dataset for Autonomous Driving in Inland Waterways
Shanliang Yao, Runwei Guan, Yi Ni, Yong Yue 0001, Ryan Wen Liu
IROS5
2025 Exploring Radar Data Representations in Autonomous Driving: A Comprehensive Review
abstract
With the rapid advancements of sensor technology and deep learning, autonomous driving systems are providing safe and efficient access to intelligent vehicles as well as intelligent transportation. Among these equipped sensors, the radar sensor plays a crucial role in providing robust perception information in diverse environmental conditions. This review focuses on exploring different radar data representations utilized in autonomous driving systems. Firstly, we introduce the capabilities and limitations of the radar sensor by examining the working principles of radar perception and signal processing of radar measurements. Then, we delve into the generation process of five radar representations, including the ADC signal, radar tensor, point cloud, grid map, and micro-Doppler signature. For each radar representation, we examine the related datasets, methods, advantages and limitations. Furthermore, we discuss the challenges faced in these data representations and propose potential research directions. Above all, this comprehensive review offers an in-depth insight into how these representations enhance autonomous system capabilities, providing guidance for radar perception researchers. To facilitate retrieval and comparison of different data representations, datasets and methods, we provide an interactive website at https://radar-camera-fusion.github.io/radar.
Shanliang Yao, Runwei Guan, Zitian Peng, Chenhang Xu, Yilu Shi, Weiping Ding 0001, Eng Gee Lim, Yong Yue 0001, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue
IEEE Trans. Intell. Transp. Syst.8
2024 ASY-VRNet: Waterway Panoptic Driving Perception Model based on Asymmetric Fair Fusion of Vision and 4D mmWave Radar
abstract
Panoptic Driving Perception (PDP) is critical for the autonomous navigation of Unmanned Surface Vehicles (USVs). A PDP model typically integrates multiple tasks, necessitating the simultaneous and robust execution of various perception tasks to facilitate downstream path planning. The fusion of visual and radar sensors is currently acknowledged as a robust and cost-effective approach. However, most existing research has primarily focused on fusing visual and radar features dedicated to object detection or utilizing a shared feature space for multiple tasks, neglecting the individual representation differences between various tasks. To address this gap, we propose a pair of Asymmetric Fair Fusion (AFF) modules with favorable explainability designed to efficiently interact with independent features from both visual and radar modalities, tailored to the specific requirements of object detection and semantic segmentation tasks. The AFF modules treat image and radar maps as irregular point sets and transform these features into a crossed-shared feature space for multitasking, ensuring equitable treatment of vision and radar point cloud features. Leveraging AFF modules, we propose a novel and efficient PDP model, ASY-VRNet, which processes image and radar features based on irregular super-pixel point sets. Additionally, we propose an effective multi-task learning method specifically designed for PDP models. Compared to other lightweight models, ASY-VRNet achieves state-of-the-art performance in object detection, semantic segmentation, and drivable-area segmentation on the WaterScenes benchmark. Our project is publicly available at https://github.com/GuanRunwei/ASY-VRNet.
Runwei Guan, Shanliang Yao, Ka Lok Man, Yong Yue 0001, Jeremy S. Smith, Eng Gee Lim, Yutao Yue
IROS5
2024 Smart Building Management System based on Digital Twin: A Case Study on Real-Time Environmental Monitoring and Thermal Comfort Prediction
abstract
As urban landscapes increasingly transform into smart cities, there is an increasing emphasis on smart buildings for sustainable and efficient management of resources. This evolution necessitates innovative approaches to efficiently manage energy resources, enhance occupant comfort, and ensure environmental sustainability. In response to these challenges, this study introduces a practical and novel smart building management system, which integrates Digital Twin (DT) technology with machine learning, primarily aimed at enhancing thermal comfort in buildings. The system, underpinned by a sensor network-based DT and real-time data visualization using the Unreal Engine, employs a Deep Neural Network (DNN) to predict thermal comfort. The empirical validation, conducted in a controlled laboratory setting, involves a comparative analysis of the DNN’s performance against traditional models and user experience evaluation through the USE Questionnaire. Results demonstrate the DNN’s superior predictive accuracy and high user satisfaction levels in usability and effectiveness. This research highlights the significant role of DT and machine learning in revolutionizing smart building operations, setting a foundation for future advancements in creating sustainable, efficient, and occupant-friendly smart cities.
Qizhong Gao, Yijie Chu, Zitian Peng, Yuhao Jin, Shuchen Ji, Songming Ping, Yong Yue 0001
ISPA10
2024 ST-GCN: A Spatiotemporal Graph Convolution Neural Network for EEG Motor Imagery Signal Decoding
abstract
Motor imagery (MI) is a mental process extensively used in the experimental paradigm for brain-computer interfaces (BCIs) across various basic science and clinical research studies. Despite its widespread use, accurately decoding intentions from MI poses significant challenges due to the complex nature of brain patterns and the limited sample sizes typically available for machine learning. This paper introduces a Spatiotemporal Graph Neural Network (ST-GCN) designed for MI classification. First, the spatial-temporal convolution layer is used to extract features from raw EEG data, where mixed depthwise convolution extracts temporal features, followed by spatial filtering convolution that decomposes the EEG signal. A graph convolution module employing the max relative aggregator is then utilized to explore the relationships between the spatially decomposed EEG components. In the final step, under the combined supervision of cross-entropy and our proposed channel selection loss, the ST-GCN achieves feature extraction that enhances interclass dispersion and intraclass compactness. We compare ST-GCN with several benchmark EEG decoding methods on two MI datasets: the BCI Competition III Dataset IVa and the BCI Competition IV Dataset 1. ST-GCN outperforms the deep learning benchmark methods by achieving an accuracy of 78.11% and 71.94%, respectively, in 10-fold cross-validation.
Jingzhou Xu, Jun Qi 0001, Junqing Zhang, Yong Yue 0001
ISPA4
2024 Toward Multi-Agent Coordination in IoT via Prompt Pool-based Continual Reinforcement Learning
abstract
The Internet of Things (IoT) represents a complex, dynamic environment where edge devices continuously optimize their policies to address a continual stream of tasks. Previous studies have typically relied on a rehearsal buffer containing data from past tasks or a known task identity to mitigate catastrophic forgetting. Our research, Prompt Pool-based Continual Reinforcement Learning (PPCRL), aims to create a more efficient memory system by expanding a single prompt into a prompt pool, allowing agents to automatically select a set of relevant prompts without needing task identity knowledge. Similar to prompt-based learning techniques, our approach utilizes a small trainable prompt pool to guide pre-trained models through sequential task learning systematically. This allows us to optimize prompts for guiding model predictions and effectively manage both shared and task-specific knowledge while maintaining model generalization. We conducted experiments on two multi-agent benchmarks where traditional methods suffer from significant performance degradation. In contrast, PPCRL demonstrates the capability to outperform baselines and exhibits high generalization ability.
Chenhang Xu, Jia Wang 0009, Yong Yue 0001, Jun Qi 0001, Jieming Ma
ISPA4
2024 CCTSDB dataset enhancement based on a cross-augmentation method for image datasets
abstract
In the digital era, the rapid advancement of artificial intelligence has put a spotlight on target detection, especially in traffic settings. This area of study is pivotal for crucial projects like autonomous vehicles, road monitoring, and traffic sign recognition. However, existing Chinese traffic datasets lack comprehensive benchmarks for traffic signs and signals, and foreign datasets do not match Chinese traffic conditions. Manually annotating a large-scale dataset tailored for Chinese traffic conditions presents a significant challenge. This study addresses this gap by proposing a cross-augmentation method for image datasets. We utilized YOLOX for target detection and trained models on the BDD100K dataset, achieving an impressive mAP of 60.25%, surpassing most algorithms. Leveraging transfer learning, we enhanced the CCTSDB dataset, creating the ACCTSDB dataset, which includes annotations for common traffic objects and Chinese traffic signs. Using YOLOX, we trained a traffic detector tailored for Chinese traffic scenarios, achieving an mAP of 75.79%. To further validate our approach, we conducted experiments on the TT100K dataset and successfully introduced the ATT100K dataset. Our methodology is poised to alleviate the limitations of manually annotating image datasets. The proposed ACCTSDB dataset and ATT100K dataset are expected to compensate for the lack of large-scale, multi-class traffic datasets in China.
Xinrui Lin, Wei Wang 0368, Yong Yue 0001
Intell. Data Anal.4
2024 Tangible and Mid-Air Interactions in Hand-Held Augmented Reality for Upper Limb Rehabilitation: An Evaluation of User Experience and Motor Performance
abstract
Hand-held augmented reality (AR) offers accessible, interactive rehabilitation options for patients with upper limb motor deficits. Incorporating hand-involved interactions (e.g., tangible and mid-air interactions) into hand-held AR provides patients with intuitive manners to perform rehabilitation exercises mimicking real-world activities. Previous work has shown the importance of user experience and motor performance in rehabilitation systems, but little was known in the literature regarding the impact of hand-involved interactions in hand-held AR on user experience and motor performance in rehabilitation exercises. Hence, this study aims to evaluate user experience and motor performance when using three types of hand-involved interactions in hand-held AR rehabilitation: (1) tangible cube (i.e., a space-multiplexed tangible interaction with a physical cube acting as a real proxy to manipulate a virtual object in the same form); (2) tangible controller (i.e., a time-multiplexed tangible interaction with a physical controller applied to manipulate a virtual object); and (3) hand motion (i.e., a form of mid-air interaction to move a virtual object with hands). Based on the findings from self-report, electroencephalography (EEG), and performance measures, this study reveals the advantages of the tangible cube over the tangible controller, both superior to the hand motion in hand-held AR rehabilitation regarding user experience and motor performance. This study offers new understanding of the advantages and disadvantages of various interaction techniques in hand-held AR rehabilitation, emphasizing crucial design considerations for these systems, with a focus on user experience and motor performance in upper limb rehabilitation.
Wenxin Sun, Mengjie Huang, Chenxin Wu, Rui Yang 0007, Yong Yue 0001, Miaomiao Jiang
Int. J. Hum. Comput. Interact.5
2024 Chain-of-thought prompting empowered generative user modeling for personalized recommendation
Fan Yang 0051, Yong Yue 0001, Gangmin Li, Terry R. Payne, Ka Lok Man
Neural Comput. Appl.2
2024 Experimental Analysis of Freehand Multi-object Selection Techniques in Virtual Reality Head-Mounted Displays
abstract
Object selection is essential in virtual reality (VR) head-mounted displays (HMDs). Prior work mainly focuses on enhancing and evaluating techniques for selecting a single object in VR, leaving a gap in the techniques for multi-object selection, a more complex but common selection scenario. To enable multi-object selection, the interaction technique should support group selection in addition to the default pointing selection mode for acquiring a single target. This composite interaction could be particularly challenging when using freehand gestural input. In this work, we present an empirical comparison of six freehand techniques, which are comprised of three mode-switching gestures (Finger Segment, Multi-Finger, and Wrist Orientation) and two group selection techniques (Cone-casting Selection and Crossing Selection) derived from prior work. Our results demonstrate the performance, user experience, and preference of each technique. The findings derive three design implications that can guide the design of freehand techniques for multi-object selection in VR HMDs.
Rongkai Shi, Yushi Wei, Xuning Hu, Yu Liu 0077, Yong Yue 0001, Lingyun Yu 0001, Hai-Ning Liang
Proc. ACM Hum. Comput. Interact.5
2024 WaterScenes: A Multi-Task 4D Radar-Camera Fusion Dataset and Benchmarks for Autonomous Driving on Water Surfaces
abstract
Autonomous driving on water surfaces plays an essential role in executing hazardous and time-consuming missions, such as maritime surveillance, survivor rescue, environmental monitoring, hydrography mapping and waste cleaning. This work presents WaterScenes, the first multi-task 4D radar-camera fusion dataset for autonomous driving on water surfaces. Equipped with a 4D radar and a monocular camera, our Unmanned Surface Vehicle (USV) proffers all-weather solutions for discerning object-related information, including color, shape, texture, range, velocity, azimuth, and elevation. Focusing on typical static and dynamic objects on water surfaces, we label the camera images and radar point clouds at pixel-level and point-level, respectively. In addition to basic perception tasks, such as object detection, instance segmentation and semantic segmentation, we also provide annotations for free-space segmentation and waterline segmentation. Leveraging the multi-task and multi-modal data, we conduct benchmark experiments on the uni-modality of radar and camera, as well as the fused modalities. Experimental results demonstrate that 4D radar-camera fusion can considerably improve the accuracy and robustness of perception on water surfaces, especially in adverse lighting and weather conditions. WaterScenes dataset is public onhttps://waterscenes.github.io.
Shanliang Yao, Runwei Guan, Zhaodong Wu, Yi Ni, Zile Huang, Ryan Wen Liu, Yong Yue 0001, Weiping Ding 0001, Eng Gee Lim, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Yutao Yue
IEEE Trans. Intell. Transp. Syst.7
2023 Tasks of a Different Color: How Crowdsourcing Practices Differ per Complex Task Type and Why This Matters
abstract
Crowdsourcing in China is a thriving industry. Among its most interesting structures, we find crowdfarms, in which crowdworkers self-organize as small organizations to tackle macrotasks. Little, however, is known as to which practices these crowdfarms use to tackle the macrotasks, and this goes hand in hand with the current practice of the HCI research community to treat all forms of complex crowdsourcing work as practically the same. However, macrotasks differ substantially regarding structure and decomposability. Treating them under one umbrella term - macrotasking - can lead to an imprecise understanding of the workforce involved. We address this gap by examining the work practices of 31 Chinese crowdfarms on the four main macrotask types, namely: modular, interlaced, wicked, and container macrotasks. Our results confirm essential differences in how these nascent crowd organizations address different macrotasks and shed light on what platforms can do to improve the uptake of such work.
Konstantinos Papangelis, Ioanna Lykourentzou, Michael Saker, Alan Chamberlain, Vassilis-Javed Khan, Hai-Ning Liang, Yong Yue 0001
CHI8
2021 Mental Workload Evaluation of Virtual Object Manipulation on WebVR: An EEG Study
abstract
Virtual object manipulation as a key feature has been studied in virtual reality (VR) environments. Previous studies highlighted user experience on three basic types of virtual object manipulation, translation, rotation and scaling. However, prior literature mainly studied task performance in manipulation modes with different degrees of freedom (DoF), and few studies assessed user experience by evaluating the psychological response, such as mental workload on these three basic manipulation types in virtual environments. This paper compared manipulation modes with 1DoF and 3DoF to assess users’ mental workload as a critical indicator of user experience by electroencephalogram (EEG) measurement and questionnaires in manipulation tasks on the webpage with VR effects (also known as WebVR). By applying signal processing and statistical methods to analyze EEG data from ten subjects, the results demonstrated that the participants generally perceive less mental workload by 1DoF manipulation modes than 3DoF on WebVR. Besides, this study also found some different results between objective and subjective data.
Wenxin Sun, Mengjie Huang, Rui Yang 0007, Yong Yue 0001
HSI5
2019 An Improved APFM for Autonomous Navigation and Obstacle Avoidance of USVs
Yong Yue 0001, Shunda Wu, MingSheng Li, Yawei Hu
ICINCO (2)2
2019 Unbalancing Pairing-Free Identity-Based Authenticated Key Exchange Protocols for Disaster Scenarios
abstract
In disaster scenarios, such as an area after a terrorist attack, security is a significant problem since communications involve information for the rescue officers, such as polices, militaries, emergency medical technicians, and the survivors. Such information is critically important for the rescue organizations; and protecting the privacy of the survivors is required. Normally, authenticated key exchange (AKE) is an underlying approach for security. However, available AKE protocols are either inconvenient or infeasible in disaster areas due to the very nature of disasters. To address the security problem in disaster scenarios, we propose two pairing-free identity-based AKE (ID-AKE) protocols that have unbalanced computational requirements on the two parties. Compared with existing AKE protocols, the proposed protocols have a number of advantages in disaster scenarios: 1) they are more convenient than symmetric cryptography-based AKE protocols since they do not require any preshared secret between the parties; 2) they are more feasible than asymmetric cryptography-based AKE protocols since they do not require any online server; and 3) they are more friendly to battery-powered and computationally limited devices than pairing-based and pairing-free ID-AKE protocols since they do not involve any bilinear pairing (a time-consuming operation), and have lower computational requirement on the limited party. Security of the proposed protocols are analyzed in detail; and prototypes of them are implemented to evaluate the performance. We also illustrate the application of the protocols through a vivid use case in a terrorist attack scenario.
Jie Zhang 0030, Xin Huang 0005, Wei Wang 0042, Yong Yue 0001
IEEE Internet Things J.4
2018 Contexts-States-Aware Access Control for Internet of Things
abstract
The more and more connected devices and rapidly developing Internet of Things (IoT) applications are the foundations of the future smart cities which provide ubiquitous services. The extension and proliferation of the technology brings huge security challenges, especially for the infrastructural IoT applications in the open environments. The traditional access model, such as Role-Based Access Control (RBAC) cannot provide the flexible fine-grained access control which is required due to dynamic changing users and environments. On the other hand, some other features of the IoT applications like constrained-resources devices and large-scale deployments make it very difficult to apply Attribute-Based Access Control (ABAC). Furthermore, the ABAC mechanism cannot control the way that the requester uses the services once the requester obtains the access permission. To address these issues, in this paper, we propose an access control model based on ABAC with Contexts-States-Awareness. The proposed model is implemented by using Semantic Web technologies with a sample ontology for the model and some access control policies in SWRL (Semantic Web Rule Language). We also give a logical architecture which is the extension from the reference architecture of XACML eXtensible Access Control Markup Language specification.
Yuji Dong, Kaiyu Wan, Xin Huang 0005, Yong Yue 0001
CSCWD4
2018 Infrared motion detection and electromyographic gesture recognition for navigating 3D environments
abstract
Abstract This research explores the suitability and effectiveness of two relatively new types of input device for navigating 3D virtual environments. These are infrared motion detection, like the Leap Motion tracker, and electromyographic gesture recognition, like the Myo Armband. Despite the introduction of a variety of new input devices intended to provide a more natural interaction experience, navigation within 3D virtual environments is still normally done on more traditional control devices such as game controllers or the keyboard–mouse combination. This study investigates the potential of new devices to support navigation in 3D environments through an experiment conducted with 27 participants using three different types of input devices to play a ball‐balancing maze‐like game. The input devices tested are a standard game controller, a Leap Motion tracker for infrared motion detection, and the Myo Armband for electromyographic gesture recognition. Results demonstrated the real potential of both types of device to support navigation interaction within 3D environments.
Hai-Ning Liang, Yong Yue 0001, Paul Craig
Comput. Animat. Virtual Worlds3
2018 Evaluating enjoyment, presence, and emulator sickness in VR games based on first- and third- person viewing perspectives
abstract
Abstract Many virtual reality (VR) games are based on a first‐person perspective (1PP). There are, however, advantages in using another perspective, such as the third‐person perspective (3PP). Although there has been some research evaluating the effect of 1PP and 3PP in gameplay experiences, it is largely unexplored for VR games played via the new generation of commercial head‐mounted display systems, such as the Oculus Rift. In this research we want to shed some light on the relationship between the different perspectives, when games are played using head‐mounted display VR, and simulator sickness, enjoyment, and presence. To do so, we perform an experiment using two different perspectives (1PP and 3PP) and displays (VR and a conventional display) with a popular game. Our findings indicate that 3PP‐VR is less likely to make people have simulator sickness when compared with 1PP‐VR. However, the former is not perceived as immersive, but this might not be a problem because our data also show that presence is not mandatory for enjoyment. Also, the data suggest that there is no clear preference between 1PP‐VR and 3PP‐VR for gameplay.
Diego Monteiro 0001, Hai-Ning Liang, Wenge Xu, Marvin Brucker, Vijayakumar Nanjappan, Yong Yue 0001
Comput. Animat. Virtual Worlds6
2016 Artificial intelligence techniques in product engineering
Kit Yan Chan, Kevin Kam Fung Yuen, Vasile Palade, Yong Yue 0001
Eng. Appl. Artif. Intell.4
2016 Video-Based Classification of Driving Behavior Using a Hierarchical Classification System with Multiple Features
abstract
Driver fatigue and inattention have long been recognized as one of the main contributing factors in traffic accidents. Therefore, the development of intelligent driver assistance systems, which provides automatic monitoring of driver's vigilance, is an urgent and challenging task. This paper presents a novel system for video-based driving behavior recognition. The fundamental idea is to monitor driver's hand movements and to use these as predictors for safe/unsafe driving behavior. In comparison to previous work, the proposed method utilizes hierarchical classification and treats driving behavior in terms of a spatio-temporal reference framework as opposed to a static image. The approach was verified using the Southeast University Driving-Posture Dataset, a dataset comprised of video clips covering aspects of driving such as: normal driving, responding to a cell phone call, eating and smoking. After pre-processing for illumination variations and motion sequence segmentation, eight classes of behavior were identified. The overall prediction accuracy obtained using the proposed approach was [Formula: see text] when using a hierarchical classification approach. The proposed approach was able to clearly identify two dangerous driving behaviors, Responding to a cellphone call and Eating, with recognition rates of 92.39% and 92.29% respectively.
Chao Yan 0003, Frans Coenen, Yong Yue 0001, Xiaosong Yang
Int. J. Pattern Recognit. Artif. Intell.3
2013 Enhancing Bayesian Estimators for Removing Camera Shake
abstract
Abstract The aim of removing camera shake is to estimate a sharp version x from a shaken image y when the blur kernel k is unknown. Recent research on this topic evolved through two paradigms called and . only solves for k by marginalizing the image prior, while recovers both x and k by selecting the mode of the posterior distribution. This paper first systematically analyses the latent limitations of these two estimators through Bayesian analysis. We explain the reason why it is so difficult for image statistics to solve the previously reported failure. Then we show that the leading methods, which depend on efficient prediction of large step edges, are not robust to natural images due to the diversity of edges. , although much more robust to diverse edges, is constrained by two factors: the prior variation over different images, and the ratio between image size and kernel size. To overcome these limitations, we introduce an inter‐scale prior prediction scheme and a principled mechanism for integrating the sharpening filter into . Both qualitative results and extensive quantitative comparisons demonstrate that our algorithm outperforms state‐of‐the‐art methods.
Chao Wang 0063, Yong Yue 0001, Feng Dong 0005, Yubo Tao, Gordon Clapworthy, Xujiong Ye
Comput. Graph. Forum2
2013 Nonedge-Specific Adaptive Scheme for Highly Robust Blind Motion Deblurring of Natural Imagess
abstract
Blind motion deblurring estimates a sharp image from a motion blurred image without the knowledge of the blur kernel. Although significant progress has been made on tackling this problem, existing methods, when applied to highly diverse natural images, are still far from stable. This paper focuses on the robustness of blind motion deblurring methods toward image diversity-a critical problem that has been previously neglected for years. We classify the existing methods into two schemes and analyze their robustness using an image set consisting of 1.2 million natural images. The first scheme is edge-specific, as it relies on the detection and prediction of large-scale step edges. This scheme is sensitive to the diversity of the image edges in natural images. The second scheme is nonedge-specific and explores various image statistics, such as the prior distributions. This scheme is sensitive to statistical variation over different images. Based on the analysis, we address the robustness by proposing a novel nonedge-specific adaptive scheme (NEAS), which features a new prior that is adaptive to the variety of textures in natural images. By comparing the performance of NEAS against the existing methods on a very large image set, we demonstrate its advance beyond the state-of-the-art.
Chao Wang 0063, Yong Yue 0001, Feng Dong 0005, Yubo Tao, Xiangyin Ma, Gordon Clapworthy, Hai Lin 0003, Xujiong Ye
IEEE Trans. Image Process.2
2012 Classification of multi-channels SEMG signals using wavelet and neural networks on assistive robot
abstract
Recently, the robot technology research is changing from manufacturing industry to non-manufacturing industry, especially the service industry related to the human life. Assistive robot is a kind of novel service robot. It can not only help the elder and disabled people to rehabilitate their impaired musculoskeletal functions, but also help healthy people to perform tasks requiring large forces. This kind of robot has a broad application prospect in many areas, such as medical rehabilitation, special military operations, special/high intensity physical labour, space, sports, and entertainment. SEMG (Surface Electromyography) of Palmaris longus, brachioradialis, flexor carpiulnaris and biceps brachii are analysed with a wavelet transform method. The absolute variance of 3-layer wavelet coefficients is distilled and regarded as signal characteristics to compose eigenvectors. The eigenvectors are input data of a neural network classifier used to identify 5 different kinds of movement patterns including wrist flexor, wrist extensor, elbow flexion, forearm pronation and forearm rotation. Experiments verify the effectiveness of the proposed method.
Yong Yue 0001, Carsten Maple, Beisheng Liu, Chengdong Wu 0001
INDIN2
2012 Fuzzy logic based symbolic grounding for best grasp pose for homecare robotics
abstract
Symbolic grounding in unstructured environments remains an important challenge in robotics [7]. Homecare robots are often required to be instructed by their human users intuitively, which means the robots are expected to take highlevel commands and execute corresponding tasks in a domestic environment. High-level commands are represented with symbolic terms such as “near” and “close” and, on the other hand, robots are controlled based on trajectories. The robots need to translate the symbolic terms to trajectories. In addition, domestic environment is unstructured where the same objects can be placed in different places over the time. This increases the difficulties in symbolic grounding. This paper presents a fuzzy logic based approach to symbolic grounding. In this approach, grounded concepts are modelled as fuzzy sets and the existing knowledge is used to deduce grounded values given real-time sensory inputs. Experiments results show that this approach works well in unstructured environment.
Beisheng Liu, Dayou Li, Yong Yue 0001, Carsten Maple, Renxi Qiu
INDIN3
2012 Fuzzy optimisation based symbolic grounding for service robots
abstract
Symbolic grounding is a bridge between high-level planning and actual robot sensing, and actuation. Uncertainties raised by the unstructured environment make a bottleneck for integrating traditional artificial intelligence with service robotics. This paper presents a fuzzy logic based approach to formalise the grounding problems into a fuzzy optimization problem, which is robust to uncertainties. Novel techniques are applied to establish the objective function, to model fuzzy constraints and to perform fuzzy optimisation. The outcome is tested with a service robot fetch and carry task, where the fuzzy optimisation approach helps the robot to determine the most comfortable position (location and orientation) for grasping objects. Experimental results show that the proposed approach improves the robustness of the task implementation in unstructured environments.
Beisheng Liu, Dayou Li, Renxi Qiu, Yong Yue 0001, Carsten Maple
IROS4
2010 Data Mapping, Matching and Loading Using Grid Services
abstract
With the proliferation of international standards for grid-enabled databases, the need for data loading and mapping is identified. It is required in a large integrated environment encompassing heterogeneous databases. This research study proposes an intermediate staging facility in order to upload and integrate data from various small to large size data repositories. The proposed facility has been employed to temporarily store and process the data before it is validated. This research expands the integration notion of a database management system (DBMS) to include the data loading, data matching and transformation processes in traditional and grid distributed environments. The data mapping is employed in the form of value correspondences by using a staging DBMS. A service-based generic strategy grid staging catalogue service (GSCATS) is introduced. It consolidates distinct catalogue schemas of federated databases to access information seamlessly. It involves data mapping and matching of structure objects of federated databases.
Ejaz Ahmed 0001, Nik Bessis, Yong Yue 0001, Muhammad Sarfraz 0001
AINA3
2010 Challenges and Perspectives of Procedural Modelling and Effects
abstract
The use of procedural modelling has risen dramatically over the last decade. This has been partly due to the increase in computing power, affording developers the opportunity to use methods that were previously infeasible. A further reason for this growth is due to the improvement of the representation of human knowledge. The development of procedural modeling has, thus far, been somewhat disjointed and ad hoc as different areas of graphics, gaming and modelling have been utilizing the technique with little reference outside of their own specialisation. This paper provides an overview of procedural modelling covering key techniques and applications and then suggests a framework for the development of a procedural modelling system that can form areas of land of both populated areas and bodies of water.
David Fletcher, Yong Yue 0001, Majid Al Kader
IV2