VLDB 2026 Research / reviewers in the wild / expert
Yunbo Rao
dblp:43/9830
· DBLP profile ↗
45ranked-venue papers
15as first author
28since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 30 · 11 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | EA-NACFE: Edge-aware nonlinear low-light image enhancement via closed-form diffusion coupling
Annicet Razafindratovolahy, Yunbo Rao, Jean Clément Tovolahy, Molla Woretaw Teshome, Subhan Uddin |
Signal Process. | 2 |
| 2026 | Smile: enhancing low-light images in lightweight networks via exposure-aware non-reference losses
Annicet Razafindratovolahy, Yunbo Rao, Shaoning Zeng, Linda Delali Fiasam, Junmin Xue, Collins Sey |
Multim. Syst. | 2 |
| 2026 | TOVO: Tone-oriented vision optimization for efficient low-light image enhancement
Annicet Razafindratovolahy, Yunbo Rao, Jean Clément Tovolahy, Mona Afanga, Albert Mutale |
Signal Process. | 2 |
| 2026 | AMS: Attention Map Seeds for enhancing interactive segmentation
Qingsong Lv, Jialong Zhu, Yunbo Rao, Zhanglin Cheng |
Signal Process. Image Commun. | 3 |
| 2025 | MID-LLM: Enhancing Medical Image Diagnostics With LLMs in a Blockchain AI FrameworkabstractThe rapid growth of medical imaging data presents significant challenges in diagnostic accuracy, data privacy, and computational efficiency. Traditional centralized AI models struggle with scalability and pose risks to patient confidentiality due to data aggregation. Moreover, heterogeneous medical data across institutions complicates the development of robust diagnostic tools. To address these issues, we propose MID-LLM, a novel framework that integrates Large Language Models (LLMs) with a blockchain-based federated learning system for medical image analysis. It also ensures the security and privacy of sensitive medical data across decentralized networks. MID-LLM uses verification mechanisms to ensure the global model’s integrity. It also employs aggregation techniques to reduce bias and improve training efficiency. Experiments on the BraTS 2020 dataset show that MID-LLM outperforms traditional federated learning, achieving higher Dice scores with improved computational efficiency. These results highlight MID-LLM’s potential to enhance diagnostic accuracy while offering a scalable, secure solution for AI in healthcare. Rajesh Kumar 0014, Yunbo Rao, Jay Kumar, Cobbinah Bernard Mawuli, Waqar Ali 0001, Shaoning Zeng |
IEEE Internet Things J. | 2 |
| 2025 | Big-LITTLE-Net: a dual-branch network for small UAV detection
Yinjie Chen, Wenyi Tang, Yunbo Rao, Shuzhen Zhu |
Multim. Syst. | 3 |
| 2025 | FSformer: fusing frequency and spatial domain transformer network for underwater image enhancement
Dalang Liu, Yunbo Rao, Jialong Zhu, Yanjin Ma |
Multim. Syst. | 2 |
| 2025 | U-net of joint spatial domains with multi-scale atrous convolution for rectal image segmentation
Yunbo Rao, Shaoning Zeng, Tingting Shao, Jihong Sun |
Multim. Tools Appl. | 1 |
| 2025 | A Simple Yet Robust Nonlinear Function for Low-Light Image Enhancement TaskabstractWe present a novel, parameter-free nonlinear transformation for low-light image enhancement that operates directly on individual pixel values. This simple yet powerful function requires no prior knowledge or external tuning, and enhances image brightness and contrast by leveraging only the input image itself. When applied iteratively, the method achieves optimal results after just three applications. Despite its minimalism, our approach outperforms recent state-of-the-art methods on benchmarks. This highlights the potential of simple signal processing operations for emergent enhancement, and suggests directions for theoretical analysis, integration with deep learning, and deployment in real-world vision systems. Annicet Razafindratovolahy, Yunbo Rao |
IEEE Signal Process. Lett. | 2 |
| 2025 | Personalized News Recommendation Towards the Era of LLMs: Review and ProspectabstractWith the prevalence of online news services, personalized news recommendation (PNR) has played an indispensable role in meeting users' needs and mitigating information overload, with the aim of providing news articles that cater to user preferences. Despite significant progress made in the field of PNR over the past few decades, their performances are still hindered by some limitations, such as insufficient news modeling, difficulties in effectively modeling diverse user interests, and ignorance of fine-grained matching signals. It is fortunate that the emergence of large language models (LLMs) provides a promising insight into empowering the capabilities of news recommendation. Known for their impressive capabilities of natural language understanding and generation, LLMs have achieved disruptive achievements in various natural language processing (NLP) tasks, which motivates us to integrate LLMs into news recommendation and benefits from them to make up existing deficiencies. In this paper, we conduct a comprehensive review of current efforts made towards utilizing LLMs for PNR, with a focus on three core modules involved in the news recommendation process, i.e., news modeling, user modeling, and accurate matching. We systematically discuss and analyze relevant works under each focus. In addition, we point out several potential research directions to provide more inspiration for future investigation in this thriving field. Linmei Hu, Yunbo Rao, Bo Fang 0007, Liqiang Nie |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2025 | LRDNet: Lightweight LiDAR Aided Cascaded Feature Pools for Free Road Space DetectionabstractHumans have long fantasized about self-driving vehicles for the sake of luxury, style, safety, and ease. Free road space detection for collision avoidance and path planning is a vital part of autonomous driving vehicles. Despite many researchers focusing on free road space detection, it remains an open and challenging problem for real-world applications. Many studies have attempted to fuse depth and LiDAR features with visual features to improve the overall performance of free road space detection. However, there is no guideline on how such features should be fused to complement the visual features. Additionally, most of the previously proposed methods are computationally expensive and not suitable for real-life applications. The main motivation of this study is to realize a lightweight model that addresses these problems without compromising performance. As the LiDAR and visual features exist in different spaces, the proposed method attempts to learn various transformation and fusion operations from LiDAR features to complement the visual features. To validate the performance of the proposed method, we conduct comprehensive experiments on prominent benchmark datasets. The results of the experiments reveal the superior performance of the proposed model while being lightweight. LRDNet ranks third overall (with a minor difference) and second among LiDAR-based methods on the KITTI road benchmark dataset. Furthermore, the proposed model is the least computationally expensive among state-of-the-art methods and can be considered an optimal trade-off between speed and accuracy. Abdullah Aman Khan, Jie Shao 0001, Yunbo Rao, Lei She, Heng Tao Shen |
IEEE Trans. Multim. | 3 |
| 2024 | Enhancing Model Interpretability Through Interactive Visual Analysis and Counterfactual Explanation Methods
Hanlin Lan, Jiansu Pu, Yulu Xia, Yilei He, Jinyue Huang, Yunbo Rao |
CDVE | 7 |
| 2024 | Robust meter reading detection via differentiable binarization
Yunbo Rao, Hangrui Guo, Dalang Liu, Shaoning Zeng |
Appl. Intell. | 1 |
| 2024 | Multi-session aware hypergraph neural network for session-based recommendation
Yunbo Rao, Tongze Mu, Shaoning Zeng, Junming Xue |
Multim. Tools Appl. | 1 |
| 2024 | RWS: Refined Weak Slice for Semantic Segmentation EnhancementabstractInterpretation of predictions made by Convolutional Neural Networks (CNNs) is a rapidly growing field of research. A common approach involves enhancing semantic segmentation predictions through the generation of heatmaps that illustrate the significance of individual pixels in the segmentation. Nevertheless, the selection of beneficial features from these heatmaps remains a challenge. This is because the introduced information often contains interfering factors such as mutual features between different objects, background, and insufficient heat map resolution which often diminish its effectiveness. To overcome these limitations, we introduce Refined Weak Slices (RWS). Our main idea is to identify low attention regions in heat maps i.e.weak slices, in conjunction with segmentation accuracy, and utilize them to select effective features across different DNN layers, to enhance segmentation. We then seamlessly integrate these features back into the CNN, thusrefiningand enhancing the semantic segmentation result with selected features. Through extensive experiments, we demonstrate that incorporating the RWS module into state-of-the-art methods yields a notable improvement in the average mIoU by 2.84% on benchmark datasets (VOC 2012, COCOStuff, ADE20K, Cityscapes) for both ResNet-101 and ResNet-50 architectures. Furthermore, we achieve a maximum improvement of 5.8% with a single CNN. Overall, the combination of RWS and CNNs exhibits excellent performance in image segmentation tasks. Yunbo Rao, Qingsong Lv, Andrei Sharf, Zhanglin Cheng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Relevance gradient descent for parameter optimization of image enhancement
Yunbo Rao, Yuling Yi, Obed Tettey Nartey, Saeed Ullah Jan 0001 |
Comput. Graph. | 1 |
| 2023 | Point completion by a Stack-Style Folding Network with multi-scaled graphical featuresabstractAbstract Point cloud completion is prevalent due to the insufficient results from current point cloud acquisition equipments, where a large number of point data failed to represent a relatively complete shape. Existing point cloud completion algorithms, mostly encoder‐decoder structures with grids transform (also presented as folding operation), can hardly obtain a persuasive representation of input clouds due to the issue that their bottleneck‐shape result cannot tell a precise relationship between the global and local structures. For this reason, this article proposes a novel point cloud completion model based on a Stack‐Style Folding Network (SSFN). Firstly, to enhance the deep latent feature extraction, SSFN enhances the exploitation of shape feature extractor by integrating both low‐level point feature and high‐level graphical feature. Next, a precise presentation is obtained from a high dimensional semantic space to improve the reconstruction ability. Finally, a refining module is designed to make a more evenly distributed result. Experimental results shows that our SSFN produces the most promising results of multiple representative metrics with a smaller scale parameters than current models. Yunbo Rao, Shaoning Zeng, Jianping Gou |
IET Comput. Vis. | 1 |
| 2023 | GAB-Net: A Robust Detector for Remote Sensing Object Detection Under Dramatic Scale Variation and Complex BackgroundsabstractDetecting objects in remote sensing images (RSIs), characterized by dramatic scale variation and complex backgrounds, has always been a challenging problem. These challenges can be further summarized into three aspects: 1) scale variation among objects; 2) feature fusion misalignment due to the semantic gap between adjacent feature layers and noise from backgrounds; and 3) boundary uncertainty under ambiguous and complex backgrounds. To alleviate these problems, we first utilize a global–local feature enhancement module (GLFEM) to capture local features with multiple receptive fields through cheap pooling operation and obtain global features through nonlocal block, thus alleviating the scale variation issues. Subsequently, attentional feature fusion alignment (AFFA) module is designed to align adjacent feature levels in the feature pyramid from pixel and channel levels. Finally, boundary-uncertainty aware head (BUAH) with distribution focal loss (DFL) is adopted to solve the boundary uncertainty problems. After fusing GLFEM, AFFA, and BUAH modules, we obtain GAB-Net. GAB-Net outperforms state-of-the-art methods on the Dior and NWPU VHR-10 datasets, achieving mAP scores of 73.8% and 89.8%, respectively, without adding high computational costs. The code is available at:https://github.com/Hong-yu-Zhang/GAB-Net. Yunbo Rao, Jie Shao 0001, Fanman Meng, Naveed Ahmad 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2023 | Heterogeneous Knowledge Network for Visual DialogabstractVisual dialog requires an agent to answer successive questions considering an image and dialog history, which is a classic vision-language task. Despite progress, there are still two key challenges: 1) parsing long or complex questions and answers and 2) dealing with the visual scene containing complicated interactions among entities. These challenges bring about the unsatisfactory consequence of current visual dialog methods. In this paper, we propose a novel Heterogeneous Knowledge Network (HKNet), which leverages textual sequence knowledge and graph knowledge to address the above issues. Specifically, the textual sequence knowledge is derived from the sentences that are retrieved from the image captions of the visual dialog dataset. The textual sequence knowledge can supplement essential common sense for parsing long or complex questions and answers. The graph knowledge is constructed via scene graph, which provides complete visual relationships for understanding the complicated interactions. These two kinds of heterogeneous knowledge complement each other and jointly improve the logical reasoning ability of the visual dialog. Extensive experimental results on two benchmark datasets: VisDial v0.9 and v1.0 demonstrate the superiority of the proposed HKNet. Ablation studies and visualization results further verify the effectiveness of our method. Lei Zhao 0017, Lianli Gao, Yunbo Rao, Jingkuan Song, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | PiCovS: Pixel-Level With Covariance Pooling Feature and Superpixel-Level Feature Fusion for Hyperspectral Image ClassificationabstractIn hyperspectral image (HSI) classification, Convolutional Neural Networks (CNNs) have exhibited exceptional performance, owing to their hierarchical nonlinear modeling. However, their fixed square receptive field constrains their ability to effectively handle irregular image regions. Graph Convolution Networks (GCNs) have been introduced to learn irregular regions through correlations between adjacent pixels modeled as superpixel-based nodes, yet they lack pixel-level information. We propose a novel approach "Pixel-level with Covariance Pooling feature and Superpixel-level feature Fusion for Hyperspectral Image Classification" (PiCovS). Our method harnesses complementary spectral-spatial features at both pixel and superpixel levels to capture characteristics of both small-scale regular and large-scale irregular regions. We introduce a hybrid network that integrates and propagates features between image-level pixels and graph-level nodes using a graph encoder-decoder, effectively reconciling the differences between regular CNN and irregular GCN data representations. To enhance superpixel boundary learning, we modify the Manifold Simple Linear Iterative Clustering (M-SLIC) algorithm by incorporating texture feature information, resulting in refined superpixel representations. Additionally, we propose a novel covariance pooling mechanism with an attention mechanism within the CNN branch, enabling the capturing and utilization of holistic HSI information along spectral and spatial dimensions by exploiting second-order statistics throughout the network. Our comprehensive experiments showcase the efficiency and robustness of the proposed framework, achieving an impressive overall accuracy of 99.84%, 99.97%, 99.98%, and 81.96% on the Indian Pines, University of Pavia, Salinas, and the Houston University datasets, respectively. Remarkably, PiCovS excels even with limited training samples, outperforming other state-of-the-art methods in accuracy. Obed Tettey Nartey, Kwabena Sarpong, Daniel Addo, Yunbo Rao, Zhiguang Qin |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Dual Projective Zero-Shot Learning Using Text DescriptionsabstractZero-shot learning (ZSL) aims to recognize image instances of unseen classes solely based on the semantic descriptions of the unseen classes. In this field, Generalized Zero-Shot Learning (GZSL) is a challenging problem in which the images of both seen and unseen classes are mixed in the testing phase of learning. Existing methods formulate GZSL as a semantic-visual correspondence problem and apply generative models such as Generative Adversarial Networks and Variational Autoencoders to solve the problem. However, these methods suffer from the bias problem since the images of unseen classes are often misclassified into seen classes. In this work, a novel model named the Dual Projective model for Zero-Shot Learning (DPZSL) is proposed using text descriptions. In order to alleviate the bias problem, we leverage two autoencoders to project the visual and semantic features into a latent space and evaluate the embeddings by a visual-semantic correspondence loss function. An additional novel classifier is also introduced to ensure the discriminability of the embedded features. Our method focuses on a more challenging inductive ZSL setting in which only the labeled data from seen classes are used in the training phase. The experimental results, obtained from two popular datasets—Caltech-UCSD Birds-200-2011 (CUB) and North America Birds (NAB)—show that the proposed DPZSL model significantly outperforms both the inductive ZSL and GZSL settings. Particularly in the GZSL setting, our model yields an improvement up to 15.2% in comparison with state-of-the-art CANZSL on datasets CUB and NAB with two splittings. Yunbo Rao, Ziqiang Yang, Shaoning Zeng, Jiansu Pu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2023 | Joint Augmented and Compressed Dictionaries for Robust Image ClassificationabstractDictionary-based Classification (DC) has been a promising learning theory in multimedia computing. Previous studies focused on learning a discriminative dictionary as well as the sparsest representation based on the dictionary, to cope with the complex conditions in real-world applications. However, robustness by learning only one single dictionary is far from the optimal level. What is worse, it cannot take advantage of the available techniques proven in modern machine learning, like data augmentation, to mitigate the same problem. In this work, we propose a novel method that utilizes joint Augmented and Compressed Dictionaries for Robust Dictionary-based Classification (ACD-RDC). For optimization under the noise model introduced by real-world conditions, the objective function of ACD-RDC incorporates only two simple, but well-designed constraints, including one enhanced sparsity constraint by the general data augmentation, which requires less case-by-case and sophisticated tuning, and another discriminative constraint solved by a jointly learned dictionary. The optimization of the objective function is then deduced theoretically to an approximate linear problem. The sparsity and discrimination enhanced by data augmentation guarantees the robustness for image classification under various conditions, which constructs the first positive case using data augmentation to obtain robust dictionary-based classification. Numerous experiments have been conducted on popular facial and object image datasets. The results demonstrate that ACD-RDC obtains more promising classification on diversely collected images than the current dictionary-based classification methods. ACD-RDC is also confirmed to be a state-of-the-art classification method when using deep features as inputs. Shaoning Zeng, Yunbo Rao, Bob Zhang 0001, Yong Xu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | LeFUNet: UNet with Learnable Feature Connections for Teeth Identification and Segmentation in Dental Panoramic X-ray ImagesabstractDeep learning methods have widely been applied to accurately identify and segment individual teeth in panoramic X-ray radiographs. However, the task becomes challenging as deep learning models grow deeper and wider. Contextual information have to pass through many layers leading to features vanishing before reaching the end of the model. This study proposes a deep learning-based three-step model to identify and segment individual teeth from panoramic X-ray radiographs to address the issues above. Firstly, an automatic binarized transformation of panoramic images to deal with computational complexity of high-dimensionality with respect to less training data is conducted. The transformed images are then trained on with a vanilla top-down learnable feature connection based UNet. Specifically, a learnable feature connection module that incorporates an improved squeeze and excitation module with dense connections to ensure feature propagation and facilitating information flow throughout the network is designed and conveniently plunged into a UNet. Finally, accurate individual tooth identification and segmentation is achieved using proposed regions via bounding detection techniques. Extensive evaluation on a publicly available dental panoramic X-ray benchmark demonstrated the effectiveness the proposed scheme by obtaining a Dice score of 96.94%, 97.93% for accuracy, 97.03% for recall, 96.81% for precision, 97.61% for specificity and 93.52% for Jaccard Coefficient Score, which are significantly superior to the recent state-of-the-art methods for individual tooth identification and segmentation. Yunbo Rao, Obed Tettey Nartey, Shaoning Zeng, Kingsley Nketia Acheampong, Charles R. Haruna, Jianxun Sun |
BIBM | 1 |
| 2022 | A Visual Analytics Approach to Understanding Gradient Boosting Tree via Click Prediction on Ads
Zhuoyue Cheng, Kehan Cheng, Yulu Xia, Jiansu Pu, Yunbo Rao |
CDVE | 5 |
| 2022 | ENet: event based highlight generation network for broadcast sports videos
Abdullah Aman Khan, Yunbo Rao, Jie Shao 0001 |
Multim. Syst. | 2 |
| 2022 | matExplorer: Visual Exploration on Predicting Ionic Conductivity for Solid-state ElectrolytesabstractLithium ion batteries (LIBs) are widely used as important energy sources for mobile phones, electric vehicles, and drones. Experts have attempted to replace liquid electrolytes with solid electrolytes that have wider electrochemical window and higher stability due to the potential safety risks, such as electrolyte leakage, flammable solvents, poor thermal stability, and many side reactions caused by liquid electrolytes. However, finding suitable alternative materials using traditional approaches is very difficult due to the incredibly high cost in searching. Machine learning (ML)-based methods are currently introduced and used for material prediction. However, learning tools designed for domain experts to conduct intuitive performance comparison and analysis of ML models are rare. In this case, we propose an interactive visualization system for experts to select suitable ML models and understand and explore the predication results comprehensively. Our system uses a multifaceted visualization scheme designed to support analysis from various perspectives, such as feature distribution, data similarity, model performance, and result presentation. Case studies with actual lab experiments have been conducted by the experts, and the final results confirmed the effectiveness and helpfulness of our system. Jiansu Pu, Boyang Gao, Zhengguo Zhu, Yanlin Zhu, Yunbo Rao |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2021 | Visual Analysis on Machine Learning Assisted Prediction of Ionic Conductivity for Solid-State ElectrolytesabstractLithium ion batteries (LIBs) are widely used as the important energy sources in our daily life such as mobile phones, electric vehicles, and drones etc. Due to the potential safety risks caused by liquid electrolytes, the experts have tried to replace liquid electrolytes with solid ones. However, it is very difficult to find suitable alternatives materials in traditional ways for its incredible high cost in searching. Machine learning (ML) based methods are currently introduced and used for material prediction. But there is rarely an assisting learning tools designed for domain experts for institutive performance comparison and analysis of ML model. In this case, we propose an interactive visualization system for experts to select suitable ML models, understand and explore the predication results comprehensively. Our system employs a multi-faceted visualization scheme designed to support analysis from the perspective of feature composition, data similarity, model performance, and results presentation. A case study with real experiments in lab has been taken by the expert and the results of confirmed the effectiveness and helpfulness of our system. Jiansu Pu, Yanlin Zhu, Boyang Gao, Zhengguo Zhu, Yunbo Rao |
PacificVis | 6 |
| 2021 | GBMVis: Visual Analytics for Interpreting Gradient Boosting Machine
Yulu Xia, Kehan Cheng, Zhuoyue Cheng, Yunbo Rao, Jiansu Pu |
CDVE | 4 |
| 2020 | Managing Massive Amounts of Small Files in All-Flash StorageabstractAll-flash array is a popular memory device available for use in modern high-performance storage systems. Compared with other types of devices such as DRAM, NVRAM, and EEPROM, flash array combines the best features: shock resistance, low cost, low power consumption, and fast access. Moreover, the ever-increasing density of flash memory has led to a dramatic increase in the capacity, which allows the storage of large volume of data. However, flash memory is not optimal for managing a large number of small files because: 1. the small and random write operation is inefficient in flash memory; 2. massive metadata information occupies a significant portion of the namespace, which is relatively limited or scarce in big data storage systems. This paper introduces a novel approach, hash partitioning-based file compaction (HFC), to improve the efficiency of storing and accessing small files in all-flash storage systems. HFC consists of a file compaction tool and an access interface. The compaction tool merges a group (usually a directory) of small files into a set of "big files" to reduce the metadata required to be maintained in the on-chip memory. The data locality and tree structure of those small files are preserved. The access interface is designed to provide transparent access to the small files in the HFC big files. Experimental results confirm that the proposed method significantly enhances the efficiency of managing massive amounts of small files in flash memory in terms of namespace usage and access speed. Rize Jin, Joon-Young Paik, Yenewondim Biadgie, Yunbo Rao, Tae-Sun Chung |
COMPSAC | 5 |
| 2020 | Multi-data UAV Images for Large Scale Reconstruction of Buildings
Menghan Zhang, Yunbo Rao, Jiansu Pu, Xun Luo, Qifei Wang |
MMM (2) | 2 |
| 2019 | Automatic Image Annotation and Deep Learning for Tooth CT Image Segmentation
Miao Gou, Yunbo Rao, Minglu Zhang, Jianxun Sun, Keyang Cheng |
ICIG (2) | 2 |
| 2019 | A generalized mean distance-based k-nearest neighbor classifier
Jianping Gou, Hongxing Ma, Weihua Ou, Shaoning Zeng, Yunbo Rao, Hebiao Yang |
Expert Syst. Appl. | 5 |
| 2018 | Roads Detection of Aerial Image with FCN-CRF ModelabstractThis paper describes a deep learning based model for roads detection in Aerial image. In general, standard CNN networks would have less ability for tiny objects detection in remote sensing image. With this regard, we propose a novel fully convolutional network, which utilizes deconvolution layers and feature map fussing to take as input intensity and pixel-wise labeling. Moreover, the class prediction are used as the input to Condition Random Field (CRF) for the final pixel prediction. The Batch Normalization (BN) algorithm and two stages training strategy were used in our model to reduce the time cost of model training. Several experimental results conducted in Massachuseets. Road dataset demonstrate the superiority of our model with respect to accuracy and time cost. Yunbo Rao, Wei Liu 0073, Jiansu Pu, Qifei Wang |
VCIP | 1 |
| 2018 | Visual Analysis of Human Motion: A Survey on Recent Advances and ApplicationsabstractThis paper summarizes the recent progress in human motion analysis and its applications. The first part of this paper reviews the motion capture systems and the representations of human's motion data. Next, the paper sketches the advanced human motion data processing technologies, including motion data filtering, temporal alignment, and segmentation. The following parts overview the state-of-the-art approaches of action recognition and dynamics measuring. The last part discusses the emerging applications of human motion analysis in healthcare and human robot interaction. The promising research topics of human motion analysis in the future are also summarized in the last part. Qifei Wang, Yunbo Rao |
VCIP | 2 |
| 2018 | A novel relevance feedback method for CBIR
Yunbo Rao, Wei Liu 0073, Bojiang Fan, Jiali Song |
World Wide Web | 1 |
| 2017 | A Multi-local Means Based Nearest Neighbor ClassifierabstractIn this paper, we propose a multi-local means based nearest neighbor classifier (MLMNN). In the MLMNN, k categorical nearest neighbors of a query sample are first found and used to calculate the corresponding k categorical multi-local mean vectors which can represent different local class-specific sample distributions. Then, the query sample is represented by a linear combination of k categorical local mean vectors and the representation coefficient of each local mean vector as the contribution to representing and classifying the query sample is obtained. Finally, the class-specific representation-based distance (i.e. reconstruction residual) between the query sample and k categorical multi-local mean vectors is adopted to determine the class label of the query sample. The experimental results on three popular face databases show that the proposed MLMNN method outperforms the related competitive KNN-based methods. Jianping Gou, Wenmo Qiu, Qirong Mao, Yongzhao Zhan 0001, Xiangjun Shen, Yunbo Rao |
ICTAI | 6 |
| 2017 | Optimization algorithm based on texture feature and frame correlation in HEVC
Yunbo Rao |
Multim. Tools Appl. | 3 |
| 2017 | Anterior cruciate ligament reconstruction model based on anatomical position locating
Yunbo Rao, Xianshu Ding, Jianping Gou, Qifei Wang |
Multim. Tools Appl. | 1 |
| 2016 | Sparse codes fusion for context enhancement of night video surveillance
Xianshu Ding, Yunbo Rao |
Multim. Tools Appl. | 3 |
| 2015 | A New Optimization Algorithm for HEVC
Yunbo Rao |
ICIG (1) | 3 |
| 2015 | Automatic vehicle recognition in multiple cameras for video surveillance
Yunbo Rao |
Vis. Comput. | 1 |
| 2014 | Improved pseudo nearest neighbor classification
Jianping Gou, Yongzhao Zhan 0001, Yunbo Rao, Xiangjun Shen, Wu He |
Knowl. Based Syst. | 3 |
| 2014 | Illumination-based nighttime video contrast enhancement using genetic algorithm
Yunbo Rao, Zhihui Wang 0001, Leiting Chen |
Multim. Tools Appl. | 1 |
| 2011 | An effecive night video enhancement algorithmabstractNight video enhancement is important for video surveillance since many objects or activities of interest occur in a dark environment which cannot be seen easily without enhancement. In this paper, we discuss several problems of existing techniques for illumination-fusion based night video enhancement, which fuses video frames from day-time backgrounds and night-time video. We then present a simple enhancement algorithm without these problems. The algorithm uses an additive enhancement term with foreground object extraction and constrained low-passed object illumination to avoid light-inversion and sensitivity problems and to reduce ghost patterns. Experimental results show the effectiveness and robustness of the proposed algorithm. Yunbo Rao, Zhong-Ho Chen, Ming-Ting Sun, Yu-Feng Hsu, Zhengyou Zhang |
VCIP | 1 |
| 2011 | Real-time control of individual agents for crowd simulation
Yunbo Rao, Leiting Chen, Qihe Liu, Weiyao Lin |
Multim. Tools Appl. | 1 |