EDBT 2026 Demo / reviewers in the wild / expert
Yiming Cui 0002
dblp:130/6308-2
· DBLP profile ↗
25ranked-venue papers
6as first author
19since 2021 · last 2025
0000-0003-2423-8972ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | All You Need is One: Capsule Prompt Tuning with a Single VectorabstractPrompt-based learning has emerged as a parameter-efficient finetuning (PEFT) approach to facilitate Large Language Model (LLM) adaptation to downstream tasks by conditioning generation with task-aware guidance. Despite its successes, current prompt-based learning methods heavily rely on laborious grid searching for optimal prompt length and typically require considerable number of prompts, introducing additional computational burden. Worse yet, our pioneer findings indicate that the task-aware prompt design is inherently limited by its absence of instance-aware information, leading to a subtle attention interplay with the input sequence. In contrast, simply incorporating instance-aware information as a part of the guidance can enhance the prompt-tuned model performance without additional fine-tuning. Moreover, we find an interesting phenomenon, namely "attention anchor", that incorporating instance-aware tokens at the earliest position of the sequence can successfully preserve strong attention to critical structural information and exhibit more active attention interaction with all input tokens. In light of our observation, we introduce Capsule Prompt-Tuning (CaPT), an efficient and effective solution that leverages off-the-shelf, informative instance semantics into prompt-based learning. Our approach innovatively integrates both instance-aware and task-aware information in a nearly parameter-free manner (i.e., one single capsule prompt).
Empirical results demonstrate that our method can exhibit superior performance across various language tasks (e.g., 84.03\% average accuracy on T5-Large), serving as an "attention anchor," while enjoying high parameter efficiency (e.g., 0.003\% of model parameters on Llama3.2-1B). Yiyang Liu 0003, James Liang, Heng Fan 0001, Yiming Cui 0002, Lifu Huang, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001 |
NeurIPS | 5 |
| 2025 | One Neuron Saved is One Neuron Earned: On Parametric Efficiency of Quadratic NetworksabstractInspired by neuronal diversity in the biological neural system, a plethora of studies proposed to design novel types of artificial neurons and introduce neuronal diversity into artificial neural networks. Recently proposed quadratic neuron, which replaces the inner-product operation in conventional neurons with a quadratic one, have achieved great success in many essential tasks. Despite the promising results of quadratic neurons, there is still an unresolved issue: Is the superior performance of quadratic networks simply due to the increased parameters or due to the intrinsic expressive capability? Without clarifying this issue, the performance of quadratic networks is always suspicious. Additionally, resolving this issue is reduced to finding killer applications of quadratic networks. In this paper, with theoretical and empirical studies, we show that quadratic networks enjoy parametric efficiency, thereby confirming that the superior performance of quadratic networks is due to the intrinsic expressive capability. This intrinsic expressive ability comes from that quadratic neurons can easily represent nonlinear interaction, while it is hard for conventional neurons. Theoretically, we derive the approximation efficiency of quadratic networks over conventional ones in terms of real space and manifolds. Moreover, from the perspective of the Barron space, we demonstrate that there exists a functional space whose functions can be approximated by quadratic networks in a dimension-free error, but the approximation error of conventional networks is dependent on dimensions. Empirically, experimental results on synthetic data, classic benchmarks, and real-world applications show that quadratic models broadly enjoy parametric efficiency, and the gain of efficiency depends on the task. Fenglei Fan, Hangcheng Dong, Zhongming Wu, Lecheng Ruan, Tieyong Zeng, Yiming Cui 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | ProMotion: Prototypes as Motion LearnersabstractIn this work, we introduce PRoMoTION, a unified proto-typical transformer-based framework engineered to model fundamental motion tasks. PRoMoTION offers a range of compelling attributes that set it apart from current task-specific paradigms. (1) We adopt a prototypical perspective, establishing a unified paradigm that harmonizes disparate motion learning approaches. This novel paradigm stream-lines the architectural design, enabling the simultaneous assimilation of diverse motion information. (2) We capitalize on a dual mechanism involving the feature denoiser and the prototypical learner to decipher the intricacies of motion. This approach effectively circumvents the pitfalls of ambiguity in pixel-wise feature matching, significantly bolstering the robustness of motion representation. (3)) We demon-strate a profound degree of transferability across distinct motion patterns. This inherent versatility reverberates robustly across a comprehensive spectrum of both 2D and 3D downstream tasks. Empirical results demonstrate that PRoMOTION outperforms various well-known specialized architectures, achieving 0.54 and 0.054$AbsRel$error on the Sintel and KITTI depth datasets, 1.04 and 2.01 average endpoint error on the clean and final pass of Sintel flow benchmark, and 4.30 F1-all error on the KITTI flow bench-mark. For its efficacy, we hope our work can catalyze a paradigm shift in universal models in computer vision. Yawen Lu, Dongfang Liu, Qifan Wang 0001, Cheng Han 0001, Yiming Cui 0002, Zhiwen Cao, Xueling Zhang, Victor Y. Chen, Heng Fan 0001 |
CVPR | 5 |
| 2024 | M²PT: Multimodal Prompt Tuning for Zero-shot Instruction LearningabstractTaowen Wang, Yiyang Liu, James Chenhao Liang, Junhan Zhao, Yiming Cui, Yuning Mao, Shaoliang Nie, Jiahao Liu, Fuli Feng, Zenglin Xu, Cheng Han, Lifu Huang, Qifan Wang, Dongfang Liu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Taowen Wang, Yiyang Liu 0003, James Liang, Junhan Zhao, Yiming Cui 0002, Yuning Mao, Shaoliang Nie, Fuli Feng, Zenglin Xu, Cheng Han 0001, Lifu Huang, Qifan Wang 0001, Dongfang Liu |
EMNLP | 5 |
| 2024 | Facing the Elephant in the Room: Visual Prompt Tuning or Full finetuning?abstractAs the scale of vision models continues to grow, the emergence of Visual Prompt Tuning (VPT) as a parameter-efficient transfer learning technique has gained attention due to its superior performance compared to traditional full-finetuning. However, the conditions favoring VPT (the "when") and the underlying rationale (the "why") remain unclear. In this paper, we conduct a comprehensive analysis across 19 distinct datasets and tasks. To understand the "when" aspect, we identify the scenarios where VPT proves favorable by two dimensions: task objectives and data distributions. We find that VPT is preferrable when there is 1) a substantial disparity between the original and the downstream task objectives ($e.g.$, transitioning from classification to counting), or 2) a notable similarity in data distributions between the two tasks ($e.g.$, both involve natural images). In exploring the "why" dimension, our results indicate VPT's success cannot be attributed solely to overfitting and optimization considerations. The unique way VPT preserves original features and adds parameters appears to be a pivotal factor. Our study provides insights into VPT's mechanisms, and offers guidance for its optimal utilization. Cheng Han 0001, Qifan Wang 0001, Yiming Cui 0002, Wenguan Wang, Lifu Huang, Siyuan Qi, Dongfang Liu |
ICLR | 3 |
| 2024 | Error-Robust and Label-Efficient Deep Learning for Understanding Tumor Microenvironment From Spatial TranscriptomicsabstractSpatial transcriptomics (ST) has become an important methodology in the analysis of the tumor microenvironment (TME) due to its ability to provide gene expression information with spatial resolution, enabling the identification and characterization of TME gene markers. Deep learning methods are proposed for analyzing spatial transcriptomic data for clustering the spatial regions of the TME based on gene expression. However, deep learning methods are often imposed by errors, which can impact the accuracy of gene expression quantification and TME gene identification. To address this issue, we propose a label-efficient method that utilizes curriculum learning and confidence learning to identify errors in graph deep learning when analyzing ST data. Our method explicitly incorporates the effect of noise in the learning process and employs probabilistic models or uncertainty estimates to represent the uncertainty in the data. Validated on human breast cancer ST data, we studied spatial gene expression in HER2-positive breast tumors using our method. The evaluation results suggest that the error quantification helps identify the noisy samples and subset the samples that results in more accurate gene expression quantification and TME gene identification. Additionally, there are biological insights obtained from the new subset formed by error samples. This error-robust deep learning method offers promising avenues for the analysis of spatial transcriptomic data, enabling accurate and label-efficient quantification of gene expression and identification of TME gene markers. Jiake Leng, Yiming Cui 0002, Junhan Zhao, Yongxin Ge |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | SAC-Net: Enhancing Spatiotemporal Aggregation in Cervical Histological Image Classification via Label-Efficient Weakly Supervised LearningabstractCervical cancer is the fourth most common cancer in women and its subtyping requires examining histopathological slides or digital images, such as whole slide images (WSIs). However, manually inspecting WSIs with gigapixel sizes can be laborious and prone to errors for pathologists. To address this issue, computer-aided approaches based on weakly-supervised learning techniques have been proposed. These methods can predict disease types directly from WSIs and highlight diagnosis-relevant regions, which can help pathologists achieve faster and more accurate diagnoses. WSIs are divided into overlapping patches using a sliding window approach, and these patches are subsequently screened in a sequential zig-zag pattern to identify spatiotemporal dependencies. These dependencies are further analyzed to generate predictions at the WSI level. Therefore, effective patch feature learning and spatiotemporal aggregation are two key issues in the weakly-supervised WSI classification (WSWC) task. In this paper, we present a label-efficient WSWC method called spatiotemporal aggregation for cervical WSIs (SAC-Net), which jointly performs online feature extraction and feature aggregation to infer the WSI-level prediction in an end-to-end manner. The online feature extractor helps to learn cervical-cancer-specific features and obtain more accurate patch representations. The feature aggregator uses an online instance clustering method to learn proper weight parameters for each cluster, which generates the WSI embedding with enhanced spatiotemporal aggregation. SAC-Net is developed and evaluated on a public cervical WSI dataset (TissueNet) containing 1015 WSIs, which are also externally tested on three independent cervical WSI datasets. Our results demonstrate that SAC-Net achieves state-of-the-art classification performance and is robust. SAC-Net has the potential to be a useful tool for clinical cervical cancer detection. De Cai, Sen Yang 0006, Yiming Cui 0002, Junyou Zhu, Kanran Wang, Junhan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Feature Aggregated Queries for Transformer-Based Video Object DetectorsabstractVideo object detection needs to solve feature degradation situations that rarely happen in the image domain. One solution is to use the temporal information and fuse the features from the neighboring frames. With Transformer-based object detectors getting a better performance on the image domain tasks, recent works began to extend those methods to video object detection. However, those existing Transformer-based video object detectors still follow the same pipeline as those used for classical object detectors, like enhancing the object feature representations by aggregation. In this work, we take a different perspective on video object detection. In detail, we improve the qualities of queries for the Transformer-based models by aggregation. To achieve this goal, we first propose a vanilla query aggregation module that weighted averages the queries according to the features of the neighboring frames. Then, we extend the vanilla module to a more practical version, which generates and aggregates queries according to the features of the input frames. Extensive experimental results validate the effectiveness of our proposed methods: On the challenging ImageNet VID benchmark, when integrated with our proposed modules, the current state-of-the-art Transformer-based object detectors can be improved by more than 2.4% on mAP and 4.2% on AP50. Code is available at https://github.com/YimingCuiCuiCui/FAQ. Yiming Cui 0002 |
CVPR | 1 |
| 2023 | E2VPT: An Effective and Efficient Approach for Visual Prompt TuningabstractAs the size of transformer-based, models continues to grow, fine-tuning these large-scale pretrained vision models for new tasks has become increasingly parameter-intensive. Parameter-efficient learning has been developed to reduce the number of tunable parameters during fine-tuning. Although these methods show promising results, there is still a significant performance gap compared to full fine-tuning. To address this challenge, we propose an Effective and Efficient Visual Prompt Tuning (E2VPT) approach for large-scale transformer-based model adaptation. Specifically, we introduce a set of learnable key-value prompts and visual prompts into self-attention and input layers, respectively, to improve the effectiveness of model fine-tuning. Moreover, we design a prompt pruning procedure to systematically prune low importance prompts while preserving model performance, which largely enhances the model’s efficiency. Empirical results demonstrate that our approach outperforms several state-of-the-art baselines on two benchmarks, with considerably low parameter usage (e.g., 0.32% of model parameters on VTAB-1k). Our code is available at https://github.com/ChengHan111/E2VPT. Cheng Han 0001, Qifan Wang 0001, Yiming Cui 0002, Zhiwen Cao, Wenguan Wang, Siyuan Qi, Dongfang Liu |
ICCV | 3 |
| 2023 | Learning Dynamic Query Combinations for Transformer-based Object Detection and SegmentationabstractTransformer-based detection and segmentation methods use a list of learned detection queries to retrieve information from the transformer network and learn to predict the location and category of one specific object from each query. We empirically find that random convex combinations of the learned queries are still good for the corresponding models. We then propose to learn a convex combination with dynamic coefficients based on the high-level semantics of the image. The generated dynamic queries, named as modulated queries, better capture the prior of object locations and categories in the different images. Equipped with our modulated queries, a wide range of DETR-based models achieve consistent and superior performance across multiple tasks (object detection, instance segmentation, panoptic segmentation) and on different benchmarks (MS COCO, CityScapes, YoutubeVIS). Yiming Cui 0002, Haichao Yu |
ICML | 1 |
| 2023 | ClusterFomer: Clustering As A Universal Visual LearnerabstractThis paper presents ClusterFormer, a universal vision model that is based on the Clustering paradigm with TransFormer. It comprises two novel designs: 1) recurrent cross-attention clustering, which reformulates the cross-attention mechanism in Transformer and enables recursive updates of cluster centers to facilitate strong representation learning; and 2) feature dispatching, which uses the updated cluster centers to redistribute image features through similarity-based metrics, resulting in a transparent pipeline. This elegant design streamlines an explainable and transferable workflow, capable of tackling heterogeneous vision tasks (i.e., image classification, object detection, and image segmentation) with varying levels of clustering granularity (i.e., image-, box-, and pixel-level). Empirical results demonstrate that ClusterFormer outperforms various well-known specialized architectures, achieving 83.41% top-1 acc. over ImageNet-1K for image classification, 54.2% and 47.0% mAP over MS COCO for object detection and instance segmentation, 52.4% mIoU over ADE20K for semantic segmentation, and 55.8% PQ over COCO Panoptic for panoptic segmentation. This work aims to initiate a paradigm shift in universal visual understanding and to benefit the broader field. James Liang, Yiming Cui 0002, Qifan Wang 0001, Tong Geng, Wenguan Wang, Dongfang Liu |
NeurIPS | 2 |
| 2022 | Dynamic Feature Aggregation for Efficient Video Object Detection
Yiming Cui 0002 |
ACCV (2) | 1 |
| 2022 | GL-RG: Global-Local Representation Granularity for Video CaptioningabstractVideo captioning is a challenging task as it needs to accurately transform visual understanding into natural language description. To date, state-of-the-art methods inadequately model global-local representation across video frames for caption generation, leaving plenty of room for improvement. In this work, we approach the video captioning task from a new perspective and propose a GL-RG framework for video captioning, namely a Global-Local Representation Granularity. Our GL-RG demonstrates three advantages over the prior efforts: 1) we explicitly exploit extensive visual representations from different video ranges to improve linguistic expression; 2) we devise a novel global-local encoder to produce rich semantic vocabulary to obtain a descriptive granularity of video contents across frames; 3) we develop an incremental training strategy which organizes model learning in an incremental fashion to incur an optimal captioning behavior. Experimental results on the challenging MSR-VTT and MSVD datasets show that our DL-RG outperforms recent state-of-the-art methods by a significant margin. Code is available at https://github.com/ylqi/GL-RG. Liqi Yan, Qifan Wang 0001, Yiming Cui 0002, Fuli Feng, Xiaojun Quan, Xiangyu Zhang 0001, Dongfang Liu |
IJCAI | 3 |
| 2022 | DG-Labeler and DGL-MOTS Dataset: Boost the Autonomous Driving PerceptionabstractMulti-object tracking and segmentation (MOTS) is a critical task for autonomous driving applications. The existing MOTS studies face two critical challenges: 1) the published datasets inadequately capture the real-world complexity for network training to address various driving settings; 2) the working pipeline annotation tool is under-studied in the literature to improve the quality of MOTS learning examples. In this work, we introduce the DG-Labeler and DGL-MOTS dataset to facilitate the training data annotation for the MOTS task and accordingly improve network training accuracy and efficiency. DG-Labeler uses the novel Depth-Granularity Module to depict the instance spatial relations and produce fine-grained instance masks. Annotated by DG-Labeler, our DGL-MOTS dataset exceeds the prior effort (i.e., KITTI MOTS and BDD100K) in data diversity, annotation quality, and temporal representations. Results on extensive cross-dataset evaluations indicate significant performance improvements for several state-of-the-art methods trained on our DGL-MOTS dataset. We believe our DGL-MOTS Dataset and DG-Labeler hold the valuable potential to boost the visual perception of future transportation. Our dataset and code are available here1. Yiming Cui 0002, Zhiwen Cao, Chloe Yixin Xie, Xingyu Jiang 0001, Feng Tao 0002, Victor Y. Chen, Dongfang Liu |
WACV | 1 |
| 2021 | DenserNet: Weakly Supervised Visual Localization Using Multi-Scale Feature AggregationabstractIn this work, we introduce a Denser Feature Network(DenserNet) for visual localization. Our work provides three principal contributions. First, we develop a convolutional neural network (CNN) architecture which aggregates feature maps at different semantic levels for image representations. Using denser feature maps, our method can produce more key point features and increase image retrieval accuracy. Second, our model is trained end-to-end without pixel-level an-notation other than positive and negative GPS-tagged image pairs. We use a weakly supervised triplet ranking loss to learn discriminative features and encourage keypoint feature repeatability for image representation. Finally, our method is computationally efficient as our architecture has shared features and parameters during forwarding propagation. Our method is flexible and can be crafted on a light-weighted backbone architecture to achieve appealing efficiency with a small penalty on accuracy. Extensive experiment results indicate that our method sets a new state-of-the-art on four challenging large-scale localization benchmarks and three image retrieval benchmarks with the same level of supervision. The code is available at https://github.com/goodproj13/DenserNet Dongfang Liu, Yiming Cui 0002, Liqi Yan, Christos Mousas, Baijian Yang 0001, Victor Y. Chen |
AAAI | 2 |
| 2021 | SG-Net: Spatial Granularity Network for One-Stage Video Instance SegmentationabstractVideo instance segmentation (VIS) is a new and critical task in computer vision. To date, top-performing VIS methods extend the two-stage Mask R-CNN by adding a tracking branch, leaving plenty of room for improvement. In contrast, we approach the VIS task from a new perspective and propose a one-stage spatial granularity network (SG-Net). Compared to the conventional two-stage methods, SG-Net demonstrates four advantages: 1) Our method has a one-stage compact architecture and each task head (detection, segmentation, and tracking) is crafted interdependently so they can effectively share features and enjoy the joint optimization; 2) Our mask prediction is dynamically performed on the sub-regions of each detected instance, leading to high-quality masks of fine granularity; 3) Each of our task predictions avoids using expensive proposal-based RoI features, resulting in much reduced runtime complexity per instance; 4) Our tracking head models objects’ centerness movements for tracking, which effectively enhances the tracking robustness to different object appearances. In evaluation, we present state-of-the-art comparisons on the YouTube-VIS dataset. Extensive experiments demonstrate that our compact one-stage method can achieve improved performance in both accuracy and inference speed. We hope our SG-Net could serve as a strong and flexible base-line for the VIS task. Our code will be available here1. Dongfang Liu, Yiming Cui 0002, Wenbo Tan, Victor Y. Chen |
CVPR | 2 |
| 2021 | Hierarchical Attention Fusion for Geo-LocalizationabstractGeo-localization is a critical task in computer vision. In this work, we cast the geo-localization as a 2D image retrieval task. Current state-of-the-art methods for 2D geo-localization are not robust to locate a scene with drastic scale variations because they only exploit features from one semantic level for image representations. To address this limitation, we introduce a hierarchical attention fusion network using multi-scale features for geo-localization. We extract the hierarchical feature maps from a convolutional neural network (CNN) and organically fuse the extracted features for image representations. Our training is self-supervised using adaptive weights to control the attention of feature emphasis from each hierarchical level. Evaluation results on the image retrieval and the large-scale geo-localization benchmarks indicate that our method outperforms the existing state-of-the-art methods. Code is available here: https://github.com/YanLiqi/HAF. Liqi Yan, Yiming Cui 0002, Victor Y. Chen, Dongfang Liu |
ICASSP | 2 |
| 2021 | TF-Blender: Temporal Feature Blender for Video Object DetectionabstractVideo objection detection is a challenging task because isolated video frames may encounter appearance deterioration, which introduces great confusion for detection. One of the popular solutions is to exploit the temporal information and enhance per-frame representation through aggregating features from neighboring frames. Despite achieving improvements in detection, existing methods focus on the selection of higher-level video frames for aggregation rather than modeling lower-level temporal relations to increase the feature representation. To address this limitation, we propose a novel solution named TF-Blender, which includes three modules: 1) Temporal relation models the relations between the current frame and its neigh-boring frames to preserve spatial information. 2). Feature adjustment enriches the representation of every neigh-boring feature map; 3) Feature blender combines outputs from the first two modules and produces stronger features for the later detection tasks. For its simplicity, TF-Blender can be effortlessly plugged into any detection network to improve detection behavior. Extensive evaluations on ImageNet VID and YouTube-VIS benchmarks indicate the performance guarantees of using TF-Blender on recent state-of-the-art methods. Code is available at https://github.com/goodproj13/TF-Blender. Yiming Cui 0002, Liqi Yan, Zhiwen Cao, Dongfang Liu |
ICCV | 1 |
| 2021 | Geometric attentional dynamic graph convolutional neural networks for point cloud analysis
Yiming Cui 0002, Xin Liu 0027, Hongmin Liu 0001, Jiyong Zhang 0001, Alina Zare, Bin Fan 0001 |
Neurocomputing | 1 |
| 2020 | Visual Localization for Autonomous Driving: Mapping the Accurate Location in the City MazeabstractAccurate localization is a foundational capacity, required for autonomous vehicles to accomplish other tasks such as navigation or path planning. It is a common practice for vehicles to use GPS to acquire location information. However, the application of GPS can result in severe challenges when vehicles run within the inner city where different kinds of structures may shadow the GPS signal and lead to inaccurate location results. To address the localization challenges of urban settings, we propose a novel feature voting technique for visual localization. Different from the conventional front-view-based method, our approach employs views from three directions (front, left, and right) and thus significantly improves the robustness of location prediction. In our work, we craft the proposed feature voting method into three state-of-the-art visual localization networks and modify their architectures properly so that they can be applied for vehicular operation. Extensive field test results indicate that our approach can predict location robustly even in challenging inner-city settings. Our research sheds light on using the visual localization approach to help autonomous vehicles to find accurate location information in a city maze, within a desirable time constraint. The source code is available at github.com/HappyDonkey13/Visual- Localization-for- Autonomous-Driving. Dongfang Liu, Yiming Cui 0002, Baijian Yang 0001, Victor Y. Chen |
ICPR | 2 |
| 2020 | A Large-scale Simulation Dataset: Boost the Detection Accuracy for Special Weather ConditionsabstractObject detection is a fundamental task for autonomous driving systems. One bottleneck hindering detection accuracy is a shortage of well-annotated image data. Virtual reality has provided a feasible low-cost way to facilitate computer vision related developments. In autonomous driving area, existing public datasets from real world generally have data biases and cannot represent a wide range of weather conditions, such as rainy or snowy roads. To address this challenge, we introduce a new large-scale simulation dataset which is generated by an automated pipeline from a high realism video game. Our dataset focuses on weather conditions, which can be adopted to train networks to effectively detect objects under such conditions. We use extensive experiments to evaluate our dataset by comparing it with public datasets. The experiment results show that networks trained with our dataset outperform the networks trained by other public datasets. Our work demonstrates the effectiveness of using simulation data to address real-world challenges in the practice of object detection. Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen |
IJCNN | 2 |
| 2020 | Indoor Navigation for Mobile Agents: A Multimodal Vision Fusion ModelabstractIndoor navigation is a challenging task for mobile agents. The latest vision-based indoor navigation methods make remarkable progress in this field but do not fully leverage visual information for policy learning and struggle to perform well in unseen scenes. To address the existing limitations, we present a multimodal vision fusion model (MVFM). We implement a joint modality of different image recognition networks for navigation policy learning. The proposed model incorporates object detection for target searching, depth estimation for distance prediction, and semantic segmentation to depict the walkable region. In design, our model provides holistic vision knowledge for navigation. Evaluation on AI2-THOR indicates that MVFM improves on the results of a strong baseline model by 3.49% for Success weighted by Path Length (SPL) and 4% for success rate respectively. In comparison with other state-of-the-art systems, MVFM performs in the lead in terms of SPL and success rate. Extensive experiments show the effectiveness of the proposed model. Dongfang Liu, Yiming Cui 0002, Zhiwen Cao, Victor Y. Chen |
IJCNN | 2 |
| 2020 | Video object detection for autonomous driving: Motion-aid feature calibration
Dongfang Liu, Yiming Cui 0002, Victor Y. Chen, Jiyong Zhang 0001, Bin Fan 0001 |
Neurocomputing | 2 |
| 2016 | Manifold learning based supervised hyperspectral data classification method using class encodingabstractManifold learning based unsupervised classification methods will be unable to obtain satisfactory results because of the lack of training samples. The employment of training samples' information makes manifold learning based classification become supervised, and thus brings the improvement on classification accuracy. In order to make full use of this information, we emphatically consider the hyperspectral data distribute by clusters. A novel supervised manifold learning method termed class encoding is proposed for hyperspectral data classification. The experimental results show that this algorithm has better classification performance than the existing supervised manifold learning algorithm. Miao Zhang 0001, Yiming Cui 0002, Yi Shen 0001 |
IGARSS | 3 |
| 2016 | Multiclassification method for hyperspectral data based on Chernoff distance and pairwise decision tree strategyabstractTo address the multi-classification problems of hyperspectral dataset, a new method with weighted kernel function based on Chernoff distance is proposed. Chernoff distance utilizes the information between categories and strengthens the separability of original dataset. The adjustable parameter in Chernoff distance can fit the hyperspectral dataset well compared with other least upper bounds. Pairwise decision tree reduces the number of subclassifiers that the dataset requires and improves the classification accuracy. The guidance of the weighed subclassifiers is global separability metric computed by Chernoff distance. Weighted subclassifiers highlight bands with more useful information and reduce accumulative error. Comparative experiment shows the effectiveness of the proposed method. Miao Zhang 0001, Zheqi Lin, Yiming Cui 0002, Yi Shen 0001 |
IGARSS | 3 |