Yaochen Li

dblp:75/10700 · DBLP profile ↗
← Back
56ranked-venue papers
13as first author
35since 2021 · last 2026
0000-0001-5741-5280ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Frequency-aware experts with multi-stage fusion for multimodal sentiment analysis
Xiaofei Zhu, Yaochen Li
J. Intell. Inf. Syst.2
2026 Key Instance-Based Spatio-Temporal Network for Group Activity Recognition
abstract
Group activity recognition involves detecting the collective actions performed by a group of individuals, where identifying the key actors and key frames is crucial for understanding the group’s behavior. To tackle this challenge, we propose a spatio-temporal reasoning framework that leverages key instances. Our key instance identification module effectively detects key roles and frames from video sequences, while a graph-based reasoning model dynamically aggregates the features of related actors. We extract joint features and RGB features from video sequences, and these are fused using our multi-modal fusion TCT module, which improves the representation power of the original features. To better understand group activity through spatio-temporal correlations, we further utilize an enhanced cross-transformer module for spatio-temporal synchronous reasoning, considering both time and space dimensions. Our method has been tested on two public datasets, showing that it achieves high accuracy and surpasses many state-of-the-art approaches.
Yaochen Li, Haoting He, Yutong Wang 0008, Gaojie Li, Yuehu Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2026 SCALE-Pose: Skeletal Correction and Language Knowledge-assisted for 3D Human Pose Estimation
abstract
Transformer-based 3D human pose estimation methods typically use 2D joint sequences as inputs, leveraging spatial and temporal transformer encoders to model the 3D human pose. However, these methods often fail to incorporate skeletal constraints to limit joint motion. The integration of prior category knowledge to enhance joint representations is also neglected. To address these challenges, a novel approach named SCALE-Pose is proposed in this article. Our method first constructs a feature extraction network based on spatiotemporal skeleton refinement, where the skeletal correction modules are designed within both the spatial and temporal skeleton encoders to enhance the backbone network’s understanding of skeletal features. Meanwhile, a weighted average joint position error loss function is applied to improve the network’s ability to represent the joints with varying difficulty levels. A new radian-based loss function for skeletal joint angles is also designed, further enhancing the model’s capability to capture subtle skeletal movements. Furthermore, a training strategy based on large language model (LLM) priors is proposed to generate category-specific prior semantic knowledge from category keywords, which is then incorporated as auxiliary information to extract motion features. The experimental results based on the Human3.6M and MPI-INF-3DHP datasets well demonstrate the effectiveness of the proposed method.
Yaochen Li, Xinnan Ma, Limeng Zhao, ChenXu Zhou, Yuehu Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Edge-Deployable Spatiotemporal Modeling Network for Vehicle Behavior Recognition
abstract
Vehicle behavior recognition is essential for autonomous vehicles to quickly perceive and respond to their driving environment. In this paper, an edge-deployable spatiotemporal modeling network for vehicle behavior recognition is developed. Firstly, an Efficient SpatioTemporal Modeling (ESTM) Block is designed to extract both long-term evolution features and short-term motion information. Secondly, a Channel-Enhanced Spatial Modeling (CESM) Block is developed to capture the interdependencies among channels in spatial modeling. The proposed network can effectively process video input from onboard cameras while minimizing computational parameters and FLOPs. The combination of the ESTM and CESM blocks can produce rich and effective features for vehicle behavior recognition on edge devices. The experimental results demonstrate the effectiveness of the proposed method in real-world driving scenarios.
Gaojie Li, Yaochen Li, Yutong Wang 0008, Sibo Hao, Yuanqi Su
IV2
2025 CATIT: Cross-Adaptive Transformer for Road Image Translation
abstract
In the high-resolution image style transfer task of road traffic scenes, how to fully transfer style information while retaining the original content structure information is a challenging problem. In this paper, a novel Transformer-based generative adversarial network for high-resolution un-paired image translation is proposed. Firstly, we design a style transformation module based on cross adaptive Transformer, which dynamically adjusts the content features to achieve statistical alignment between content features and the target style. Meanwhile, an image frequency-domain enhancement module is designed based on cross-attention, which fuses the global information of the low-frequency style with the local details of the high-frequency content information. The detailed texture is then enhanced while the image style consistency is maintained. Furthermore, we design a threshold-guided negative sample screening strategy based on contrastive learning, which can improves the model's transfer effect. The experimental results well demonstrate the effectiveness of the proposed method.
Jinhuo Yang, Yaochen Li, Yi Han 0009, Sitong Li, Peijun Chen, Jintao Chang, Yuanqi Su
IV2
2025 Adaptive Semantic Segmentation of Traffic Scenes via Frequency Domain Analysis
abstract
High-precision semantic segmentation is an important research topic in the communities of computer vision and intelligent transportation. The existing unsupervised domain adaptation methods based on image translation often lead to artifacts and structural distortions. To overcome this problem, a novel adaptive semantic segmentation method of traffic scenes via frequency domain analysis is proposed. Firstly, we leverage the frequency domain space to decouple style and semantic features. The Fast Fourier Transform is applied to achieve structural-preserving style alignment. Subsequently, a content enhancement module is proposed based on the Wavelet transform, which utilizes the original source images to correct and enhance high-frequency structural and semantic details. Furthermore, a convolutional enhancement attention module is proposed, which utilizes depthwise separable convolution to capture more local details. The experimental results based on the GTA5→Cityscapes and SYNTHIA→Cityscapes tasks have respectively attained state-of-the-art mIoU of 76.4 and 67.7, convincingly demonstrating the effectiveness of the methods.
Tengwen Zhang, Yaochen Li, Runlin Zou, Chao Qiu, Ziyuan He, Hong Ni
IV2
2025 Image Captioning with Multimodal Guidance and Search Space Optimization
abstract
Image captioning bridges the gap between visual perception and natural language understanding by transforming image content into descriptive text. While existing methods have made significant progress in visual feature extraction, encoding, and cross-modal semantic alignment, challenges remain in terms of fine-grained feature representation, cross-modal alignment efficiency, and suboptimal search strategies. To address these issues, a multimodal-guided and search space-optimized image captioning model is proposed. In the visual encoding stage, we construct a hierarchical network that integrates regional and grid features through a geometry-constrained multi-layer feature aggregation mechanism, which enhances the model's capability to jointly capture global semantics and local details. In the decoding stage, we introduce a dynamic grouped beam width adjustment strategy to improve semantic path exploration. Additionally, a diversity-driven scoring function is designed to enforce intra-group diversity rewards and inter-group similarity penalties, encouraging the generation of more diverse captions. Finally, we incorporate a two-level pruning algorithm based on syntactic and spatial logic constraints to refine the search space from both hard and soft constraint perspectives, improving both the accuracy and diversity of generated captions. A 3% improvement in CIDEr is achieved by the proposed method over state-of-the-art (SOTA) models, as demonstrated by experiments on the COCO and Flickr30k datasets.
Yimou Guo, Yaochen Li, Jingze Liu, Haoyi Lou, Yuanqi Su
ACM Multimedia2
2025 SVDGNet: Shapley Value-Based Weight Adjustment for Unsupervised Image Style Transfer
abstract
With the advancement of autonomous driving technology, there is an increasing demand for high-quality and diverse images of road traffic scenes. Style transfer techniques can be employed to synthesize large-scale datasets. However, existing image style transfer methods often exhibit suboptimal performance in transferring styles for road scenes, frequently struggling to maintain structural consistency. In this paper, we propose a novel network architecture for unsupervised image style transfer named SVDGNet. This architecture dynamically adjusts the weights of different image regions during model training by calculating the Shapley values for the source and target domain images. We also employ a pre-trained diffusion model to generate better-stylized images. The experimental results demonstrate that the proposed method achieves better performance compared to the existing methods, which can preserve the structural consistency of the source domain images while providing impressive style transfer results.
Yi Han 0009, Yaochen Li, Peijun Chen, Jinhuo Yang, Jintao Chang
ACM Multimedia2
2025 Domain adaptation for semantic segmentation of road scenes via two-stage alignment of traffic elements
Yaochen Li, Hao Liao, Tenweng Zhang, Chao Qiu
Neurocomputing2
2025 Prototype-based multi-domain self-distillation for unbiased scene graph generation
Yaochen Li, Yujie Zang, Jingze Liu, Yuehu Liu
Neurocomputing2
2025 PFENet: Towards precise feature extraction from sparse point cloud for 3D object detection
Yaochen Li, Shengjing Gao
Neural Networks1
2025 Dynamic Fusion Label Assignment Network for Remote Sensing Object Detection
abstract
Object detection in remote sensing images faces significant challenges posed by tiny objects, which are often overwhelmed by background noise. These tiny objects exhibit significant scale differences compared to larger objects, making it difficult to simultaneously consider both large and small objects during label assignment. In this paper, a novel dynamic fusion label assignment network (DFLAN) is proposed to address these problems. Firstly, to effectively extract features of tiny objects in background noise, we introduce a novel feature selection interactive pyramid network. Secondly, a novel dynamic fusion label assignment algorithm is developed, which achieves a collaborative approach to label assignment for both tiny and large objects. Finally, a new decoupling detection head is proposed to prevent task coupling from interfering with the already weak features of tiny objects. The proposed DFLAN method achieves state-of-the-art performance on two widely-used datasets: DOTA-v1.0 (79.42% mAP), HRSC2016 (98.86% mAP) and DIOR-R (67.68% mAP).
Chaoshi Lu, Yaochen Li, Yuanbo Kou, Dinghao Li, Mingtao He
IEEE Trans. Geosci. Remote. Sens.2
2025 Vision-Based Driving Decision Making Using Multi-Action Deep Q Network
Sheng Yuan, Yaochen Li, Li Zhu 0003, Xinnan Ma, Yuncheng Xu
IEEE Trans. Intell. Transp. Syst.2
2025 ESE-GAN: Zero-Shot Food Image Classification Based on Low Dimensional Embedding of Visual Features
abstract
Existing zero-shot learning based image classification methods transform the zero-shot learning problem into supervised learning by applying generative adversarial network (GAN) to synthesize visual features of unseen classes. However, the visual features generated by the generator tend to be biased towards seen classes, and the discriminator is too weak to generate high-quality image features. To solve these problems, we propose a novel zero-shot food image classification method based on low dimensional embedding of visual features. Our method applies reinforced semantic guidance to increase the discriminative ability of the model by enhancing the strong distribution of input features. Moreover, the visual space is utilized as the embedding space to reduce the bias towards seen classes by reducing the distance between semantic information and visual features in the embedding space. Finally, the feature distribution of unseen classes is further specified by improving the prototype similarity function. Extensive experiments on three food datasets and four general benchmark datasets demonstrate the effectiveness of the proposed method.
Gaojie Li, Yaochen Li, Jingle Liu, Wenneng Tang, Yuehu Liu
IEEE Trans. Multim.2
2024 SEIT: Structural Enhancement for Unsupervised Image Translation in Frequency Domain
abstract
For the task of unsupervised image translation, transforming the image style while preserving its original structure remains challenging. In this paper, we propose an unsupervised image translation method with structural enhancement in frequency domain named SEIT. Specifically, a frequency dynamic adaptive (FDA) module is designed for image style transformation that can well transfer the image style while maintaining its overall structure by decoupling the image content and style in frequency domain. Moreover, a wavelet-based structure enhancement (WSE) module is proposed to improve the intermediate translation results by matching the high-frequency information, thus enriching the structural details. Furthermore, a multi-scale network architecture is designed to extract the domain-specific information using image-independent encoders for both the source and target domains. The extensive experimental results well demonstrate the effectiveness of the proposed method.
Zhifeng Zhu, Yaochen Li, Jinhuo Yang, Peijun Chen, Yuehu Liu
AAAI2
2024 Group Activity Recognition via Spatio-Temporal Reasoning of Key Instances
Haoting He, Yaochen Li, Yutong Wang 0008, Gaojie Li, Runlin Zou
BMVC2
2024 Template-Guided Data Augmentation for Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to identify objects within an image and infer relationships among them, providing a comprehensive description of image content. However, current methods are heavily impacted by a severe long-tailed problem, making it challenging to adequately train fine-grained predicates and resulting in inaccurate content understanding. To address this issue, we propose a template-guided data augmentation (TGDA) strategy that effectively balances data distribution and conducts secondary training for classifiers. Initially, we employ self-driven distillation learning to transfer advanced representation capabilities across all categories, extracting unique templates for each predicate. Furthermore, we apply centroid radiating and gate filtering on these learnable templates to construct reliable instances, thus providing a rich source of supplementary data for fine-grained predicates. We conduct extensive experiments to validate the effectiveness of the proposed method, which demonstrates state-of-the-art performance on the VG dataset.
Yujie Zang, Yaochen Li, Luguang Cao, Ruitao Lu
ICASSP2
2024 Diffusion-Based Generative Self-Supervised Model for Few-Shot PolSAR Image Classification
abstract
With the emergence of deep neural networks, polarimetric synthetic aperture radar (PolSAR) image classification has seen significant advancements in recent years. However, most deep learning-based PolSAR image classification methods heavily rely on a large amount of labeled data which is difficult and expensive to collect for PolSAR. Generative self-supervised learning, which aims to learn representations from unlabeled data via generative auxiliary tasks, is a promising solution to address this issue. Specifically, diffusion models have shown favorable generative capabilities in recent years. Inspired by this, we propose a diffusion-based generative self-supervised model for PolSAR image classification in this work. Via simulating the noise-adding and denoising process, the proposed method is able to learn discriminative and robust polarimetric representations without any annotations, which greatly facilitates the downstream classification. Experiments on the benchmark Flevoland dataset demonstrate the effectiveness of our proposed model.
Zuzheng Kuang, Shirou Jing, Haixia Bi, Yaochen Li
IGARSS5
2024 SCALE-Pose: Skeletal Correction and Language Knowledge-assisted for 3D Human Pose Estimation
Xinnan Ma, Yaochen Li, Limeng Zhao, ChenXu Zhou, Yuncheng Xu
PRCV (11)2
2024 Refine and Redistribute: Multi-Domain Fusion and Dynamic Label Assignment for Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) plays an important role in enhancing visual image comprehension. However, existing approaches often struggle to represent implicit relationship features, resulting in a limited ability to distinguish predicates. Meanwhile, they are vulnerable to skewed instance distributions, which impairs effective training for fine-grained predicates. To address these problems, we propose a novel feature refinement and data redistribution framework (RAR). Specifically, a multi-domain fusion (MDF) module is designed to acquire comprehensive predicate representations, integrating global knowledge from the contextual domain and local details in the spatial-frequency domains. Then, we introduce a dynamic label assignment (DLA) strategy to tackle the long-tailed problem. Different predicate categories are adaptively grouped, accommodating varying training conditions. Guided by this strategy, we leverage a hierarchical auto-encoder to generate siamese samples, expanding the label cardinality. Furthermore, we explore the updated sample space to derive reliable samples and assign tailored labels, ultimately achieving the data rebalancing. Experiments on VG and GQA demonstrate that our model contributes to correcting prediction bias and achieves a significant improvement of approximately 10% in mean recall compared to baseline models.
Yujie Zang, Yaochen Li, Yimou Guo, Wenneng Tang, Yanxue Li, Meklit Atlaw
WACV2
2024 Driver fatigue detection and human-machine cooperative decision-making for road scenarios
Anna Li, Xinnan Ma, Jingyue Zhang, Yaochen Li
Multim. Tools Appl.7
2023 BEV-LaneDet: An Efficient 3D Lane Detection Based on Virtual Camera via Key-Points
abstract
3D lane detection which plays a crucial role in vehicle routing, has recently been a rapidly developing topic in autonomous driving. Previous works struggle with practicality due to their complicated spatial transformations and inflexible representations of 3D lanes. Faced with the issues, our work proposes an efficient and robust monocular 3D lane detection called BEV-LaneDet with three main contributions. First, we introduce the Virtual Camera that unifies the in/extrinsic parameters of cameras mounted on different vehicles to guarantee the consistency of the spatial relationship among cameras. It can effectively promote the learning procedure due to the unified visual space. We secondly propose a simple but efficient 3D lane representation called Key-Points Representation. This module is more suitable to represent the complicated and diverse 3D lane structures. At last, we present a light-weight and chip-friendly spatial transformation module named Spatial Transformation Pyramid to transform multiscale front-view features into BEV features. Experimental results demonstrate that our work outperforms the state-of-the-art approaches in terms of F-Score, being 10.6% higher on the OpenLane dataset and 4.0% higher on the Apollo 3D synthetic dataset, with a speed of 185 FPS. Code is released at https://github.com/gigo-team/bev_lane_det.
Kaiying Li, Yaochen Li, Jintao Xu 0001
CVPR4
2023 Swin-UNIT: Transformer-based GAN for High-resolution Unpaired Image Translation
abstract
The transformer model has gained a lot of success in various computer vision tasks owing to its capacity of modeling long-range dependencies. However, its application has been limited in the area of high-resolution unpaired image translation using GANs due to the quadratic complexity with the spatial resolution of input features. In this paper, we propose a novel transformer-based GAN for high-resolution unpaired image translation named Swin-UNIT. A two-stage generator is designed which consists of a global style translation (GST) module and a recurrent detail supplement (RDS) module. The GST module focuses on translating low-resolution global features using the ability of self-attention. The RDS module offers quick information propagation from the global features to the detail features at a high resolution using cross-attention. Moreover, we customize a dual-branch discriminator to guide the generator. Extensive experiments demonstrate that our model achieves state-of-the-art results on the unpaired image translation tasks.
Yaochen Li, Wenneng Tang, Zhifeng Zhu, Jinhuo Yang, Yuehu Liu
ACM Multimedia2
2023 Efficient Point-Based Single Scale 3D Object Detection from Traffic Scenes
Wenneng Tang, Yaochen Li
PRCV (2)2
2023 Spatiotemporal Analysis of Static and Dynamic Traffic Elements From Road Scenes
abstract
Spatiotemporal analysis of road scenes is a hot research topic in the communities of computer vision and intelligent transportation systems. In this paper, we propose a new framework for spatiotemporal analysis of static and dynamic traffic elements from road scenes. In the first stage, a bottom-up analysis method for static traffic elements is proposed based on a hierarchical spatiotemporal model using hidden conditional random fields (HCRF). The bottom-level features are extracted from sub-regions in the hierarchical model, and the local and global features of the image sequence are then fully combined for spatial and temporal layers. In the second stage, a lightweight multi-stream 3DCNN network is developed for the behavior classification of dynamic traffic elements, which is composed of three parts. Firstly, a SELayer-3DCNN is designed to extract the appearance, motion and edge information from the image sequences. Secondly, the channel attention fusion strategy (CAF) is introduced to enhance the feature fusion ability. Finally, the 3D-RFB module is incorporated to expand the receptive field of the convolution kernel. The experimental results well demonstrate the effectiveness of the proposed framework.
Yaochen Li, Haochuan Hou, Zikun Dong, Yujie Zang, Yonghong Song
IEEE Trans. Intell. Transp. Syst.1
2022 RITNet: A Rotation Invariant Transformer based Network for Point Cloud Registration
abstract
Conventional point cloud registration methods usually employ an encoder-decoder architecture, where mid-level features are locally aggregated to extract geometric information. However, the over-reliance on local features may raise the boundary points cannot be adequately matched for two point clouds. To address this issue, we argue that the boundary features can be further enhanced by the rotation information, and propose a rotation invariant representation to replace common 3D Cartesian coordinates as the network inputs that enhances generalization to arbitrary orientations. Based on this technique, we propose rotation invariant Transformer for point cloud registration, which utilizes insensitivity to arrangement and quantity of data in the Transformer module to capture global structural knowledge within local parts for overall comprehension of each point clouds. Extensive quantitative and qualitative experimental on ModelNet40 evaluations show the effectiveness of the proposed method.
Yaochen Li, Shaohan Yang, Hujun Liu
ICTAI2
2022 Attention-Based Deep Driving Model for Autonomous Vehicles with Surround-View Cameras
abstract
Experienced human drivers always make safe driving decisions by selectively observing the front, rear and side- view mirrors. Several end - to-end methods have been pro-posed to learn driving models with multi-view visual infor-mation. However, these benchmark methods lack semantic understanding of multi-view image contents, where human drivers usually reason these information for decision making with different visual region of interests. In this paper, we propose an attention-based deep learning method to learn a driving model with input of surround-view visual information and the route planner, in which a multi-view attention module is designed for obtaining region of interests from human drivers. We evaluate our model on the Drive360 dataset with comparison of benchmarking deep driving models. Results demonstrate that our model achieves a competitive accuracy in both steering angle and speed prediction than benchmarking methods. Code is available at https://githuh.com/jet-uestc/MVA-Net.
Yang Zhao 0024, Rui Huang 0008, Boqi Li 0001, Ao Luo, Yaochen Li, Hong Cheng 0002
IROS6
2022 Driver Behavior Decision Making Based on Multi-Action Deep Q Network in Dynamic Traffic Scenes
Yaochen Li, Hujun Liu, Anna Li, Yuehu Liu
PRCV (1)2
2022 SiamPolar: Semi-supervised realtime video object segmentation with polar representation
Yaochen Li, Yuhui Hong, Yonghong Song
Neurocomputing1
2022 Caption Generation From Road Images for Traffic Scene Modeling
abstract
In this traffic-scene-modeling study, we propose an image-captioning network which incorporates element attention into an encoder-decoder mechanism to generate more reasonable scene captions. A visual-relationship-detecting network is also developed to detect the relative positions of object pairs. Firstly, the traffic scene elements are detected and segmented according to their clustered locations. Then, the image-captioning network is applied to generate the corresponding description of each traffic scene element. The visual-relationship-detecting network is utilized to detect the position relations of all object pairs in the subregion. The static and dynamic traffic elements are appropriately selected and organized to construct a 3D model according to the captions and the position relations. The reconstructed 3D traffic scenes can be utilized for the offline test of unmanned vehicles. The evaluations and comparisons based on the TSD-max, KITTI and Microsoft’s COCO datasets demonstrate the effectiveness of the proposed framework.
Yaochen Li, Yuehu Liu, Jihua Zhu
IEEE Trans. Intell. Transp. Syst.1
2021 Point Cloud Segmentation via Edge-fused Local Graph Learning
abstract
Traditional convolution for capturing local structures and relationships remains a key technical limit in 3D semantic segmentation, which neglects the certain influence of the adjacent points on the central point in the disordered local point clouds. In this paper, we propose a novel joint-edge graph convolution neural network (JEGCN), which can extract the dynamic features of each local area and transfer the edge information between the vertex pairs to the adjacent vertices. In the proposed graph convolution module, the adjacent vertices are selected with high classification confidence which can guide the central vertex, and then reweight these vertices. Considering the lack of texture features in 3D point clouds, we incorporate 2D image features to adjacent feature propagation to effectively extract the local and global features of point clouds. The experimental results based on ScanNet and S3DIS datasets demonstrate the effectiveness of the proposed method.
Mengtao Han, Yaochen Li, Liangyu Zuo, Chi Zhang 0020, Yuanqi Su
ICRA2
2021 Robust Motion Averaging under Maximum Correntropy Criterion
abstract
Recently, the motion averaging method has been introduced as an effective means to solve the multi-view registration problem. This method aims to recover global motions from a set of relative motions, where the original method is sensitive to outliers due to using the Frobenius norm error in the optimization. Accordingly, this paper proposes a novel robust motion averaging method based on the maximum correntropy criterion (MCC). Specifically, the correntropy measure is used instead of utilizing Frobenius norm error to improve the robustness of motion averaging against outliers. According to the half-quadratic technique, the correntropy measure based optimization problem can be solved by the alternating minimization procedure, which includes operations of weight assignment and weighted motion averaging. Further, we design a selection strategy of adaptive kernel width to take advantage of correntropy. Experimental results on benchmark data sets illustrate that our method has superior performance on accuracy and robustness for multi-view registration. What’s more, it can be applied to robot mapping.
Jihua Zhu, Huimin Lu 0001, Badong Chen, Zhongyu Li 0002, Yaochen Li
ICRA6
2021 Multiview spectral clustering via complementary information
abstract
Abstract In this article, multiview spectral clustering via complementary information (MSCC) is proposed, in which both the consensus information and the complementary information are explored for multiview clustering. In contrast to most multiview spectral clustering methods, the proposed MSCC considers the differences among multiple views and constructs a similarity matrix for clustering. Furthermore, a convex relaxation is employed and an algorithm that is based on the augmented Lagrange multiplier is proposed for optimizing the objective function of MSCC. In extensive experiments on five real‐world benchmark datasets, our proposed method outperforms two baselines and has significantly improved to several state‐of‐the‐art multiview clustering methods.
Shuangxun Ma, Yuehu Liu, Qinghai Zheng, Yaochen Li, Zhichao Cui
Concurr. Comput. Pract. Exp.4
2021 Geometric and semantic analysis of road image sequences for traffic scene construction
Yaochen Li, Yuehu Liu, Yuhui Hong, Jianji Wang 0001
Neurocomputing1
2021 Merging Grid Maps in Diverse Resolutions by the Context-based Descriptor
abstract
Building an accurate map is essential for autonomous robot navigation in the environment without GPS. Compared with single-robot, the multiple-robot system has much better performance in terms of accuracy, efficiency and robustness for the simultaneous localization and mapping (SLAM). As a critical component of multiple-robot SLAM, the problem of map merging still remains a challenge. To this end, this article casts it into point set registration problem and proposes an effective map merging method based on the context-based descriptors and correspondence expansion. It first extracts interest points from grid maps by the Harris corner detector. By exploiting neighborhood information of interest points, it automatically calculates the maximum response radius as scale information to compute the context-based descriptor, which includes eigenvalues and normals computed from local structures of each interest point. Then, it effectively establishes origin matches with low precision by applying the nearest neighbor search on the context-based descriptor. Further, it designs a scale-based corresponding expansion strategy to expand each origin match into a set of feature matches, where one similarity transformation between two grid maps can be estimated by the Random Sample Consensus algorithm. Subsequently, a measure function formulated from the trimmed mean square error is utilized to confirm the best similarity transformation and accomplish the coarse map merging. Finally, it utilizes the scaling trimmed iterative closest point algorithm to refine initial similarity transformation so as to achieve accurate merging. As the proposed method considers scale information in the context-based descriptor, it is able to merge grid maps in diverse resolutions. Experimental results on real robot datasets demonstrate its superior performance over other related methods on accuracy and robustness.
Zhiyang Lin, Jihua Zhu, Zutao Jiang, Yujie Li 0001, Yaochen Li, Zhongyu Li 0002
ACM Trans. Internet Techn.5
2020 Detail Fusion GAN: High-Quality Translation for Unpaired Images with GAN-based Data Augmentation
abstract
Image-to-image translation, a task to learn the mapping relation between two different domains, is a rapid-growing research field in deep learning. Although existing Generative Adversarial Network (GAN)-based methods have achieved decent results in this field, there are still some limitations in generating high-quality images for practical applications (e.g., data augmentation and image inpainting). In this work, we aim to propose a GAN-based network for data augmentation which can generate translated images with more details and less artifacts. The proposed Detail Fusion Generative Adversarial Network (DFGAN) consists of a detail branch, a transfer branch, a filter module, and a reconstruction module. The detail branch is trained by a super-resolution loss and its intermediate features can be used to introduce more details to the transfer branch by the filter module. Extensive evaluations demonstrate that our model generates more satisfactory images against the state-of-the-art approaches for data augmentation.
Yaochen Li, Hang Dong 0001, Peilin Jiang, Fei Wang 0008
ICPR2
2020 Meta Generalized Network for Few-Shot Classification
abstract
Few-shot classification aims to learn a well generalized model with very limited labeled examples. There are mainly two directions for this aim, namely, meta- and metric-learning. Meta learning trains models in a particular way to fast adapt to new tasks, but it neglects variational features of images. Metric learning considers relationships among same or different classes, however on the downside, it usually fails to achieve competitive performance on unseen boundary examples. In this paper, we propose a Meta Generalized Network (MGNet) that aims to combine advantages of both meta- and metric-learning. There are two novel components in MGNet. Specifically, we first develop a meta backbone training method that learns a flexible feature extractor and a classifier initializer efficiently, delightedly leading to fast adaption to unseen few-shot tasks without overfitting. Second, we design a trainable adaptive interval model to improve the cosine classifier, which increases the recognition accuracy of hard examples. We train the meta backbone in the training stage by all classes, and fine-tune the meta-backbone as well as train the adaptive classifier in the testing stage. We evaluate MGNet on three standard image recognition benchmarks, and experimental results validate the superiority over recent competitive methods.
Shanmin Pang, Yaochen Li
ICPR4
2020 Caption Generation from Road Images for Traffic Scene Construction
abstract
In this paper, an image captioning network is proposed for traffic scene modeling, which incorporates element attention into the encoder-decoder mechanism to generate more reasonable scene captions. Firstly, the traffic scene elements are detected and segmented according to their clustered locations. Then, the image captioning network is applied to generate the corresponding caption of each subregion. The static and dynamic traffic elements are appropriately organized to construct a 3D corridor scene model. The semantic relationships between the traffic elements are specified according to the captions. The constructed 3D scene model can be utilized for the offline test of unmanned vehicles. The evaluations and comparisons based on the TSD-max and COCO datasets prove the effectiveness of the proposed framework.
Yaochen Li, Le Wang 0003, Yuehu Liu
IV2
2020 Spatial-Content Image Search in Complex Scenes
abstract
Although the topic of image search has been heavily studied in the last two decades, many works have focused on either instance-level retrieval or semantic-level retrieval. In this work, we develop a novel visually similar spatial-semantic method, namely spatial-content image search, to search images that not only share the same spatial-semantics but also enjoy visual consistency as the query image in complex scenes. We achieve the goal by capturing spatial-semantic concepts as well as the visual representation of each concept contained in an image. Specifically, we first generate a set of bounding boxes and their category labels representing spatial-semantic constraints with YOLOV3, and then obtain visual content of each bounding box with deep features extracted from a convolutional neural network. After that, we customize a similarity computation method that evaluates the relevance between dataset images and input queries according to the developed image representations. Experimental results on two large-scale benchmark retrieval datasets with images consisting of multiple objects demonstrate that our method provides an effective way to query image databases. Our code is available at https://github.com/MaJinWakeUp/spatial-content.
Shanmin Pang, Bo Yang 0041, Jihua Zhu, Yaochen Li
WACV5
2020 Coarse-to-fine 3D road model registration for traffic video augmentation
abstract
This study addresses the problem of non‐perspective pose estimation from line correspondences in the traffic scenarios. A coarse‐to‐fine 3D road registration method is proposed for this problem in two stages. Firstly, the iterative closest point algorithm is exploited to estimate the pose coarsely. An objective function is then established to incorporate the feature correspondences for refining the coarse pose. Besides, the framework including road registration is employed for traffic video augmentation. The framework begins with the inputs of traffic videos, road information from Geographic Information Systems and 3D models of traffic elements (e.g. vehicles, pedestrians). Subsequently, 3D road model generation and point‐to‐line correspondence establishment are achieved in the preprossessing stage. After road and viewpoint registration, the 3D graphic engine is employed to simulate the traffic scene with the road, viewpoints and traffic elements. The augmented videos are generated by fusing the original frames and newly projected traffic elements. The authors demonstrate the superiority of the proposed registration method by the comparison to state‐of‐the‐arts in both quantitative and qualitative experiments. In addition, the frames of the augmented videos validate the proposed method in the application.
Zhichao Cui, Yaochen Li, Chi Zhang 0020, Yuehu Liu, Fuji Ren
IET Image Process.2
2020 Feature concatenation multi-view subspace clustering
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Shanmin Pang, Jun Wang 0024, Yaochen Li
Neurocomputing6
2020 Spatiotemporal road scene reconstruction using superpixel-based Markov random field
Yaochen Li, Yuehu Liu, Jihua Zhu, Shiqi Ma, Zhenning Niu
Inf. Sci.1
2020 Multi-view point cloud registration with adaptive convergence threshold and its application in 3D model retrieval
Yaochen Li, Li Zhu 0003
Multim. Tools Appl.1
2020 Self-Weighting and Hypergraph Regularization for Multi-view Spectral Clustering
abstract
Leveraging the consensus and complementary principle to find a common representation for different views is an essential problem of multi-view clustering. To address the problem, many Low-Rank Representation (LRR) based methods have been proposed. However, existing LRR based methods have two common limitations: 1) they adopt graph regularization that only considers simple pairwise similarities among data points, and 2) they do not generally characterize the importance of each view. In this letter, we correspondingly utilize hypergraph regularization and a self-weighting strategy to handle the limitations with an LRR based model. Specifically, in our model, we construct hypergraph Laplacian matrices of each view that explicitly contain high order relations among data points, to improve the usage of complementary information. Meanwhile, the self-weighting strategy that preserves view specific information and assigns adaptive weights to each view is leveraged to take full advantage of multi-view consensus information. Based on the Augmented Lagrangian Multiplier (ALM) scheme, we design an effective alternating iterative strategy to optimize the model. Extensive experiments conducted on four benchmark datasets validate the superiority of our method.
Wenyu Hao, Shanmin Pang, Jihua Zhu, Yaochen Li
IEEE Signal Process. Lett.4
2019 Jointly Detecting and Retrieving Vehicles from Road Image Sequences based on CNN
abstract
In this paper, a CNN-based vehicle detection and retrieval framework is proposed for the intelligent transportation system. Firstly, the vehicle target is detected from the traffic scene. The proposed object detection method uses a fully convolutional neural network (CNN) based on SqueezeNet, which has the characteristics of real-time, high accuracy and has small model size. Secondly, an intra-class image retrieval method is presented to search vehicles which are similar to the target vehicle in the dataset. The image retrieval results can be used for traffic scenes simulation and modeling. The experiments and comparisons prove the effectiveness of our framework.
Yaochen Li, Yuehu Liu, Shanmin Pang, Le Wang 0003, Huihui Huo
IV2
2019 Road Scene Layout Reconstruction based on CNN and its Application in Traffic Simulation
abstract
In this paper, we propose a road scene prediction framework based on the control points of road boundaries using CNN. Firstly, the image features are extracted and the heatmaps are generated by CNN to locate the control points of road boundaries. The input images are then segmented to specify the scene layout based on the control points. Furthermore, the 3D traffic scene models are constructed. The applications for traffic simulation are then developed. The evaluations and comparisons based on TSD-max dataset prove the effectiveness of the proposed method.
Yaochen Li, Yuehu Liu, Zhichao Cui, Chi Zhang 0020
IV2
2019 Multi-view registration based on weighted LRS matrix decomposition of motions
abstract
Recently, the low‐rank and sparse (LRS) matrix decomposition has been introduced as an effective mean to solve the multi‐view registration. It views each available relative motion as a block element to reconstruct one sparse matrix, which then is used to approximate the low‐rank matrix, where global motions can be recovered for multi‐view registration. However, this approach is sensitive to the sparsity of the reconstructed matrix and it treats all block elements equally in spite of their varied reliabilities. Therefore, this study proposes an effective approach for multi‐view registration by weighted LRS matrix decomposition. On the basis of the inverse symmetry property of relative motions, it first proposes a completion method to reduce the sparsity of the reconstructed matrix. The reduced sparsity of the reconstructed matrix can improve the robustness and efficiency of LRS matrix decomposition. Then, it proposes the weighted LRS matrix decomposition, where each block element is assigned with one estimated weight to denote its reliability. By introducing the weight, more accurate registration results can be efficiently recovered from the estimated low‐rank matrix. Experimental results tested on public datasets illustrate the superiority of the proposed approach over the state‐of‐the‐art approaches on robustness, accuracy and efficiency.
Congcong Jin, Jihua Zhu, Yaochen Li, Shanmin Pang, Lei Chen 0011, Jun Wang 0024
IET Comput. Vis.3
2019 Co-weighting semantic convolutional features for object retrieval
Jihua Zhu, Shanmin Pang, Weili Guan, Zhongyu Li 0002, Yaochen Li, Xueming Qian
J. Vis. Commun. Image Represent.6
2018 Adaptive Co-Weighting Deep Convolutional Features for Object Retrieval
abstract
Aggregating deep convolutional features into a global image vector has attracted sustained attention in image retrieval. In this paper, we propose an efficient unsupervised aggregation method that uses an adaptive Gaussian filter and an element-value sensitive vector to co-weight deep features. Specifically, the Gaussian filter assigns large weights to features of region-of-interests (RoI) by adaptively determining the RoI's center, while the element-value sensitive channel vector suppresses burstiness phenomenon by assigning small weights to feature maps with large sum values of all locations. Experimental results on benchmark datasets validate the proposed two weighting schemes both effectively improve the discrimination power of image vectors. Furthermore, with the same experimental setting, our method outperforms other very recent aggregation approaches by a considerable margin.
Jihua Zhu, Shanmin Pang, Zhongyu Li 0002, Yaochen Li, Xueming Qian
ICME5
2018 Weighted motion averaging for the registration of multi-view range scans
Jihua Zhu, Yaochen Li, Dapeng Chen, Zhongyu Li 0002, Yongqin Zhang
Multim. Tools Appl.3
2016 Fast two-cycle curve evolution with narrow perception of background for object tracking and contour refinement
Yaochen Li, Yuanqi Su, Yuehu Liu
Signal Process. Image Commun.1
2016 Three-Dimensional Traffic Scenes Simulation From Road Image Sequences
abstract
In this paper, we present a novel framework to allow users to tour simulated traffic scenes from the first-person view. Constructing 3-D scenes from road image sequences is in general difficult, due to the intrinsic complexity of dynamic road scenes, which are composed of a drastically moving background, not to mention numerous other surrounding vehicles. With the definitions of the traffic scene models, we first introduce the construction process of the simple traffic scenes. After the detection of road boundaries by a semantic fast two-cycle (FTC) level set method, we generate the control points on road sides to construct the “floor-wall” background scene that is subsequently propagated to each frame. Furthermore, we approach the cluttered traffic scenes through a three-component processing pipeline as follows: 1) traffic elements segmentation; 2) background images inpainting; and 3) traffic scenes construction. The traffic elements in the cluttered images are segmented by the semantic FTC level set method first. A Gaussian mixture model is then employed to inpaint the occluded background utilizing the optical flows. The cluttered traffic scenes can be constructed after the segmentation and inpainting components. The foreground polygons such as vehicles and traffic signs are then modeled. Users can change their viewpoints according to their own interpretations. We present the evaluations of each technical component, followed by our findings from comprehensive user studies, which well demonstrate the effectiveness of the proposed framework in delivering good touring experience to users.
Yaochen Li, Yuehu Liu, Yuanqi Su, Gang Hua 0001, Nanning Zheng 0001
IEEE Trans. Intell. Transp. Syst.1
2015 Fast Two-Cycle level set tracking with narrow perception of background
abstract
The problem of tracking foreground objects in a video sequence with moving background remains challenging. In this paper, we propose the Fast Two-Cycle level set method with Narrow band Background (FTCNB) to automatically extract the foreground objects in such video sequences. The level set curve evolution process consists of two successive cycles: one cycle for data dependent term and a second cycle for smoothness regularization. The curve evolution is implemented by computing the signs of region competition terms on two linked lists of contour pixels rather than solving any Partial Differential Equations (PDEs). Maximum A Posterior (MAP) optimization is applied in the FTCNB method for curve refinement with the assistance of optical flows. The comparison with other level set methods demonstrate the tracking accuracy of our method. The tracking speed of the proposed method also outperforms the traditional level set methods.
Yaochen Li, Yuanqi Su, Yuehu Liu
ICME1
2015 Autonomous Driving Simulation for Unmanned Vehicles
abstract
Human can judge driver's driving ability by observing the vehicle motion in different traffic scenes. Identically, driving behavior can be the main basis for evaluating the performance of an unmanned vehicle in both field test and simulation test. Although simulation test avoids disadvantages of field test, existing simulation systems lack traffic scene data with perception granularity of visual sensors. In order for realizing vehicle-in-loop simulation, simulation technique of driving behaviors must be able to exhibit actual motion of unmanned vehicles. In this paper, we propose an automatic approach of simulating autonomous driving behaviors of vehicles in traffic scene represented by image sequences. Different from general simulation systems, we use actual traffic environment data to build the traffic scene and simulate the driving behaviors. After the proposed method was embedded in scene browser, a typical traffic scene including the intersections was chosen for virtual vehicle to execute the driving tasks of lane change, overtaking, slowing down and stop, right turn and U-Turn. The experimental results show that different driving behaviors of vehicles in typical traffic scene can be exhibited smoothly and realistically. Our method can also be used for generating simulation data of traffic scenes that are difficult to collect.
Danchen Zhao, Yuehu Liu, Chi Zhang 0020, Yaochen Li
WACV4
2011 Fractal image coding using SSIM
abstract
Since Jacquin proposed original fractal image compression technique in 1990, fractal coding method has been developed into various schemes. Traditionally, fractal coding uses mean square error (MSE) to evaluate similarity of image blocks, but the similarity evaluated by MSE usually differs from human visual system (HVS). Compared with MSE, structural similarity (SSIM) is an image measure index which is more appropriate for the HVS. This paper proposes a new fractal coding scheme which uses structural similarity to measure the similarity between image blocks and compute these blocks' coefficients. The experiment results show that the proposed method generates higher quality images for the HVS than MSE scheme.
Jianji Wang 0001, Yuehu Liu, Ping Wei 0001, Yaochen Li, Nanning Zheng 0001
ICIP5
2011 3D facial mesh detection using geometric saliency of surface
abstract
This paper proposes a 3D facial mesh detection algorithm based on the geometric saliency of surface. Specifically, the geometric saliency of each vertex on 3D triangle mesh is measured by the combination of Gaussian-weighted curvature and spin-image correlation. Salient vertices with similar properties are clustered into regions on the saliency map, and represented as nodes by the graph model. To detect a 3D facial mesh, initialization and registration steps are applied to match each triangle in the graph model with a reference graph, corresponding to a 3D reference facial mesh. Furthermore, the match error between the graph model of the testing 3D mesh and the reference facial mesh is computed to classify face and non-face meshes. Experimental results demonstrate that the proposed algorithm is effective to detect 3D facial meshes and robust to facial expressions and geometric noises.
Yaochen Li, Yuehu Liu, Yuanchun Wang 0003, Zhengwang Wu, Yang Yang 0066
ICME1