Michael Ying Yang

dblp:24/7671 · DBLP profile ↗
← Back
49ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-0649-9987ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 1 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-authorSystems, architecture and hardware · 6 · 3 since 2021
YearPublicationVenuePosition
2026 4DSTR: Advancing Generative 4D Gaussians with Spatial-Temporal Rectification for High-Quality and Consistent 4D Generation
abstract
Remarkable advances in recent 2D image and 3D shape generation have induced a significant focus on dynamic 4D content generation. However, previous 4D generation methods commonly struggle to maintain spatial-temporal consistency and adapt poorly to rapid temporal variations, due to the lack of effective spatial-temporal modeling. To address these problems, we propose a novel 4D generation network called 4DSTR, which modulates generative 4D Gaussian Splatting with spatial-temporal rectification. Specifically, temporal correlation across generated 4D sequences is designed to rectify deformable scales and rotations and guarantee temporal consistency. Furthermore, an adaptive spatial densification and pruning strategy is proposed to address significant temporal variations by dynamically adding or deleting Gaussian points with the awareness of their pre-frame movements. Extensive experiments demonstrate that our 4DSTR achieves state-of-the-art performance in video-to-4D generation, excelling in reconstruction quality, spatial-temporal consistency, and adaptation to rapid temporal movements.
Jiuming Liu, Michael Ying Yang, Francesco Nex, Hao Cheng 0008
AAAI5
2026 SPAN: Learning Similarity Between Scene Graphs and Images With Transformers
abstract
Learning similarity between scene graphs and images aims to estimate a similarity score given a scene graph and an image. There is currently no research dedicated to this task, although it is critical for scene graph generation and downstream applications. Scene graph generation is conventionally evaluated by Recall$@K$@K and mean Recall$@K$@K, which measure the ratio of predicted triplets that appear in the human-labeled triplet set. However, such triplet-oriented metrics fail to demonstrate the overall semantic difference between a scene graph and an image and are sensitive to annotation bias and noise. Using generated scene graphs in the downstream applications is therefore limited. To address this issue, for the first time, we propose a Scene graPh-imAge coNtrastive learning framework, SPAN, that can measure the similarity between scene graphs and images. Our novel framework consists of a graph Transformer and an image Transformer to align scene graphs and their corresponding images in the shared latent space. We introduce a novel graph serialization technique that transforms a scene graph into a sequence with structural encodings. Based on our framework, we propose R-Precision measuring image retrieval accuracy as a new evaluation metric for scene graph generation. We establish new benchmarks on the Visual Genome and Open Images datasets. Extensive experiments are conducted to verify the effectiveness of SPAN, which shows great potential as a scene graph encoder.
Yuren Cong, Wentong Liao, Bodo Rosenhahn, Michael Ying Yang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 DVLO4D: Deep Visual-Lidar Odometry with Sparse Spatial-Temporal Fusion
abstract
Visual-LiDAR odometry is a critical component for autonomous system localization, yet achieving high accuracy and strong robustness remains a challenge. Traditional approaches commonly struggle with sensor misalignment, fail to fully leverage temporal information, and require extensive manual tuning to handle diverse sensor configurations. To address these problems, we introduce DVLO4D, a novel visual-LiDAR odometry framework that leverages sparse spatial-temporal fusion to enhance accuracy and robustness. Our approach proposes three key innovations: (1) Sparse Query Fusion, which utilizes sparse LiDAR queries for effective multi-modal data fusion; (2) a Temporal Interaction and Update module that integrates temporally-predicted positions with current frame data, providing better initialization values for pose estimation and enhancing model's robustness against accumulative errors; and (3) a Temporal Clip Training strategy combined with a Collective Average Loss mechanism that aggregates losses across multiple frames, enabling global optimization and reducing the scale drift over long sequences. Extensive experiments on the KITTI and Argoverse Odometry dataset demonstrate the superiority of our proposed DVLO4D, which achieves state-of-the-art performance in terms of both pose accuracy and robustness. Additionally, our method has high efficiency, with an inference time of 82 ms, possessing the potential for the real-time deployment.
Michael Ying Yang, Jiuming Liu, Sander Oude Elberink, George Vosselman, Hao Cheng 0008
ICRA2
2025 Attribute-Centric Compositional Text-to-Image Generation
abstract
Abstract Despite the recent impressive breakthroughs in text-to-image generation, generative models have difficulty in capturing the data distribution of underrepresented attribute compositions while over-memorizing overrepresented attribute compositions, which raises public concerns about their robustness and fairness. To tackle this challenge, we propose ACTIG, an attribute-centric compositional text-to-image generation framework. We present an attribute-centric feature augmentation and a novel image-free training scheme, which greatly improves model’s ability to generate images with underrepresented attributes. We further propose an attribute-centric contrastive loss to avoid overfitting to overrepresented attribute compositions. We validate our framework on the CelebA-HQ and CUB datasets. Extensive experiments show that the compositional generalization of ACTIG is outstanding, and our framework outperforms previous works in terms of image quality and text-image consistency. The source code and trained models are publicly available at https://github.com/yrcong/ACTIG .
Yuren Cong, Martin Renqiang Min, Li Erran Li, Bodo Rosenhahn, Michael Ying Yang
Int. J. Comput. Vis.5
2025 Guest Editorial: Special Issue on Multimodal Learning
Michael Ying Yang, Paolo Rota, Massimiliano Mancini, Pietro Morerio, Bodo Rosenhahn, Vittorio Murino
Int. J. Comput. Vis.1
2024 Robust Shape Fitting for 3D Scene Abstraction
abstract
Humans perceive and construct the world as an arrangement of simple parametric models. In particular, we can often describe man-made environments using volumetric primitives such as cuboids or cylinders. Inferring these primitives is important for attaining high-level, abstract scene descriptions. Previous approaches for primitive-based abstraction estimate shape parameters directly and are only able to reproduce simple objects. In contrast, we propose a robust estimator for primitive fitting, which meaningfully abstracts complex real-world environments using cuboids. A RANSAC estimator guided by a neural network fits these primitives to a depth map. We condition the network on previously detected parts of the scene, parsing it one-by-one. To obtain cuboids from single RGB images, we additionally optimise a depth estimation CNN end-to-end. Naively minimising point-to-primitive distances leads to large or spurious cuboids occluding parts of the scene. We thus propose an improved occlusion-aware distance metric correctly handling opaque scenes. Furthermore, we present a neural network based cuboid solver which provides more parsimonious scene abstractions while also reducing inference time. The proposed algorithm does not require labour-intensive labels, such as cuboid annotations, for training. Results on the NYU Depth v2 dataset demonstrate that the proposed algorithm successfully abstracts cluttered real-world 3D scene layouts.
Florian Kluger, Eric Brachmann, Michael Ying Yang, Bodo Rosenhahn
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 RelTR: Relation Transformer for Scene Graph Generation
abstract
Different objects in the same scene are more or less related to each other, but only a limited number of these relationships are noteworthy. Inspired by Detection Transformer, which excels in object detection, we view scene graph generation as a set prediction problem. In this article, we propose an end-to-end scene graph generation model Relation Transformer (RelTR), which has an encoder-decoder architecture. The encoder reasons about the visual feature context while the decoder infers a fixed-size set of triplets subject-predicate-object using different types of attention mechanisms with coupled subject and object queries. We design a set prediction loss performing the matching between the ground truth and predicted triplets for the end-to-end training. In contrast to most existing scene graph generation methods, RelTR is a one-stage method that predicts sparse scene graphs directly only using visual appearance without combining entities and labeling all possible predicates. Extensive experiments on the Visual Genome, Open Images V6, and VRD datasets demonstrate the superior performance and fast inference of our model.
Yuren Cong, Michael Ying Yang, Bodo Rosenhahn
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Flow-based GAN for 3D Point Cloud Generation from a Single Image
George Vosselman, Michael Ying Yang
BMVC3
2022 Text to Image Generation with Semantic-Spatial Aware GAN
abstract
Text-to-image synthesis (T2I) aims to generate photorealistic images which are semantically consistent with the text descriptions. Existing methods are usually built upon conditional generative adversarial networks (GANs) and initialize an image from noise with sentence embedding, and then refine the features with fine-grained word embedding iteratively. A close inspection of their generated images reveals a major limitation: even though the generated image holistically matches the description, individual image regions or parts of somethings are often not recognizable or consistent with words in the sentence, e.g. “a white crown”. To address this problem, we propose a novel framework Semantic-Spatial Aware GAN for synthesizing images from input text. Concretely, we introduce a simple and effective Semantic-Spatial Aware block, which (1) learns semantic-adaptive transformation conditioned on text to effectively fuse text features and image features, and (2) learns a semantic mask in a weakly-supervised way that depends on the current text-image fusion process in order to guide the transformation spatially. Experiments on the challenging COCO and CUB bird datasets demonstrate the advantage of our method over the recent state-of-the-art approaches, regarding both visual fidelity and alignment with input text description. Code available at https://github.com/wtliao/text2image.
Wentong Liao, Michael Ying Yang, Bodo Rosenhahn
CVPR3
2022 Object Detection in Aerial Images: A Large-Scale Benchmark and Challenges
abstract
In he past decade, object detection has achieved significant progress in natural images but not in aerial images, due to the massive variations in the scale and orientation of objects caused by the bird's-eye view of aerial images. More importantly, the lack of large-scale benchmarks has become a major obstacle to the development of object detection in aerial images (ODAI). In this paper, we present a large-scale Dataset of Object deTection in Aerial images (DOTA) and comprehensive baselines for ODAI. The proposed DOTA dataset contains 1,793,658 object instances of 18 categories of oriented-bounding-box annotations collected from 11,268 aerial images. Based on this large-scale and well-annotated dataset, we build baselines covering 10 state-of-the-art algorithms with over 70 configurations, where the speed and accuracy performances of each model have been evaluated. Furthermore, we provide a code library for ODAI and build a website for evaluating different algorithms. Previous challenges run on DOTA have attracted more than 1300 teams worldwide. We believe that the expanded large-scale DOTA dataset, the extensive baselines, the code library and the challenges can facilitate the designs of robust algorithms and reproducible research on the problem of object detection in aerial images.
Jian Ding 0001, Nan Xue 0001, Gui-Song Xia, Xiang Bai, Wen Yang 0001, Michael Ying Yang, Serge J. Belongie, Jiebo Luo 0001, Mihai Datcu, Marcello Pelillo, Liangpei Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Context-Aware Layout to Image Generation With Enhanced Object Appearance
abstract
A layout to image (L2I) generation model aims to generate a complicated image containing multiple objects (things) against natural background (stuff), conditioned on a given layout. Built upon the recent advances in generative adversarial networks (GANs), existing L2I models have made great progress. However, a close inspection of their generated images reveals two major limitations: (1) the object-to-object as well as object-to-stuff relations are often broken and (2) each object’s appearance is typically distorted lacking the key defining characteristics associated with the object class. We argue that these are caused by the lack of context-aware object and stuff feature encoding in their generators, and location-sensitive appearance representation in their discriminators. To address these limitations, two new modules are proposed in this work. First, a context-aware feature transformation module is introduced in the generator to ensure that the generated feature encoding of either object or stuff is aware of other coexisting objects/stuff in the scene. Second, instead of feeding location-insensitive image features to the discriminator, we use the Gram matrix computed from the feature maps of the generated object images to preserve location-sensitive information, resulting in much enhanced object appearance. Extensive experiments show that the proposed method achieves state-of-the-art performance on the COCO-Thing-Stuff and Visual Genome benchmarks. Code available at: https://github.com/wtliao/layout2img.
Sen He 0001, Wentong Liao, Michael Ying Yang, Yongxin Yang, Yi-Zhe Song, Bodo Rosenhahn, Tao Xiang 0002
CVPR3
2021 Cuboids Revisited: Learning Robust 3D Shape Fitting to Single RGB Images
abstract
Humans perceive and construct the surrounding world as an arrangement of simple parametric models. In particular, man-made environments commonly consist of volumetric primitives such as cuboids or cylinders. Inferring these primitives is an important step to attain high-level, abstract scene descriptions. Previous approaches directly estimate shape parameters from a 2D or 3D input, and are only able to reproduce simple objects, yet unable to accurately parse more complex 3D scenes. In contrast, we propose a robust estimator for primitive fitting, which can meaningfully abstract real-world environments using cuboids. A RANSAC estimator guided by a neural network fits these primitives to 3D features, such as a depth map. We condition the network on previously detected parts of the scene, thus parsing it one-by-one. To obtain 3D features from a single RGB image, we additionally optimise a feature extraction CNN in an end-to-end manner. However, naively minimising point-to-primitive distances leads to large or spurious cuboids occluding parts of the scene behind. We thus propose an occlusion-aware distance metric correctly handling opaque scenes. The proposed algorithm does not require labour-intensive labels, such as cuboid annotations, for training. Results on the challenging NYU Depth v2 dataset demonstrate that the proposed algorithm successfully abstracts cluttered real-world 3D scene layouts.
Florian Kluger, Hanno Ackermann, Eric Brachmann, Michael Ying Yang, Bodo Rosenhahn
CVPR4
2021 Spatial-Temporal Transformer for Dynamic Scene Graph Generation
abstract
Dynamic scene graph generation aims at generating a scene graph of the given video. Compared to the task of scene graph generation from images, it is more challenging because of the dynamic relationships between objects and the temporal dependencies between frames allowing for a richer semantic interpretation. In this paper, we propose Spatial-temporal Transformer (STTran), a neural network that consists of two core modules: (1) a spatial encoder that takes an input frame to extract spatial context and reason about the visual relationships within a frame, and (2) a temporal decoder which takes the output of the spatial encoder as input in order to capture the temporal dependencies between frames and infer the dynamic relationships. Furthermore, STTran is flexible to take varying lengths of videos as input without clipping, which is especially important for long videos. Our method is validated on the benchmark dataset Action Genome (AG). The experimental results demonstrate the superior performance of our method in terms of dynamic scene graphs. Moreover, a set of ablative studies is conducted and the effect of each proposed module is justified. Code available at: https://github.com/yrcong/STTran.
Yuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn, Michael Ying Yang
ICCV5
2021 Disentangled Lifespan Face Synthesis
abstract
A lifespan face synthesis (LFS) model aims to generate a set of photo-realistic face images of a person’s whole life, given only one snapshot as reference. The generated face image given a target age code is expected to be age-sensitive reflected by bio-plausible transformations of shape and texture, while being identity preserving. This is extremely challenging because the shape and texture characteristics of a face undergo separate and highly nonlinear transformations w.r.t. age. Most recent LFS models are based on generative adversarial networks (GANs) whereby age code conditional transformations are applied to a latent face representation. They benefit greatly from the recent advancements of GANs. However, without explicitly disentangling their latent representations into the texture, shape and identity factors, they are fundamentally limited in modeling the nonlinear age-related transformation on texture and shape whilst preserving identity. In this work, a novel LFS model is proposed to disentangle the key face characteristics including shape, texture and identity so that the unique shape and texture age transformations can be modeled effectively. This is achieved by extracting shape, texture and identity features separately from an encoder. Critically, two transformation modules, one conditional convolution based and the other channel attention based, are designed for modeling the nonlinear shape and texture feature transformations respectively. This is to accommodate their rather distinct aging processes and ensure that our synthesized images are both age-sensitive and identity preserving. Extensive experiments show that our LFS model is clearly superior to the state-of-the-art alternatives. Codes and demo are available on our project website: https://senhe.github.io/projects/iccv_2021_lifespan_face.
Sen He 0001, Wentong Liao, Michael Ying Yang, Yi-Zhe Song, Bodo Rosenhahn, Tao Xiang 0002
ICCV3
2021 Exploring Dynamic Context for Multi-path Trajectory Prediction
abstract
To accurately predict future positions of different agents in traffic scenarios is crucial for safely deploying intelligent autonomous systems in the real-world environment. However, it remains a challenge due to the behavior of a target agent being affected by other agents dynamically and there being more than one socially possible paths the agent could take. In this paper, we propose a novel framework, named Dynamic Context Encoder Network (DCENet). In our framework, first, the spatial context between agents is explored by using self-attention architectures. Then, the two-stream encoders are trained to learn temporal context between steps by taking the respective observed trajectories and the extracted dynamic spatial context as input. The spatial-temporal context is encoded into a latent space using a Conditional Variational Auto-Encoder (CVAE) module. Finally, a set of future trajectories for each agent is predicted conditioned on the learned spatial-temporal context by sampling from the latent space, repeatedly. DCENet is evaluated on one of the most popular challenging benchmarks for trajectory forecasting Trajnet and reports a new state-of-the-art performance. It also demonstrates superior performance evaluated on the benchmark inD for mixed traffic at intersections. A series of ablation studies is conducted to validate the effectiveness of each proposed module. Our code is available at https://github.com/wtliao/DCENet.
Hao Cheng 0008, Wentong Liao, Xuejiao Tang, Michael Ying Yang, Monika Sester, Bodo Rosenhahn
ICRA4
2021 CABiNet: Efficient Context Aggregation Network for Low-Latency Semantic Segmentation
abstract
With the increasing demand of autonomous machines, pixel-wise semantic segmentation for visual scene understanding needs to be not only accurate but also efficient for any potential real-time applications. In this paper, we propose CABiNet (Context Aggregated Bi-lateral Network), a dual branch convolutional neural network (CNN), with significantly lower computational costs as compared to the state-of-the-art, while maintaining a competitive prediction accuracy. Building upon the existing multi-branch architectures for high-speed semantic segmentation, we design a cheap high resolution branch for effective spatial detailing and a context branch with light-weight versions of global aggregation and local distribution blocks, potent to capture both long-range and local contextual dependencies required for accurate semantic segmentation, with low computational overheads. Specifically, we achieve 76.6% and 75.9% mIOU on Cityscapes validation and test sets respectively, at 76 FPS on an NVIDIA RTX 2080Ti and 8 FPS on a Jetson Xavier NX.
Saumya Kumaar, Ye Lyu, Francesco Nex, Michael Ying Yang
ICRA4
2020 Image Captioning Through Image Transformer
Sen He 0001, Wentong Liao, Hamed Rezazadegan Tavakoli, Michael Ying Yang, Bodo Rosenhahn, Nicolas Pugeault
ACCV (4)4
2020 CONSAC: Robust Multi-Model Fitting by Conditional Sample Consensus
abstract
We present a robust estimator for fitting multiple parametric models of the same form to noisy measurements. Applications include finding multiple vanishing points in man-made scenes, fitting planes to architectural imagery, or estimating multiple rigid motions within the same sequence. In contrast to previous works, which resorted to hand-crafted search strategies for multiple model detection, we learn the search strategy from data. A neural network conditioned on previously detected models guides a RANSAC estimator to different subsets of all measurements, thereby finding model instances one after another. We train our method supervised, as well as, self-supervised. For supervised training of the search strategy, we contribute a new dataset for vanishing point estimation. Leveraging this dataset, the proposed algorithm is superior with respect to other robust estimators, as well as, to designated vanishing point estimation algorithms. For self-supervised learning of the search, we evaluate the proposed algorithm on multi-homography estimation and demonstrate an accuracy that is superior to state-of-the-art methods.
Florian Kluger, Eric Brachmann, Hanno Ackermann, Carsten Rother, Michael Ying Yang, Bodo Rosenhahn
CVPR5
2020 FairNN - Conjoint Learning of Fair Representations for Fair Decisions
Tongxin Hu, Vasileios Iosifidis, Wentong Liao, Michael Ying Yang, Eirini Ntoutsi, Bodo Rosenhahn
DS5
2020 NODIS: Neural Ordinary Differential Scene Understanding
Yuren Cong, Hanno Ackermann, Wentong Liao, Michael Ying Yang, Bodo Rosenhahn
ECCV (20)4
2020 Temporally Consistent Horizon Lines
abstract
The horizon line is an important geometric feature for many image processing and scene understanding tasks in computer vision. For instance, in navigation of autonomous vehicles or driver assistance, it can be used to improve 3D reconstruction as well as for semantic interpretation of dynamic environments. While both algorithms and datasets exist for single images, the problem of horizon line estimation from video sequences has not gained attention. In this paper, we show how convolutional neural networks are able to utilise the temporal consistency imposed by video sequences in order to increase the accuracy and reduce the variance of horizon line estimates. A novel CNN architecture with an improved residual convolutional LSTM is presented for temporally consistent horizon line estimation. We propose an adaptive loss function that ensures stable training as well as accurate results. Furthermore, we introduce an extension of the KITTI dataset which contains precise horizon line labels for 43699 images across 72 video sequences. A comprehensive evaluation shows that the proposed approach consistently achieves superior performance compared with existing methods.
Florian Kluger, Hanno Ackermann, Michael Ying Yang, Bodo Rosenhahn
ICRA3
2019 Deep Learning for Semantic Segmentation of UAV Videos
abstract
As one of the key problems in both remote sensing and computer vision, video semantic segmentation has been attracting increasing amounts of attention. Using video segmentation technique for Unmanned Aerial Vehicle (UAV) data processing is also a popular application. Previous methods extended single image segmentation approaches to multiple frames. The temporal dependencies are ignored in these methods. This paper proposes a novel segmentation method to solve this problem. Combining the fully convolutional networks (FCN) and the Convolution Long Short Term Memory (Conv-LSTM) together, we segment the sequence of the video frames instead of segmenting each individual frame separately. FCN serves as the frame-based segmentation method. Conv-LSTM makes use of the temporal information between consecutive frames. Experimental results show the superiority of this method especially in some classes compared to the single image segmentation model using video dataset from UAV.
Ye Lyu, Yanpeng Cao, Michael Ying Yang
IGARSS4
2019 Fusing Airborne Laser Scanning and Rapideye Sensor Parameters for Tropical Forest Biomass Estimation of Nepal
abstract
This study aimed to integrate and optimise multi-sensor data consisting of airborne laser scanning (ALS) and RapidEye image-derived parameters with field-measured biomass using Random forests (RF) to predict the tropical forest biomass (FB). The multi-sensor parameters which include combined spectral, textural and ALS derived 119 variables provided a better result, R2= 0.95 and relative RMSE (relRMSE) = 17.25%, as compared to any single sensor parameter for the FB estimation. The top 20 variables were used to compute FB and its distribution map with the R2, RMSE and relRMSE values of 0.93, 35.46 Mg ha-1and 17.40% respectively. Furthermore, its validation was checked with the in-situ biomass to obtain satisfactory accuracy (R2= 0.72, RMSE = 47.71 Mg ha-1and relRMSE = 23.41%). The result showed that the integration of multi-sensor data using RF regression algorithm proves to be a reliable algorithm for accurate FB estimation.
Kashi Ram Yadav, Subrata Nandy, Ritika Srinet, Raja Ram Aryal, Michael Ying Yang
IGARSS5
2019 Accurate salient object detection via dense recurrent connections and residual-based hierarchical feature integration
Yanpeng Cao, Guizhong Fu, Jiangxin Yang, Yanlong Cao, Michael Ying Yang
Signal Process. Image Commun.5
2019 Cascaded Deep Networks With Multiple Receptive Fields for Infrared Image Super-Resolution
abstract
Infrared images have a wide range of military and civilian applications, including night vision, surveillance, and robotics. However, high-resolution infrared detectors are difficult to fabricate and their manufacturing cost is expensive. In this paper, we present a cascaded architecture of deep neural networks with multiple receptive fields to increase the spatial resolution of infrared images by a large scale factor (x8). Instead of reconstructing a high-resolution image from its low-resolution version using a single complex deep network, the key idea of our approach is to set up a mid-point (scale x2) between scale x1 and x8 such that lost information can be divided into two components. Lost information within each component contains similar patterns thus can be more accurately recovered even using a simpler deep network. In our proposed cascaded architecture, two consecutive deep networks with different receptive fields are jointly trained through a multi-scale loss function. The first network with a large receptive field is applied to recover largescale structure information, while the second one uses a relatively smaller receptive field to reconstruct small-scale image details. Our proposed method is systematically evaluated using realistic infrared images. Compared with state-of-the-art super-resolution methods, our proposed cascaded approach achieves improved reconstruction accuracy using significantly fewer parameters.
Zewei He, Siliang Tang, Jiangxin Yang, Yanlong Cao, Michael Ying Yang, Yanpeng Cao
IEEE Trans. Circuits Syst. Video Technol.5
2018 Deep Learning for Vehicle Detection in Aerial Images
abstract
The detection of vehicles in aerial images is widely applied in many domains. In this paper, we propose a novel double focal loss convolutional neural network framework (DFL-CNN). In the proposed framework, the skip connection is used in the CNN structure to enhance the feature learning. Also, the focal loss function is used to substitute for conventional cross entropy loss function in both of the region proposed network and the final classifier. We further introduce the first large-scale vehicle detection dataset ITCVD with ground truth annotations for all the vehicles in the scene. The experimental results show that our DFL-CNN outperforms the baselines on vehicle detection.
Michael Ying Yang, Wentong Liao, Xinbo Li, Bodo Rosenhahn
ICIP1
2018 Object Recognition from very few Training Examples for Enhancing Bicycle Maps
abstract
In recent years, data-driven methods have shown great success for extracting information about the infrastructure in urban areas. These algorithms are usually trained on large datasets consisting of thousands or millions of labeled training examples. While large datasets have been published regarding cars, for cyclists very few labeled data is available although appearance, point of view, and positioning of even relevant objects differ. Unfortunately, labeling data is costly and requires a huge amount of work. In this paper, we thus address the problem of learning with very few labels. The aim is to recognize particular traffic signs in crowdsourced data to collect information which is of interest to cyclists. We propose a system for object recognition that is trained with only 15 examples per class on average. To achieve this, we combine the advantages of convolutional neural networks and random forests to learn a patch-wise classifier. In the next step, we map the random forest to a neural network and transform the classifier to a fully convolutional network. Thereby, the processing of full images is significantly accelerated and bounding boxes can be predicted. Finally, we integrate data of the Global Positioning System (GPS) to localize the predictions on the map. In comparison to Faster R-CNN and other networks for object recognition or algorithms for transfer learning, we considerably reduce the required amount of labeled data. We demonstrate good performance on the recognition of traffic signs for cyclists as well as their localization in maps.
Christoph Reinders, Hanno Ackermann, Michael Ying Yang, Bodo Rosenhahn
Intelligent Vehicles Symposium3
2017 Analyzing modular CNN architectures for joint depth prediction and semantic segmentation
abstract
This paper addresses the task of designing a modular neural network architecture that jointly solves different tasks. As an example we use the tasks of depth estimation and semantic segmentation given a single RGB image. The main focus of this work is to analyze the cross-modality influence between depth and semantic prediction maps on their joint refinement. While most of the previous works solely focus on measuring improvements in accuracy, we propose a way to quantify the cross-modality influence. We show that there is a relationship between final accuracy and cross-modality influence, although not a simple linear one. Hence a larger cross-modality influence does not necessarily translate into an improved accuracy. We find that a beneficial balance between the cross-modality influences can be achieved by network architecture and conjecture that this relationship can be utilized to understand different network design choices. Towards this end we propose a Convolutional Neural Network (CNN) architecture that fuses the state-of-the-art results for depth estimation and semantic labeling. By balancing the cross-modality influences between depth and semantic prediction, we achieve improved results for both tasks using the NYU-Depth v2 benchmark.
Omid Hosseini Jafari, Oliver Groth, Alexander Kirillov, Michael Ying Yang, Carsten Rother
ICRA4
2016 Mapping Auto-context Decision Forests to Deep ConvNets for Semantic Segmentation
David L. Richmond, Dagmar Kainmüller, Michael Ying Yang, Eugene W. Myers, Carsten Rother
BMVC3
2016 Uncertainty-Driven 6D Pose Estimation of Objects and Scenes from a Single RGB Image
abstract
In recent years, the task of estimating the 6D pose of object instances and complete scenes, i.e. camera localization, from a single input image has received considerable attention. Consumer RGB-D cameras have made this feasible, even for difficult, texture-less objects and scenes. In this work, we show that a single RGB image is sufficient to achieve visually convincing results. Our key concept is to model and exploit the uncertainty of the system at all stages of the processing pipeline. The uncertainty comes in the form of continuous distributions over 3D object coordinates and discrete distributions over object labels. We give three technical contributions. Firstly, we develop a regularized, auto-context regression framework which iteratively reduces uncertainty in object coordinate and object label predictions. Secondly, we introduce an efficient way to marginalize object coordinate distributions over depth. This is necessary to deal with missing depth information. Thirdly, we utilize the distributions over object labels to detect multiple objects simultaneously with a fixed budget of RANSAC hypotheses. We tested our system for object pose estimation and camera localization on commonly used data sets. We see a major improvement over competing systems.
Eric Brachmann, Frank Michel 0002, Alexander Krull, Michael Ying Yang, Stefan Gumhold, Carsten Rother
CVPR4
2016 Real-time RGB-D based template matching pedestrian detection
abstract
Pedestrian detection is one of the most popular topics in computer vision and robotics. Considering challenging issues in multiple pedestrian detection, we present a real-time depth-based template matching people detector. In this paper, we propose different approaches for training the depth-based template. We train multiple templates for handling issues due to various upper-body orientations of the pedestrians and different levels of detail in depth-map of the pedestrians with various distances from the camera. And, we take into account the degree of reliability for different regions of sliding window by proposing the weighted template approach. Furthermore, we combine the depth-detector with an appearance based detector as a verifier to take advantage of the appearance cues for dealing with the limitations of depth data. We evaluate our method on the challenging ETH dataset sequence. We show that our method outperforms the state-of-the-art approaches.
Omid Hosseini Jafari, Michael Ying Yang
ICRA2
2016 Node-Grained Incremental Community Detection for Streaming Networks
abstract
Community detection has been one of the key research topics in the analysis of networked data, which is a powerful tool for understanding organizational structures of complex networks. One major challenge in community detection is to analyze community structures for streaming networks in real-time in which changes arrive sequentially and frequently. The existing incremental algorithms are often designed for edge-grained sequential changes, which are sensitive to the processing sequence of edges. However, there exist many real-world net-works that changes occur on node-grained, i.e., node with its connecting edges is added into network simultaneously and all edges arrive at the same time. In this paper, we propose a novel incremental community detection method based on modularity optimization for node-grained streaming networks. This method takes one vertex and its connecting edges as a processing unit, and equally treats edges involved by same node. Our algorithm is evaluated on a set of real-world networks, and is compared with several representative incremental and non-incremental algorithms. The experimental results show that our method is highly effective for discovering communities in an incremental way. In addition, our algorithm even got better results than Louvain method (the famous modularity optimization algorithm using global information) in some test networks, e.g., citation networks, which are more likely to be node-grained. This may further indicate the significance of the node-grained incremental algorithms.
Siwen Yin, Shizhan Chen, Zhiyong Feng 0002, Keman Huang, Dongxiao He, Michael Ying Yang
ICTAI7
2016 Bi-layer dictionary learning for remote sensing image classification
abstract
With the widely application of high-resolution remote sensing images, its classification has attracted a lot of attention. Usually, some different categories share common patterns, which make these categories look similar. This makes the classification of such categories a challenging task. In this paper, we propose a novel dictionary learning based bilayer classification algorithm to solve this problem. Using SIFT descriptor, instead of directly classifying an image, we separate the classification in two steps. In the first step, the similar categories are clustered to be as a new category for the first classification layer. In this step, the inter-class variation are maximized. The second layer is designed to classify the similar categories clustered in the same group. Experimental results show the superiority of our method compared to the state-of-the-art methods using UCMerced LandUse dataset.
Michael Ying Yang, Saif Dawood Salman Al-Shaikhli, Yanpeng Cao, Bodo Rosenhahn
IGARSS1
2016 Effective Strip Noise Removal for Low-Textured Infrared Images Based on 1-D Guided Filtering
abstract
Infrared images typically contain obvious strip noise. It is a challenging task to eliminate such noise without blurring fine image details in low-textured infrared images. In this paper, we introduce an effective single-image-based algorithm to accurately remove strip-type noise present in infrared images without causing blurring effects. First, a 1-D row guided filter is applied to perform edge-preserving image smoothing in the horizontal direction. The extracted high-frequency image part contains both strip noise and a significant amount of image details. Through a thermal calibration experiment, we discover that a local linear relationship exists between infrared data and strip noise of pixels within a column. Based on the derived strip noise behavioral model, strip noise components are accurately decomposed from the extracted high-frequency signals by applying a 1-D column guided filter. Finally, the estimated noise terms are subtracted from the raw infrared images to remove strips without blurring image details. The performance of the proposed technique is thoroughly investigated and is compared with the state-of-the-art 1-D and 2-D denoising algorithms using captured infrared images.
Yanpeng Cao, Michael Ying Yang, Christel-Loïc Tisse
IEEE Trans. Circuits Syst. Video Technol.2
2015 Pose Estimation of Kinematic Chain Instances via Object Coordinate Regression
abstract
Accurate pose estimation of object instances is a key aspect in many applications, including augmented reality or robotics. For example, a task of a domestic robot could be to fetch an item from an open drawer. The poses of both, the drawer and the item have to be known by the robot in order to fulfil the task. 6D pose estimation of rigid objects has been addressed with great success in recent years. In large part, this has been due to the advent of consumer-level RGB-D cameras, which provide rich, robust input data. However, the practical use of state-of-the-art pose estimation approaches is limited by the assumption that objects are rigid. In cluttered, domestic environments this assumption does often not hold. Examples are doors, many types of furniture, certain electronic devices and toys. A robot might encounter these items in any state of articulation. This work considers the task of one-shot pose estimation of articulated object instances from an RGB-D image. In particular, we address objects with the topology of a kinematic chain of any length, i.e. objects are composed of a chain of parts interconnected by joints. We restrict joints to either revolute joints with 1 DOF (degrees of freedom) rotational movement or prismatic joints with 1 DOF translational movement. This topology covers a wide range of common objects (see our dataset for examples). However, our approach can easily be expanded to any topology, and to joints with higher degrees of freedom.
Frank Michel 0002, Alexander Krull, Eric Brachmann, Michael Ying Yang, Stefan Gumhold, Carsten Rother
BMVC4
2015 Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images
abstract
Analysis-by-synthesis has been a successful approach for many tasks in computer vision, such as 6D pose estimation of an object in an RGB-D image which is the topic of this work. The idea is to compare the observation with the output of a forward process, such as a rendered image of the object of interest in a particular pose. Due to occlusion or complicated sensor noise, it can be difficult to perform this comparison in a meaningful way. We propose an approach that "learns to compare", while taking these difficulties into account. This is done by describing the posterior density of a particular object pose with a convolutional neural network (CNN) that compares observed and rendered images. The network is trained with the maximum likelihood paradigm. We observe empirically that the CNN does not specialize to the geometry or appearance of specific objects. It can be used with objects of vastly different shapes and appearances, and in different backgrounds. Compared to state-of-the-art, we demonstrate a significant improvement on two different datasets which include a total of eleven objects, cluttered background, and heavy occlusion.
Alexander Krull, Eric Brachmann, Frank Michel 0002, Michael Ying Yang, Stefan Gumhold, Carsten Rother
ICCV4
2015 A novel dictionary learning method for remote sensing image classification
abstract
With the widely application of high-resolution remote sensing images, its classification has attracted a lot of attention. Most classification methods focus on various combination of features and ignore the similarities between different categories. In this paper we present a modification by combining ScSPM [1] with a dictionary learning method DL-COPAR [2], which separates the particularity and commonality atoms of class-specific sub-dictionaries. With this over-complete dictionary, the sparse representation of a query image can be specified to capture salient and unique properties. Experimental results on two remote sensing datasets show that, this modification achieves state-of-the-art classification accuracy, when merely SIFT feature is applied.
Michael Ying Yang, Saif Dawood Salman Al-Shaikhli, Bodo Rosenhahn
IGARSS1
2015 Hyperspectral image classification using Gaussian process models
abstract
Hyperspectral image processing has been a very dynamic area in remote sensing and other applications since last decades. Hyperspectral images provide abundant spectral information to identify and distinguish spectrally similar materials. Recent advances in kernel machines promote the novel use of Gaussian processes (GP) for classifying hyper-spectral images. Many sophisticated kernel functions have been provided for kernel-based methods. However, different kernel functions has different performance in different applications. This paper introduces GP models with different kernel functions for classifying hyperspectral images. We first provided the mathematical formulation of GP models for classification. Then, several popular kernel functions and their hyperparaeters selection for GP models are introduced. The experiment are performed on three benchmark datasets to evaluate the performances of different kernel functions in terms of classification accuracy. Their performances are compared with each other and discussed in detailed.
Michael Ying Yang, Wentong Liao, Bodo Rosenhahn
IGARSS1
2015 A Global-to-Local Framework for Infrared and Visible Image Sequence Registration
abstract
Based on the development of image registration, sequence registration can be done by computing the transformations between consecutive frames. To take into account the accumulated error, global registration method is usually employed as a global error minimizing approach. However, in real surveillance applications, the visible sequence and infrared sequence may be taken at different times, or from different viewpoints, and may have different dynamic contents. Therefore, global registration is only an approximate estimation for two sequences, resulting in inferior local contents. In this paper we present a novel integrated global-to-local framework that addresses the problems of dynamic infrared and visible image sequence registration. We propose to maximize the sum of the mutual information of two sequences for the global homography estimation. Then, frame-to-frame registration is performed to estimate the per-frame local homography. Finally, a smoothing strategy is adopted to smooth the local homographies in the temporal domain to enforce temporal consistency. We evaluate our proposed framework by comparing it to the state-of-the art sequence registration algorithm. Our method achieves improved performance on the public benchmark dataset.
Michael Ying Yang, Yu Qiang, Bodo Rosenhahn
WACV1
2015 Descriptor evaluation and feature regression for multimodal image analysis
Xuanzi Yong, Michael Ying Yang, Yanpeng Cao, Bodo Rosenhahn
Mach. Vis. Appl.2
2015 Joint Object Segmentation and Depth Upsampling
abstract
With the advent of powerful ranging and visual sensors, nowadays, it is convenient to collect sparse 3-D point clouds and aligned high-resolution images. Benefitted from such convenience, this letter proposes a joint method to perform both depth assisted object-level image segmentation and image guided depth upsampling. To this end, we formulate these two tasks together as a bi-task labeling problem, defined in a Markov random field. An alternating direction method (ADM) is adopted for the joint inference, solving each sub-problem alternatively. More specifically, the sub-problem of image segmentation is solved by Graph Cuts, which attains discrete object labels efficiently. Depth upsampling is addressed via solving a linear system that recovers continuous depth values. By this joint scheme, robust object segmentation results and high-quality dense depth maps are achieved. The proposed method is applied to the challenging KITTI vision benchmark suite, as well as the Leuven dataset for validation. Comparative experiments show that our method outperforms stand-alone approaches.
Wenqi Huang 0002, Xiaojin Gong, Michael Ying Yang
IEEE Signal Process. Lett.3
2014 Brain tumor classification using sparse coding and dictionary learning
abstract
Brain tumor classification is considered as one of the most challenging tasks in medical imaging. In this paper, a novel approach for multi-class brain tumor classification based on sparse coding and dictionary learning is proposed. We propose an individual (per-class) dictionary learning and sparse coding classification using K-SVD algorithm. This approach combines topological and texture features to build and learn a dictionary. Experimental results demonstrate that the sparse coding based classification outperforms other state-of-the-art methods.
Saif Dawood Salman Al-Shaikhli, Michael Ying Yang, Bodo Rosenhahn
ICIP2
2014 Simultaneous remote sensing image classification and annotation based on the spatial coherent topic model
abstract
The traditional LDA models to solve the problem of scene classification lack the spatial relationship between the fragments of images or the parts of targets and linkages between the global and local information, so their performance is usually poor in stability for the images with clutter background. In this paper, a novel method for the simultaneous classification and annotation of remote sensing images with complex scenes is proposed. The Spatially Consistent Topic Model is defined by making full use of the correlation between image classification and annotation. We choose SIFT features, hue features and texture features as the visual words, which help to endow pixels of similar appearance region with the same hidden topic. Competitive results on remote sensing images demonstrate the precision and robustness of the proposed method.
Michael Ying Yang, Mei Zhou, Xiang-zhao Zeng
IGARSS2
2014 Improved trihedral corner reflector for high-precision SAR calibration and validation
abstract
Trihedral corner reflectors have been widely used as standard point targets for synthetic aperture radar (SAR) calibration. The RCS accuracy of the trihedral corner reflector is vital especially for future high-precision SAR radiometric calibration. In order to reduce the interaction between corner reflector and the ground on which the reflector deployed, and also reduce the edge diffraction through minimizing the panel external edge length, improved trihedral corner reflector was developed and it had smaller edge length than the trihedral corner reflector with the polygonous shape for a given panel area. General expression for the panel area and panel external edge length of arbitrarily self-illuminating corner reflectors was presented by parameter equation firstly. Then the edge of reflector panel was assumed as circular arc and through a numerical approximation approach, improved corner reflector of circular arc panel geometry with smaller edge length was obtained.
Yongsheng Zhou, Chuanrong Li, Lingling Ma 0001, Michael Ying Yang
IGARSS4
2014 Video segmentation with joint object and trajectory labeling
abstract
Unsupervised video object segmentation is a challenging problem because it involves a large amount of data and object appearance may significantly change over time. In this paper, we propose a bottom-up approach for the combination of object segmentation and motion segmentation using a novel graphical model, which is formulated as inference in a conditional random field (CRF) model. This model combines object labeling and trajectory clustering in a unified probabilistic framework. The CRF contains binary variables representing the class labels of image pixels as well as binary variables indicating the correctness of trajectory clustering, which integrates dense local interaction and sparse global constraint. An optimization scheme based on a coordinate ascent style procedure is proposed to solve the inference problem. We evaluate our proposed framework by comparing it to other video and motion segmentation algorithms. Our method achieves improved performance on state-of-the-art benchmark datasets.
Michael Ying Yang, Bodo Rosenhahn
WACV1
2014 Estimating layout of cluttered indoor scenes using trajectory-based priors
Muhammad Shoaib 0007, Michael Ying Yang, Bodo Rosenhahn, Jörn Ostermann
Image Vis. Comput.2
2013 Slice Sampling Particle Belief Propagation
abstract
Inference in continuous label Markov random fields is a challenging task. We use particle belief propagation (PBP) for solving the inference problem in continuous label space. Sampling particles from the belief distribution is typically done by using Metropolis-Hastings (MH) Markov chain Monte Carlo (MCMC) methods which involves sampling from a proposal distribution. This proposal distribution has to be carefully designed depending on the particular model and input data to achieve fast convergence. We propose to avoid dependence on a proposal distribution by introducing a slice sampling based PBP algorithm. The proposed approach shows superior convergence performance on an image denoising toy example. Our findings are validated on a challenging relational 2D feature tracking application.
Oliver Müller 0003, Michael Ying Yang, Bodo Rosenhahn
ICCV2
2011 Robust alignment of wide baseline terrestrial laser scans via 3D viewpoint normalization
abstract
The complexity of natural scenes and the amount of information acquired by terrestrial laser scanners turn the registration among scans into a complex problem. This problem becomes even more challenging when two individual scans captured at significantly changed viewpoints (wide baseline). Since laser-scanning instruments nowadays are often equipped with an additional image sensor, it stands to reason making use of the image content to improve the registration process of 3D scanning data. In this paper, we present a novel improvement to the existing feature techniques to enable automatic alignment between two widely separated 3D scans. The key idea consists of extracting dominant planar structures from 3D point clouds and then utilizing the recovered 3D geometry to improve the performance of 2D image feature extraction and matching. The resulting features are very discriminative and robust to perspective distortions and viewpoint changes due to exploiting the underlying 3D structure. Using this novel viewpoint invariant feature, the corresponding 3D points are automatically linked in terms of wide baseline image matching. Initial experiments with real data demonstrate the potential of the proposed method for the challenging wide baseline 3D scanning data alignment tasks.
Yanpeng Cao, Michael Ying Yang, John McDonald 0001
WACV2
2009 Multiregion level-set segmentation of synthetic aperture radar images
abstract
Due to the presence of speckle, segmentation of SAR images is generally acknowledged as a difficult problem. A large effort has been done in order to cope with the influence of speckle noise on image segmentation such as edge detection or direct global segmentation. Recent works address this problem by using statistical image representation and deformable models. We suggest a novel variational approach to SAR image segmentation, which consists of minimizing a functional containing an original observation term derived from maximum a posteriori (MAP) estimation framework and a Gamma image representation. The minimization is carried out efficiently by a new multiregion method which embeds a simple partition assumption directly in curve evolution to guarantee a partition of the image domain from an arbitrary initial partition. Experiments on both synthetic and real images show the effectiveness of the proposed method.
Michael Ying Yang
ICIP1