VLDB 2026 Research / reviewers in the wild / expert
Bin Yang 0009
dblp:77/377-9
· DBLP profile ↗
107ranked-venue papers
15as first author
46since 2021 · last 2026
0000-0002-8322-117XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 14 first-author · 15 since 2021Artificial intelligence and machine learning · 33 · 28 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Systems, architecture and hardware · 8 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DOS: Distilling Observable Softmaps of Zipfian Prototypes for Self-Supervised Point RepresentationabstractRecent advances in self-supervised learning (SSL) have shown tremendous potential for learning 3D point cloud representations without human annotations. However, SSL for 3D point clouds still faces critical challenges due to irregular geometry, shortcut-prone reconstruction, and unbalanced semantics distribution. In this work, we propose DOS (Distilling Observable Softmaps), a novel SSL framework that self-distills semantic relevance softmaps only at observable (unmasked) points. This strategy prevents information leakage from masked regions and provides richer supervision than discrete token-to-prototype assignments. To address the challenge of unbalanced semantics in an unsupervised setting, we introduce Zipfian prototypes and incorporate them using a modified Sinkhorn-Knopp algorithm, Zipf-Sinkhorn, which enforces a power-law prior over prototype usage and modulates the sharpness of the target softmap during training. DOS outperforms current state-of-the-art methods on semantic segmentation and 3D object detection across multiple benchmarks, including nuScenes, Waymo, SemanticKITTI, ScanNet, and ScanNet200, without relying on extra data or annotations. Our results demonstrate that observable-point softmaps distillation offers a scalable and effective paradigm for learning robust 3D representations. Mohamed Abdelsamad, Michael Ulrich, Bin Yang 0009, Miao Zhang 0043, Yakov Miron, Abhinav Valada |
AAAI | 3 |
| 2026 | An interpretable data-driven framework for variable group selection in high-dimensional manufacturing test data
Bin Yang 0009, Yiwen Liao Data |
ETS | 2 |
| 2026 | GroupEnsemble: Efficient Uncertainty Estimation for DETR-based Object Detection
Yutong Yang, Katarina Popovic, Julian Wiederer, Markus Braun 0003, Vasileios Belagiannis, Bin Yang 0009 |
IV | 6 |
| 2026 | SVS-GAN for Semantic Synthesis of Traffic Videos for Autonomous DrivingabstractAutonomous driving demands robust perception modules trained on diverse scenarios, yet collecting and annotating real-world datasets is both expensive and often lacks sufficient coverage of all possible driving conditions. Semantic Image Synthesis (SIS)—the process of generating realistic images from semantic label maps—has proven effective for producing large-scale labeled data. However, extending SIS to the video domain as Semantic Video Synthesis (SVS), where entire sequences are generated from semantic maps, remains underexplored. We introduce SVS-GAN, a framework specifically tailored for SVS that generates high-quality, temporally coherent videos at a resolution of 1024×512 in real-time (45 FPS). Our approach leverages a deformable motion triple-pyramid generator and a segmentation-aware discriminator to ensure strong semantic alignment and visual fidelity. Through this combination of tailored architecture and loss design, we bridge the gap between SIS and SVS, outperforming state-of-the-art GAN- and diffusion-based baselines on both Cityscapes and KITTI-360. When combined with a semantic-map generator, SVS-GAN enables controllable generation of diverse driving scenarios, providing a scalable source of labeled video for data augmentation and closed-loop testing. Code is available at https://github.com/KhaledSerry/SVS-GAN. Khaled M. Seyam, Julian Wiederer, Markus Braun 0003, Bin Yang 0009 |
WACV | 4 |
| 2025 | Class-Aware PillarMix: Can Mixed Sample Data Augmentation Enhance 3D Object Detection with Radar Point Clouds?abstractDue to the significant effort required for data collection and annotation in 3D perception tasks, mixed sample data augmentation (MSDA) has been widely studied to generate diverse training samples by mixing existing data. Among these methods, MixUp is a prominent approach that generates new samples by linearly combining two existing ones, using a mix ratio sampled from a β distribution. This simple yet powerful method has inspired numerous variations and applications in 2D and 3D data domains. Recently, many MSDA techniques have been developed for point clouds, but they mainly target LiDAR data, leaving their application to radar point clouds largely unexplored. In this paper, we examine the feasibility of applying existing MSDA methods to radar point clouds and identify several challenges in adapting these techniques. These obstacles stem from the radar’s irregular angular distribution, deviations from a single-sensor polar layout in multi-radar setups, and point sparsity. To address these issues, we propose Class-Aware PillarMix (CAPMix), a novel MSDA approach that applies MixUp at the pillar level in 3D point clouds, guided by class labels. Unlike methods that rely a single mix ratio to the entire sample, CAPMix assigns an independent ratio to each pillar, boosting sample diversity. To account for the density of different classes, we use class-specific distributions: for dense objects (e.g., large vehicles), we skew ratios to favor points from another sample, while for sparse objects (e.g., pedestrians), we sample more points from the original. This class-aware mixing retains critical details and enriches each sample with new information, ultimately generating more diverse training data. Experimental results demonstrate that our method not only significantly boosts performance but also outperforms existing MSDA approaches across two datasets (Bosch Street and K-Radar). We believe that this straightforward yet effective approach will spark further investigation into MSDA techniques for radar data. Miao Zhang 0043, Sherif Abdulatif, Benedikt Loesch, Marco Altmann, Bin Yang 0009 |
IROS | 5 |
| 2025 | Memory-Efficient Pseudo-Labeling for Online Source-Free Universal Domain Adaptation using a Gaussian Mixture ModelabstractIn practice, domain shifts are likely to occur between training and test data, necessitating domain adaptation (DA) to adjust the pre-trained source model to the target domain. Recently, universal domain adaptation (UniDA) has gained attention for addressing the possibility of an additional category (label) shift between the source and target domain. This means new classes can appear in the target data, some source classes may no longer be present, or both at the same time. For practical applicability, UniDA methods must handle both source-free and online scenarios, enabling adaptation without access to the source data and performing batch-wise updates in parallel with prediction. In an online setting, preserving knowledge across batches is crucial. However, existing methods often require substantial memory, which is impractical because memory is limited and valuable, in particular on embedded systems. Therefore, we consider memory-efficiency as an additional constraint. To achieve memory-efficient online source-free universal domain adaptation (SF-UniDA), we propose a novel method that continuously captures the distribution of known classes in the feature space using a Gaussian mixture model (GMM). This approach, combined with entropy-based out-of-distribution detection, allows for the generation of reliable pseudo-labels. Finally, we combine a contrastive loss with a KL divergence loss to perform the adaptation. Our approach not only achieves state-of-the-art results in all experiments on the DomainNet and Office-Home datasets but also significantly outperforms the existing methods on the challenging VisDA-C dataset, setting a new benchmark for online SF-UniDA. Our code is available at https://github.com/pascalschlachter/GMM. Pascal Schlachter, Simon Wagner, Bin Yang 0009 |
WACV | 3 |
| 2024 | CNN Mixture-of-Depths
Rinor Cakaj, Jens Mehnert, Bin Yang 0009 |
ACCV (7) | 3 |
| 2024 | Towards calibration-free online EEG motor imagery decoding using Deep LearningabstractThe prevalence of stroke-induced disability drives research in motor imagery Brain-Computer Interfaces (BCIs) for rehabilitation.Closed-loop systems using traditional decoding models prevail but deep learning advances in single-trial offline decoding offer promises.However, transferring methods from offline to online decoding poses challenges.To address this, we propose a new approach to tune existing offline deep learning models towards online decoding, outperforming traditional pipelines without the need for subject-specific calibration data.Our proposed method is a step towards calibration-free BCIs that enable immediate feedback and user learning. Martin Wimpff, Jan Zerfowski, Bin Yang 0009 |
ESANN | 3 |
| 2024 | Deep Regression for Biological Age Estimation in Multiple Organs: Investigations on 40, 000 Subjects of the UK BiobankabstractAge plays an important role in shaping medical decisions, but the biological changes associated with aging do not solely depend on the chronological age. Genetics, lifestyle, and environment cause variations in age-related characteristics, even within the same chronological age group. Biological age (BA) was introduced to better capture an individual’s actual aging process, although its imprecise definition remains a challenge. Organ systems can age at different rates, necessitating organ-specific BA evaluation. To our knowledge, there have been no studies assessing age regionally across multiple organ systems. These age estimations, reflecting various body parts, enable a comprehensive patient-based analysis. We conducted brain, heart, kidney, liver, spleen, pancreas, and retinal fundus age estimations using MRI and OCT scans in 40,000 subjects of the UK Biobank with an uncertainty-aware ResNet-based network. Our results demonstrate the feasibility of organ-specific age estimation with cross-organ correlations of age-related changes. We achieve a mean age difference between predicted and chronological age of 2.72 years across all organs and an averaged Pearson correlation coefficient of 0.87. Veronika Ecker, Marcel Frueh, Bin Yang 0009, Sergios Gatidis, Thomas Kustner |
ICASSP | 3 |
| 2024 | Squeeze-and-Remember BlockabstractConvolutional Neural Networks (CNNs) are important for many machine learning tasks. They are built with different types of layers: convolutional layers that detect features, dropout layers that help to avoid over-reliance on any single neuron, and residual layers that allow the reuse of features. However, CNNs lack a dynamic feature retention mechanism similar to the human brain's memory, limiting their ability to use learned information in new contexts. To bridge this gap, we introduce the “Squeeze-and-Remember” (SR) block, a novel architectural unit that gives CNNs dynamic memory-like functionalities. The SR block selectively memorizes important features during training, and then adaptively re-applies these features during inference. This improves the network's ability to make contextually informed predictions. Empirical results on ImageNet and Cityscapes datasets demonstrate the SR block's efficacy: integration into ResNet50 improved top-1 validation accuracy on ImageNet by 0.52% over dropout2d alone, and its application in DeepLab v3 increased mean Intersection over Union in Cityscapes by 0.20%. These improvements are achieved with minimal computational overhead. This show the SR block's potential to enhance the capabilities of CNNs in image processing tasks. Rinor Cakaj, Jens Mehnert, Bin Yang 0009 |
ICMLA | 3 |
| 2024 | Spectral Wavelet Dropout: Regularization in the Wavelet DomainabstractRegularization techniques help prevent overfitting and therefore improve the ability of convolutional neural net-works (CNNs) to generalize. One reason for overfitting is the complex co-adaptations among different parts of the network, which make the CNN dependent on their joint response rather than encouraging each part to learn a useful feature representation independently. Frequency domain manipulation is a powerful strategy for modifying data that has temporal and spatial coherence by utilizing frequency decomposition. This work intro-duces Spectral Wavelet Dropout (SWD), a novel regularization method that includes two variants: ID-SWD and 2D-SWD. These variants improve CNN generalization by randomly dropping detailed frequency bands in the discrete wavelet decomposition of feature maps. Our approach distinguishes itself from the pre-existing Spectral “Fourier” Dropout (2D-SFD), which eliminates coefficients in the Fourier domain. Notably, SWD requires only a single hyperparameter, unlike the two required by SFD. We also extend the literature by implementing a one-dimensional version of Spectral “Fourier” Dropout (lD-SFD), setting the stage for a comprehensive comparison. Our evaluation shows that both ID and 2D SWD variants have competitive performance on CIFAR-IO/IOO benchmarks relative to both ID-SFD and 2D-SFD. Specifically, ID-SWD has a significantly lower computational complexity compared to ID/2D-SFD. In the Pascal VOC Object Detection benchmark, SWD variants surpass ID-SFD and 2D-SFD in performance and demonstrate lower computational complexity during training. Rinor Cakaj, Jens Mehnert, Bin Yang 0009 |
ICMLA | 3 |
| 2024 | Introducing Intermediate Domains for Effective Self-Training during Test-TimeabstractExperiencing domain shifts during test-time is nearly inevitable in practice and likely results in a severe performance degradation. To overcome this issue, test-time adaptation continues to update the initial source model after deployment. A promising direction are methods based on self-training which have been shown to be well suited for gradual domain adaptation, since reliable pseudo-labels can be provided. In this work, we address two problems that exist when applying self-training in the setting of test-time adaptation. First, adapting a model to long test sequences that contain multiple domains can lead to error accumulation. Second, naturally, not all shifts are gradual in practice. To tackle these challenges, we introduce GTTA. By creating artificial intermediate domains that divide the current domain shift into a more gradual one, effective self-training through high quality pseudo-labels can be performed. To create the intermediate domains, we propose two independent variations: mixup and light-weight style transfer. We demonstrate the effectiveness of our approach on the continual and gradual corruption benchmarks, as well as ImageNet-R. To further investigate gradual shifts in the context of urban scene segmentation, we publish a benchmark: CarlaTTA. It enables the exploration of several non-stationary domain shifts.1 Robert A. Marsden, Mario Döbler, Bin Yang 0009 |
IJCNN | 3 |
| 2024 | COMET: Contrastive Mean Teacher for Online Source-Free Universal Domain AdaptationabstractIn real-world applications, there is often a domain shift (distribution change) from training to test data. This observation recently resulted in the development of test-time adaptation (TTA). It aims to adapt a pre-trained source model to the test data without requiring access to the source data. Thereby, most existing works are so far limited to the closed-set assumption, i.e. there is no category shift (class change) between source and target domain. We argue that in a realistic open-world setting a category shift can appear in addition to a domain shift. This means, individual source classes may not appear in the target domain anymore, samples of new unknown classes may be part of the target domain or even both at the same time. Moreover, in many real-world scenarios the test data is not accessible all at once in form of a dataset but arrives sequentially as a stream of batches which require an immediate prediction. Hence, TTA must be applied in an online manner. To the best of our knowledge, the combination of these aspects, i.e. online source-free universal domain adaptation (online SF-UniDA), has not been studied yet despite its practical relevance. In this paper, we are the first ones to tackle this challenging task. We introduce a Contrastive Mean Teacher (COMET) tailored to this novel scenario. It applies a contrastive loss to rebuild a feature space where the samples of known classes build distinct clusters and the samples of new classes separate well from them. It is complemented by an entropy loss which ensures that the classifier output has a small entropy for samples of known classes and a large entropy for samples of new classes to be easily detected and rejected as unknown. To provide the losses with reliable pseudo labels, they are embedded into a mean teacher (MT) framework. We evaluate our method across two datasets and all category shifts to set an initial benchmark for online SF-UniDA. Thereby, COMET yields state-of-the-art performance and proves to be consistent and robust across a variety of different scenarios. Our code is available at https://github.com/pascalschlachter/COMET. Pascal Schlachter, Bin Yang 0009 |
IJCNN | 2 |
| 2024 | Universal Test-time Adaptation through Weight Ensembling, Diversity Weighting, and Prior CorrectionabstractSince distribution shifts are likely to occur during testtime and can drastically decrease the model’s performance, online test-time adaptation (TTA) continues to update the model after deployment, leveraging the current test data. Clearly, a method proposed for online TTA has to perform well for all kinds of environmental conditions. By introducing the variable factors domain non-stationarity and temporal correlation, we first unfold all practically relevant settings and define the entity as universal TTA. We want to highlight that this is the first work that covers such a broad spectrum, which is indispensable for the use in practice. To tackle the problem of universal TTA, we identify and highlight several challenges a self-training based method has to deal with: 1) model bias and the occurrence of trivial solutions when performing entropy minimization on varying sequence lengths with and without multiple domain shifts, 2) loss of generalization which exacerbates the adaptation to multiple domain shifts and the occurrence of catastrophic forgetting, and 3) performance degradation due to shifts in class prior. To prevent the model from becoming biased, we leverage a dataset and model-agnostic certainty and diversity weighting. In order to maintain generalization and prevent catastrophic forgetting, we propose to continually weightaverage the source and adapted model. To compensate for disparities in the class prior during test-time, we propose an adaptive prior correction scheme that reweights the model’s predictions. We evaluate our approach, named ROID, on a wide range of settings, datasets, and models, setting new standards in the field of universal TTA. Code is available at: https://github.com/mariodoebler/testtime-adaptation Robert A. Marsden, Mario Döbler, Bin Yang 0009 |
WACV | 3 |
| 2024 | Prompt tuning for parameter-efficient medical image segmentation
Marc Fischer 0003, Alexander Bartler, Bin Yang 0009 |
Medical Image Anal. | 3 |
| 2024 | CMGAN: Conformer-Based Metric-GAN for Monaural Speech EnhancementabstractIn this work, we further develop the conformerbased metric generative adversarial network (CMGAN) model1for speech enhancement (SE) in the time-frequency (TF) domain. This paper builds on our previous work but takes a more indepth look by conducting extensive ablation studies on model inputs and architectural design choices. We rigorously tested the generalization ability of the model to unseen noise types and distortions. We have fortified our claims through DNSMOS measurements and listening tests. Rather than focusing exclusively on the speech denoising task, we extend this work to address the dereverbration and super-resolution tasks. This necessitated exploring various architectural changes, specifically metric discriminator scores and masking techniques. It is essential to highlight that this is among the earliest works that attempted complex TF-domain super-resolution. Our findings show that CMGAN outperforms existing state-of-the-art methods in the three major speech enhancement tasks: denoising, dereverberation, and super-resolution. For example, in the denoising task using the Voice Bank+DEMAND dataset, CMGAN notably exceeded the performance of prior models, attaining a PESQ score of 3.41 and an SSNR of 11.10 dB. Audio samples and CMGAN implementations are available online2. Sherif Abdulatif, Ruizhe Cao, Bin Yang 0009 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Robust Mean Teacher for Continual and Gradual Test-Time AdaptationabstractSince experiencing domain shifts during test-time is inevitable in practice, test-time adaption (TTA) continues to adapt the model after deployment. Recently, the area of continual and gradual test-time adaptation (TTA) emerged. In contrast to standard TTA, continual TTA considers not only a single domain shift, but a sequence of shifts. Gradual TTA further exploits the property that some shifts evolve gradually over time. Since in both settings long test sequences are present, error accumulation needs to be addressed for methods relying on self-training. In this work, we propose and show that in the setting of TTA, the symmetric cross-entropy is better suited as a consistency loss for mean teachers compared to the commonly used cross-entropy. This is justified by our analysis with respect to the (symmetric) cross-entropy's gradient properties. To pull the test feature space closer to the source domain, where the pre-trained model is well posed, contrastive learning is leveraged. Since applications differ in their requirements, we address several settings, including having source data available and the more challenging source-free setting. We demonstrate the effectiveness of our proposed method “robust mean teacher” (RMT) on the continual and gradual corruption benchmarks CIFAR10C, CIFAR100C, and Imagenet-C. We further consider ImageNet-R and propose a new continual DomainNet-126 benchmark. State-of-the-art results are achieved on all benchmarks.11Code is available at: https://github.com/mariodoebler/test-time-adaptation Mario Döbler, Robert A. Marsden, Bin Yang 0009 |
CVPR | 3 |
| 2023 | NIFF: Alleviating Forgetting in Generalized Few-Shot Object Detection via Neural Instance Feature ForgingabstractPrivacy and memory are two recurring themes in a broad conversation about the societal impact of AI. These con-cerns arise from the need for huge amounts of data to train deep neural networks. A promise of Generalized Few-shot Object Detection (G-FSOD), a learning paradigm in AI, is to alleviate the need for collecting abundant training samples of novel classes we wish to detect by leveraging prior knowledge from old classes (i.e., base classes). G-FSOD strives to learn these novel classes while alleviating catas-trophic forgetting of the base classes. However, existing approaches assume that the base images are accessible, an assumption that does not hold when sharing and storing data is problematic. In this work, we propose the first data-free knowledge distillation (DFKD) approach for G-FSOD that leverages the statistics of the region of interest (RoI) features from the base model to forge instance-level features without accessing the base images. Our contribution is three-fold: (1) we design a standalone lightweight generator with (2) class-wise heads (3) to generate and replay diverse instance-level base features to the RoI head while finetuning on the novel data. This stands in contrast to standard DFKD approaches in image classification, which invert the entire network to generate base images. Moreover, we make careful design choices in the novel finetuning pipeline to regularize the model. We show that our approach can dramatically reduce the base memory requirements, all while setting a new standard for G-FSOD on the challenging MS-COCO and PASCAL-VOC benchmarks. Karim Guirguis, Johannes Meier, George Eskandar, Matthias Kayser, Bin Yang 0009, Jürgen Beyerer |
CVPR | 5 |
| 2023 | A Semi-Paired Approach for Label-to-Image TranslationabstractData efficiency, or the ability to generalize from a few labeled data, remains a major challenge in deep learning. Semi-supervised learning has thrived in traditional recognition tasks alleviating the need for large amounts of labeled data, yet it remains understudied in image-to-image translation (I2I) tasks. In this work, we introduce the first semi-supervised (semi-paired) framework for label-to-image translation, a challenging subtask of I2I which generates photorealistic images from semantic label maps. In the semi-paired setting, the model has access to a small set of paired data and a larger set of unpaired images and labels. Instead of using geometrical transformations as a pretext task like previous works, we leverage an input reconstruction task by exploiting the conditional discriminator on the paired data as a reverse generator. We propose a training algorithm for this shared network, and we present a rare classes sampling algorithm to focus on under-represented classes. Experiments on 3 standard benchmarks show that the proposed model outperforms state-of-the-art unsupervised and semi-supervised approaches, as well as some fully supervised approaches while using a much smaller number of paired samples. George Eskandar, Mohamed Abdelsamad, Mark Youssef, Diandian Guo, Bin Yang 0009 |
ICIP | 6 |
| 2023 | Weight Compander: A Simple Weight Reparameterization for RegularizationabstractRegularization is a set of techniques that are used to improve the generalization ability of deep neural networks. In this paper, we introduce weight compander (WC), a novel effective method to improve generalization by reparameterizing each weight in deep neural networks using a nonlinear function. It is a general, intuitive, cheap and easy to implement method, which can be combined with various other regularization techniques. Large weights in deep neural networks are a sign of a more complex network that is overfitted to the training data. Moreover, regularized networks tend to have a greater range of weights around zero with fewer weights centered at zero. We introduce a weight reparameterization function which is applied to each weight and implicitly reduces overfitting by restricting the magnitude of the weights while forcing them away from zero at the same time. This leads to a more democratic decision-making in the network. Firstly, individual weights cannot have too much influence in the prediction process due to the restriction of their magnitude. Secondly, more weights are used in the prediction process, since they are forced away from zero during the training. This promotes the extraction of more features from the input data and increases the level of weight redundancy, which makes the network less sensitive to statistical differences between training and test data. From an optimizational point of view, the second effect of WC can be seen as a reactivation of “dead” (near zero) weights to participate in the training. This increases the probability to find an ensemble of weights which performs better in the given task. We extend our method to learn the hyperparameters of the introduced weight reparameterization function. This avoids hyperparameter search and gives the network the opportunity to align the weight reparameterization with the training progress. We show experimentally that using weight compander in addition to standard regularization methods improves the performance of neural networks. Furthermore, we empirically analyze the weight distribution with and without weight compander after training to confirm the companding effects of our method on the weights. Rinor Cakaj, Jens Mehnert, Bin Yang 0009 |
IJCNN | 3 |
| 2023 | Spectral Batch Normalization: Normalization in the Frequency DomainabstractRegularization is a set of techniques that are used to improve the generalization ability of deep neural networks. In this paper, we introduce spectral batch normalization (SBN), a novel effective method to improve generalization by normalizing feature maps in the frequency (spectral) domain. The activations of residual networks without batch normalization (BN) tend to explode exponentially in the depth of the network at initialization. This leads to extremely large feature map norms even though the parameters are relatively small. These explosive dynamics can be very detrimental to learning. BN makes weight decay regularization on the scaling factors$\gamma,\beta$approximately equivalent to an additive penalty on the norm of the feature maps, which prevents extremely large feature map norms to a certain degree. It was previously shown that preventing explosive growth at the final layer at initialization and during training in ResNets can recover a large part of Batch Normalization's generalization boost. However, we show experimentally that, despite the approximate additive penalty of BN, feature maps in deep neural networks (DNNs) tend to explode at the beginning of the training and that feature maps of DNNs contain large values during the whole training. This phenomenon also occurs in a weakened form in non-residual networks. Intuitively, it is not preferred to have large values in feature maps since they have too much influence on the prediction in contrast to other parts of the feature map. SBN addresses large feature maps by normalizing them in the frequency domain. In our experiments, we empirically show that SBN prevents exploding feature maps at initialization and large feature map values during the training. Moreover, the normalization of feature maps in the frequency domain leads to more uniform distributed frequency components. This discourages the DNNs to rely on single frequency components of feature maps. These, together with other effects (e.g. noise injection, scaling and shifting of the feature map) of SBN, have a regularizing effect on the training of residual and non-residual networks. We show experimentally that using SBN in addition to standard regularization methods improves the performance of DNNs by a relevant margin, e.g. ResNet50 on CIFAR-100 by 2.31%, on ImageNet by 0.71% (from 76.80% to 77.51%) and VGG19 on CIFAR-100 by 0.66%. Rinor Cakaj, Jens Mehnert, Bin Yang 0009 |
IJCNN | 3 |
| 2023 | Indoor Positioning Based on Active Radar Sensing and Passive Reflectors: Reflector Placement OptimizationabstractWe extend our work on a novel indoor positioning system (IPS) for autonomous mobile robots (AMRs) based on radar sensing of local, passive radar reflectors. Through the combination of simple reflectors and a single-channel frequency modulated continuous wave (FMCW) radar, high positioning accuracy at low system cost can be achieved. Further, a multi-objective (MO) particle swarm optimization (PSO) algorithm is presented that optimizes the 2D placement of radar reflectors in complex room settings. Sven Hinderer, Pascal Schlachter, Zhibin Yu 0004, Bin Yang 0009 |
IPIN | 5 |
| 2023 | Urban-StyleGAN: Learning to Generate and Manipulate Images of Urban ScenesabstractA promise of Generative Adversarial Networks (GANs) is to provide cheap photorealistic data for training and validating AI models in autonomous driving. Despite their huge success, their performance on complex images featuring multiple objects is understudied. While some frameworks produce high-quality street scenes with little to no control over the image content, others offer more control at the expense of high-quality generation. A common limitation of both approaches is the use of global latent codes for the whole image, which hinders the learning of independent object distributions. Motivated by SemanticStyleGAN (SSG), a recent work on latent space disentanglement in human face generation, we propose a novel framework, Urban-StyleGAN, for urban scene generation and manipulation. We find that a straightforward application of SSG leads to poor results because urban scenes are more complex than human faces. To provide a more compact yet disentangled latent representation, we develop a class grouping strategy wherein individual classes are grouped into super-classes. Moreover, we employ an unsupervised latent exploration algorithm in the $\mathcal{S}$-space of the generator and show that it is more efficient than the conventional ${\mathcal{W}^ + }$-space in controlling the image content. Results on the Cityscapes and Mapillary datasets show the proposed approach achieves significantly more controllability and improved image quality than previous approaches on urban scenes and is on par with general-purpose non-controllable generative models (like StyleGAN2) in terms of quality. George Eskandar, Youssef Farag, Tarun Yenamandra, Daniel Cremers, Karim Guirguis, Bin Yang 0009 |
IV | 6 |
| 2023 | Towards Pragmatic Semantic Image Synthesis for Urban ScenesabstractThe need for large amounts of training and validation data is a huge concern in scaling AI algorithms for autonomous driving. Semantic Image Synthesis (SIS), or label-to-image translation, promises to address this issue by translating semantic layouts to images, providing a controllable generation of photorealistic data. However, they require a large amount of paired data, incurring extra costs. In this work, we present a new task: given a dataset with synthetic images and labels and a dataset with unlabeled real images, our goal is to learn a model that can generate images with the content of the input mask and the appearance of real images. This new task reframes the well-known unsupervised SIS task in a more practical setting, where we leverage cheaply available synthetic data from a driving simulator to learn how to generate photorealistic images of urban scenes. This stands in contrast to previous works, which assume that labels and images come from the same domain but are unpaired during training. We find that previous unsupervised works underperform on this task, as they do not handle distribution shifts between two different domains. To bypass these problems, we propose a novel framework with two main contributions. First, we leverage the synthetic image as a guide to the content of the generated image by penalizing the difference between their high-level features on a patch level. Second, in contrast to previous works which employ one discriminator that overfits the target domain semantic distribution, we employ a discriminator for the whole image and multiscale discriminators on the image patches. Extensive comparisons on the benchmarks benchmarks GTA-V → Cityscapes and GTA-V → Mapillary show the superior performance of the proposed model against state-of-the-art on this task. George Eskandar, Diandian Guo, Karim Guirguis, Bin Yang 0009 |
IV | 4 |
| 2023 | Towards Discriminative and Transferable One-Stage Few-Shot Object DetectorsabstractRecent object detection models require large amounts of annotated data for training a new classes of objects. Few-shot object detection (FSOD) aims to address this problem by learning novel classes given only a few samples. While competitive results have been achieved using two-stage FSOD detectors, typically one-stage FSODs under-perform compared to them. We make the observation that the large gap in performance between two-stage and one-stage FSODs are mainly due to their weak discriminability, which is explained by a small post-fusion receptive field and a small number of foreground samples in the loss function. To address these limitations, we propose the Few-shot RetinaNet (FSRN) that consists of: a multi-way support training strategy to augment the number of foreground samples for dense meta-detectors, an early multi-level feature fusion providing a wide receptive field that covers the whole anchor area and two augmentation techniques on query and source images to enhance transferability. Extensive experiments show that the proposed approach addresses the limitations and boosts both discriminability and transferability. FSRN is almost two times faster than two-stage FSODs while remaining competitive in accuracy, and it outperforms the state-of-the-art of one-stage meta-detectors and also some two-stage FSODs on the MS-COCO and PASCAL VOC benchmarks. Karim Guirguis, Mohamed Abdelsamad, George Eskandar, Ahmed Hendawy, Matthias Kayser, Bin Yang 0009, Jürgen Beyerer |
WACV | 6 |
| 2023 | USIS: Unsupervised Semantic Image Synthesis
George Eskandar, Mohamed Abdelsamad, Karim Armanious, Bin Yang 0009 |
Comput. Graph. | 4 |
| 2022 | MT3: Meta Test-Time Training for Self-Supervised Test-Time AdaptionabstractAn unresolved problem in Deep Learning is the ability of neural networks to cope with domain shifts during test-time, imposed by commonly fixing network parameters after training. Our proposed method Meta Test-Time Training (MT3), however, breaks this paradigm and enables adaption at test-time. We combine meta-learning, self-supervision and test-time training to learn to adapt to unseen test distributions. By minimizing the self-supervised loss, we learn task-specific model parameters for different tasks. A meta-model is optimized such that its adaption to the different task-specific models leads to higher performance on those tasks. During test-time a single unlabeled image is sufficient to adapt the meta-model parameters. This is achieved by minimizing only the self-supervised loss component resulting in a better prediction for that image. Our approach significantly improves the state-of-the-art results on the CIFAR-10-Corrupted image classification benchmark. Alexander Bartler, Andre Bühler, Felix Wiewel, Mario Döbler, Bin Yang 0009 |
AISTATS | 5 |
| 2022 | Intelligent Methods for Test and ReliabilityabstractTest methods that can keep up with the ongoing increase in complexity of semiconductor products and their underlying technologies are an essential prerequisite for maintaining quality and safety of our daily lives and for continued success of our economies and societies. There is a huge potential how test methods can benefit from recent breakthroughs in domains such as artificial intelligence, data analytics, virtual/augmented reality, and security. The Graduate School on “Intelligent Methods for Semiconductor Test and Reliability” (GS-IMTR) at the University of Stuttgart is a large-scale, radically interdisciplinary effort to address the scientific-technological challenges in this domain. It is funded by Advantest, one of the world leaders in automatic test equipment. In this paper, we describe the overall philosophy of the Graduate School and the specific scientific questions targeted by its ten projects. Hussam Amrouch, Jens Anders, Steffen Becker 0001, Maik Betka, Gerd Bleher, Peter Domanski, Nourhan Elhamawy, Thomas Ertl, Athanasios Gatzastras, Paul R. Genssler, Sebastian Hasler, Martin Heinrich, André van Hoorn, Hanieh Jafarzadeh, Ingmar Kallfass, Florian Klemme, Steffen Koch 0001, Ralf Küsters, Andrés Lalama, Raphaël Latty, Yiwen Liao, Natalia Lylina, Zahra Paria Najafi-Haghi, Dirk Pflüger, Ilia Polian, Jochen Rivoir, Matthias Sauer 0002, Denis Schwachhofer, Steffen Templin, Christian Volmer, Stefan Wagner 0001, Daniel Weiskopf, Hans-Joachim Wunderlich, Bin Yang 0009 |
DATE | 34 |
| 2022 | Wavelet-Based Unsupervised Label-to-Image TranslationabstractSemantic Image Synthesis (SIS) is a subclass of image-to-image translation where a semantic layout is used to generate a photorealistic image. State-of-the-art conditional Generative Adversarial Networks (GANs) need a huge amount of paired data to accomplish this task while generic un-paired image-to-image translation frameworks underperform in comparison, because they color-code semantic layouts and learn correspondences in appearance instead of semantic content. Starting from the assumption that a high quality generated image should be segmented back to its semantic layout, we propose a new Unsupervised paradigm for SIS (USIS) that makes use of a self-supervised segmentation loss and whole image wavelet based discrimination. Furthermore, in order to match the high-frequency distribution of real images, a novel generator architecture in the wavelet domain is proposed. We test our methodology on 3 challenging datasets and demonstrate its ability to bridge the performance gap between paired and unpaired models. George Eskandar, Mohamed Abdelsamad, Karim Armanious, Bin Yang 0009 |
ICASSP | 5 |
| 2022 | On-Board Pedestrian Trajectory Prediction Using Behavioral FeaturesabstractThis paper presents a novel approach to pedestrian trajectory prediction for on-board camera systems, which utilizes behavioral features of pedestrians that can be inferred from visual observations. Our proposed method, called Behavior-Aware Pedestrian Trajectory Prediction (BA-PTP), processes multiple input modalities, i.e. bounding boxes, body and head orientation of pedestrians as well as their pose, with independent encoding streams. The encodings of each stream are fused using a modality attention mechanism, resulting in a final embedding that is used to predict future bounding boxes in the image.In experiments on two datasets for pedestrian behavior prediction, we demonstrate the benefit of using behavioral features for pedestrian trajectory prediction and evaluate the effectiveness of the proposed encoding strategy. Additionally, we investigate the relevance of different behavioral features on the prediction performance based on an ablation study. Phillip Czech, Markus Braun 0003, Ulrich Kressel, Bin Yang 0009 |
ICMLA | 4 |
| 2022 | TTAPS: Test-Time Adaption by Aligning Prototypes using Self-SupervisionabstractNowadays, deep neural networks outperform humans in many tasks. However, if the input distribution drifts away from the one used in training, their performance drops significantly. Recently published research has shown that adapting the model parameters to the test sample can mitigate this performance degradation. In this paper, we therefore propose a novel modification of the self-supervised training algorithm SwAV that adds the ability to adapt to single test samples. Using the provided prototypes of SwAV and our derived test-time loss, we align the representation of unseen test samples with the self-supervised learned prototypes. We show the success of our method on the common benchmark dataset CIFAR10-C. Alexander Bartler, Florian Bender, Felix Wiewel, Bin Yang 0009 |
IJCNN | 4 |
| 2022 | To Generalize or Not to Generalize: Towards Autoencoders in One-Class ClassificationabstractIn One-Class Classification (OCC), data affiliated with only one given class are accessible during training, while the trained algorithm must be capable of distinguishing the given class from all other unknown classes during test. OCC can be considered as a general case for many anomaly detection and novelty discovery problems and is more challenging than conventional binary and multi-class classification due to the absence of unknown classes. In literature, one of the most popular techniques towards OCC is autoencoder because autoencoders can capture the intrinsic structure of training data and are expected to have smaller reconstruction errors for the given class than those of unknown classes. As a result, during test, data from the known and unknown classes can be discriminated by thresholds. Despite the wide application of autoencoders in OCC, we have recently discovered that autoencoders can easily generalize to unknown classes, although they are trained on one given class only. The unexpected behavior reduces the overall OCC performance and leads to performance degradation during training. This paper proposes a novel solution to regularize the generalization ability of autoencoders and reduce performance degradation in OCC by introducing a feature weighting block in the latent space of autoencoders. Intensive experiments show that our method enables a training without performance degradation and significantly improves the OCC ability in comparison with contemporary state-of-the-art approaches. Yiwen Liao, Bin Yang 0009 |
IJCNN | 2 |
| 2022 | Contrastive Learning and Self-Training for Unsupervised Domain Adaptation in Semantic SegmentationabstractDeep convolutional neural networks have considerably improved state-of-the-art results for semantic segmentation. Nevertheless, even modern architectures lack the ability to generalize well to a test dataset that originates from a different domain. To avoid the costly annotation of training data for unseen domains, unsupervised domain adaptation (UDA) attempts to provide efficient knowledge transfer from a labeled source domain to an unlabeled target domain. Previous work has mainly focused on minimizing the discrepancy between the two domains by using adversarial training or self-training. While adversarial training may fail to align the correct semantic categories as it minimizes the discrepancy between the global distributions, self-training raises the question of how to provide reliable pseudolabels. To align the correct semantic categories across domains, we propose a contrastive learning approach that adapts category-wise centroids across domains. Furthermore, we extend our method with self-training, where we use a memory-efficient temporal ensemble to generate consistent and reliable pseudo-labels. Although both contrastive learning and self-training (CLST) through temporal ensembling enable knowledge transfer between two domains, it is their combination that leads to a symbiotic structure. We validate our approach on two domain adaptation benchmarks: GTA5 → Cityscapes and SYNTHIA → Cityscapes. Our method achieves better results than the state-of-the-art. Robert A. Marsden, Alexander Bartler, Mario Döbler, Bin Yang 0009 |
IJCNN | 4 |
| 2022 | Continual Unsupervised Domain Adaptation for Semantic Segmentation using a Class-Specific TransferabstractIn recent years, there has been tremendous progress in the field of semantic image segmentation. However, one remaining challenging problem is that segmentation models do not generalize to unseen domains. To overcome this problem, one either has to label lots of data covering the whole variety of possible domains, which is often infeasible in practice, or apply unsupervised domain adaptation (UDA), only requiring labeled source data. In this work, we focus on UDA and additionally address the case of adapting not only to a single domain, but to a sequence of target domains. This requires mechanisms preventing the model from forgetting its previously learned knowledge. To adapt a segmentation model to a target domain, we follow the idea of utilizing light-weight style transfer to convert the style of labeled source images into the style of the target domain, while retaining the source content. To mitigate the distributional shift between the source and the target domain, the model is fine-tuned on the transferred source images in a second step. Existing light-weight style transfer approaches relying on adaptive instance normalization (AdaIN) or Fourier transformation (FDA) still lack performance and do not substantially improve upon common data augmentation, such as color jittering. The reason for this is that these methods do not focus on region- or class-specific differences, but mainly capture the most salient style. Therefore, we propose a simple and light-weight framework that incorporates two class-conditional AdaIN layers. To extract the class-specific target moments needed for the transfer layers, we use unfiltered pseudo-labels, which we show to be an effective approximation compared to real labels. We extensively validate our approach (CACE) on a synthetic sequence and further propose a challenging sequence consisting of real domains. CACE outperforms existing methods visually and quantitatively. Robert A. Marsden, Felix Wiewel, Mario Döbler, Bin Yang 0009 |
IJCNN | 5 |
| 2022 | Dirichlet Prior Networks for Continual LearningabstractDeep Neural Networks (DNNs) suffer from the long known phenomenon of catastrophic forgetting when trained on a sequence of tasks without appropriate counter measures. Overcoming this is of great interest as it would enable DNNs to accumulate knowledge over a potentially long sequence of tasks without forgetting. In this paper, we study the commonly used method of rehearsal for mitigating catastrophic forgetting and show that it can cause an unwanted distribution shift that negatively affects performance. Building on recently introduced Dirichlet Prior Networks (DPNs), we propose a novel method that incorporates prior knowledge of known distribution shifts into its predictions in order to reduce their negative influence. The proposed method is evaluated on commonly used benchmark datasets and compared to related methods. Felix Wiewel, Alexander Bartler, Bin Yang 0009 |
IJCNN | 3 |
| 2022 | CMGAN: Conformer-based Metric GAN for Speech EnhancementabstractRecently, convolution-augmented transformer (Conformer) has achieved promising performance in automatic speech recognition (ASR) and time-domain speech enhancement (SE), as it can capture both local and global dependencies in the speech signal. In this paper, we propose a conformer-based metric generative adversarial network (CMGAN) for SE in the time-frequency (TF) domain. In the generator, we utilize two-stage conformer blocks to aggregate all magnitude and complex spectrogram information by modeling both time and frequency dependencies. The estimation of magnitude and complex spectrogram is decoupled in the decoder stage and then jointly incorporated to reconstruct the enhanced speech. In addition, a metric discriminator is employed to further improve the quality of the enhanced estimated speech by optimizing the generator with respect to a corresponding evaluation score. Quantitative analysis on Voice Bank+DEMAND dataset indicates the capability of CMGAN in outperforming various previous models with a margin, i.e., PESQ of 3.41 and SSNR of 11.10 dB. Ruizhe Cao, Sherif Abdulatif, Bin Yang 0009 |
INTERSPEECH | 3 |
| 2022 | An Unsupervised Domain Adaptive Approach for Multimodal 2D Object Detection in Adverse Weather ConditionsabstractIntegrating different representations from complementary sensing modalities is crucial for robust scene interpretation in autonomous driving. While deep learning architectures that fuse vision and range data for 2D object detection have thrived in recent years, the corresponding modalities can degrade in adverse weather or lighting conditions, ultimately leading to a drop in performance. Although domain adaptation methods attempt to bridge the domain gap between source and target domains, they do not readily extend to heterogeneous data distributions. In this work, we propose an unsupervised domain adaptation framework, which adapts a 2D object detector for RGB and LiDAR sensors to one or more target domains featuring adverse weather conditions. Our proposed approach consists of three components. First, a data augmentation scheme that simulates weather distortions is devised to add domain confusion and prevent overfitting on the source data. Second, to promote cross-domain foreground object alignment, we leverage the complementary features of multiple modalities through a multi-scale entropy-weighted domain discriminator. Finally, we use carefully designed pretext tasks to learn a more robust representation of the target domain data. Experiments performed on the DENSE dataset show that our method can substantially alleviate the domain gap under the single-target domain adaptation setting and the less explored yet more general multi-target domain adaptation setting. George Eskandar, Robert A. Marsden, Pavithran Pandiyan, Mario Döbler, Karim Guirguis, Bin Yang 0009 |
IROS | 6 |
| 2022 | Wafer Map Defect Classification Based on the Fusion of Pattern and Pixel InformationabstractWith the dramatically increasing requirements on semiconductor products, improving the yield is one of the major tasks for semiconductor manufacturers. To minimize losses, automatic and efficient wafer testing tools are required to quickly notify the engineers of potential problems. One such technique is wafer map defect pattern classification, which has inspired and motivated extensive research over the last decades. Many popular studies often design novel wafer map defect identification algorithms based on manual feature extraction, statistical learning and deep neural networks, having achieved significant advancement and success. However, these methods often face challenges of training large-scale networks and few of them have noticed the full usage of the information within each wafer map. Based on the concerns above, this paper proposes a multi-task learning framework based on neural networks that fuses the information of the entire wafer map as well as the state of each individual die to enhance the defect pattern classification capability. Extensive experiments on a public real-world dataset have been conducted to justify the effectiveness of our method. Specifically, our method achieved an classification accuracy of 96.3%, which was better or comparable to other state-of-the-art approaches that required notably larger network sizes and heavy data augmentation. Yiwen Liao, Raphaël Latty, Paul R. Genssler, Hussam Amrouch, Bin Yang 0009 |
ITC | 5 |
| 2022 | Efficient and Robust Resistive Open Defect Detection Based on Unsupervised Deep LearningabstractBoth process variations and defects in cells can lead to additional small delays within specifications, while the latter must be identified because they may degrade soon into critical faults for circuits and result in threat to reliability. Therefore, discriminating small delays due to defects from those due to variations has drawn increasingly attention in the test community over the recent years. One promising research direction is to formulate the task into binary classification by using delays under a few supply voltages as the only variables for data-driven algorithms. However, many approaches often assume the availability of delay information from both defective and non-defective cells or combinational circuits. This assumption implies a large time consumption for simulation, and considerable costs for manufactured defective devices. To address the issues above, this paper proposes to use unsupervised deep learning techniques to train an recognizer on non-defective data only but still can identify defects during inference. Specifically, we have proposed to use a weighted autoencoder with a novel data augmentation technique to solve this problem. Experiments show that our approach has comparable detection capability as supervised learning schemes, while our method does not require any defective data. Moreover, in practice, our approach is more robust to unbalanced datasets and to non-target defects than other methods. Yiwen Liao, Zahra Paria Najafi-Haghi, Hans-Joachim Wunderlich, Bin Yang 0009 |
ITC | 4 |
| 2021 | Uncertainty-Based Biological Age Estimation of Brain MRI ScansabstractAge is an essential factor in modern diagnostic procedures. However, assessment of the true biological age (BA) remains a daunting task due to the lack of reference ground-truth labels. Current BA estimation approaches are either restricted to skeletal images or rely on non-imaging modalities that yield a whole-body BA assessment. However, various organ systems may exhibit different aging characteristics due to lifestyle and genetic factors. In this initial study, we propose a new framework for organ-specific BA estimation utilizing 3D magnetic resonance image (MRI) scans. As a first step, this framework predicts the chronological age (CA) together with the corresponding patient-dependent aleatoric uncertainty. An iterative training algorithm is then utilized to segregate atypical aging patients from the given population based on the predicted uncertainty scores. In this manner, we hypothesize that training a new model on the remaining population should approximate the true BA behavior. We apply the proposed methodology on a brain MRI dataset containing healthy individuals as well as Alzheimer’s patients. We demonstrate the correlation between the predicted BAs and the expected cognitive deterioration in Alzheimer’s patients. Karim Armanious, Sherif Abdulatif, Wenbin Shi, Tobias Hepp 0002, Sergios Gatidis, Bin Yang 0009 |
ICASSP | 6 |
| 2021 | Automated Multi-Organ Segmentation in Pet Images Using Cascaded Training of a 3d U-Net and Convolutional AutoencoderabstractPET imaging is an important tool in clinical diagnostics, especially in oncology as it is able to visualize ongoing metabolic processes, e.g. caused by a tumor. Due to the low spatial resolution, a corresponding CT or MRI scan is normally necessary to gain knowledge about the physiological structures of a patient and especially to perform some computer-aided diagnostics methods, e.g. segmentation of structures of interest. As this transfer of information from CT/MRI to the PET domain is not always feasible, e.g. when the corresponding CT or MRI images are unavailable or corrupted by artifacts, we propose a novel approach to perform organ segmentation on the PET images directly. We utilize a CNN architecture based on a 3D U-Net combined with a convolutional autoencoder and train our model purely on PET images and corresponding ground truth masks. Our resulting Dice scores of 0.88, 0.82 an 0.59 for liver, spleen and spine, respectively, show that standalone PET organ segmentation is generally feasible. Annika Liebgott, Charlotte Lorenz, Sergios Gatidis, Viet Chau Vu, Konstantin Nikolaou, Bin Yang 0009 |
ICASSP | 6 |
| 2021 | Multi-Class Uncertainty Calibration via Mutual Information Maximization-based Binning
Kanil Patel, William Beluch, Bin Yang 0009, Michael Pfeiffer 0001, Dan Zhang 0003 |
ICLR | 3 |
| 2021 | Feature Selection Using Batch-Wise Attenuation and Feature Mask NormalizationabstractFeature selection is generally used as one of the most important preprocessing techniques in machine learning, as it helps to reduce the dimensionality of data and assists researchers and practitioners in understanding data. Thereby, by utilizing feature selection, better performance and reduced computational consumption, memory complexity and even data amount can be expected. Although there exist approaches leveraging the power of deep neural networks to carry out feature selection, many of them often suffer from sensitive hyperparameters. This paper proposes a feature mask module (FM-module) for feature selection based on a novel batch-wise attenuation and feature mask normalization. The proposed method is almost free from hyperparameters and can be easily integrated into common neural networks as an embedded feature selection method. Experiments on popular image, text and speech datasets have shown that our approach is easy to use and has superior performance in comparison with other state-of-the-art deep-learning-based feature selection methods. Yiwen Liao, Raphaël Latty, Bin Yang 0009 |
IJCNN | 3 |
| 2021 | Condensed Composite Memory Continual LearningabstractDeep Neural Networks (DNNs) suffer from a rapid decrease in performance when trained on a sequence of tasks where only data of the most recent task is available. This phenomenon, known as catastrophic forgetting, prevents DNNs from accumulating knowledge over time. Overcoming catastrophic forgetting and enabling continual learning is of great interest since it would enable the application of DNNs in settings where unrestricted access to all the training data at any time is not always possible, e.g. due to storage limitations or legal issues. While many recently proposed methods for continual learning use some training examples for rehearsal, their performance strongly depends on the number of stored examples. In order to improve performance of rehearsal for continual learning, especially for a small number of stored examples, we propose a novel way of learning a small set of synthetic examples which capture the essence of a complete dataset. Instead of directly learning these synthetic examples, we learn a weighted combination of shared components for each example that enables a significant increase in memory efficiency. We demonstrate the performance of our method on commonly used datasets and compare it to recently proposed related methods and baselines. Felix Wiewel, Bin Yang 0009 |
IJCNN | 2 |
| 2021 | Age-Net: An MRI-Based Iterative Framework for Brain Biological Age EstimationabstractThe concept of biological age (BA) - although important in clinical practice - is hard to grasp mainly due to the lack of a clearly defined reference standard. For specific applications, especially in pediatrics, medical image data are used for BA estimation in a routine clinical context. Beyond this young age group, BA estimation is mostly restricted to whole-body assessment using non-imaging indicators such as blood biomarkers, genetic and cellular data. However, various organ systems may exhibit different aging characteristics due to lifestyle and genetic factors. Thus, a whole-body assessment of the BA does not reflect the deviations of aging behavior between organs. To this end, we propose a new imaging-based framework for organ-specific BA estimation. In this initial study we focus mainly on brain MRI. As a first step, we introduce a chronological age (CA) estimation framework using deep convolutional neural networks (Age-Net). We quantitatively assess the performance of this framework in comparison to existing state-of-the-art CA estimation approaches. Furthermore, we expand upon Age-Net with a novel iterative data-cleaning algorithm to segregate atypical-aging patients (BA [Formula: see text] CA) from the given population. We hypothesize that the remaining population should approximate the true BA behavior. We apply the proposed methodology on a brain magnetic resonance image (MRI) dataset containing healthy individuals as well as Alzheimer's patients with different dementia ratings. We demonstrate the correlation between the predicted BAs and the expected cognitive deterioration in Alzheimer's patients. A statistical and visualization-based analysis has provided evidence regarding the potential and current challenges of the proposed methodology. Karim Armanious, Sherif Abdulatif, Wenbin Shi, Shashank Salian, Thomas Kustner, Daniel Weiskopf, Tobias Hepp 0002, Sergios Gatidis, Bin Yang 0009 |
IEEE Trans. Medical Imaging | 9 |
| 2021 | LAPNet: Non-Rigid Registration Derived in k-Space for Magnetic Resonance ImagingabstractPhysiological motion, such as cardiac and respiratory motion, during Magnetic Resonance (MR) image acquisition can cause image artifacts. Motion correction techniques have been proposed to compensate for these types of motion during thoracic scans, relying on accurate motion estimation from undersampled motion-resolved reconstruction. A particular interest and challenge lie in the derivation of reliable non-rigid motion fields from the undersampled motion-resolved data. Motion estimation is usually formulated in image space via diffusion, parametric-spline, or optical flow methods. However, image-based registration can be impaired by remaining aliasing artifacts due to the undersampled motion-resolved reconstruction. In this work, we describe a formalism to perform non-rigid registration directly in the sampled Fourier space, i.e. k-space. We propose a deep-learning based approach to perform fast and accurate non-rigid registration from the undersampled k-space data. The basic working principle originates from the Local All-Pass (LAP) technique, a recently introduced optical flow-based registration. The proposed LAPNet is compared against traditional and deep learning image-based registrations and tested on fully-sampled and highly-accelerated (with two undersampling strategies) 3D respiratory motion-resolved MR images in a cohort of 40 patients with suspected liver or lung metastases and 25 healthy subjects. The proposed LAPNet provided consistent and superior performance to image-based approaches throughout different sampling trajectories and acceleration factors. Thomas Kustner, Jiazhen Pan, Haikun Qi, Gastão Cruz, Christopher Gilliam, Thierry Blu, Bin Yang 0009, Sergios Gatidis, René M. Botnar, Claudia Prieto |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Supervised Canonical Correlation Analysis of Data on Symmetric Positive Definite Manifolds by Riemannian Dimensionality ReductionabstractMost computer vision problems entail data that reside on Riemannian manifolds. Canonical correlation analysis (CCA) is a powerful method that captures correlations between any two sets of matrices. In this paper, we propose a framework for a supervised CCA of manifold-based data. This framework aims to find the optimal dimensionality reduction maps that maximize the discriminative power of any classifier in the reduced dimensional space and the correlation between the projected sets. This allows to incorporate the CCA into a classifier that analyzes multichannel or multimodal data on separate manifolds. The proposed method is evaluated on the challenging task of segmenting cardiac adipose tissues on fat-water (2-channel) magnetic resonance images. Faezeh Fallah, Bin Yang 0009 |
ICASSP | 2 |
| 2020 | Continual Learning Through One-Class Classification Using VAEabstractArtificial neural networks (ANNs) suffer from catastrophic forgetting, a sharp decrease in performance on previously learned tasks, when trained on a new task without constant rehearsal. In this paper, we propose a new method for overcoming this phenomenon based on one-class classification. It is not only able to incrementally learn new but also detect unknown classes. This is a desirable property, since it enables the detection of new and unknown classes in a stream of data and adaption to a changing environment. Experiments on commonly used continual learning setups show competitive results and verify the concept. Felix Wiewel, Andreas Brendle, Bin Yang 0009 |
ICASSP | 3 |
| 2020 | ipA-MedGAN: Inpainting of Arbitrary Regions in Medical ImagingabstractLocal deformations in medical modalities are common phenomena due to a multitude of factors such as metallic implants or limited field of views in magnetic resonance imaging (MRI). Completion of the missing or distorted regions is of special interest for automatic image analysis frameworks to enhance post-processing tasks such as segmentation or classification. In this work, we propose a new generative framework for medical image inpainting, titled ipA-MedGAN. It bypasses the limitations of previous frameworks by enabling inpainting of arbitrary shaped regions without a prior localization of the regions of interest. Thorough qualitative and quantitative comparisons with other inpainting and translational approaches have illustrated the superior performance of the proposed framework for the task of brain MR inpainting. Karim Armanious, Vijeth Kumar, Sherif Abdulatif, Tobias Hepp 0002, Sergios Gatidis, Bin Yang 0009 |
ICIP | 6 |
| 2020 | On-manifold Adversarial Data Augmentation Improves Uncertainty CalibrationabstractUncertainty estimates help to identify ambiguous, novel, or anomalous inputs, but the reliable quantification of uncertainty has proven to be challenging for modern deep networks. To improve uncertainty estimation, we propose On-Manifold Adversarial Data Augmentation or OMADA, which specifically attempts to generate challenging examples by following an on-manifold adversarial attack path in the latent space of an autoencoder that closely approximates the decision boundaries between classes. On a variety of datasets and for multiple network architectures, OMADA consistently yields more accurate and better calibrated classifiers than baseline models, and outperforms competing approaches such as Mixup, as well as achieving similar performance to (at times better than) postprocessing calibration methods such as temperature scaling. Variants of OMADA can employ different sampling schemes for ambiguous on-manifold examples based on the entropy of their estimated soft labels, which exhibit specific strengths for generalization, calibration of predicted uncertainty, or detection of out-of-distribution inputs. Kanil Patel, William Beluch, Dan Zhang 0003, Michael Pfeiffer 0001, Bin Yang 0009 |
ICPR | 5 |
| 2019 | Adversarial Inpainting of Medical Image ModalitiesabstractNumerous factors could lead to partial deteriorations of medical images. For example, metallic implants will lead to localized perturbations in MRI scans. This will affect further post-processing tasks such as attenuation correction in PET/MRI or radiation therapy planning. In this work, we propose the inpainting of medical images via Generative Adversarial Networks (GANs). The proposed framework incorporates two patch-based discriminator networks with additional style and perceptual losses for the inpainting of missing information in realistically detailed and contextually consistent manner. The proposed framework outperformed other natural image inpainting techniques both qualitatively and quantitatively on two different medical modalities. Karim Armanious, Youssef Mecky, Sergios Gatidis, Bin Yang 0009 |
ICASSP | 4 |
| 2019 | Continual Learning for Anomaly Detection with Variational AutoencoderabstractDetecting anomalies using a variational autoencoder (VAE) suffers from catastrophic forgetting when trained on a continually growing set of normal data where only the most recently added data is available. Solving this problem would allow the use of the VAE for anomaly detection in settings where it is difficult or even impossible to retain all normal data at the same time. We propose an efficient extension of a method for continual learning which alleviates catastrophic forgetting for anomaly detection using a VAE. We show on some anomaly detection problems that the definition of normal data can be continually expanded without requiring all previously seen data. Felix Wiewel, Bin Yang 0009 |
ICASSP | 2 |
| 2019 | Simultaneous Volumetric Segmentation of Vertebral Bodies and Intervertebral Discs on Fat-Water MR ImagesabstractFat-water magnetic resonance (MR) images allow automated noninvasive analysis of morphological properties and fat fractions of vertebral bodies (VBs) and intervertebral discs (IVDs) that constitute an important part of human biomechanical systems. In this paper, we propose a fully automated approach for simultaneously segmenting multiple VBs and IVDs on fat-water MR images without prior localization or geometry estimation. This method involved a hierarchical random forest (HRF) classifier and a hierarchical conditional random field (HCRF) that encoded a multi-resolution image pyramid based on a set of multiscale local and contextual features. The HRF classifier employed penalized multivariate linear discriminants and SMOTEBagging to handle limited and imbalanced training data with large feature dimension. The HCRF estimated optimum labels according to their spatial and hierarchical consistencies by using the layer-wise significant features determined over the trained HRF classifier. To handle variable sample numbers at different resolutions, resolution-specific hyperparameters were used. This method was trained and evaluated for segmenting 15 thoracic and lumbar VBs and their IVDs on fat-water MR images of a subset of a large cohort data set. It was further evaluated for segmenting seven IVDs of the lower spine on fat-water images of a public grand challenge. These evaluations revealed the comparable accuracy of this method with the state-of-the-art while requiring less computational burden due to a simultaneous localization and segmentation. Faezeh Fallah, Sven S. Walter, Fabian Bamberg, Bin Yang 0009 |
IEEE J. Biomed. Health Informatics | 4 |
| 2018 | Person Recognition Based on Micro-Doppler and Thermal Infrared Camera Fusion for FirefightingabstractThis paper examines the recognition of real persons, mirrored persons and other objects using thermal infrared (TIR) images and radar micro-Doppler (μ-D). Mirrored persons lead to confusion of firefighters, when only a TIR camera is used. However, mirrored persons exhibit the μ-D of the mirroring objects, hence radar can resolve this ambiguity. In this paper, multiple sensor fusion architectures are investigated for this classification task. The first approach uses an attention stage, where bounding boxes of candidates for real/mirrored persons are determined in TIR images. These bounding boxes are associated to the radar targets and subsequently classified. A joint classification of the radar μ-D and TIR image at measurement level is compared to a separate classification with subsequent combination (object level). Furthermore, a classification of the complete scene is proposed, omitting the TIR attention stage and data association. Experiments with real measurements are used for an evaluation of the presented approaches. Michael Ulrich, Thomas Hess, Sherif Abdulatif, Bin Yang 0009 |
FUSION | 4 |
| 2018 | Automatic Motion Artifact Detection for Whole-Body Magnetic Resonance ImagingabstractMagnetic resonance (MR) plays an important role in medical imaging. It can be flexibly tuned towards different applications for deriving a meaningful diagnosis. However, its long acquisition times and flexible parametrization make it on the other hand prone to artifacts which obscure the underlying image content or can be misinterpreted as anatomy. Patient-induced motion artifacts are still one of the major extrinsic factors which degrade image quality. In this work, an automatic reference-free motion artifact detection, including localization and quantification, is proposed which can be used prospectively as quality control (e.g. scan adjustment) or retrospectively as quality control (e.g. supported diagnosis). The detection is achieved via trained convolutional neural networks (CNN). This study focuses on investigating the optimal CNN architecture and required training set composition to derive a general and robust network for MR motion artifact detection. In a volunteer cohort an average accuracy of 91% was achieved. Thomas Kustner, Marvin Jandt, Annika Liebgott, Lukas Mauch, Petros Martirosian, Fabian Bamberg, Konstantin Nikolaou, Sergios Gatidis, Fritz Schick, Bin Yang 0009 |
ICASSP | 10 |
| 2018 | Automated Detection of High FDG Uptake Regions in CT ImagesabstractCombined PET-CT scan is an important diagnostic tool in modern medicine, e.g. for staging or treatment planning in the field of oncology. Especially in small structures, like a tumour, textural variations visible in a PET image are not visually recognizable within a CT scan from the same region. Thus, both modalities are necessary for diagnosis. Since both techniques expose the patient to radiation, it would be desirable to get the same information about metabolic activity contained in the PET image from a CT scan only. To investigate the relationship between both imaging modalities, we propose a machine learning approach to automatically identify regions in a CT scan corresponding to areas with high FDG uptakes in a PET image. Annika Liebgott, Florian Liebgott, Bin Yang 0009, Sergios Gatidis, Konstantin Nikolaou |
ICASSP | 3 |
| 2018 | Semantic Organ Segmentation in 3D Whole-Body MR ImagesabstractAutomated organ segmentation is a prerequisite for efficient analysis of MR data in large cohorts with thousands of participants. The feasibility and generalizability of previously proposed methods has mostly been demonstrated in smaller cohorts. The aim of this work is to implement and validate automated semantic 3D segmentation of liver and spleen on multi-contrast MR data of the body trunk which were acquired in a large epidemiological imaging study with the objective to provide a robust and general setup in a setting of limited training data. Liver and spleen were manually segmented in 173 MR images by an experienced radiologist, providing labeled ground-truth. Varying amount of training datasets were randomly chosen to train a convolutional neural network (CNN)-based segmentation with 4-fold patient-leave-out cross-validation and compared against a Random Forest (RF)-based segmentation. Validation amongst participants revealed high accuracies of 99.7%/99.9% for liver/spleen-segmentation with superiority of CNN to RF. In conclusion, automated semantic organ segmentation is feasible in a robust and general setup. Thomas Kustner, Sarah Müller, Marc Fischer 0003, Jakob Weiß, Konstantin Nikolaou, Fabian Bamberg, Bin Yang 0009, Fritz Schick, Sergios Gatidis |
ICIP | 7 |
| 2018 | Binary Segmentation Based Class Extension in Semantic Image Segmentation Using Convolutional Neural NetworksabstractWe deal with semantic image segmentation using deep convolutional neural networks (CNNs) and propose to extent a well-trained model to capture more classes. Because ground truth is very expensive in such a pixel-wise classification task, we avoid the manual annotation of the new classes by using a binary segmentation model to support the class extension. We use soft targets (probabilities), reuse and distill knowledge from the old segmentation model, and fuse information from the binary model to regularize the training of a new model with extended classes. In the experiments, we show that our method outperforms two other methods and improves the accuracy of small object classes. Moreover, our method is robust and more capable of tolerating bad binary models. Chunlai Wang, Lukas Mauch, Bin Yang 0009 |
ICIP | 4 |
| 2018 | Ensemble Learning to EEG-Based Brain Computer Interfaces with Applications on P300-SpellersabstractBrain-Computer Interfaces (BCI) are systems in which the electrical activity of an animal brain becomes the main controller of an external electronic device capable of reading and processing electroencephalographic (EEG) signals. One early application of such systems is the attention-based spellers utilizing the P300 visually evoked potential in a framework known as the oddball paradigm. In this paper, we propose novel variants of machine learning model ensembles in addressing the task of P300 detection and attended target recognition in attention-based speller systems. Proposed ensembles adopted Bootstrap aggregation (Bagging) of calibrated Support Vector Machines (SVM) as well as data-driven learners. The latter is dominantly represented by Convolutional Neural Networks (CNN) with several variants, including what is referred to as Inception, Xception, and Interleaved Group Convolutions (IGC) modules. The proposed models are evaluated on a publicly available EEG dataset developed specifically for BCI applications and published in public contests, namely the 2nd dataset of the 3rd BCI competitions. The proposed models consistently outperform all previous works on the same dataset and show the highest 5- and 15-trial recognition rates of 76.5% and 98.5%, respectively, for both subjects in the dataset jointly. Additionally, we introduce in this work the first study on inter-subject evaluation under similar training protocols, reaching between 30%-40% recognition rates for either subject. We further investigate the effect of reducing the training data on the performance of the proposed models showing possibilities of reduced training time for a target recognition rate. Karim Said Barsim, Wangbo Zheng, Bin Yang 0009 |
SMC | 3 |
| 2017 | Selecting optimal layer reduction factors for model reduction of deep neural networksabstractDeep neural networks (DNN) achieve very good performance in many machine learning tasks, but are computationally very demanding. Hence, there is a growing interest on model reduction methods for DNN. Model reduction allows to reduce the number of computations needed to evaluate a trained DNN without a significant performance degradation. In this paper, we study layerwise reduction methods that reduce the number of computations in each layer independently. We consider the pruning and low-rank approximation method for model reduction. Up to now, often a constant reduction factor is used in all layers. In this paper, we show that a non-uniform allocation of reduction factors to different layers can greatly improve the performance of the reduced DNN. For this purpose, we select the optimal layer reduction factors in terms of an optimization problem. Experiments on three different benchmark datasets demonstrate the superior performance of our method. Lukas Mauch, Bin Yang 0009 |
ICASSP | 2 |
| 2017 | A novel layerwise pruning method for model reduction of fully connected deep neural networksabstractDeep neural networks (DNN) are powerful models for many pattern recognition tasks, yet they tend to have many layers and many neurons resulting in a high computational complexity. This limits their application to high-performance computing platforms. In order to evaluate a trained DNN on a lower-performance computing platform like a mobile or embedded device, model reduction techniques which shrink the network size and reduce the number of parameters without considerable performance degradation performance are highly desirable. In this paper, we start with a trained fully connected DNN and show how to reduce the network complexity by a novel layerwise pruning method. We show that if some neurons are pruned and the remaining parameters (weights and biases) are adapted correspondingly to correct the errors introduced by pruning, the model reduction can be done almost without performance loss. The main contribution of our pruning method is a closed-form solution that only makes use of the first and second order moments of the layer outputs and, therefore, only needs unlabeled data. Using three benchmark datasets, we compare our pruning method with the low-rank approximation approach. Lukas Mauch, Bin Yang 0009 |
ICASSP | 2 |
| 2017 | Unsupervised image segmentation using convolutional autoencoder with total variation regularization as preprocessingabstractConventional unsupervised image segmentation methods use color and geometric information and apply clustering algorithms over pixels. They preserve object boundaries well but often suffer from over-segmentation due to noise and artifacts in the images. In this paper, we contribute on a preprocessing step for image smoothing, which alleviates the burden of conventional unsupervised image segmentation and enhance their performance. Our approach relies on a convolutional autoencoder (CAE) with the total variation loss (TVL) for unsupervised learning. We show that, after our CAE-TVL preprocessing step, the over-segmentation effect is significantly reduced using the same unsupervised image segmentation methods. We evaluate our approach using the BSDS500 image segmentation benchmark dataset and show the performance enhancement introduced by our approach in terms of both increased segmentation accuracy and reduced computation time. We examine the robustness of the trained CAE and show that it is directly applicable to other natural scene images. Chunlai Wang, Bin Yang 0009, Yiwen Liao |
ICASSP | 2 |
| 2017 | MR-based respiratory and cardiac motion correction for PET imaging
Thomas Kustner, Martin Schwartz, Petros Martirosian, Sergios Gatidis, Ferdinand Seith, Christopher Gilliam, Thierry Blu, Hadi Fayad, Dimitris Visvikis, Fritz Schick, Bin Yang 0009, Nina F. Schwenzer |
Medical Image Anal. | 11 |
| 2016 | A novel feedforward noise shaping for word-length reductionabstractIn this paper, we present a novel approach to shape the quantization noise during word-length reduction. In comparison to the traditional feedback noise shaping, our approach is feedforward and thus inherently stable. It can achieve one or multiple frequency notches in the quantization noise spectrum with controlled notch width and notch depth while keeping the out-of-band noise level lower than the feedback noise shaping. Results from a digital communication example demonstrate the performance of this new noise shaping method. Mohamed Ibrahim 0001, Bin Yang 0009, Andreas Menkhoff |
ICASSP | 2 |
| 2016 | Active learning for magnetic resonance image quality assessmentabstractIn medical imaging, the acquired images are usually analyzed by a human observer and rated with respect to a diagnostic question. However, this procedure is time-demanding and expensive. Further more, the lack of a reference image makes this task challenging. In order to support the human observer in assessing image quality and to ensure an objective evaluation, we extend in this paper our previous no-reference magnetic resonance (MR) image quality assessment system with an active learning loop to reduce the amount of necessary labeled training data. We employ two different active learning query strategies based on uncertainty sampling. Since the classification task is performed on 2D image slices, but the human observer labels complete 3D image volumes, we present a method to select representative 3D images instead of independant 2D image slices. The performance is evaluated on in-vivo MR image data. Annika Liebgott, Thomas Kustner, Sergios Gatidis, Fritz Schick, Bin Yang 0009 |
ICASSP | 5 |
| 2016 | A novel DNN-HMM-based approach for extracting single loads from aggregate power signalsabstractThis paper presents a new supervised approach to extract the power trace of individual loads from single channel aggregate power signals in non-intrusive load monitoring (NILM) systems. Recent approaches to this source separation problem are based on factorial hidden Markov models (FHMM). Drawbacks are the needed knowledge of HMM models for all loads, what is infeasible for large buildings, and the large combinatorial complexity. Our approach trains HMM with two emission probabilities, one for the single load to be extracted and the other for the aggregate power signal. A Gaussian distribution is used to model observations of the single load whereas observations of the aggregate signal are modeled with a Deep Neural Network (DNN). By doing so, a single load can be extracted from the aggregate power signal without knowledge of the remaining loads. The performance of the algorithm is evaluated on the Reference Energy Disaggregation (REDD) dataset. Lukas Mauch, Bin Yang 0009 |
ICASSP | 2 |
| 2016 | Saliency-guided object proposal for refined salient region detectionabstractAutomatic detection of visually salient regions across images is useful in many applications. Traditional methods predict saliency values of pixels in a bottom-up fashion and use low-level features. Recent researches demonstrate that high-level information are also important for salient region detection. In this paper, we propose a novel approach of integrating object-level information and bottom-up saliency model. By using saliency-guided object proposals, we implicitly remove the noisy salient regions and produce a refined saliency map. We experimentally show that our approach boosted the performance of bottom-up saliency models and performs favorably against the state-of-the-art methods. Chunlai Wang, Bin Yang 0009 |
VCIP | 2 |
| 2016 | MR Image Reconstruction Using a Combination of Compressed Sensing and Partial Fourier Acquisition: ESPReSSoabstractA Cartesian subsampling scheme is proposed incorporating the idea of PF acquisition and variable-density Poisson Disc (vdPD) subsampling by redistributing the sampling space onto a smaller region aiming to increase k-space sampling density for a given acceleration factor. Especially the normally sparse sampled high-frequency components benefit from this sampling redistribution, leading to improved edge delineation. The prospective subsampled and compacted k-space can be reconstructed by a seamless combination of a CS-algorithm with a Hermitian symmetry constraint accounting for the missing part of the k-space. This subsampling and reconstruction scheme is called Compressed Sensing Partial Subsampling (ESPReSSo) and was tested on in-vivo abdominal MRI datasets. Different reconstruction methods and regularizations are investigated and analyzed via global (intensity-based) and local (region-of-interest and line evaluation) image metrics, to conclude a clinical feasible setup. Results substantiate that ESPReSSo can provide improved edge delineation and regional homogeneity for multidimensional and multi-coil MRI datasets and is therefore useful in applications depending on well-defined tissue boundaries, such as image registration and segmentation or detection of small lesions in clinical diagnostics. Thomas Kustner, Christian Würslin, Sergios Gatidis, Petros Martirosian, Konstantin Nikolaou, Nina F. Schwenzer, Fritz Schick, Bin Yang 0009 |
IEEE Trans. Medical Imaging | 8 |
| 2015 | Combining Compressed Sensing with motion correction in acquisition and reconstruction for PET/MRabstractIn the field of oncology, simultaneous Positron-Emission-Tomography/Magnetic Resonance (PET/MR) scanners offer a great potential for improving diagnostic accuracy. However, to achieve a high Signal-to-Noise Ratio (SNR) for an accurate lesion detection and quantification in the PET/MR images, one has to overcome the induced respiratory motion artifacts. The simultaneous acquisition allows performing a MR-based non-rigid motion correction of the PET data. It is essential to acquire a 4D (3D + time) motion model as accurate and fast as possible to minimize additional MR scan time overhead. Therefore, a Compressed Sensing (CS) acquisition by means of a variable-density Gaussian subsampling is employed to achieve high accelerations. Reformulating the sparse reconstruction as a combination of the inverse CS problem with a non-rigid motion correction improves the accuracy by alternately projecting the reconstruction results on either the motion-compensated CS reconstruction or on the motion model optimization. In-vivo patient data substantiates the diagnostic improvement. Thomas Kustner, Christian Würslin, Bin Yang 0009 |
ICASSP | 4 |
| 2015 | On the spectral growth of the polar representation of communication signalsabstractThe traditional modulator in digital communication systems is the quadrature (IQ) modulator using the Cartesian representation of a complex baseband signal. In the last years, the polar transmitter has been becoming an attractive alternative due to its significantly increased energy efficiency. It uses a polar representation of the baseband signal before transmission. The result of the changed signal representation is an internal spectral growth of the polar signals whose understanding is fundamental for the design of modern polar transmitters. This papers gives a mathematical analysis of the spectral growth of the polar signals. Bin Yang 0009 |
ICASSP | 1 |
| 2014 | A theoretical study of the statistical and spectral properties of polar transmitter signalsabstractThe polar transmitter is an increasingly interesting architecture for modern wireless communication devices. Several studies have been done investigating optimal hardware architectures with minimal power consumption and space. In this paper, we present a mathematical analysis of the signals in a polar transmitter. We study the statistical and spectral properties of the amplitude and phase of the baseband signal and their relationship to influencing parameters in the transmitter like modulation and pulse shaping. This knowledge helps us to better understand the potential and limitation of a polar transmitter and facilitates future research. Mohamed Ibrahim 0001, Bin Yang 0009 |
ISCAS | 2 |
| 2014 | Eyelid-based driver state classification under simulated and real driving conditionsabstractOn account of the increase in vehicle accidents due to driver drowsiness over the last years, the development of reliable drowsiness assistant systems by a reference drowsiness measure is highlighted. Since eyelid features have shown acceptable correlation with driver vigilance in driving simulators, this study focuses on 18 blink features of 43 subjects collected by electrooculography under both simulated and real driving conditions during 67 hours of driving. We have assessed the driver state by artificial neural network, support vector machine and k-nearest neighbors classifiers for both binary and multi-class cases. There, binary classifiers are trained both subject-independent and subject-dependent to address the generalization aspects of the results for unseen data. The drawback of driving simulators in comparison to real driving is also discussed and to this end we have performed a data reduction approach as a remedy. For the binary driver state prediction (awake vs. drowsy) by eyelid features, we have attained an average detection rate of 82% by each classifier separately. For 3-class classification (awake vs. medium vs. drowsy), however, the result was only 66%, possibly due to inaccurate self-rated vigilance states. Parisa Ebrahim, Amira Abdellaoui, Wolfgang Stolzmann, Bin Yang 0009 |
SMC | 4 |
| 2013 | Fast and reliable TDOA assignment in multi-source reverberant environmentsabstractThe localization of acoustic sources based on Time Difference of Arrivals (TDOA) is very vulnerable in reverberant environments. In this paper, we propose a method to synthesize fully and partially consistent TDOA combinations. They fulfill the zero cyclic sum condition along all loops, which is a necessary condition for assigning TDOAs to an acoustic source. Our method is based on an efficient search of all sets of maximally connected compatible fundamental loops in a compatibility-conflict graph. We both prove the correctness of our algorithm and show some experimental results. Martin Kreißig, Bin Yang 0009 |
ICASSP | 2 |
| 2013 | Segmentation of magnetic resonance images in presence of severe intensity inhomogeneitiesabstractIn high-field whole body magnetic resonance imaging (MRI), images usually suffer from intensity inhomogeneities. The BC-FAT (bias correction by fitting of adipose tissue intensity) algorithm can compensate for this; however, it is limited to images containing only one object, e.g. the torso. In this paper, we present a method, which extends the BC-FAT algorithm to images containing multiple objects and thus to cross-sectional images of the whole body. This is achieved by an algorithm for the robust and fully automated object detection in MR images using the Hough transform and a modified k-means clustering. We also present a two-scale approach for active contours in order to eliminate the need of object size dependent parametrization for BC-FAT. Florian Liebgott, Christian Würslin, Bin Yang 0009 |
ICASSP | 3 |
| 2013 | Colocated MIMO radar: Cramer-Rao bound and optimal time division multiplexing for DOA estimation of moving targetsabstractMultiple-Input-Multiple-Output (MIMO) radars with colocated transmit and receive antennas offer the advantage of a larger (virtual) aperture compared to a conventional Single-Input-Multiple-Output (SIMO) radar. Hence a higher accuracy of the estimated direction of arrival (DOA) of a target can be achieved. In general, the accuracy of DOA estimators decreases in a MIMO radar if the target moves relative to the radar, because the motion causes an unknown phase change of the baseband signal due to the Doppler effect. We compute the Cramer-Rao bound (CRB) of DOA estimation of a non-stationary target for a MIMO radar with colocated antennas for a general time division multiplexing (TDM) scheme. This allows a quantitative comparison of different MIMO and SIMO radars. Moreover, we derive an optimal TDM scheme such that the CRB is as small as in the stationary case. The results are confirmed by simulations. Kilian Rambach, Bin Yang 0009 |
ICASSP | 2 |
| 2013 | Maximum discriminant margin transform of discriminant functionsabstractMany current multiclass classification approaches can be described by a set of discriminant functions, where the label of the class with the largest discriminant function is chosen as the best prediction. We develop an affine transform called maximum discriminant margin (MDM), which can use independently estimated discriminant functions to solve multiclass classification problems. Fabian Schmieder, Bin Yang 0009 |
ICASSP | 2 |
| 2013 | Eye Movement Detection for Assessing Driver Drowsiness by ElectrooculographyabstractMany studies show that driver drowsiness is one of the main reasons for road accidents. To prevent such car crashes, systems are needed to monitor and characterize the driver based on the driving information. In order to have highly reliable assistant systems, reference drowsiness measurements are required. Among different physiological measures, previous studies have introduced driver eye movements, particularly blinking, as a measure with high correlation to drowsiness. Hence, in this study, eye movements of 14 drivers have been observed using electrooculography (EOG) at the moving-base driving simulator of Mercedes Benz to assess driver drowsiness. Based on the measured signals, an adaptive detection approach is introduced to simultaneously detect not only eye blinks, but also other driving-relevant eye movements such as saccades and micro sleep events. Moreover, in spite of the fact that drowsiness influences eye movement patterns, the proposed algorithm distinguishes between the often-confused driving-related saccades and decreased amplitude blinks of a drowsy driver. The evaluation of results shows that the presented detection algorithm outperforms common methods so that eye movements are detected correctly during both awake and drowsy phases. Parisa Ebrahim, Wolfgang Stolzmann, Bin Yang 0009 |
SMC | 3 |
| 2012 | An efficient algorithm for the synthesis of fully consistent graphsabstractIn this paper we present an efficient algorithm for the synthesis of fully consistent graphs. A consistent graph is a graph whose cyclic sum of edge weights along all loops is zero. It plays an important role in many sensor array processing applications like Time Difference of Arrival (TDOA) based source localization. By applying the concept of fundamental loops, a linearly independent basis of the loop space of the graph, our algorithm is able to find all consistent sets of edge weights for the full graph efficiently. Martin Kreißig, Bin Yang 0009 |
ICASSP | 2 |
| 2012 | Identification and validation of lateral driver models on experimentally induced driving behaviorabstractThis paper presents a real-world driving experiment with aim on controlled variation of steering and lane keeping behavior and investigates the ability of three common driver models to distinguish variations in driving performance. Nine drivers executed a lane keeping task with visual occlusion of the upper or lower field of view restraining them to near or far road scene information. Three common driver models are applied to replicate driving behavior. An autoregressive model with exogeneous input (ARX) is identified using vehicle lateral lane deviation as input and steering wheel angle as output. Two output error models are identified using vehicle heading deviation angles with respect to near and far preview points as respective inputs and steering wheel angle as output. The results show that the driving behaviors induced in the experiment are significantly different in terms of lane keeping performance. In simulations, the output error models exhibit advantages over the ARX model in capturing driving behavior. However, the model natural frequency and the model simulation error show weak performance in discerning this varying driving behavior and are largely determined by track effects. Peter Hermannstadter, Bin Yang 0009 |
SMC | 2 |
| 2011 | On the relation between ICA and MMSE based source separationabstractThis paper aims at deriving a relationship between minimum mean square error (MMSE) based source separation and independent component analysis (ICA) based on the Kullback-Leibler divergence (KLD) for a linear noisy mixing model. Starting from a description of the demixing task and two well-known solutions, inverse mixing matrix and MMSE solution, we derive an analytic expression for the demixing matrix of KLD-based ICA in the presence of noise. The derivation is done by using a perturbation analysis valid for small noise variance. Furthermore, we provide an analytic expression for the mean square error (MSE) of the demixed signals using KLD-based ICA. We show that for a wide range of the shape parameter of the generalized Gaussian distribution (GGD), the MSE of KLD-based ICA is very close to the MMSE. Simulations verify this and show that in practice the variance of the ICA estimation due to limited amount of data also influences the achievable performance. Benedikt Loesch, Bin Yang 0009 |
ICASSP | 2 |
| 2011 | Recursive estimation of room impulse responses with energy conservation constraintsabstractThis paper considers the problem of constrained tracking the time-varying room impulse response of a source/microphone pair. The constraint which is used to improve the performance stems from the energy conservation that has to hold for real-world impulse responses. We consider three different recursive estimators and compare their performance with the recursive weighted least squares algorithm which does not take the constraint into account. The simulation results show that exploiting this constraint decreases the mean squared error and is thus interesting for applications, especially in the low SNR regime. Stefan Uhlich, Bin Yang 0009 |
ICASSP | 2 |
| 2011 | An introduction to consistent graphs and their signal processing applicationsabstractIn this paper, we present an introduction into the synthesis of consistent graphs. A consistent graph is a directed weighted graph whose cyclic sum of edge weights along all loops is zero. We show that this novel graph theoretical framework is a useful tool for various signal processing tasks, in particular sensor fusion applications. We introduce the concept and notation of consistent graphs, discuss their properties, address different synthesis issues, and describe some signal processing applications. Bin Yang 0009, Martin Kreißig |
ICASSP | 1 |
| 2010 | On the robustness of the multidimensional state coherence transform for solving the permutation problem of frequency-domain ICAabstractA common problem in frequency domain independent component analysis (ICA) is the so called permutation problem which arises due to the independent demixing in each frequency bin. This paper evaluates the robustness of a an extension of a recently proposed method for permutation correction based on the time difference of arrival (TDOA) of the sources. First, we discuss the permutation problem, review the proposed method, and give an intuitive model to predict the number of permutations. Then the theoretical performance using perfect knowledge of the TDOAs of the sources as well as the practical performance using the TDOAs estimated from a multidimensional state coherence transform (SCT) are evaluated through extensive simulations. In our experiments, ICA with SCT based permutation correction outperforms Independent Vector Analysis (IVA). Benedikt Loesch, Francesco Nesta, Bin Yang 0009 |
ICASSP | 3 |
| 2010 | Camera-based drowsiness reference for driver state classification under real driving conditionsabstractExperts assume that accidents caused by drowsiness are significantly under-reported in police crash investigations (1-3%). They estimate that about 24-33% of the severe accidents are related to drowsiness. In order to develop warning systems that detect reduced vigilance based on the driving behavior, a reliable and accurate drowsiness reference is needed. Studies have shown that measures of the driver's eyes are capable to detect drowsiness under simulator or experiment conditions. In this study, the performance of the latest eye tracking based in-vehicle fatigue prediction measures are evaluated. These measures are assessed statistically and by a classification method based on a large dataset of 90 hours of real road drives. The results show that eye-tracking drowsiness detection works well for some drivers as long as the blinks detection works properly. Even with some proposed improvements, however, there are still problems with bad light conditions and for persons wearing glasses. As a summary, the camera based sleepiness measures provide a valuable contribution for a drowsiness reference, but are not reliable enough to be the only reference. Fabian Friedrichs, Bin Yang 0009 |
Intelligent Vehicles Symposium | 2 |
| 2010 | Automatic extrinsic camera self-calibration based on homography and epipolar geometryabstractIn this paper we present a method to calibrate the extrinsic parameters of a monocular camera on a moving vehicle. The method is based on a homography between two camera shots. Therefore, only the road surface has to be visible in the pair of images. A reasonable definition of the vehicle coordinate system in combination with the use of epipolar geometry reduces the complexity to parameterize the underlying homography matrix. The extrinsic parameters are determined analytically by two correctly matched feature points located on the road surface. The final parameter set is determined by a recursive filter which considers various estimates over time. Results with a real-world video sequence indicate that the method is comparable to classical offline calibration techniques using objects of known geometry. Michael Miksch, Bin Yang 0009, Klaus Zimmermann |
Intelligent Vehicles Symposium | 2 |
| 2010 | Motion compensation for obstacle detection based on homography and odometric data with virtual camera perspectivesabstractIn this paper we present a method to compensate the image motion of a monocular camera on a moving vehicle in order to detect obstacles. Due to the camera motion, the road surface induces a characteristic image motion between two camera shots. The motion of the camera is determined by the use of odometric data received from the CAN-bus, and the position and orientation of the road is continuously estimated with camera self-calibration. This all leads to a motion field which is predicted based on homography. To prevent the drawbacks of the real camera perspective, different virtual camera perspectives are presented in combination with motion compensation. Possible virtual perspectives are the bird's eye view and image rectification. In addition, a non-linear camera model is used which does not limit the range of obstacle detection to a certain distance and efficiently uses the available image information. Michael Miksch, Bin Yang 0009, Klaus Zimmermann |
Intelligent Vehicles Symposium | 2 |
| 2009 | Online blind source separation based on time-frequency sparsenessabstractRecently, blind source separation (BSS) has been proposed to separate signals recorded by a microphone array in a reverberant environment. This paper deals with BSS of a time-varying number of moving sources, which often occurs in practical situations. We develop two online algorithms based on time-frequency (TF) sparseness that are able to deal with moving sources: A block online algorithm that estimates the number of sources and a gradient based online algorithm with prespecified maximum number of sources. Both algorithms are evaluated in simulations and real-world scenarios and show good separation performance. Benedikt Loesch, Bin Yang 0009 |
ICASSP | 2 |
| 2009 | MMSE estimation in a linear signal model with ellipsoidal constraintsabstractThe estimation of an unknown parameter vector in a Gaussian linear model is studied in this paper. Two different cases are analyzed: the parameter vector is assumed to lie either in or on a given ellipsoid. The best estimator in terms of the mean squared error is derived. The performance of this estimator is analyzed and compared with the ordinary least squares, the constrained least squares and the linear minimax approach. Stefan Uhlich, Bin Yang 0009 |
ICASSP | 2 |
| 2008 | A generalized optimal correlating transform for multiple description coding and its theoretical analysisabstractThis paper considers a coding scheme for data transmission over erasure channels which is also known as multiple description coding. The LMMSE prefilter method of Romano [1] is reviewed and generalized to allow three different operational modes of the prefilter. They include the possibility to decrease or increase the number of descriptions to be transmitted. We derive explicitly the Hessian matrix for an efficient calculation of the prefilter. We also study the properties of the distortion measure theoretically. Stefan Uhlich, Bin Yang 0009 |
ICASSP | 2 |
| 2008 | A study of inverse short-time fourier transformabstractIn this paper, we study the inverse short-time Fourier transform (STFT). We propose a new vector formulation of STFT. We derive a family of inverse STFT estimators and a least squares one. We discuss their relationship and compare their performance with respect to both additive and multiplicative modifications to STFT. The influence of window, overlap, and zero-padding are investigated as well. Bin Yang 0009 |
ICASSP | 1 |
| 2008 | Disambiguation of TDOA Estimation for Multiple Sources in Reverberant EnvironmentsabstractThis paper presents a novel approach to estimate the time difference of arrival (TDOA) for multiple sources in reverberant environments. It resolves ambiguities in TDOA estimation caused by multipath propagation and multiple sources. By exploiting two TDOA constraints, the raster condition and the zero cyclic sum condition, we are able to identify and reject the echo path TDOAs and to assign the direct path TDOAs correctly to different sources. For the latter purpose, an efficient algorithm for the synthesis of approximately consistent TDOA graphs has been developed. A real experiment demonstrates the superior performance of our algorithms. Jan Scheuing, Bin Yang 0009 |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Efficient Synthesis of Approximately Consistent Graphs for Acoustic Multi-Source LocalizationabstractDue to ambiguities and estimation errors, combining time differences of arrival (TDOAs) for simultaneous localization of multiple acoustic sources is a challenging task. This paper studies this problem under the framework of consistent graphs and proposes an efficient algorithm to determine TDOAs originating from the source. Jan Scheuing, Bin Yang 0009 |
ICASSP (4) | 2 |
| 2007 | Different Sensor Placement Strategies for TDOA Based LocalizationabstractThis paper studies different optimization strategies for the sensor placement in source localization by using time differences of arrival. It continues the works in (B. Yang et al., 2005, 2006) and gives answers to some open questions there. In particular, we discuss the relationship between maximum Fisher information matrix, minimum Cramer-Rao bound, spherical codes, uniform angular arrays, and Platonic solids as well as their roles in optimizing the sensor placement. Various new optimum sensor array geometries are given. Bin Yang 0009 |
ICASSP (2) | 1 |
| 2006 | Disambiguation of TDOA Estimates in Multi-Path Multi-Source Environments (DATEMM)abstractOne major problem of time delay estimation for acoustic localization in multi-source reverberant environments is the ambiguity in identifying out of many peaks of generalized cross-correlation the desired time differences of arrival (TDOAs) caused by direct paths and in assigning them correctly to individual sources. In this paper, we propose a novel geometrically motivated approach "Disambiguation of TDOA estimates in multi-path multi-source environments" (DATEMM). It utilizes additional information from the auto-correlation of sensor signals and a zero TDOA sum condition to suppress spurious TDOA estimates. Furthermore, this method can be used as an add-on module to improve the robustness of any existing TDOA estimation method. Jan Scheuing, Bin Yang 0009 |
ICASSP (4) | 2 |
| 2006 | A Theoretical Analysis of 2D Sensor Arrays for TDOA Based LocalizationabstractIn source localization from time difference of arrival, the impact of the sensor array geometry to the localization accuracy is not well understood yet. A first rigorous analysis can be found in [1]. It derived sufficient and necessary conditions for optimum array geometry in terms of minimum CramerRao bound. This paper continues the above work and studies theoretically the localization accuracy of two-dimensional sensor arrays. It addresses different issues: a) optimum vs. uniform angular array b) near-field vs. far-field array c) using all sensor pairs vs. those with a common reference sensor as required from spherical position estimators. The paper ends up with some new insights into the sensor placement problem. Bin Yang 0009, Jan Scheuing |
ICASSP (4) | 1 |
| 2005 | Cramer-Rao bound and optimum sensor array for source localization from time differences of arrivalabstractThis paper presents a theoretical analysis of the Cramer-Rao lower bound for source localization from time differences of arrival. We derive properties of the Cramer-Rao bound and design optimum sensor arrays which minimize the bound. Bin Yang 0009, Jan Scheuing |
ICASSP (4) | 1 |
| 1996 | Convergence analysis of the subspace tracking algorithms PAST and PASTdabstractWe prove the asymptotic convergence of the subspace tracking (stochastic approximation) algorithms PAST and PASTd. We also present new results about their asymptotic convergence rate. First we review the algorithms. Then we derive the corresponding ordinary differential equation (ODE). The convergence behaviour established by studying the asymptotically stable equilibrium states of the ODE, and the computer simulation results are shown. Bin Yang 0009 |
ICASSP | 1 |
| 1996 | Asymptotic distribution of recursive subspace estimatorsabstractWe derive the asymptotic distribution of recursive subspace estimators. In particular, we study the PAST algorithm for tracking the signal subspace and the Oja (1982) rule for updating the eigenvector corresponding to the largest eigenvalue. Both the decreasing gain and the constant gain case are considered. It turns out that their asymptotic distributions differ from that of the batch eigenvalue decomposition. The asymptotic rate of convergence is also addressed. Bin Yang 0009, Frank Gersemsky |
ICASSP | 1 |
| 1996 | Asymptotic convergence analysis of the projection approximation subspace tracking algorithms
Bin Yang 0009 |
Signal Process. | 1 |
| 1995 | A robust and efficient algorithm for source parameter estimationabstractA large number of array processing applications such as radar, sonar, etc. require the estimation of some parameters given the output of an array of sensors. Many high resolution methods for source parameter estimation are based on the eigen decomposition of the covariance matrix of the sensor output. The PASTd (projection approximation subspace tracking with deflation) algorithm [Yang, 1994] has been published for tracking both the signal subspace and its rank at a computational cost of order O(nr), where n is the number of sensors and r the number of sources to be detected. The present authors address the problem of tracking the physical parameters as direction, distance, etc. given the estimated signal subspace. All known parameter estimation methods as MUSIC, MinNorm or WSF are based on a different cost function which is minimized with respect to the desired parameters. Standard minimization methods as gradient or Newton's method fail to converge to the global minimum if the starting value is not close enough to the desired solution [Viberg, 1991]. The present authors introduce a new cost function which has to be minimized with respect to the parameters and an algorithm of low computational cost which is able to find the global minimum, starting from any initial value in all the experiments. Frank Gersemsky, Bin Yang 0009 |
ICASSP | 2 |
| 1995 | An extension of the PASTd algorithm to both rank and subspace trackingabstractIn this letter, we present an extension of the PASTd algorithm to both rank and signal subspace tracking. It has a low computational complexity O(nr), where n is the input vector length, and r denotes the signal subspace dimension. Its performance in tracking time-varying direction of arrival is comparable with that of the expensive eigenvalue decomposition and more robust than the O(n/sup 2/) rank revealing URV updating algorithm proposed by Stewart.> Bin Yang 0009 |
IEEE Signal Process. Lett. | 1 |
| 1994 | A new efficient subspace tracking algorithm based on singular value decompositionabstractA new algorithm for signal subspace tracking is presented. It is based on an approximated singular value decomposition using interlaced QR-updating and Jacobi plane rotations. By forcing the noise subspace to be spherical, the computational complexity of the algorithm is brought down to O(nr), where n is the problem dimension and r is the desired number of signal components. The algorithm lends itself for a very efficient systolic array implementation, resulting in a throughput of O(n/sup 0/). Simulations show that the frequency tracking capabilities of the new method are at least as good as those of the computationally much more expensive exact singular value decomposition.> Aleksandar Kavcic, Bin Yang 0009 |
ICASSP (4) | 2 |
| 1994 | An adaptive algorithm of linear computational complexity for both rank and subspace trackingabstractRank and subspace estimation is important in a variety of modern signal processing applications. In this paper we present a new approach for tracking both the rank and the signal subspace recursively. At arrival of each new sample, we update the eigenvectors spanning the signal subspace plus one or a fixed number of auxiliary eigenvectors, their corresponding eigenvalues, and an averaged noise eigenvalue. Then we apply information theoretic criteria to estimate the number of signals. The resulting adaptive algorithm has a computational complexity which is linearly proportional to the sample vector size n. In comparison to the URV based subspace tracking requiring O(n/sup 2/) operations, our approach is computationally simpler, easier to implement, and does not need user supplied tolerances. Simulation results show similar tracking performance of our algorithm to the URV updating and the exact eigenvalue decomposition.> Bin Yang 0009, Frank Gersemsky |
ICASSP (4) | 1 |
| 1993 | Subspace tracking based on the projection approach and the recursive least squares method
Bin Yang 0009 |
ICASSP (4) | 1 |
| 1993 | A rotation based multichannel least squares lattice algorithm for adaptive nonlinear filters
Bin Yang 0009 |
ISCAS | 1 |
| 1989 | On a systolic implementation and the numerical properties of a multiple constrained adaptive beamformerabstractAn efficient algorithm for the multiple linearly constrained adaptive beamformer is developed. It is based on rank-one updating and downdating Cholesky factorizations. These operations can be realized by a sequence of complex Givens and hyperbolic rotations. A linear systolic array using CORDIC processors as the processing elements for implementing the algorithm is presented. The authors also investigate the numerical properties of the algorithm.> Bin Yang 0009, Johann F. Böhme |
ICASSP | 1 |
| 1988 | Systolic implementation of a general adaptive array processing algorithmabstractThe authors present a general adaptive array processing algorithm based on the MVDR (minimum variance distortionless response) beamformer for multiple steering vectors and multiple frequencies. They show how they use a sequence of Givens rotations to update the beamformer output, and how they implement the algorithm by means of two linear systolic arrays. The pipelined CORDIC processor is proposed as the processing element to achieve a high data throughput and a straightforward implementation. The problem of complex arithmetic for frequency-domain data is discussed.> Bin Yang 0009, Johann F. Böhme |
ICASSP | 1 |