EDBT 2026 Demo / reviewers in the wild / expert
Min Bai
dblp:157/1666
· DBLP profile ↗
26ranked-venue papers
7as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Proposer-Agent-Evaluator (PAE): Autonomous Skill Discovery For Foundation Model Internet AgentsabstractA generalist foundation model agent needs to have a large and diverse skill repertoire, such as finding directions between two travel locations and buying specific items from the Internet. If each skill needs to be specified manually through a fixed set of human-annotated instructions, the agent’s skill repertoire will necessarily be limited due to the scalability of human-annotated instructions. In this work, we address this challenge by proposing Proposer-Agent-Evaluator (PAE), an effective learning system that enables foundation model agents to autonomously discover and practice skills in the wild. After a context-aware task proposer generates instructions based on website information, the agent policy attempts those tasks in the real world with resulting trajectories evaluated by an autonomous VLM-based success evaluator. The success evaluation serves as the reward signal for the agent to refine its policies through RL. We validate PAE on challenging vision-based web navigation, using both real-world and selfhosted websites from WebVoyager and WebArena. Our results show that PAE significantly improves the zero-shot generalization capability of VLM Internet agents (around 50% relative improvement) to both unseen tasks and websites. Qianlan Yang, Kaixiang Lin, Min Bai, Yu-Xiong Wang, Sergey Levine, Li Erran Li |
ICML | 4 |
| 2024 | ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
Siming Yan, Min Bai, Qixing Huang, Li Erran Li |
ECCV (61) | 2 |
| 2024 | Snapper: Accelerating Bounding Box Annotation in Object Detection Tasks with Find-and-Snap ToolingabstractObject detection tasks are central to the development of datasets and algorithms in computer vision and machine learning. Despite its centrality, object detection remains tedious and time-consuming due to the inherent interactions that are often associated with drawing precise annotations. In this paper, we introduce Snapper, an interactive and intelligent annotation tool that intercepts bounding box annotations as they’re drawn and “snaps” them to the nearby object edges in real-time. Through a mixed-design user study with 18 full-time annotators, we compare Snapper’s annotation mode to alternative modes of annotation and find that Snapper enables participants to complete object detection tasks 39% more quickly without diminishing annotation quality. Further, we find that participants perceive Snapper as a tool that is interactively intuitive, trustworthy, and helpful. We conclude by discussing the implications of our findings as they relate to augmenting annotators’ conventions for drawing annotations in practice. Alex C. Williams, Min Bai, Jonathan Buck, Tristan McKinney, Amy Rechkemmer, Koushik Kalyanaraman, Matthew Lease, Patrick Haffner, Li Erran Li |
IUI | 2 |
| 2024 | Efficient Convolutional Sparse Coding Based on Penalized Weighted Least Squares for Seismic Data DenoisingabstractRecently, convolutional sparse coding (CSC) has been successfully applied to seismic data denoising. CSC differs from traditional dictionary learning methods based on patching schemes in that it can directly process the whole data and capture correlations between local neighborhoods. However, the learned filters by CSC may contain inaccurate features resulting in the structure loss of data, and solving CSC problems has a heavy computational burden. To optimize these problems, we investigate the denoising accuracy of the CSC model and introduce CSC as a regularization term into the penalized weighted least-squares (PWLSs) framework. In particular, we design an effective method to solve the problem of updating sparse feature maps, which improves computational efficiency. Combining the above two points, we propose an efficient CSC (ECSC) model for random noise attenuation of seismic data. The numerical experiments on synthetic data and field data demonstrate that ECSC performs better than the K-singular value decomposition (K-SVD) algorithm, sequential generalized K-means (SGK) algorithm, and fast and flexible CSC (FF-CSC) in seismic data denoising performance and computational efficiency. Min Bai |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | One-Dimensional Dictionary Learning With Variational Sparse Representation for Single-Channel Seismic DenoisingabstractSeismic data acquired from the field inevitably suffer from noise pollution, which covers the useful signals and affects the reliability of subsequent seismic data processing and interpretation. Many two-dimensional (2D) multi-channel seismic denoising methods depend on the assumption that the receiver array is spatial coherent in field microseismic data acquisition, which limits their performance when dealing with field data. However, the single-channel methods are more flexible when faced with real microseismic data because they do not require any assumptions regarding spatial coherency. Therefore, we propose a one-dimensional (1D) dictionary learning (DL) framework based on variational sparse representation to suppress background noise in seismic data. Compared with the 2D multi-channel denoising method, the proposed method takes into account the waveform characteristics and requires no spatial coherency of single-channel seismic data, thus achieving better denoising performance. Additionally, the 1D dictionary learning method requires fewer training samples than the 2D method to reach a promising result, thereby causing less consumption time. Numerical results show that compared with the bandpass (BP) filtering, structure-oriented filtering (SOF), and K-SVD methods, the proposed DL framework can significantly improve the signal-to-noise ratio (SNR) of seismic data and protect effective signals better without causing extra computation. Furthermore, we discuss how to further suppress the residual horizontal noise and erratic noise based on the proposed method. Min Bai, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Fast Dictionary Learning Based on Data-Driven Tight Frame for 3-D Seismic Data DenoisingabstractSeismic denoising is a fundamental and critical task in seismic data processing. Aiming at solving the computational complexity of in 3D seismic data processing, we propose a novel data-driven tight frame (DDTF) dictionary learning method with over-complete dictionary constructed by discrete cosine transform for 3D seismic data denoising. The advantage of the DDTF algorithm is that only one singular value decomposition is required to update the entire dictionary, so as to accelerate the computational efficiency of 3D seismic data denoising. First, the seismic data is divided into patches to form matrix samples, and discrete cosine transform is selected according to preset parameters to initialize the dictionary. Then, the initial dictionary is trained by DDTF algorithm to update the dictionary. After that, the updated dictionary is used to denoise the block samples of seismic data. Finally, the proposed method is tested with synthetic data and field data. The results show that this method can significantly reduce the computational burden of state-of-the-art method, such as damped rank-reduction (DRR) method in 3D seismic data denoising, and the denoising performance is better than traditional DDTF method, which is conducive to the application of field data. Min Bai |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Implicit Surface Contrastive Clustering for LiDAR Point CloudsabstractSelf-supervised pretraining on large unlabeled datasets has shown tremendous success in improving the task performance of many 2D and small scale 3D computer vision tasks. However, the popular pretraining approaches have not been impactfully applied to outdoor LiDAR point cloud perception due to the latter's scene complexity and wide range. We propose a new self-supervised pretraining method ISCC with two novel pretext tasks for LiDAR point clouds. The first task uncovers semantic information by sorting local groups of points in the scene into a globally consistent set of semantically meaningful clusters using contrastive learning, complemented by a second task which reasons about precise surfaces of various parts of the scene through implicit surface reconstruction to learn geometric structures. We demonstrate their effectiveness through transfer learning on 3D object detection and semantic segmentation in real world LiDAR scenes. We further design an unsupervised semantic grouping task to show that our approach learns highly semantically meaningful features. Zaiwei Zhang, Min Bai, Li Erran Li |
CVPR | 2 |
| 2023 | HiSec: Towards Cyber Threat Correlation and Discovery Based on Hierarchical Graph Neural NetworksabstractA multi-step targeted cyberattack refers to a sophisticated, systematic, and persistent form of attack aiming to compromise the security and integrity of a complex network system. These attacks exhibit spatial and temporal correlations in their execution sequences. However, conventional anomaly detection methods, which often focus on single facets such as network traffic or host behavior, lack the capacity to correlate and validate these steps. To address this deficiency, we introduce HiSec, a cyber threat correlation and discovery framework that leverages dynamic graph modeling and hierarchical graph neural networks. HiSec enhances the modeling of complex network systems and the analysis of spatio-temporal characteristics of multi-step targeted cyberattacks. Specifically, we introduce a novel dynamic graph modeling algorithm that employs overlapping samplers and sliding windows to establish long-term correlations among system activities. Aided by graph attention networks and the Transformer, HiSec uniquely exploits the spatio-temporal correlated edge feature representation, a capability inaccessible to traditional algorithms. Evaluation results reveal that HiSec surpasses existing benchmarks in unsupervised detection, achieving a high degree of accuracy (93.10%) and recall (94.77%). When deployed in our intranet for two rounds of evaluation, HiSec demonstrated remarkable efficiency, taking only 0.15 seconds to model 10,000 system activities occurring within approximately an hour. Min Bai |
TrustCom | 4 |
| 2023 | Coherent Noise Attenuation by Kurtosis-Guided Adaptive Dictionary Learning Based on Variational Sparse RepresentationabstractIn seismic data denoising, coherent noise is usually challenging to remove because it has similar features with signal. Therefore, attenuating coherent noise is of great importance for seismic data. In order to obtain a good denoising result, it is necessary to fully consider the features of coherent noise and choose the appropriate denoising methods. Here, we propose a kurtosis-guided adaptive dictionary learning (KGADL) algorithm based on variational sparse representation model. First, the variational sparse representation is to construct initial dictionary that depends on the seismic data, so that it can accurately contain the features of the seismic data, so as to improve the accuracy of sparse representation. Then, we use the K-singular value decomposition (K-SVD) algorithm to update the dictionary and introduce kurtosis to measure each atom after updating. Atoms with high kurtosis values usually exhibit strong irregularities and are considered noise, while atoms with low kurtosis values have more feature of valid waves and are considered signal. The atoms with low kurtosis values characterizing valid signal are retained to get the special dictionary that adapts to the complexity of the input seismic data, and this dictionary is used to suppress the coherent noise in seismic data. Finally, the denoising performance of the proposed method is illustrated by examples of both synthetic and field seismic data. Min Bai |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Self-Supervised Pretraining for Large-Scale Point CloudsabstractPretraining on large unlabeled datasets has been proven to improve the down-stream task performance on many computer vision tasks, such as 2D object detection and video classification. However, for large-scale 3D scenes, such as outdoor LiDAR point clouds, pretraining is not widely used. Due to the special data characteristics of large 3D point clouds, 2D pretraining frameworks tend to not generalize well. In this paper, we propose a new self-supervised pretraining method that targets large-scale 3D scenes. We pretrain commonly used point-based and voxel-based model architectures and show the transfer learning performance on 3D object detection and also semantic segmentation. We demonstrate the effectiveness of our approach on both dense 3D indoor point clouds and also sparse outdoor lidar point clouds. Zaiwei Zhang, Min Bai, Li Erran Li |
NeurIPS | 2 |
| 2022 | Frequency-Space-Dependent Smoothing Regularized Nonstationary Predictive FilteringabstractPredictive filtering is one of the most widely used denoising algorithms in the seismic data processing community because of its high efficiency and stability in different situations. The traditional predictive filtering, however, is not able to deal with structurally complex data set unless applied in local windows. We develop a novel noncausal predictive filtering method that is free of the windowing step but is able to denoise complicated data set. We extend the stationary predictive filtering method to its nonstationary version, where the predictive filter coefficients vary across the frequency-space domain. The nonstationary predictive filtering (NPF) model requires solving a highly underdetermined inverse problem using an iterative shaping regularization method. The traditional shaping regularization method solves an inverse problem by applying a constant smoothing operator and thus does not consider the heterogeneity of the filter coefficients in the frequency–space domain. We propose to apply a nonstationary smoothing operator to constrain the model in the shaping regularization framework. The smoothing radius in the nonstationary smoothing operator is chosen based ona prioriinformation of the model, e.g., the nonstationarity of the data in the frequency–space domain. The proposed NPF method offers the flexibility in controlling the smoothness and sharpness of the calculated filter coefficients in both frequency and space dimensions. Several synthetic data sets and complicated real data examples are used to demonstrate the advantages of the new method. Guangtan Huang, Min Bai, Xingye Liu, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Directional Total Variation Regularized High-Resolution Prestack AVA InversionabstractPrestack seismic inversion has emerged as a powerful technique for reconstructing parameters attribute to the subsurface properties and building the geophysical parameter models. However, the inversion algorithms always suffer from spatial blur and low resolution. Total variation (TV) regularization preserves the spatial variation boundary of data by highlighting the sparsity of the first-order difference, which is regarded as an important technical means for image restoration. However, when the data do not change along the spatial grid direction, TV regularization is prone to a staircase effect. In this article, a directional TV (DTV) method is proposed to conduct the prestack amplitude variation with offset/angle (AVO/AVA) inversion. The method consists of three essential steps: estimating the seismic slope attribute from the seismic data, introducing seismic slope attribute to the TV regularization to establish the objective function, and optimizing the objective function by the split-Bregman algorithm. Finally, the conventional and proposed methods are applied to the synthetic and the real seismic data. The comparison of different methods demonstrates that the proposed method is applicable to reveal the detailed subsurface models, alleviate the staircase effect or artifact substantially, and further upgrade the quality of prestack inversion results. Guangtan Huang, Xiaohong Chen 0003, Shan Qu, Min Bai, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Self-Attention Deep Image Prior Network for Unsupervised 3-D Seismic Data EnhancementabstractWe develop a deep learning framework based on deep image prior (DIP) and attention networks for 3-D seismic data enhancement. First, the 3-D noisy data are divided into several overlapped patches. Second, the DIP network has a U-NET architecture, where the input patches are encoded to extract the significant latent features, while the decoder tries to reconstruct the input patches using these extracted features. Besides, the attention network is used to scale the extracted features from the encoder and the decoder. Third, the attention network output of the encoder is concatenated with that of the decoder to obtain high-order features and guide the network to extract the most significant information related to the seismic signals and discard the others. Finally, the 3-D seismic data are reconstructed using the output patches obtained by the DIP network. The proposed algorithm is an iterative and unsupervised approach, which does not require labeled data. We evaluate the proposed algorithm using several synthetic and field data examples. As a result, the proposed algorithm shows the ability to enhance the 3-D seismic data by attenuating the random noise and preserving the 3-D seismic signal with minimal signal leakage. Moreover, the proposed algorithm shows good denoising performance when tested using various types of events, e.g., linear, hyperbolic, low and high dominant frequencies, and weak amplitude. In addition, the proposed method outperforms the predictive filtering (PF) and damped rank-reduction (DRR) methods. To further understand the principle of the proposed method inside the DIP network, we analyze the weighting matrices in the encoder and decoder parts in detail. We attribute the denoising ability of the DIP network to the improvement of the extracted basis features from the encoder to the decoder layers through a deep network. Omar M. Saad, Yapo Abolé Serge Innocent Oboué, Min Bai, Lotfy Samy, Liuqing Yang 0004, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | An Unsplit CFS-PML Scheme for the Second-Order Wave Equation With Its Application in Fractional Viscoacoustic SimulationabstractThe unsplit complex frequency-shifted perfectly matched layer (CFS-PML) has been widely used in the first-order wave equation in velocity and stress while rarely formulated for the wave equation recast as a second-order system in displacement. Among different variants of PML, the unsplit CFS-PML for the second-order wave equation enjoys better absorbing performance and numerical stability, compared to the traditional PML, due to the presence of the general form of CFS stretching factors, as well as higher computational efficiency over the split PMLs since it avoids wave equation order reduction and splitting the state variables into multiple directional components. This study aims to develop an unsplit CFS-PML scheme for the second-order wave equation and devote specific attention to fractional viscoacoustic simulation where fractional time derivatives are involved and hard to be reformulated into a first-order form. In the complex space, PML is typically regarded as an analytical continuation of the real coordinates; thus, we define an explicit coordinate stretching operator acting on the Laplacian operator. This stretching operator consists of several convolution terms; each of them can be efficiently resolved by a recursive convolution updating strategy. Viscoacoustic simulations on homogeneous Pierre Shale, Marmousi model, and 3-D SEG/EAGE overthrust model verify the feasibility and absorbing the performance of our proposed scheme. Yufeng Wang 0009, Min Bai, Liuqing Yang 0004, Xuebin Zhao, Omar M. Saad, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Least-Squares Gaussian Beam Transform for Deblending Distance-Separated Simultaneous SourcesabstractSimultaneous-source acquisition saves enormous acquisition costs and greatly enhances the quality of seismic data. However, it brings challenges to subsequent data processing due to the intense interferences of neighbor sources. Deblending is one of the methods to solve the problem of crosstalk noise. The deblending applied in the shot domain does not require the random scheduling that is used in the conventional simultaneous-source acquisition, so it has more flexibility. However, because the records of different sources show continuous traces in common shot gathers, they have the same characteristics and are difficult to separate directly. In this article, we use the least-squares Gaussian beam transform (LSGBT) to separate the useful seismic signals from crosstalk noise in the distance-separated simultaneous-sourcing (DSSS) survey and propose a novel deblending framework based on the LSGBT in the shot domain. Unlike most state-of-the-art Gaussian beam approaches that construct Gaussian beams in the frequency domain, the LSGBT constructs time-domain Gaussian beams, which can be considered as functions of amplitude, position, dip field, and so on. The essence of the proposed algorithm is that the blended data can be decomposed into the Gaussian beams that represent different dip-angle components. Thus, the single-source seismic records can be reconstructed in terms of a carefully selected combination of dip-oriented Gaussian beams. Two synthetic examples and one field data example show that the iterative deblending scheme based on LSGBT obtains better performance than the conventional frequency-wavenumber-based methods. Min Bai, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Time-Lapse Seismic Difference-and-Joint Prestack AVA InversionabstractTime-lapse (4-D) seismic exploration is one of the essential means for accurately reconstructing the underground geological model, enhancing oil/gas recovery (EOR), and predicting remaining oil distribution. Under the assumption that the rock skeleton is almost unchanged in the process of development, inversion of time-lapse seismic difference data can be used to characterize the dynamic reservoir parameter. However, the inaccuracy of the forward operator, lack of constraints, and inappropriate regularization weight may lead to ill-posedness. Thus, the ill-posedness is still a key factor affecting the accuracy and stability of time-lapse seismic inversion. In this article, an improved difference-and-joint inversion strategy is proposed using the modified linear approximation as the forward operator. First, based on the modified approximation as the forward operator, time-lapse seismic difference inversion is exploited to obtain the variations of elastic parameter reflectance caused by the production-induced model perturbation. Then, we innovatively proposed to take the production-induced model perturbation as constraints, thereby achieving more accurate geological models for multiperiod seismic data joint inversion. Moreover, the L-curve method is exploited in this article to acquire the optimal regularization weight adaptively in each iteration step. Finally, the difference inversion results are used as constraints to precisely invert the geological model by combining multiperiod seismic data. Guangtan Huang, Xiaohong Chen 0003, Min Bai, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2019 | UPSNet: A Unified Panoptic Segmentation NetworkabstractIn this paper, we propose a unified panoptic segmentation network (UPSNet) for tackling the newly proposed panoptic segmentation task. On top of a single backbone residual network, we first design a deformable convolution based semantic segmentation head and a Mask R-CNN style instance segmentation head which solve these two subtasks simultaneously. More importantly, we introduce a parameter-free panoptic head which solves the panoptic segmentation via pixel-wise classification. It first leverages the logits from the previous two heads and then innovatively expands the representation for enabling prediction of an extra unknown class which helps better resolving the conflicts between semantic and instance segmentation. Besides, it handles the challenge caused by the varying number of instances and permits back propagation to the bottom modules in an end-to-end manner. Extensive experimental results on Cityscapes, COCO and our internal dataset demonstrate that our UPSNet achieves state-of-the-art performance with much faster inference. Code has been made available at: https://github.com/uber-research/UPSNet. Yuwen Xiong, Renjie Liao 0001, Hengshuang Zhao, Rui Hu 0001, Min Bai, Ersin Yumer, Raquel Urtasun |
CVPR | 5 |
| 2019 | Exploiting Sparse Semantic HD Maps for Self-Driving Vehicle LocalizationabstractIn this paper we propose a novel semantic localization algorithm that exploits multiple sensors and has precision on the order of a few centimeters. Our approach does not require detailed knowledge about the appearance of the world, and our maps require orders of magnitude less storage than maps utilized by traditional geometry- and LiDAR intensity-based localizers. This is important as self-driving cars need to operate in large environments. Towards this goal, we formulate the problem in a Bayesian filtering framework, and exploit lanes, traffic signs, as well as vehicle dynamics to localize robustly with respect to a sparse semantic map. We validate the effectiveness of our method on a new highway dataset consisting of 312km of roads. Our experiments show that the proposed approach is able to achieve 0.05m lateral accuracy and 1.12m longitudinal accuracy on average while taking up only 0.3% of the storage required by previous LiDAR intensity-based approaches. Wei-Chiu Ma, Raquel Urtasun, Ignacio Tartavull, Ioan Andrei Barsan, Shenlong Wang, Min Bai, Gellért Máttyus, Namdar Homayounfar, Shrinidhi Kowshika Lakshmikanth, Andrei Pokrovsky |
IROS | 6 |
| 2019 | Nonstationary Least-Squares Decomposition With Structural Constraint for Denoising Multi-Channel Seismic DataabstractThe seismic data usually contains strong random noise, which impedes the effective usage of the seismic signals for imaging and inversion. We propose an effective seismic denoising method based on a least-squares decomposition model. We assume that each trace in the multi-channel seismic data can be decomposed into several smoothly variable components. Since the decomposition is basically an inverse problem, we apply the temporal smoothness to constrain the inversion and control the stability. Considering the spatial coherency in a multi-channel seismic data, we also apply the spatial smoothness constraint to the decomposition. The space constraint is applied along the structural direction of the seismic events to preserve the dipping energy. The structural constraint is equivalent to applying a structure-oriented smoothing that requires the estimation of the local slope from the input noisy data. We validate the effectiveness of the proposed algorithm via several synthetic and real seismic data. The proposed method outperforms the state-of-the-art single-channel and multi-channel algorithms even in the case of strong random noise. Min Bai, Wei Chen 0031, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Least-Squares Gaussian Beam Transform for Seismic Noise AttenuationabstractWe propose a novel seismic noise attenuation approach based on least-squares Gaussian beam transform (LSGBT). Gaussian beam transform uses time-domain Gaussian beam (TGB), which can be characterized by a particular location, arrival time, amplitude, slope, curvature, and width. We implement the local attributes such as beam center, spacing, and width to perform Gaussian beam decomposition. In this approach, we first introduce the plane-wave decomposition (PWD) theory to implement TGB decomposition of noisy seismic data and then apply data reconstruction. Unlike most state-of-the-art algorithms, random noise is attenuated in the process of Gaussian beam reconstruction. In the reconstruction records, the useful events are well preserved simultaneously removing random noise. Comparisons of experimental results on field data using traditional$f$-$x$deconvolution (FX Decon) and median filter (MF) methods are also provided, which suggest that our method achieves better denoising performance than FX Decon and MF methods. Taking into account that signal loss is sometimes unavoidable in almost all existing denoising methods. In addition to the signal-to-noise ratio (SNR) measurement, we also use local similarity as an efficient tool to evaluate denoising performance. A group of synthetic and field examples demonstrates the effectiveness of the proposed approach. Min Bai, Mi Zhang 0005, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Seismic Noise Attenuation Using Unsupervised Sparse Feature LearningabstractNoise attenuation plays an important role in seismic data processing. We propose a novel denoising method for seismic data based on unsupervised sparse feature learning. Our goal is to obtain the identifiable feature of the noisy seismic data and then to represent the effective signals. By preprocessing the raw data and training the autoencoder neural network with sparse constraint, the sparse feature of the seismic data can be learned and stored in the neural network. We use the adaptive moment estimation as a backpropagation algorithm to minimize the cost function with a sparse penalty term and combine the dropout technique in the training process to improve the feature extraction and generalization capability of the neural network. Then, the test data set can be reconstructed by the most important sparse features. The final denoising result can be obtained by rearranging the output test data set. Compared with three commonly used state-of-the-art denoising methods, the proposed method performs well in applications to denoising for synthetic and real seismic data. Mi Zhang 0005, Yang Liu 0143, Min Bai, Yangkang Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2018 | Learning Deep Structured Active Contours End-to-EndabstractThe world is covered with millions of buildings, and precisely knowing each instance's position and extents is vital to a multitude of applications. Recently, automated building footprint segmentation models have shown superior detection accuracy thanks to the usage of Convolutional Neural Networks (CNN). However, even the latest evolutions struggle to precisely delineating borders, which often leads to geometric distortions and inadvertent fusion of adjacent building instances. We propose to overcome this issue by exploiting the distinct geometric properties of buildings. To this end, we present Deep Structured Active Contours (DSAC), a novel framework that integrates priors and constraints into the segmentation process, such as continuous boundaries, smooth edges, and sharp corners. To do so, DSAC employs Active Contour Models (ACM), a family of constraint- and prior-based polygonal models. We learn ACM parameterizations per instance using a CNN, and show how to incorporate all components in a structured output model, making DSAC trainable end-to-end. We evaluate DSAC on three challenging building instance segmentation datasets, where it compares favorably against state-of-the-art. Code will be made available on https://github.com/dmarcosg/DSAC. Diego Marcos, Devis Tuia, Benjamin Kellenberger, Lisa Zhang 0003, Min Bai, Renjie Liao 0001, Raquel Urtasun |
CVPR | 5 |
| 2018 | Deep Multi-Sensor Lane DetectionabstractReliable and accurate lane detection has been a long-standing problem in the field of autonomous driving. In recent years, many approaches have been developed that use images (or videos) as input and reason in image space. In this paper we argue that accurate image estimates do not translate to precise 3D lane boundaries, which are the input required by modern motion planning algorithms. To address this issue, we propose a novel deep neural network that takes advantage of both LiDAR and camera sensors and produces very accurate estimates directly in 3D space. We demonstrate the performance of our approach on both highways and in cities, and show very accurate estimates in complex scenarios such as heavy traffic (which produces occlusion), fork, merges and intersections. Min Bai, Gellért Máttyus, Namdar Homayounfar, Shenlong Wang, Shrinidhi Kowshika Lakshmikanth, Raquel Urtasun |
IROS | 1 |
| 2017 | Deep Watershed Transform for Instance SegmentationabstractMost contemporary approaches to instance segmentation use complex pipelines involving conditional random fields, recurrent neural networks, object proposals, or template matching schemes. In this paper, we present a simple yet powerful end-to-end convolutional neural network to tackle this task. Our approach combines intuitions from the classical watershed transform and modern deep learning to produce an energy map of the image where object instances are unambiguously represented as energy basins. We then perform a cut at a single energy level to directly yield connected components corresponding to object instances. Our model achieves more than double the performance over the state-of-the-art on the challenging Cityscapes Instance Level Segmentation task. Min Bai, Raquel Urtasun |
CVPR | 1 |
| 2017 | TorontoCity: Seeing the World with a Million EyesabstractIn this paper we introduce the TorontoCity benchmark, which covers the full greater Toronto area (GTA) with 712.5km2 of land, 8439km of road and around 400, 000 buildings. Our benchmark provides different perspectives of the world captured from airplanes, drones and cars driving around the city. Manually labeling such a large scale dataset is infeasible. Instead, we propose to utilize different sources of high-precision maps to create our ground truth. Towards this goal, we develop algorithms that allow us to align all data sources with the maps while requiring minimal human supervision. We have designed a wide variety of tasks including building height estimation (reconstruction), road centerline and curb extraction, building instance segmentation, building contour extraction (reorganization), semantic labeling and scene type classification (recognition). Our pilot study shows that most of these tasks are still difficult for modern convolutional neural networks. Shenlong Wang, Min Bai, Gellért Máttyus, Hang Chu, Wenjie Luo 0002, Bin Yang 0021, Justin Liang 0001, Joel Cheverie, Sanja Fidler, Raquel Urtasun |
ICCV | 2 |
| 2016 | Exploiting Semantic Information and Deep Matching for Optical Flow
Min Bai, Wenjie Luo 0002, Kaustav Kundu, Raquel Urtasun |
ECCV (6) | 1 |