EDBT 2026 Demo / reviewers in the wild / expert
Xin Yang 0008
dblp:44/1152-8
· DBLP profile ↗
129ranked-venue papers
29as first author
72since 2021 · last 2026
0000-0001-6252-1061ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 75 · 20 first-author · 39 since 2021Artificial intelligence and machine learning · 50 · 4 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 7 first-author · 23 since 2021Systems, architecture and hardware · 10 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedRNC: Addressing Spatio-Temporal Label Misalignment in Federated Noisy Class-Incremental LearningabstractFederated class-incremental learning (FCIL) aims to incrementally learn new classes across decentralized clients under non-IID data distributions. However, the pervasive challenge of label noise in FCIL has been completely overlooked. In this work, we introduce federated noisy class-incremental learning (FNCIL) and, for the first time, identify a novel form of label noise—spatio-temporal label misalignment—where samples from unseen classes are entirely mislabeled as known classes, with their correctly labeled counterparts appearing in latter tasks or other clients. This phenomenon undermines the effectiveness of existing centralized denoising strategies and creates a clear requirement for noise-robust methods in real-world FNCIL scenarios. To tackle this issue, we propose FedRNC, a dual-phase framework that leverages feature-space associations to establish spatio-temporal correspondences between clean global prototypes and noisy cached samples for progressive label correction. Experiments on standard benchmarks demonstrate FedRNC's superiority against existing baselines, along with its plug-and-play capability to upgrade FCIL systems for FNCIL. Xingwei Huang, Zhaobin Sun, Xin Yang 0008, Zengqiang Yan |
AAAI | 4 |
| 2026 | Generalized Geometry Encoding Volume for Real-time Stereo MatchingabstractReal-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In contrast, recent stereo foundation models leverage monocular foundation models (MFMs) to improve generalization, but typically suffer from substantial inference latency. To address this trade-off, we propose Generalized Geometry Encoding Volume (GGEV), a novel real-time stereo matching network that achieves strong generalization. We first extract depth-aware features that encode domain-invariant structural priors as guidance for cost aggregation. Subsequently, we introduce a Depth-aware Dynamic Cost Aggregation (DDCA) module that adaptively incorporates these priors into each disparity hypothesis, effectively enhancing fragile matching relationships in unseen scenes. Both steps are lightweight and complementary, leading to the construction of a generalized geometry encoding volume with strong generalization capability. Experimental results demonstrate that our GGEV surpasses all existing real-time methods in zero-shot generalization capability, and achieves state-of-the-art performance on the KITTI 2012, KITTI 2015, and ETH3D benchmarks. Gangwei Xu, Xianqi Wang 0001, Chengliang Zhang, Xin Yang 0008 |
AAAI | 5 |
| 2026 | BAT: Learning Event-based Optical Flow with Bidirectional Adaptive Temporal CorrelationabstractEvent cameras deliver visual information characterized by a high dynamic range and high temporal resolution, offering significant advantages in estimating optical flow for complex lighting conditions and fast-moving objects. Current advanced optical flow methods for event cameras largely adopt established image-based frameworks. However, the spatial sparsity of event data limits their performance. In this paper, we present BAT, an innovative framework that estimates event-based optical flow using bidirectional adaptive temporal correlation. BAT includes three novel designs: 1) a bidirectional temporal correlation that transforms bidirectional temporally dense motion cues into spatially dense ones, enabling accurate and spatially dense optical flow estimation; 2) an adaptive temporal sampling strategy for maintaining temporal consistency in correlation; 3) spatially adaptive temporal motion aggregation to efficiently and adaptively aggregate consistent target motion features into adjacent motion features while suppressing inconsistent ones. Our BAT achieves state-of-the-art performance on the DSEC-Flow benchmark, outperforming existing methods by a large margin while also exhibiting sharp edges and high-quality details. Our BAT can accurately predict future optical flow using only past events, significantly outperforming E-RAFT’s warm-start approach. Gangwei Xu, Haotong Lin, Zhaoxing Zhang, Hongcheng Luo, Xin Yang 0008 |
AAAI | 6 |
| 2026 | Addressing Imbalanced Modal Incompleteness in Realistic Multi-Modal Medical Image Segmentation via Hierarchical Gradient AlignmentabstractDespite the promising potential of multi-modal learning in medical image segmentation, real-world applications often encounter modal incompleteness sourced from diverse domains and institutions, sparking significant discussions on incomplete multi-modal learning. Existing approaches either train a unified model for all or develop individual models for specific multi-modal combinations to ensure model fairness and robustness during inference. However, the assumption of complete multi-modal data for training is unrealistic and infeasible in clinical practice. In this paper, we thoroughly formulate such a challenging setting and propose hierarchical gradient alignment (HGA) to address uni- and multi-modal imbalance. Specifically, gradient direction is aligned through sequential meta learning for multi-modal combinations and multi-level self-distillation for uni-modals within each combination. Gradient magnitude is aligned based on relative preference estimation to balance the dominance of each modal during training. Extensive experiments on five public benchmarks (BraTS2018, BraTS2020, BraTS2023, MyoPS2020, and MSSEG2016) demonstrate that HGA consistently outperforms state-of-the-art incomplete and imbalanced multi-modal learning methods, as well as representative multi-task learning optimization techniques. More importantly, HGA is validated to work as plug-and-play modules for consistent performance improvement across different backbones. Code is available at https://github.com/Jun-Jie-Shi/HGA. Zhaobin Sun, Li Yu 0003, Xin Yang 0008, Zengqiang Yan |
IEEE Trans. Medical Imaging | 4 |
| 2025 | FlowMamba: Learning Point Cloud Scene Flow with Global Motion PropagationabstractScene flow methods based on deep learning have achieved impressive performance. However, current top-performing methods still struggle with ill-posed regions, such as extensive flat regions or occlusions, due to insufficient local evidence. In this paper, we propose a novel global-aware scene flow estimation network with global motion propagation, named FlowMamba. The core idea of FlowMamba is a novel Iterative Unit based on the State Space Model (ISU), which first propagates global motion patterns and then adaptively integrates the global motion information with previously hidden states. As the irregular nature of point clouds limits the performance of ISU in global motion propagation, we propose a feature-induced ordering strategy (FIO). The FIO leverages semantic-related and motion-related features to order points into a sequence characterized by spatial continuity. Extensive experiments demonstrate the effectiveness of FlowMamba, with 21.9% and 20.5% EPE3D reduction from the best published results on FlyingThings3D and KITTI datasets. Specifically, our FlowMamba is the first method to achieve millimeter-level prediction accuracy in FlyingThings3D and KITTI. Furthermore, the proposed ISU can be seamlessly embedded into existing iterative networks as a plug-and-play module, improving their estimation accuracy significantly. Min Lin 0006, Gangwei Xu, Yun Wang 0013, Xianqi Wang 0001, Xin Yang 0008 |
AAAI | 5 |
| 2025 | SVDC: Consistent Direct Time-of-Flight Video Depth Completion with Frequency Selective FusionabstractLightweight direct Time-of-Flight (dToF) sensors are ideal for 3D sensing on mobile devices. However, due to the manufacturing constraints of compact devices and the inherent physical principles of imaging, dToF depth maps are sparse and noisy. In this paper, we propose a novel video depth completion method, called SVDC, by fusing the sparse dToF data with the corresponding RGB guidance. Our method employs a multi-frame fusion scheme to mitigate the spatial ambiguity resulting from the sparse dToF imaging. Misalignment between consecutive frames during multi-frame fusion could cause blending between object edges and the background, which results in a loss of detail. To address this, we introduce an adaptive frequency selective fusion (AFSF) module, which automatically selects convolution kernel sizes to fuse multi-frame features. Our AFSF utilizes a channel-spatial enhancement attention (CSEA) module to enhance features and generates an attention map as fusion weights. The AFSF ensures edge detail recovery while suppressing high-frequency noise in smooth regions. To further enhance temporal consistency, We propose a cross-window consistency loss to ensure consistent predictions across different windows, effectively reducing flickering. Our proposed SVDC achieves optimal accuracy and consistency on the TartanAir and Dynamic Replica datasets. Code is available at https://github.com/Lan1eve/SVDC. Xuan Zhu 0009, Jijun Xiang, Xianqi Wang 0001, Longliang Liu, Xin Yang 0008 |
CVPR | 8 |
| 2025 | MonSter: Marry Monodepth to Stereo Unleashes PowerabstractStereo matching recovers depth from image correspondences. Existing methods struggle to handle ill-posed regions with limited matching cues, such as occlusions and textureless areas. To address this, we propose MonSter, a novel method that leverages the complementary strengths of monocular depth estimation and stereo matching. MonSter integrates monocular depth and stereo matching into a dual-branch architecture to iteratively improve each other. Confidence-based guidance adaptively selects reliable stereo cues for monodepth scale-shift recovery. The refined monodepth is in turn guides stereo effectively at ill-posed regions. Such iterative mutual enhancement enables MonSter to evolve monodepth priors from coarse object-level structures to pixel-level geometry, fully unlocking the potential of stereo matching. As shown in Fig. 2, MonSter ranks 1stacross five most commonly used leaderboards — SceneFlow, KITTI 2012, KITTI 2015, Middlebury, and ETH3D. Achieving up to 49.5% improvements (Bad 1.0 on ETH3D) over the previous best method. Comprehensive analysis verifies the effectiveness of MonSter in ill-posed regions. In terms of zero-shot generalization, MonSter significantly and consistently outperforms state-of-the-art across the board. The code is publicly available at: https://github.com/Junda24/MonSter. Junda Cheng, Longliang Liu, Gangwei Xu, Xianqi Wang 0001, Zhaoxing Zhang, Jinliang Zang, Yurui Chen, Zhipeng Cai 0003, Xin Yang 0008 |
CVPR | 10 |
| 2025 | PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal ViewabstractPanoramic optical flow enables a comprehensive understanding of temporal dynamics across wide fields of view. However, severe distortions caused by sphere-to-plane projections, such as the equirectangular projection (ERP), significantly degrade the performance of conventional perspective-based optical flow methods, especially in polar regions. To address this challenge, we propose PriOr-Flow, a novel dual-branch framework that leverages the low-distortion nature of the orthogonal view to enhance optical flow estimation in these regions. Specifically, we introduce the Dual-Cost Collaborative Lookup (DCCL) operator, which jointly retrieves correlation information from both the primitive and orthogonal cost volumes, effectively mitigating distortion noise during cost volume construction. Furthermore, our Ortho-Driven Distortion Compensation (ODDC) module iteratively refines motion features from both branches, further suppressing polar distortions. Extensive experiments demonstrate that PriOr-Flow is compatible with various perspective-based iterative optical flow methods and consistently achieves state-of-the-art performance on publicly available panoramic optical flow datasets, setting a new benchmark for wide-field motion estimation. The code is publicly available at: https://github.com/longliangLiu/PriOr-Flow. Longliang Liu, Miaojie Feng, Junda Cheng, Jijun Xiang, Xuan Zhu 0009, Xin Yang 0008 |
ICCV | 6 |
| 2025 | ZeroStereo: Zero-Shot Stereo Matching from Single Images
Xianqi Wang 0001, Gangwei Xu, Junda Cheng, Min Lin 0006, Jinliang Zang, Yurui Chen, Xin Yang 0008 |
ICCV | 9 |
| 2025 | DEPTHOR: Depth Enhancement from a Practical Light-Weight dToF Sensor and RGB Image
Jijun Xiang, Xuan Zhu 0009, Xianqi Wang 0001, Xin Yang 0008 |
ICCV | 7 |
| 2025 | BANet: Bilateral Aggregation Network for Mobile Stereo MatchingabstractState-of-the-art stereo matching methods typically use costly 3D convolutions to aggregate a full cost volume, but their computational demands make mobile deployment challenging. Directly applying 2D convolutions for cost aggregation often results in edge blurring, detail loss, and mismatches in textureless regions. Some complex operations, like deformable convolutions and iterative warping, can partially alleviate this issue; however, they are not mobile-friendly, limiting their deployment on mobile devices. In this paper, we present a novel bilateral aggregation network (BANet) for mobile stereo matching that produces high-quality results with sharp edges and fine details using only 2D convolutions. Specifically, we first separate the full cost volume into detailed and smooth volumes using a spatial attention map, then perform detailed and smooth aggregations accordingly, ultimately fusing both to obtain the final disparity map. Experimental results demonstrate that our BANet-2D significantly outperforms other mobile-friendly methods, achieving 35.3\% higher accuracy on the KITTI 2015 leaderboard than MobileStereoNet-2D, with faster runtime on mobile devices. Code: \textcolor{magenta}{https://github.com/gangweix/BANet}. Gangwei Xu, Xianqi Wang 0001, Junda Cheng, Jinliang Zang, Yurui Chen, Xin Yang 0008 |
ICCV | 8 |
| 2025 | Semi-Elastic LiDAR-Inertial OdometryabstractThis work proposes a semi-elastic optimizationbased LiDAR-inertial state estimation method, which balances the constraints from LiDAR, IMU and consistency according to their unique characteristics, thereby imparts appropriate elasticity for current state to be optimized to the correct value, and ensure the accuracy, consistency, and robustness of state estimation. We incorporate the proposed LiDAR-inertial state estimation method into a self-developed optimizationbased LiDAR-inertial odometry (LIO) framework. Experimental results on four public datasets demonstrate that the proposed method enhances the performance of optimizationbased LiDAR-inertial state estimation. We have released the source code of this work for the development of the community. Zikang Yuan, Fengtian Lang, Tianle Xu, Ruiye Ming, Xin Yang 0008 |
ICRA | 6 |
| 2025 | ACP-MVS: Efficient Multi-View Stereo with Attention-based Context PerceptionabstractThe core of Multi-View Stereo (MVS) is to find corresponding pixels in neighboring images. However, due to challenging regions in input images such as untextured areas, repetitive patterns, or reflective surfaces, existing methods struggle to find precise pixel correspondence therein, resulting in inferior reconstruction quality. In this paper, we present an efficient context-perception MVS network, termed ACP-MVS. The ACP-MVS constructs a context-aware cost volume that can enhance pixels containing essential context information while suppressing irrelevant or noisy information via our proposed Context-stimulated Weighting Fusion module. Furthermore, we introduce a new Context-Guided Global Aggregation module, based on the insight that similar-looking pixels tend to have similar depths, which exploits global contextual cues to implicitly guide depth detail propagation from high-confidence regions to low-confidence ones. These two modules work in synergy to substantially improve reconstruction quality of ACP-MVS without incurring significant additional computational and time cost. Extensive experiments demonstrate that our approach not only achieves state-of-the-art performance but also offers the fastest inference speed and minimal GPU memory usage, providing practical value for practitioners working with high-resolution MVS image sets. Notably, our method ranks 2nd on the challenging Tanks and Temples advanced benchmark among all published methods. Code is available at https://github.com/HaoJia-mongh/ACP-MVS. Gangwei Xu, Miaojie Feng, Xianqi Wang 0001, Junda Cheng, Min Lin 0006, Xin Yang 0008 |
IROS | 7 |
| 2025 | A Multi-Branch Framework for Cross-Domain Vessel Segmentation via the Few-Shot Paradigm
Tianyu Zhao 0005, Xin Yang 0008 |
MICCAI (5) | 4 |
| 2025 | SR-SAM: Subspace Regularization for Domain Generalization of Segment Anything Model
Xixi Jiang, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (10) | 5 |
| 2025 | TemSAM: Temporal-Aware Segment Anything Model for Cerebrovascular Segmentation in Digital Subtraction Angiography Sequences
Xixi Jiang, Xiaohuan Ding, Tianyu Zhao 0005, Xin Yang 0008 |
MICCAI (1) | 6 |
| 2025 | ISAC: Redefining the Vascular Segmentation Paradigm Through Mask Completion for Cross-Domain Generalization
Tianyu Zhao 0005, Xixi Jiang, Xiaohuan Ding, Xin Yang 0008 |
MICCAI (7) | 6 |
| 2025 | Pixel-Perfect Depth with Semantics-Prompted Diffusion TransformersabstractThis paper presents **Pixel-Perfect Depth**, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds from estimated depth maps. Current generative depth estimation models fine-tune Stable Diffusion and achieve impressive performance. However, they require a VAE to compress depth maps into the latent space, which inevitably introduces flying pixels at edges and details. Our model addresses this challenge by directly performing diffusion generation in the pixel space, avoiding VAE-induced artifacts. To overcome the high complexity associated with pixel-space generation, we introduce two novel designs: 1) **Semantics-Prompted Diffusion Transformers** (**SP-DiT**), which incorporate semantic representations from vision foundation models into DiT to prompt the diffusion process, thereby preserving global semantic consistency while enhancing fine-grained visual details; and 2) **Cascade DiT Design** that progressively increases the number of tokens to further enhance efficiency and accuracy. Our model achieves the best performance among all published generative models across five benchmarks, and significantly outperforms all other models in edge-aware point cloud evaluation. Project page: https://pixel-perfect-depth.github.io/. Gangwei Xu, Haotong Lin, Hongcheng Luo, Xianqi Wang 0001, Jingfeng Yao, Lianghui Zhu, Yuechuan Pu, Hangjun Ye, Sida Peng, Xin Yang 0008 |
NeurIPS | 14 |
| 2025 | Labeled-to-unlabeled distribution alignment for partially-supervised multi-organ medical image segmentation
Xixi Jiang, Kangyi Liu, Kwang-Ting Cheng, Xin Yang 0008 |
Medical Image Anal. | 6 |
| 2025 | Toward high-quality pseudo masks from noisy or weak annotations for robust medical image segmentation
Zhiwei Wang 0002, Tianyu Zhao 0005, Xiaohuan Ding, Xin Yang 0008 |
Neural Networks | 5 |
| 2025 | IGEV++: Iterative Multi-Range Geometry Encoding Volumes for Stereo MatchingabstractStereo matching is a core component in many computer vision and robotics systems. Despite significant advances over the last decade, handling matching ambiguities in ill-posed regions and large disparities remains an open challenge. In this paper, we propose a new deep network architecture, called IGEV++, for stereo matching. The proposed IGEV++ constructs Multi-range Geometry Encoding Volumes (MGEV), which encode coarse-grained geometry information for ill-posed regions and large disparities, while preserving fine-grained geometry information for details and small disparities. To construct MGEV, we introduce an adaptive patch matching module that efficiently and effectively computes matching costs for large disparity ranges and/or ill-posed regions. We further propose a selective geometry feature fusion module to adaptively fuse multi-range and multi-granularity geometry features in MGEV. Then, we input the fused geometry features into ConvGRUs to iteratively update the disparity map. MGEV allows to efficiently handle large disparities and ill-posed regions, such as occlusions and textureless regions, and enjoys rapid convergence during iterations. Our IGEV++ achieves the best performance on the Scene Flow test set across all disparity ranges, up to 768px. Our IGEV++ also achieves state-of-the-art accuracy on the Middlebury, ETH3D, KITTI 2012, and 2015 benchmarks. Specifically, IGEV++ achieves a 3.23% 2-pixel outlier rate (Bad 2.0) on the large disparity benchmark, Middlebury, representing error reductions of 31.9% and 54.8% compared to RAFT-Stereo and GMStereo, respectively. We also present a real-time version of IGEV++ that achieves the best performance among all published real-time methods on the KITTI benchmarks. Gangwei Xu, Xianqi Wang 0001, Zhaoxing Zhang, Junda Cheng, Chunyuan Liao, Xin Yang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Multi-Granularity Topology-Aware Cell Localization and Counting in Pathological ImagesabstractCell localization and counting in pathological images play an important role in the diagnosis and treatment of life-threatening diseases (e.g., tumor). However, they still remain a challenging work, due to cell clustering and adhesion, blurred boundaries, deformation, and difficulty of annotation. In this work, we address these problems by introducing multi-granularity topological constraints in model training. First, a loss function of topological structure constraint for single cells is proposed, which encourages the trained model to avoid the wrong prediction of multiple cells within an instance (false positives). Second, a loss function of constraint of spatial topological structure distribution is proposed for clustered cells, which helps the trained model to reduce the wrong prediction of some crowded cells as one (false negative). Third, a loss is proposed from the expert check of annotation and inference errors, which enables positioning of difficult samples and facilitates the correction of errors. The multi-granularity loss under topological feature constraints enables a significant enhancement in the performance of the trained model. Experimental results on a self-collected COVID-19 pathological dataset and two public pathological datasets validate the performance advantages of the proposed method over some state-of-the-art methods. Our code will be available at https://github.com/MedicalYajieChen/MGTopology. Yajie Chen, Shujuan Wang, Boshuai Zhang, Lihua Lin, Qianqian Chai, Jiazheng Yang, Xin Yang 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | MC-Stereo: Multi-Peak Lookup and Cascade Search Range for Stereo MatchingabstractStereo matching is a fundamental task in scene comprehension. In recent years, the method based on iterative optimization has shown promise in stereo matching. However, the current iteration framework employs a single-peak lookup, which struggles to handle the multi-peak problem effectively. Additionally, the fixed search range used during the iteration process limits the final convergence effects. To address these issues, we present a novel iterative optimization architecture called MC-Stereo. This architecture mitigates the multi-peak distribution problem in matching through the multi-peak lookup strategy, and integrates the coarse-to-fine concept into the iterative framework via the cascade search range. Furthermore, given that feature representation learning is crucial for successful learn-based stereo matching, we introduce a pre-trained network to serve as the feature extractor, enhancing the front end of the stereo matching pipeline. Based on these improvements, MC-Stereo ranks first among all publicly available methods on the KITTI-2012 and KITTI-2015 benchmarks, and also achieves state-of-the-art performance on ETH3D. Code is available at https://github.com/MiaoJieF/MC-Stereo Miaojie Feng, Junda Cheng, Longliang Liu, Gangwei Xu, Xin Yang 0008 |
3DV | 6 |
| 2024 | Adaptive Fusion of Single-View and Multi-View Depth for Autonomous DrivingabstractMulti- view depth estimation has achieved impressive performance over various benchmarks. However, almost all current multi-view systems rely on given ideal camera poses, which are unavailable in many real-world scenarios, such as autonomous driving. In this work, we propose a new robustness benchmark to evaluate the depth estimation system under various noisy pose settings. Surprisingly, we find current multi-view depth estimation methods or single-view and multi-view fusion methods will fail when given noisy pose settings. To address this challenge, we propose a single-view and multi-view fused depth estimation system, which adaptively integrates high-confident multi-view and single-view results for both robust and accurate depth es-timations. The adaptive fusion module performs fusion by dynamically selecting high-confidence regions between two branches based on a wrapping confidence map. Thus, the system tends to choose the more reliable branch when facing textureless scenes, inaccurate calibration, dynamic ob-jects, and other degradation or challenging conditions. Our method outperforms state-of-the-art multi-view and fusion methods under robustness testing. Furthermore, we achieve state-of-the-art performance on challenging benchmarks (KITTI and DDAD) when given accurate pose estimations. Project website: https://github.com/Junda24/Afnet/. Junda Cheng, Wei Yin 0006, Xiaozhi Chen, Xin Yang 0008 |
CVPR | 6 |
| 2024 | Selective-Stereo: Adaptive Frequency Information Selection for Stereo MatchingabstractStereo matching methods based on iterative optimization, like RAFT-Stereo and IGEV-Stereo, have evolved into a cornerstone in the field of stereo matching. However, these methods struggle to simultaneously capture high-frequency information in edges and low-frequency information in smooth regions due to the fixed receptive field. As a result, they tend to lose details, blur edges, and produce false matches in textureless areas. In this paper, we propose Selective Recurrent Unit (SRU), a novel iterative update operator for stereo matching. The SRU module can adaptively fuse hidden disparity information at multiple frequencies for edge and smooth regions. To perform adaptive fusion, we introduce a new Contextual Spatial Attention (CSA) module to generate attention maps as fusion weights. The SRU empowers the network to aggregate hidden disparity information across multiple frequencies, mitigating the risk of vital hidden disparity information loss during iterative processes. To verify SRU’s universality, we apply it to representative iterative stereo matching methods, collectively referred to as Selective-Stereo. Our Selective-Stereo ranks 1ston KITTI 2012, KITTI 2015, ETH3D, and Middle- bury leaderboards among all published methods. Code is available at https://github.com/Windsrain/Selective-Stereo. Xianqi Wang 0001, Gangwei Xu, Xin Yang 0008 |
CVPR | 4 |
| 2024 | HDRFlow: Real-Time HDR Video Reconstruction with Large MotionsabstractReconstructing High Dynamic Range (HDR) video from image sequences captured with alternating exposures is challenging, especially in the presence of large camera or object motion. Existing methods typically align low dynamic range sequences using optical flow or attention mechanism for deghosting. However, they often struggle to handle large complex motions and are computation-ally expensive. To address these challenges, we propose a robust and efficient flow estimator tailored for real-time HDR video reconstruction, named HDRFlow. HDRFlow has three novel designs: an HDR-domain alignment loss (HALoss), an efficient flow network with a multi-size large kernel (MLK), and a new HDR flow training scheme. The HALoss supervises our flow network to learn an HDR-oriented flow for accurate alignment in saturated and dark regions. The MLK can effectively model large motions at a negligible cost. In addition, we incorporate synthetic data, Sintel, into our training dataset, utilizing both its provided forward flow and backward flow generated by us to super-vise our flow network, enhancing our performance in large motion regions. Extensive experiments demonstrate that our HDRFlow outperforms previous methods on standard benchmarks. To the best of our knowledge, HDRFlow is the first real-time HDR video reconstruction method for video sequences captured with alternating exposures, capable of processing 720p resolution inputs at 25ms. Project website: https: https://openimaginglab.github.io/HDRFlow/. Gangwei Xu, Yujin Wang 0001, Jinwei Gu, Tianfan Xue, Xin Yang 0008 |
CVPR | 5 |
| 2024 | A Noise Robust Framework via Uncertainty Guidance for Medical Image Segmentation with Noisy LabelabstractIn medical image segmentation, acquiring sufficient accurate pixel-level annotations demands substantial manual labor and expertise, making it difficult to be satisfied and yielding noisy annotations. Visually distinguishable random label noises and boundary uncertainties are two primary aspects of annotation noises, but most existing methods fail to handle their co-occurrence issue. In this paper, we present a novel frame-work that primarily consists of an adaptive uncertainty-based label revision method and a contrastive learning approach to address the above challenge. The label revision method can correct random label noises by utilizing low uncertainty predictions, while enhancing the model’s tolerance to boundary uncertainties through the use of soft labels. The contrastive learning method can assist the model in reliably learning intrinsic relationships between pixels in the presence of noisy annotations. Experiments on two public datasets demonstrate the superiority of our method. Tianyu Zhao 0005, Xin Yang 0008 |
ICME | 4 |
| 2024 | SR-LIO: LiDAR-Inertial Odometry with Sweep ReconstructionabstractThis paper proposes a novel LiDAR-Inertial odometry (LIO), named SR-LIO, based on an error state iterated Kalman filter (ESIKF) framework. We adapt the sweep reconstruction method, which segments and reconstructs raw input sweeps from spinning LiDAR to obtain reconstructed sweeps with higher frequency. We found that such method can effectively reduce the time interval for each iterated state update, improving the state estimation accuracy and enabling the usage of ESIKF framework for fusing high-frequency IMU and low-frequency LiDAR. To prevent inaccurate trajectory caused by multiple distortion correction to a particular point, we further propose to perform distortion correction for each segment. Experimental results on four public datasets demonstrate that our SR-LIO outperforms all existing state-of-the-art methods on accuracy, and reducing the time interval of iterated state update via the proposed sweep reconstruction can improve the accuracy and frequency of estimated states. The source code of SR-LIO is publicly available for the development of the community. Zikang Yuan, Fengtian Lang, Tianle Xu, Xin Yang 0008 |
IROS | 4 |
| 2024 | FedIA: Federated Medical Image Segmentation with Heterogeneous Annotation Completeness
Yangyang Xiang, Li Yu 0003, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (10) | 4 |
| 2024 | Masked Snake Attention for Fundus Image Restoration with Vessel PreservationabstractRestoring low-quality fundus images, especially the recovery of vessel structures, is crucial for clinical observation and diagnosis. Existing state-of-the-art methods use standard convolution and window based self-attention block to recover low-quality fundus images, but these feature capturing approaches do not effectively match the slender and tortuous structure of retinal vessels. Therefore, these methods struggle to accurately restore vessel structures. To overcome this challenge, we propose a novel low-quality fundus image restoration method called Masked Snake Attention Network (MSANet). It is designed specifically for accurately restoring vessel structures. Specifically, we introduce the Snake Attention module (SA) to adaptively aggregate vessel features based on the morphological structure of the vessels. Due to the small proportion of vessel pixels in the image, we further present the Masked Snake Attention module (MSA) to more efficiently capture vessel features. MSA enhances vessel features by constraining snake attention within regions predicted by segmentation methods. Extensive experimental results demonstrate that our MSANet outperforms the state-of-the-art methods in enhancement evaluation and downstream segmentation tasks. Xiaohuan Ding, Yangrui Gong, Gangwei Xu, Xin Yang 0008 |
ACM Multimedia | 6 |
| 2024 | PASSION: Towards Effective Incomplete Multi-Modal Medical Image Segmentation with Imbalanced Missing Rates
Caozhi Shang, Zhaobin Sun, Li Yu 0003, Xin Yang 0008, Zengqiang Yan |
ACM Multimedia | 5 |
| 2024 | Coatrsnet: Fully Exploiting Convolution and Attention for Stereo Matching by Region Separation
Junda Cheng, Gangwei Xu, Peng Guo 0001, Xin Yang 0008 |
Int. J. Comput. Vis. | 4 |
| 2024 | Accurate and Efficient Stereo Matching via Attention Concatenation VolumeabstractStereo matching is a fundamental building block for many vision and robotics applications. An informative and concise cost volume representation is vital for stereo matching of high accuracy and efficiency. In this article, we present a novel cost volume construction method, named attention concatenation volume (ACV), which generates attention weights from correlation clues to suppress redundant information and enhance matching-related information in the concatenation volume. The ACV can be seamlessly embedded into most stereo matching networks, the resulting networks can use a more lightweight aggregation network and meanwhile achieve higher accuracy. We further design a fast version of ACV to enable real-time performance, named Fast-ACV, which generates high likelihood disparity hypotheses and the corresponding attention weights from low-resolution correlation clues to significantly reduce computational and memory cost and meanwhile maintain a satisfactory accuracy. Furthermore, we design a highly accurate network ACVNet and a real-time network Fast-ACVNet based on our ACV and Fast-ACV respectively, which achieve state-of-the-art performance on several benchmarks. Gangwei Xu, Yun Wang 0013, Junda Cheng, Jinhui Tang 0001, Xin Yang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | UCTNet: Uncertainty-guided CNN-Transformer hybrid networks for medical image segmentation
Xiayu Guo, Xian Lin, Xin Yang 0008, Li Yu 0003, Kwang-Ting Cheng, Zengqiang Yan |
Pattern Recognit. | 3 |
| 2024 | LENAS: Learning-Based Neural Architecture Search and Ensemble for 3-D Radiotherapy Dose PredictionabstractRadiation therapy treatment planning requires balancing the delivery of the target dose while sparing normal tissues, making it a complex process. To streamline the planning process and enhance its quality, there is a growing demand for knowledge-based planning (KBP). Ensemble learning has shown impressive power in various deep learning tasks, and it has great potential to improve the performance of KBP. However, the effectiveness of ensemble learning heavily depends on the diversity and individual accuracy of the base learners. Moreover, the complexity of model ensembles is a major concern, as it requires maintaining multiple models during inference, leading to increased computational cost and storage overhead. In this study, we propose a novel learning-based ensemble approach named LENAS, which integrates neural architecture search with knowledge distillation for 3-D radiotherapy dose prediction. Our approach starts by exhaustively searching each block from an enormous architecture space to identify multiple architectures that exhibit promising performance and significant diversity. To mitigate the complexity introduced by the model ensemble, we adopt the teacher-student paradigm, leveraging the diverse outputs from multiple learned networks as supervisory signals to guide the training of the student network. Furthermore, to preserve high-level semantic information, we design a hybrid loss to optimize the student network, enabling it to recover the knowledge embedded within the teacher networks. The proposed method has been evaluated on two public datasets: 1) OpenKBP and 2) AIMIS. Extensive experimental results demonstrate the effectiveness of our method and its superior performance to the state-of-the-art methods. Code: github.com/hust-linyi/LENAS. Yi Lin 0009, Hao Chen 0011, Xin Yang 0008, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
IEEE Trans. Cybern. | 4 |
| 2024 | A Discrepancy Aware Framework for Robust Anomaly DetectionabstractDefect detection is a critical research area in artificial intelligence. Recently, synthetic data-based self-supervised learning has shown great potential on this task. Although many sophisticated synthesizing strategies exist, little research has been done to investigate the robustness of models when faced with different strategies. In this article, we focus on this issue and find that existing methods are highly sensitive to them. To alleviate this issue, we present a discrepancy aware framework (DAF), which demonstrates robust performance consistently with simple and cheap strategies across different anomaly detection benchmarks. We hypothesize that the high sensitivity to synthetic data of existing self-supervised methods arises from their heavy reliance on the visual appearance of synthetic data during decoding. In contrast, our method leverages an appearance-agnostic cue to guide the decoder in identifying defects, thereby alleviating its reliance on synthetic appearance. To this end, inspired by existing knowledge distillation methods, we employ a teacher-student network, which is trained based on synthesized outliers, to compute the discrepancy map as the cue. Extensive experiments on two challenging datasets prove the robustness of our method. Under the simple synthesis strategies, it outperforms existing methods by a large margin. Furthermore, it also achieves the state-of-the-art localization performance. Dingkang Liang, Dongliang Luo, Xinwei He 0001, Xin Yang 0008, Xiang Bai |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | MFTrans: Modality-Masked Fusion Transformer for Incomplete Multi-Modality Brain Tumor SegmentationabstractBrain tumor segmentation is a fundamental task and existing approaches usually rely on multi-modality magnetic resonance imaging (MRI) images for accurate segmentation. However, the common problem of missing/incomplete modalities in clinical practice would severely degrade their segmentation performance, and existing fusion strategies for incomplete multi-modality brain tumor segmentation are far from ideal. In this work, we propose a novel framework named M$^{2}$FTrans to explore and fuse cross-modality features through modality-masked fusion transformers under various incomplete multi-modality settings. Considering vanilla self-attention is sensitive to missing tokens/inputs, both learnable fusion tokens and masked self-attention are introduced to stably build long-range dependency across modalities while being more flexible to learn from incomplete modalities. In addition, to avoid being biased toward certain dominant modalities, modality-specific features are further re-weighted through spatial weight attention and channel-wise fusion transformers for feature redundancy reduction and modality re-balancing. In this way, the fusion strategy in M$^{2}$FTrans is more robust to missing modalities. Experimental results on the widely-used BraTS2018, BraTS2020, and BraTS2021 datasets demonstrate the effectiveness of M$^{2}$FTrans, outperforming the state-of-the-art approaches with large margins under various incomplete modalities for brain tumor segmentation. Li Yu 0003, Qimin Cheng, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | APCAFlow: All-Pairs Cost Volume Aggregation for Optical Flow EstimationabstractOptical flow estimation is a fundamental task in computer vision. The all-pairs correlation volume has enabled state-of-the-art performance in many optical flow estimation methods. However, all-pairs correlations provide only local matching clues, and lack global context, which could lead to mismatches in textureless and occluded regions. In this paper, we propose a novel all-pairs correlation volume aggregation (APCA) method which includes two key innovations. The first is a cost volume splitting and reassembling approach which partitions the full cost volume into smaller blocks and re-arranges those blocks to allow the use of 2D and 3D convolutions for cost volume aggregation. The second is hierarchical aggregation which performs 2D convolutions within blocks for local matching aggregation and 3D convolutions across blocks for global matching aggregation. We further design a novel optical flow estimation network APCAFlow based on APCA. APCAFlow achieves comparable performance to the most advanced approach, FlowFormer, but with significantly lower complexity. Specifically, APCAFlow reduces the model parameters, inference time, and memory consumption by 24.1%, 35.5%, and 21.6%, respectively, compared to FlowFormer. Furthermore, APCA can be easily integrated into several existing all-pairs cost volume-based methods for performance improvement. Code is available athttps://github.com/MiaoJieF/APCAFlow. Miaojie Feng, Zengqiang Yan, Xin Yang 0008 |
IEEE Trans. Multim. | 4 |
| 2023 | Iterative Geometry Encoding Volume for Stereo MatchingabstractRecurrent All-Pairs Field Transforms (RAFT) has shown great potentials in matching tasks. However, all-pairs correlations lack non-local geometry knowledge and have difficulties tackling local ambiguities in ill-posed regions. In this paper, we propose Iterative Geometry Encoding Volume (IGEV-Stereo), a new deep network architecture for stereo matching. The proposed IGEV-Stereo builds a combined geometry encoding volume that encodes geometry and context information as well as local matching details, and iteratively indexes it to update the disparity map. To speed up the convergence, we exploit GEV to regress an accurate starting point for ConvGRUs iterations. Our IGEV-Stereo ranks 1st on KITTI 2015 and 2012 (Reflective) among all published methods and is the fastest among the top 10 methods. In addition, IGEV-Stereo has strong cross-dataset generalization as well as high inference efficiency. We also extend our IGEV to multi-view stereo (MVS), i.e. IGEV-MVS, which achieves competitive accuracy on DTU benchmark. Code is available at https://github.com/gangweiX/IGEV. Gangwei Xu, Xianqi Wang 0001, Xiaohuan Ding, Xin Yang 0008 |
CVPR | 4 |
| 2023 | IHNet: Iterative Hierarchical Network Guided by High-Resolution Estimated Information for Scene Flow EstimationabstractScene flow estimation, which predicts the 3D displacements of point clouds, is a fundamental task in autonomous driving. Most methods have adopted a coarse-to-fine structure to balance computational efficiency with accuracy, particularly when handling large displacements. However, inaccuracies in the initial coarse layer’s scene flow estimates may accumulate, leading to incorrect final estimates. To alleviate this, we introduce a novel Iterative Hierarchical Network——IHNet. This approach circulates high-resolution estimated information (scene flow and feature) from the preceding iteration back to the low-resolution layer of the current iteration. Serving as a guide, the high-resolution estimated scene flow, instead of initializing the scene flow from zero, provides a more precise center for low-resolution layer to identify matches. Meanwhile, the decoder’s feature at the high-resolution layer can contribute essential movement information. Furthermore, based on the recurrent structure, we design a resampling scheme to enhance the correspondence between points across two consecutive frames. By employing the previous estimated scene flow to fine-tune the target frame’s coordinates, we can significantly reduce the correspondence discrepancy between two frame points, a problem often caused by point sparsity. Following this adjustment, we continue to estimate the scene flow using the newly updated coordinates, along with the reencoded feature. Our approach outperforms the recent state-of-the-art method WSAFlowNet by 20.1% on FlyingThings3D and 56.0% on KITTI scene flow datasets according to EPE3D metric. The code is available at https://github.com/wangyunlhr/IHNet. Yun Wang 0013, Min Lin 0006, Xin Yang 0008 |
ICCV | 4 |
| 2023 | Dual-view Correlation Hybrid Attention Network for Robust Holistic Mammogram ClassificationabstractMammogram image is important for breast cancer screening, and typically obtained in a dual-view form, i.e., cranio-caudal (CC) and mediolateral oblique (MLO), to provide complementary information for clinical decisions. However, previous methods mostly learn features from the two views independently, which violates the clinical knowledge and ignores the importance of dual-view correlation in the feature learning. In this paper, we propose a dual-view correlation hybrid attention network (DCHA-Net) for robust holistic mammogram classification. Specifically, DCHA-Net is carefully designed to extract and reinvent deep feature maps for the two views, and meanwhile to maximize the underlying correlations between them. A hybrid attention module, consisting of local relation and non-local attention blocks, is proposed to alleviate the spatial misalignment of the paired views in the correlation maximization. A dual-view correlation loss is introduced to maximize the feature similarity between corresponding strip-like regions with equal distance to the chest wall, motivated by the fact that their features represent the same breast tissues, and thus should be highly-correlated with each other. Experimental results on the two public datasets, i.e., INbreast and CBIS-DDSM, demonstrate that the DCHA-Net can well preserve and maximize feature correlations across views, and thus outperforms previous state-of-the-art methods for classifying a whole mammogram as malignant or not. Zhiwei Wang 0002, Junlin Xian, Kangyi Liu, Xin Li 0001, Qiang Li 0018, Xin Yang 0008 |
IJCAI | 6 |
| 2023 | LIWO: LiDAR-Inertial-Wheel OdometryabstractLiDAR-inertial odometry (LIO), which fuses complementary information of a LiDAR and an Inertial Measurement Unit (IMU), is an attractive solution for state estimation. In LIO, both pose and velocity are regarded as state variables that need to be solved. However, the widely-used Iterative Closest Point (ICP) algorithm can only provide constraint for pose, while the velocity can only be constrained by IMU pre-integration. As a result, the velocity estimates inclined to be updated accordingly with the pose results. In this paper, we propose LIWO, an accurate and robust LiDAR-inertial-wheel (LIW) odometry, which fuses the measurements from LiDAR, IMU and wheel encoder in a bundle adjustment (BA) based optimization framework. The involvement of a wheel encoder could provide velocity measurement as an important observation, which assists LIO to provide a more accurate state prediction. In addition, con-straining the velocity variable by the observation from wheel encoder in optimization can further improve the accuracy of state estimation. Experiment results on two public datasets demonstrate that our system outperforms all state-of-the-art LIO systems in terms of smaller absolute trajectory error (ATE), and embedding a wheel encoder can greatly improve the performance of LIO based on the BA framework. Zikang Yuan, Fengtian Lang, Tianle Xu, Xin Yang 0008 |
IROS | 4 |
| 2023 | FedIIC: Towards Robust Federated Learning for Class-Imbalanced Medical Image Classification
Li Yu 0003, Xin Yang 0008, Kwang-Ting Cheng, Zengqiang Yan |
MICCAI (2) | 3 |
| 2023 | Confidence-weighted mutual supervision on dual networks for unsupervised cross-modality image segmentation
Yajie Chen, Xin Yang 0008, Xiang Bai |
Sci. China Inf. Sci. | 2 |
| 2023 | Dual-distribution discrepancy with self-supervised refinement for anomaly detection in medical images
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
Medical Image Anal. | 3 |
| 2023 | SDV-LOAM: Semi-Direct Visual-LiDAR Odometry and MappingabstractVisual-LiDAR odometry and mapping (V-LOAM), which fuses complementary information of a camera and a LiDAR, is an attractive solution for accurate and robust pose estimation and mapping. However, existing systems could suffer nontrivial tracking errors arising from 1) association between 3D LiDAR points and sparse 2D features (i.e., 3D-2D depth association) and 2) obvious drifts in the vertical direction in the 6-degree of freedom (DOF) sweep-to-map optimization. In this paper, we present SDV-LOAM which incorporates a semi-direct visual odometry and an adaptive sweep-to-map LiDAR odometry to effectively avoid the above-mentioned errors and in turn achieve high tracking accuracy. The visual module of our SDV-LOAM directly extracts high-gradient pixels where 3D LiDAR points project on for tracking. To avoid the problem of large scale difference between matching frames in the VO, we design a novel point matching with propagation method to propagate points of a host frame to an intermediate keyframe which is closer to the current frame to reduce scale differences. To reduce the pose estimation drifts in the vertical direction, our LiDAR module employs an adaptive sweep-to-map optimization method which automatically choose to optimize 3 horizontal DOF or 6 full DOF pose according to the richness of geometric constraints in the vertical direction. In addition, we propose a novel sweep reconstruction method which can increase the input frequency of LiDAR point clouds to the same frequency as the camera images, and in turn yield a high frequency output of the LiDAR odometry in theory. Experimental results demonstrate that our SDV-LOAM ranks 8th on the KITTI odometry benchmark which outperforms most LiDAR/visual-LiDAR odometry systems. In addition, our visual module outperforms state-of-the-art visual odometry and our adaptive sweep-to-map optimization can improve the performance of several existing open-sourced LiDAR odometry systems. Moreover, we demonstrate our SDV-LOAM on a custom-built hardware platform in large-scale environments which achieves both a high accuracy and output frequency. We have released the source code of our SDV-LOAM for the development of the community. Zikang Yuan, Qingjie Wang, Ken Cheng, Tianyu Hao, Xin Yang 0008 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Vision Transformer With Hybrid Shifted Windows for Gastrointestinal Endoscopy Image ClassificationabstractAutomated classification of gastrointestinal endoscope images can help reduce the workload of doctors and improve the accuracy of diagnoses. The rapidly developed vision Transformer, represented by Swin Transformer, has become an impressive technique for medical image classification. However, Swin Transformer cannot capture the long-range dependency well in complex gastrointestinal endoscopy images. As a result, it fails to represent features of some widely-spread targets in digestive tract images, such as normal-z-line and esophagitis, effectively. To solve this problem, we propose a novel vision Transformer model based on hybrid shifted windows for digestive tract image classification, which can obtain both short-range and long-range dependency concurrently. Extensive experiments demonstrate the superiority of our method to the state-of-the-art methods with a classification accuracy of 95.42% on the Kvasir v2 dataset and a classification accuracy of 86.81% on the HyperKvasir dataset. Wei Wang 0355, Xin Yang 0008, Jinhui Tang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Accurate Cobb Angle Estimation on Scoliosis X-Ray Images via Deeply-Coupled Two-Stage Network With Differentiable Cropping and Random PerturbationabstractAutomated Cobb angle estimation on X-ray images is crucial to scoliosis diagnosis. The existing efforts are typically two extremes, which either laboriously detect the raw vertebral landmarks or directly regress Cobb angles from the entire image. In this paper, we propose a novel two-stage end-to-end method as a balanced solution, to avoid vulnerability to false landmarks, and to preserve flexibility in clinical usages. Concretely, we cascade two stages sequentially for detecting vertebrae and then regressing their bending directions instead of raw landmarks. In the detection stage, we combine two networks called LocNet and SegNet to robustly localize vertebrae, and meanwhile to suppress the false positives by additionally segmenting the whole spine. In the subsequent stage, we introduce a regression network named RegNet to accurately regress bending directions of localized vertebrae. Furthermore, the vertebra-aligned local regions on LocNet's intermediate features are cropped via RoIAlign-pooling, and RegNet inherits the cropped regions to learn only feature residuals. By doing so, the regression difficulty can be dramatically alleviated, and the two stages are deeply coupled and mutually guided in an end-to-end training. Moreover, a random perturbation on the inherited features further enhances RegNet's robustness. We benchmark our method on both public and private datasets, and the errors are 2.92 $\pm$ 2.34$^{\circ }$ and 6.87 $\pm$ 6.26% in terms of CMAE and SMAPE on the widely-employed AASCE dataset, outperforming other state-of-the-arts by at least 16.81% and 6.15%, respectively. Also, a clinical user study verifies our promising flexibility for allowing convenient rectifications to further decrease errors by a large marge. Yuanhuai Liang, Jinxin Lv, Dun Li, Xin Yang 0008, Zhiwei Wang 0002, Qiang Li 0018 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Affinity Feature Strengthening for Accurate, Complete and Robust Vessel SegmentationabstractVessel segmentation is crucial in many medical image applications, such as detecting coronary stenoses, retinal vessel diseases and brain aneurysms. However, achieving high pixel-wise accuracy, complete topology structure and robustness to various contrast variations are critical and challenging, and most existing methods focus only on achieving one or two of these aspects. In this paper, we present a novel approach, the affinity feature strengthening network (AFN), which jointly models geometry and refines pixel-wise segmentation features using a contrast-insensitive, multiscale affinity approach. Specifically, we compute a multiscale affinity field for each pixel, capturing its semantic relationships with neighboring pixels in the predicted mask image. This field represents the local geometry of vessel segments of different sizes, allowing us to learn spatial- and scale-aware adaptive weights to strengthen vessel features. We evaluate our AFN on four different types of vascular datasets: X-ray angiography coronary vessel dataset (XCAD), portal vein dataset (PV), digital subtraction angiography cerebrovascular vessel dataset (DSA) and retinal vessel dataset (DRIVE). Extensive experimental results demonstrate that our AFN outperforms the state-of-the-art methods in terms of both higher accuracy and topological metrics, while also being more robust to various contrast changes. Xiaohuan Ding, Wei Zhou 0068, Zengqiang Yan, Xiang Bai, Xin Yang 0008 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Bidirectional Semi-Supervised Dual-Branch CNN for Robust 3D Reconstruction of Stereo Endoscopic Images via Adaptive Cross and Parallel SupervisionsabstractSemi-supervised learning via teacher-student network can train a model effectively on a few labeled samples. It enables a student model to distill knowledge from the teacher's predictions of extra unlabeled data. However, such knowledge flow is typically unidirectional, having the accuracy vulnerable to the quality of teacher model. In this paper, we seek to robust 3D reconstruction of stereo endoscopic images by proposing a novel fashion of bidirectional learning between two learners, each of which can play both roles of teacher and student concurrently. Specifically, we introduce two self-supervisions, i.e., Adaptive Cross Supervision (ACS) and Adaptive Parallel Supervision (APS), to learn a dual-branch convolutional neural network. The two branches predict two different disparity probability distributions for the same position, and output their expectations as disparity values. The learned knowledge flows across branches along two directions: a cross direction (disparity guides distribution in ACS) and a parallel direction (disparity guides disparity in APS). Moreover, each branch also learns confidences to dynamically refine its provided supervisions. In ACS, the predicted disparity is softened into a unimodal distribution, and the lower the confidence, the smoother the distribution. In APS, the incorrect predictions are suppressed by lowering the weights of those with low confidence. With the adaptive bidirectional learning, the two branches enjoy well-tuned mutual supervisions, and eventually converge on a consistent and more accurate disparity estimation. The experimental results on four public datasets demonstrate our superior accuracy over other state-of-the-arts with a relative decrease of averaged disparity error by at least 9.76%. Hongkuan Shi, Zhiwei Wang 0002, Dun Li, Xin Yang 0008, Qiang Li 0018 |
IEEE Trans. Medical Imaging | 5 |
| 2023 | FedMix: Mixed Supervised Federated Learning for Medical Image SegmentationabstractThe purpose of federated learning is to enable multiple clients to jointly train a machine learning model without sharing data. However, the existing methods for training an image segmentation model have been based on an unrealistic assumption that the training set for each local client is annotated in a similar fashion and thus follows the same image supervision level. To relax this assumption, in this work, we propose a label-agnostic unified federated learning framework, named FedMix, for medical image segmentation based on mixed image labels. In FedMix, each client updates the federated model by integrating and effectively making use of all available labeled data ranging from strong pixel-level labels, weak bounding box labels, to weakest image-level class labels. Based on these local models, we further propose an adaptive weight assignment procedure across local clients, where each client learns an aggregation weight during the global model update. Compared to the existing methods, FedMix not only breaks through the constraint of a single level of image supervision but also can dynamically adjust the aggregation weight of each local client, achieving rich yet discriminative feature representations. Experimental results on multiple publicly-available datasets validate that the proposed FedMix outperforms the state-of-the-art methods by a large margin. In addition, we demonstrate through experiments that FedMix is extendable to multi-class medical image segmentation and much more feasible in clinical scenarios. The code is available at: https://github.com/Jwicaksana/FedMix. Jeffry Wicaksana, Zengqiang Yan, Xijie Huang, Huimin Wu 0001, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Unsupervised Cross-Modality Adaptation via Dual Structural-Oriented Guidance for 3D Medical Image SegmentationabstractDeep convolutional neural networks (CNNs) have achieved impressive performance in medical image segmentation; however, their performance could degrade significantly when being deployed to unseen data with heterogeneous characteristics. Unsupervised domain adaptation (UDA) is a promising solution to tackle this problem. In this work, we present a novel UDA method, named dual adaptation-guiding network (DAG-Net), which incorporates two highly effective and complementary structural-oriented guidance in training to collaboratively adapt a segmentation model from a labelled source domain to an unlabeled target domain. Specifically, our DAG-Net consists of two core modules: 1) Fourier-based contrastive style augmentation (FCSA) which implicitly guides the segmentation network to focus on learning modality-insensitive and structural-relevant features, and 2) residual space alignment (RSA) which provides explicit guidance to enhance the geometric continuity of the prediction in the target modality based on a 3D prior of inter-slice correlation. We have extensively evaluated our method with cardiac substructure and abdominal multi-organ segmentation for bidirectional cross-modality adaptation between MRI and CT images. Experimental results on two different tasks demonstrate that our DAG-Net greatly outperforms the state-of-the-art UDA approaches for 3D medical image segmentation on unlabeled target images. Junlin Xian, Dandan Tu, Senhua Zhu, Changzheng Zhang, Xiaowu Liu, Xin Li 0001, Xin Yang 0008 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Region Separable Stereo MatchingabstractConvolutional neural networks (CNNs) have shown attractive performance for stereo matching. However, spatially shared convolution weights of CNN-based methods usually face a dilemma that the convolution weights suitable for aggregating contextual information in smooth regions often blur local matching details of textured regions and vice versa. This paper tries to find a way out of the dilemma via a novel region separable stereo matching (RSSM) method, which is universally applicable to CNN stereo models based on 4D cost volumes and can greatly improve the accuracy and efficiency of existing models. The key idea of our method is to automatically group image pixels into regions according to the gradients, and then construct and process the respective cost volume of each region separately. To perform cost aggregation, we propose a two-stage network consisted of regional grouping aggregation (RGA) and regional fusion aggregation (RFA). In RGA, convolutions are grouped in channel-wise, and each group of convolutions learn dedicated weights for the corresponding region via regional supervision. Through RGA, each group of convolutions can extract the most representative features from the corresponding region. In RFA, we combine matching clues of all convolution groups from RGA to output the final prediction map. We further extend the idea of regional grouping to feature extraction and modify the skip connection in aggregation networks to better adapt our method to stereo matching models. Experimental results on five public datasets show that our method can significantly improve several state-of-the-art 3D CNN based stereo models. Junda Cheng, Xin Yang 0008, Yuechuan Pu, Peng Guo 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Path-Analysis-Based Reinforcement Learning Algorithm for Imitation FilmingabstractImitation filming has been applied to autonomous filming by mimicking human operators. To imitate the operation of cameramen when filming multiple human actions, existing methods plan the camera motion through time series prediction or train multiple models to handle a particular style in a specific situation. As a result, these methods require various settings to adapt to different scenarios. In this work, we overcome such limitations and propose an end-to-end imitation learning framework for drone cinematography systems. The framework consists of two main components: (1) an efficient motion feature extraction module for generating a compact motion feature space, (2) a path-analysis-based reinforcement learning (PABRL) algorithm for imitating multiple filming styles from demonstrations and incorporating aesthetical features for improved perspective shots. Our PABRL method is based on the actor–critic network, which regards multiple human motion variables, camera translations, and image composition as inputs and then outputs an aesthetical filming strategy related to the subject motion. In addition, we propose an attention mechanism and a long–short-term rewarding function to enhance the motion feature space and the integrity of the generated trajectory, respectively. Extensive experimental results in simulated and real outdoor environments demonstrate that compared with state-of-the-art methods, our method can achieve 69.8% higher performance in terms of trajectory planning accuracy while successfully incorporating aesthetical features into the captured videos. Yuanjie Dang, Chong Huang 0005, Peng Chen 0008, Ronghua Liang, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 5 |
| 2023 | CR-LDSO: Direct Sparse LiDAR-Assisted Visual Odometry With Cloud ReusingabstractLiDAR-assisted visual odometry (VO) is a widely-used solution for pose estimation and mapping. However, most existing LiDAR-assisted VO systems could suffer from the problems of 1) lacking distinctive and evenly distributed pixels for tracking due to the sparsity of LiDAR points and limited FOV overlap between a camera and LiDAR, and 2) nontrivial errors when processing LiDAR point clouds. To address above problems, we present CR-LDSO, a direct sparse LiDAR-assisted VO with the core parts being: 1) a novel cloud reusing method with point extraction/re-extraction to increase both the camera-LiDAR FOV overlap and the number of high-quality tracking pixels and 2) an occlusion removal method to exclude mismatching pixels due to occluded 3D object from sliding-window optimization and a point extraction strategy without depth interpolation. Extensive experimental results on public datasets demonstrates the superiority of our method to the existing state-of-the-art methods. Zikang Yuan, Junda Cheng, Xin Yang 0008 |
IEEE Trans. Multim. | 3 |
| 2022 | Attention Concatenation Volume for Accurate and Efficient Stereo MatchingabstractStereo matching is a fundamental building block for many vision and robotics applications. An informative and concise cost volume representation is vital for stereo matching of high accuracy and efficiency. In this paper, we present a novel cost volume construction method which generates attention weights from correlation clues to suppress redundant information and enhance matching-related information in the concatenation volume. To generate reliable attention weights, we propose multi-level adaptive patch matching to improve the distinctiveness of the matching cost at different disparities even for textureless regions. The proposed cost volume is named attention concatenation volume (ACV) which can be seamlessly embedded into most stereo matching networks, the resulting networks can use a more lightweight aggregation network and meanwhile achieve higher accuracy, e.g. using only 1/25 parameters of the aggregation network can achieve higher accuracy for GwcNet. Furthermore, we design a highly accurate network (ACVNet) based on our ACV, which achieves state-of-the-art performance on several benchmarks. The code is available at https://github.com/gangweiX/ACVNet. Gangwei Xu, Junda Cheng, Peng Guo 0001, Xin Yang 0008 |
CVPR | 4 |
| 2022 | PUA-MOS: End-to-End Point-wise Uncertainty Weighted Aggregation for Moving Object SegmentationabstractSegmenting moving objects in the 3D LiDAR point cloud can provide important guidance to localization, mapping and decision-making for self-driving vehicles. As for the conventional approaches to point cloud segmentation, they rely on semantic-level information, which makes it inevitable for long-tail problems to arise as there are always unseen types of objects on the road. To achieve moving segmentation while avoiding the reliance on the object category, the point motion is identified in this paper by fully exploring and aggregating the point-level geometric consistency in sequential point clouds. More specifically, an end-to-end point-wise uncertainty weighted aggregation approach known as PUA-MOS is proposed to segment the moving points in 3D LiDAR Data. Our method is applicable to estimate point-wise moving mask, scene flow and rigid-body transformation simultaneously in a coarse- to-fine network, where the relations between each prediction are implicitly learned. To explicitly model the inner and inter relations across these predictions among all points, the point- wise estimation and the average value of the same motion points are aggregated according to a predicted uncertainty. Then, the aggregated estimation is fed again into the next-level fusion, where the points will be re-segmented using the aggregated mask from the last level. Through iterative joint aggregation, our PUA-MOS outperforms the previous methods significantly on both KITTI [4] and Waymo [26] datasets. The code will be provided to generate the moving segmentation labels on both datasets for reproduction. Peiliang Li 0001, Xiaozhi Chen, Xin Yang 0008 |
IROS | 4 |
| 2022 | Dual-Distribution Discrepancy for Anomaly Detection in Chest X-Rays
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
MICCAI (3) | 3 |
| 2022 | Subspace-PnP: A Geometric Constraint Loss for Mutual Assistance of Depth and Optical Flow Estimation
Tianyu Hao, Qingjie Wang, Peng Guo 0001, Xin Yang 0008 |
Int. J. Comput. Vis. | 5 |
| 2022 | Convolutional-capsule network for gastrointestinal endoscopy image classificationabstractAutomated diagnosis of digestive tract diseases from gastrointestinal endoscopy images is of high importance for improving the diagnosis accuracy and efficiency. The current mainstream methods for image classification of digestive tract endoscopy images are based on Convolutional Neural Networks (CNNs). However, due to their inherent defects, CNNs are not strong enough in learning deformation-invariant global features which is essential in gastrointestinal endoscopic image classification. To solve this problem, in this paper we present a two-stage endoscopic image classification method which can effectively combine complementary advantages of midlevel CNN features and a capsule network. Specifically, the core of our method is a lesion-aware CNN feature extraction module which can encode sufficiently detailed information of lesions in midlevel CNN features and in turn enable the subsequent capsule classification network to effectively learn deformation-invariant relationships between image entities. Extensive experiments demonstrate the superiority of our method to the state-of-the-art methods with the classification accuracy of 94.83% on the Kvasir v2 data set and the classification accuracy of 85.99% on the HyperKvasir data set. Wei Wang 0355, Xin Yang 0008, Xin Li 0001, Jinhui Tang 0001 |
Int. J. Intell. Syst. | 2 |
| 2022 | One-Shot Imitation Drone Filming of Human Motion VideosabstractImitation learning has recently been applied to mimic the operation of a cameraman in existing autonomous camera systems. To imitate a certain demonstration video, existing methods require users to collect a significant number of training videos with a similar filming style. Because the trained model is style-specific, it is challenging to generalize the model to imitate other videos with a different filming style. To address this problem, we propose a framework that we term "one-shot imitation filming", which can imitate a filming style by "seeing" only a single demonstration video of the target style without style-specific model training. This is achieved by two key enabling techniques: 1) filming style feature extraction, which encodes sequential cinematic characteristics of a variable-length video clip into a fixed-length feature vector; and 2) camera motion prediction, which dynamically plans the camera trajectory to reproduce the filming style of the demo video. We implemented the approach with a deep neural network and deployed it on a 6 degrees of freedom (DOF) drone system by first predicting the future camera motions, and then converting them into the drone's control commands via an odometer. Our experimental results on comprehensive datasets and showcases exhibit that the proposed approach achieves significant improvements over conventional baselines, and our approach can mimic the footage of an unseen style with high fidelity. Chong Huang 0005, Yuanjie Dang, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Cell Localization and Counting Using Direction Field MapabstractAutomatic cell counting in pathology images is challenging due to blurred boundaries, low-contrast, and overlapping between cells. In this paper, we train a convolutional neural network (CNN) to predict a two-dimensional direction field map and then use it to localize cell individuals for counting. Specifically, we define a direction field on each pixel in the cell regions (obtained by dilating the original annotation in terms of cell centers) as a two-dimensional unit vector pointing from the pixel to its corresponding cell center. Direction field for adjacent pixels in different cells have opposite directions departing from each other, while those in the same cell region have directions pointing to the same center. Such unique property is used to partition overlapped cells for localization and counting. To deal with those blurred boundaries or low contrast cells, we set the direction field of the background pixels to be zeros in the ground-truth generation. Thus, adjacent pixels belonging to cells and background will have an obvious difference in the predicted direction field. To further deal with cells of varying density and overlapping issues, we adopt geometry adaptive (varying) radius for cells of different densities in the generation of ground-truth direction field map, which guides the CNN model to separate cells of different densities and overlapping cells. Extensive experimental results on three widely used datasets (i.e., VGG Cell, CRCHistoPhenotype2016, and MBM datasets) demonstrate the effectiveness of the proposed approach. Yajie Chen, Dingkang Liang, Xiang Bai, Yongchao Xu, Xin Yang 0008 |
IEEE J. Biomed. Health Informatics | 5 |
| 2022 | Customized Federated Learning for Multi-Source Decentralized Medical Image ClassificationabstractThe performance of deep networks for medical image analysis is often constrained by limited medical data, which is privacy-sensitive. Federated learning (FL) alleviates the constraint by allowing different institutions to collaboratively train a federated model without sharing data. However, the federated model is often suboptimal with respect to the characteristics of each client's local data. Instead of training a single global model, we propose Customized FL (CusFL), for which each client iteratively trains a client-specific/private model based on a federated global model aggregated from all private models trained in the immediate previous iteration. Two overarching strategies employed by CusFL lead to its superior performance: 1) the federated model is mainly for feature alignment and thus only consists of feature extraction layers; 2) the federated feature extractor is used to guide the training of each private model. In that way, CusFL allows each client to selectively learn useful knowledge from the federated model to improve its personalized model. We evaluated CusFL on multi-source medical image datasets for the identification of clinically significant prostate cancer and the classification of skin lesions. Jeffry Wicaksana, Zengqiang Yan, Xin Yang 0008, Yang Liu 0165, Lixin Fan, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | RGB-D DSO: Direct Sparse Odometry With RGB-D Cameras for Indoor ScenesabstractVisual odometry (VO) is a fundamental technique for many robotics and augmented reality (AR) applications. However, most existing RGB-D VO systems suffer from large performance degradation when large occlusions are present and/or a large portion of depth values are invalid due to the limited range of an RGB-D camera, prohibiting the usage of most systems in practical applications. To address above two problems, we present RGB-D DSO, an RGB-D direct sparse odometry with the core part being sliding-window optimization with occlusion removal and a depth refinement module. Occlusion removal excludes negative effects arising from occluded objects when minimizing the final energy function for camera pose tracking. Depth refinement ensures sufficient valid depth values uniformly distributed for the depth map of a keyframe. Experimental results on three public datasets demonstrate that our method achieves smaller tracking error than most existing state-of-the-art methods. Meanwhile, our system takes only 21.93 ms to track a frame, which is faster than most existing methods. Zikang Yuan, Ken Cheng, Jinhui Tang 0001, Xin Yang 0008 |
IEEE Trans. Multim. | 4 |
| 2021 | Feature-Level Collaboration: Joint Unsupervised Learning of Optical Flow, Stereo Depth and Camera MotionabstractPrecise estimation of optical flow, stereo depth and camera motion are important for the real-world 3D scene understanding and visual perception. Since the three tasks are tightly coupled with the inherent 3D geometric constraints, current studies have demonstrated that the three tasks can be improved through jointly optimizing geometric loss functions of several individual networks. In this paper, we show that effective feature-level collaboration of the networks for the three respective tasks could achieve much greater performance improvement for all three tasks than only loss-level joint optimization. Specifically, we propose a single network to combine and improve the three tasks. The network extracts the features of two consecutive stereo images, and simultaneously estimates optical flow, stereo depth and camera motion. The whole network mainly contains four parts: (I) a feature-sharing encoder to extract features of input images, which can enhance features’ representation ability; (II) a pooled decoder to estimate both optical flow and stereo depth; (III) a camera pose estimation module which fuses optical flow and stereo depth information; (IV) a cost volume complement module to improve the performance of optical flow in static and occluded regions. Our method achieves state-of-the-art performance among the joint unsupervised methods, including optical flow and stereo depth estimation on KITTI 2012 and 2015 benchmarks, and camera motion estimation on KITTI VO dataset. Qingjie Wang, Tianyu Hao, Peng Guo 0001, Xin Yang 0008 |
CVPR | 5 |
| 2021 | Towards Robust Dual-View Transformation via Densifying Sparse Supervision for Mammography Lesion Matching
Junlin Xian, Zhiwei Wang 0002, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (5) | 4 |
| 2021 | Comprehensive Linguistic-Visual Composition Network for Image RetrievalabstractComposing text and image for image retrieval (CTI-IR) is a new yet challenging task, for which the input query is not the conventional image or text but a composition, i.e., a reference image and its corresponding modification text. The key of CTI-IR lies in how to properly compose the multi-modal query to retrieve the target image. In a sense, pioneer studies mainly focus on composing the text with either the local visual descriptor or global feature of the reference image. However, they overlook the fact that the text modifications are indeed diverse, ranging from the concrete attribute changes, like "change it to long sleeves", to the abstract visual property adjustments, e.g., "change the style to professional". Thus, simply emphasizing the local or global feature of the reference image for the query composition is insufficient. In light of the above analysis, we propose a Comprehensive Linguistic-Visual Composition Network (CLVC-Net) for image retrieval. The core of CLVC-Net is that it designs two composition modules: fine-grained local-wise composition module and fine-grained global-wise composition module, targeting comprehensive multi-modal compositions. Additionally, a mutual enhancement module is designed to promote local-wise and global-wise composition processes by forcing them to share knowledge with each other. Extensive experiments conducted on three real-world datasets demonstrate the superiority of our CLVC-Net. We released the codes to benefit other researchers. Haokun Wen, Xuemeng Song, Xin Yang 0008, Yibing Zhan, Liqiang Nie |
SIGIR | 3 |
| 2021 | DENAO: Monocular Depth Estimation Network With Auxiliary Optical FlowabstractEstimating depth from multi-view images captured by a localized monocular camera is an essential task in computer vision and robotics. In this study, we demonstrate that learning a convolutional neural network (CNN) for depth estimation with an auxiliary optical flow network and the epipolar geometry constraint can greatly benefit the depth estimation task and in turn yield large improvements in both accuracy and speed. Our architecture is composed of two tightly-coupled encoder-decoder networks, i.e., an optical flow net and a depth net, the core part being a list of exchange blocks between the two nets and an epipolar feature layer in the optical flow net to improve predictions of both depth and optical flow. Our architecture allows to input arbitrary number of multiview images with a linearly growing time cost for optical flow and depth estimation. Experimental result on five public datasets demonstrates that our method, named DENAO, runs at 38.46fps on a single Nvidia TITAN Xp GPU which is 5.15X ∼ 142X faster than the state-of-the-art depth estimation methods Meanwhile, our DENAO can concurrently output predictions of both depth and optical flow, and performs on par with or outperforms the state-of-the-art depth estimation methods and optical flow methods. Xin Yang 0008, Qizeng Jia, Chunyuan Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Variation-Aware Federated Learning With Multi-Source Decentralized Medical Image DataabstractPrivacy concerns make it infeasible to construct a large medical image dataset by fusing small ones from different sources/institutions. Therefore, federated learning (FL) becomes a promising technique to learn from multi-source decentralized data with privacy preservation. However, the cross-client variation problem in medical image data would be the bottleneck in practice. In this paper, we propose a variation-aware federated learning (VAFL) framework, where the variations among clients are minimized by transforming the images of all clients onto a common image space. We first select one client with the lowest data complexity to define the target image space and synthesize a collection of images through a privacy-preserving generative adversarial network, called PPWGAN-GP. Then, a subset of those synthesized images, which effectively capture the characteristics of the raw images and are sufficiently distinct from any raw image, is automatically selected for sharing with other clients. For each client, a modified CycleGAN is applied to translate its raw images to the target image space defined by the shared synthesized images. In this way, the cross-client variation problem is addressed with privacy preservation. We apply the framework for automated classification of clinically significant prostate cancer and evaluate it using multi-source decentralized apparent diffusion coefficient (ADC) image data. Experimental results demonstrate that the proposed VAFL framework stably outperforms the current horizontal FL framework. As VAFL is independent of deep learning architectures for classification, we believe that the proposed framework is widely applicable to other medical image classification tasks. Zengqiang Yan, Jeffry Wicaksana, Zhiwei Wang 0002, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Fast Depth Prediction and Obstacle Avoidance on a Monocular Drone Using Probabilistic Convolutional Neural NetworkabstractRecent studies employ advanced deep convolutional neural networks (CNNs) for monocular depth perception, which can hardly run efficiently on small drones that rely on low/middle-grade GPU(e.g. TX2 and 1050Ti) for computation. In addition, the methods which can effectively and efficiently produce probabilistic depth prediction with a measure of model confidence have not been well studied. The lack of such a method could yield erroneous, sometimes fatal, decisions in drone applications (e.g. selecting a waypoint in a region with a large depth yet a low estimation confidence). This paper presents a real-time onboard approach for monocular depth prediction and obstacle avoidance with a lightweight probabilistic CNN (pCNN), which will be ideal for use in a lightweight energy-efficient drone. For each video frame, our pCNN can efficiently predict its depth map and the corresponding confidence. The accuracy of our lightweight pCNN is greatly boosted by integrating sparse depth estimation from a visual odometry into the network for guiding dense depth and confidence inference. The estimated depth map is transformed into Ego Dynamic Space (EDS) by embedding both dynamic motion constraints of a drone and the confidence values into the spatial depth map. Traversable waypoints are automatically computed in EDS based on which appropriate control inputs for the drone are produced. Extensive experimental results on public datasets demonstrate that our depth prediction method runs at 12Hz and 45Hz on TX2 and 1050Ti GPU respectively, which is 1.8X~5.6X faster than the state-of-the-art methods and achieves better depth estimation accuracy. We also conducted experiments of obstacle avoidance in both simulated and real environments to demonstrate the superiority of our method to the baseline methods. Xin Yang 0008, Yuanjie Dang, Hongcheng Luo, Yuesheng Tang, Chunyuan Liao, Peng Chen 0008, Kwang-Ting Cheng |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Robust and Efficient RGB-D SLAM in Dynamic EnvironmentsabstractSimultaneous localization and mapping (SLAM) using an RGB-D camera is a key enabling technique for many augmented reality (AR) applications. However, most existing RGB-D SLAM methods could fail in dynamic scenarios due to non-trivial pose estimation errors arising from moving objects. In this study, we present an accurate and robust RGB-D SLAM system for dynamic scenarios which can run real-time on a single dual-core CPU. The core of our system is a robust and efficient dynamic keypoint exclusion method which consists of three steps: 1) grouping spatially and appearance related pixels of a keyframe into regions; 2) identifying dynamic regions by checking motion consistency of keypoints in every region; 3) excluding keypoints in the identified dynamic regions as well as the matching points in the 3D local map. The dynamic keypoint exclusion method can be easily integrated into any keypoint based RGB-D SLAM system for improving the accuracy and robustness in dynamic scenes with trivial time increase (16.6ms per frame). Experimental results on the TUM dataset demonstrates that our method which runs on an Intel i7-4900 CPU is even 2.3X faster than the state-of-the-art method DS-SLAM [1] which runs parallel on a P4000 GPU and a comparable CPU. In addition, our system outperforms the state-of-the-art methods [1]–[4] in terms of smaller absolute trajectory errors (ATE). We also apply our system to a real AR application and live experiments with a hand-held RGB-D camera demonstrate the robustness and generalizability of our method in practical scenarios.11A demo video is provided onhttps://github.com/cc-qy/Dynamic-RGB-D-SLAM Xin Yang 0008, Zikang Yuan, Dongfu Zhu, Chunyuan Liao |
IEEE Trans. Multim. | 1 |
| 2021 | Attribute-wise Explainable Fashion Compatibility ModelingabstractWith the boom of the fashion market and people’s daily needs for beauty, clothing matching has gained increased research attention. In a sense, tackling this problem lies in modeling the human notions of the compatibility between fashion items, i.e., Fashion Compatibility Modeling (FCM), which plays an important role in a wide bunch of commercial applications, including clothing recommendation and dressing assistant. Recent advances in multimedia processing have shown remarkable effectiveness in accurate compatibility evaluation. However, these studies work like a black box and cannot provide appropriate explanations, which are indeed of importance for gaining users’ trust and improving their experience. In fact, fashion experts usually explain the compatibility evaluation through the matching patterns between fashion attributes (e.g., a silk tank top cannot go with a knit dress). Inspired by this, we devise an attribute-wise explainable FCM solution, named ExFCM , which can simultaneously generate the item-level compatibility evaluation for input fashion items and the attribute-level explanations for the evaluation result. In particular, ExFCM consists of two key components: attribute-wise representation learning and attribute interaction modeling. The former works on learning the region-aware attribute representation for each item with the threshold global average pooling. Besides, the latter is responsible for compiling the attribute-level matching signals into the overall compatibility evaluation adaptively with the attentive interaction mechanism. Note that ExFCM is trained without any attribute-level compatibility annotations, which facilitates its practical applications. Extensive experiments on two real-world datasets validate that ExFCM can generate more accurate compatibility evaluations than the existing methods, together with reasonable explanations. Xin Yang 0008, Xuemeng Song, Fuli Feng, Haokun Wen, Ling-Yu Duan, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Celeb-DF: A Large-Scale Challenging Dataset for DeepFake ForensicsabstractAI-synthesized face-swapping videos, commonly known as DeepFakes, is an emerging problem threatening the trustworthiness of online information. The need to develop and evaluate DeepFake detection algorithms calls for datasets of DeepFake videos. However, current DeepFake datasets suffer from low visual quality and do not resemble DeepFake videos circulated on the Internet. We present a new large-scale challenging DeepFake video dataset, Celeb-DF, which contains 5,639 high-quality DeepFake videos of celebrities generated using improved synthesis process. We conduct a comprehensive evaluation of DeepFake detection methods and datasets to demonstrate the escalated level of challenges posed by Celeb-DF. Yuezun Li, Xin Yang 0008, Pu Sun 0001, Honggang Qi, Siwei Lyu |
CVPR | 2 |
| 2020 | D2VO: Monocular Deep Direct Visual OdometryabstractIn this paper, we present a novel deep learning and direct method based monocular visual odometry system named D2VO. Our system reconstructs the dense depth map of each keyframe and tracks camera poses based on these keyframes. Combining direct method and deep learning, both tracking and mapping of the system could benefit from the geometric measurement and semantic information. For each input frame, a feature pyramid is built and shared by both tracking and mapping process. The depth map of keyframe is efficiently estimated from coarse to fine with the followed multi-view hierarchical depth estimation network. We optimize the camera pose by minimizing photometric error between re-projected features of each frame and its reference keyframe with bundle adjustment. Experimental results on TUM dataset demonstrate that our approach outperforms the state-of-the-art methods on both tracking and mapping. Qizeng Jia, Yuechuan Pu, Junda Cheng, Chunyuan Liao, Xin Yang 0008 |
IROS | 6 |
| 2020 | Multi-phase and Multi-level Selective Feature Fusion for Automated Pancreas Segmentation from CT Images
Xixi Jiang, Qingqing Luo, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (4) | 8 |
| 2020 | Generative Attribute Manipulation Scheme for Flexible Fashion SearchabstractIn this work, we aim to investigate the practical task of flexible fashion search with attribute manipulation, where users can retrieve the target fashion items by replacing the unwanted attributes of an available query image with the desired ones (e.g., changing the collar attribute from v-neck to round). Although several pioneer efforts have been dedicated to fulfilling the task, they mainly ignore the potential of generative models in enhancing the visual understanding of target fashion items. To this end, we propose an end-to-end generative attribute manipulation scheme, which consists of a generator and a discriminator. The generator works on producing the prototype image that meets the user's requirement of attribute manipulation over the query image with the regularization of visual-semantic consistency and pixel-wise consistency. Besides, the discriminator aims to jointly fulfill the semantic learning towards correct attribute manipulation and adversarial metric learning for fashion search. Pertaining to the adversarial metric learning, we provide two general paradigms: the pair-based scheme and the triplet-based scheme, where the fake generated prototype images that closely resemble the ground truth images of target items are incorporated as hard negative samples to boost the model performance. Extensive experiments on two real-world datasets verify the effectiveness of our scheme. Xin Yang 0008, Xuemeng Song, Xianjing Han, Haokun Wen, Jie Nie, Liqiang Nie |
SIGIR | 1 |
| 2020 | Bi-Real Net: Binarizing Deep Network Towards Real-Network Performance
Zechun Liu, Wenhan Luo, Baoyuan Wu, Xin Yang 0008, Wei Liu 0005, Kwang-Ting Cheng |
Int. J. Comput. Vis. | 4 |
| 2020 | Semi-supervised mp-MRI data synthesis with StitchLayer and auxiliary distance maximization
Zhiwei Wang 0002, Yi Lin 0009, Kwang-Ting Cheng, Xin Yang 0008 |
Medical Image Anal. | 4 |
| 2020 | Bi-Modality Medical Image Synthesis Using Semi-Supervised Sequential Generative Adversarial NetworksabstractIn this paper, we propose a bi-modality medical image synthesis approach based on sequential generative adversarial network (GAN) and semi-supervised learning. Our approach consists of two generative modules that synthesize images of the two modalities in a sequential order. A method for measuring the synthesis complexity is proposed to automatically determine the synthesis order in our sequential GAN. Images of the modality with a lower complexity are synthesized first, and the counterparts with a higher complexity are generated later. Our sequential GAN is trained end-to-end in a semi-supervised manner. In supervised training, the joint distribution of bi-modality images are learned from real paired images of the two modalities by explicitly minimizing the reconstruction losses between the real and synthetic images. To avoid overfitting limited training images, in unsupervised training, the marginal distribution of each modality is learned based on unpaired images by minimizing the Wasserstein distance between the distributions of real and fake images. We comprehensively evaluate the proposed model using two synthesis tasks based on three types of evaluate metrics and user studies. Visual and quantitative results demonstrate the superiority of our method to the state-of-the-art methods, and reasonable visual quality and clinical significance. Code is made publicly available at https://github.com/hust- linyi/Multimodal-Medical-Image-Synthesis. Xin Yang 0008, Yi Lin 0009, Zhiwei Wang 0002, Xin Li 0001, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 1 |
| 2020 | Multi-Task Siamese Network for Retinal Artery/Vein Separation via Deep Convolution Along VesselabstractVascular tree disentanglement and vessel type classification are two crucial steps of the graph-based method for retinal artery-vein (A/V) separation. Existing approaches treat them as two independent tasks and mostly rely on ad hoc rules (e.g. change of vessel directions) and hand-crafted features (e.g. color, thickness) to handle them respectively. However, we argue that the two tasks are highly correlated and should be handled jointly since knowing the A/V type can unravel those highly entangled vascular trees, which in turn helps to infer the types of connected vessels that are hard to classify based on only appearance. Therefore, designing features and models isolatedly for the two tasks often leads to a suboptimal solution of A/V separation. In view of this, this paper proposes a multi-task siamese network which aims to learn the two tasks jointly and thus yields more robust deep features for accurate A/V separation. Specifically, we first introduce Convolution Along Vessel (CAV) to extract the visual features by convolving a fundus image along vessel segments, and the geometric features by tracking the directions of blood flow in vessels. The siamese network is then trained to learn multiple tasks: i) classifying A/V types of vessel segments using visual features only, and ii) estimating the similarity of every two connected segments by comparing their visual and geometric features in order to disentangle the vasculature into individual vessel trees. Finally, the results of two tasks mutually correct each other to accomplish final A/V separation. Experimental results demonstrate that our method can achieve accuracy values of 94.7%, 96.9%, and 94.5% on three major databases (DRIVE, INSPIRE, WIDE) respectively, which outperforms recent state-of-the-arts. Zhiwei Wang 0002, Xixi Jiang, Jingen Liu, Kwang-Ting Cheng, Xin Yang 0008 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Enabling a Single Deep Learning Model for Accurate Gland Instance Segmentation: A Shape-Aware Adversarial Learning FrameworkabstractSegmenting gland instances in histology images is highly challenging as it requires not only detecting glands from a complex background but also separating each individual gland instance with accurate boundary detection. However, due to the boundary uncertainty problem in manual annotations, pixel-to-pixel matching based loss functions are too restrictive for simultaneous gland detection and boundary detection. State-of-the-art approaches adopted multi-model schemes, resulting in unnecessarily high model complexity and difficulties in the training process. In this paper, we propose to use one single deep learning model for accurate gland instance segmentation. To address the boundary uncertainty problem, instead of pixel-to-pixel matching, we propose a segment-level shape similarity measure to calculate the curve similarity between each annotated boundary segment and the corresponding detected boundary segment within a fixed searching range. As the segment-level measure allows location variations within a fixed range for shape similarity calculation, it has better tolerance to boundary uncertainty and is more effective for boundary detection. Furthermore, by adjusting the radius of the searching range, the segment-level shape similarity measure is able to deal with different levels of boundary uncertainty. Therefore, in our framework, images of different scales are down-sampled and integrated to provide both global and local contextual information for training, which is helpful in segmenting gland instances of different sizes. To reduce the variations of multi-scale training images, by referring to adversarial domain adaptation, we propose a pseudo domain adaptation framework for feature alignment. By constructing loss functions based on the segment-level shape similarity measure, combining with the adversarial loss function, the proposed shape-aware adversarial learning framework enables one single deep learning model for gland instance segmentation. Experimental results on the 2015 MICCAI Gland Challenge dataset demonstrate that the proposed framework achieves state-of-the-art performance with one single deep learning model. As the boundary uncertainty problem widely exists in medical image segmentation, it is broadly applicable to other applications. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Learning to Film From Professional Human Motion VideosabstractWe investigate the problem of 6 degrees of freedom (DOF) camera planning for filming professional human motion videos using a camera drone. Existing methods either plan motions for only a pan-tilt-zoom (PTZ) camera, or adopt ad-hoc solutions without carefully considering the impact of video contents and previous camera motions on the future camera motions. As a result, they can hardly achieve satisfactory results in our drone cinematography task. In this study, we propose a learning-based framework which incorporates the video contents and previous camera motions to predict the future camera motions that enable the capture of professional videos. Specifically, the inputs of our framework are video contents which are represented using subject-related feature based on 2D skeleton and scene-related features extracted from background RGB images, and camera motions which are represented using optical flows. The correlation between the inputs and output future camera motions are learned via a sequence-to-sequence convolutional long short-term memory (Seq2Seq ConvLSTM) network from a large set of video clips. We deploy our approach to a real drone cinematography system by first predicting the future camera motions, and then converting them to the drone's control commands via an odometer. Our experimental results on extensive datasets and showcases exhibit significant improvements in our approach over conventional baselines and our approach can successfully mimic the footage of a professional cameraman. Chong Huang 0005, David Chuan-En Lin, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
CVPR | 6 |
| 2019 | Exposing Deep Fakes Using Inconsistent Head PosesabstractIn this paper, we propose a new method to expose AI-generated fake face images or videos (commonly known as the Deep Fakes). Our method is based on the observations that Deep Fakes are created by splicing synthesized face region into the original image, and in doing so, introducing errors that can be revealed when 3D head poses are estimated from the face images. We perform experiments to demonstrate this phenomenon and further develop a classification method based on this cue. Using features based on this cue, an SVM classifier is evaluated using a set of real face images and Deep Fakes. Xin Yang 0008, Yuezun Li, Siwei Lyu |
ICASSP | 1 |
| 2019 | MetaPruning: Meta Learning for Automatic Neural Network Channel PruningabstractIn this paper, we propose a novel meta learning approach for automatic channel pruning of very deep neural networks. We first train a PruningNet, a kind of meta network, which is able to generate weight parameters for any pruned structure given the target network. We use a simple stochastic structure sampling method for training the PruningNet. Then, we apply an evolutionary procedure to search for good-performing pruned networks. The search is highly efficient because the weights are directly generated by the trained PruningNet and we do not need any finetuning at search time. With a single PruningNet trained for the target network, we can search for various Pruned Networks under different constraints with little human participation. Compared to the state-of-the-art pruning methods, we have demonstrated superior performances on MobileNet V1/V2 and ResNet. Codes are available on https://github.com/liuzechun/MetaPruning. Zechun Liu, Haoyuan Mu, Xiangyu Zhang 0005, Zichao Guo, Xin Yang 0008, Kwang-Ting Cheng, Jian Sun 0001 |
ICCV | 5 |
| 2019 | Learning to Capture a Film-Look Video with a Camera DroneabstractThe development of intelligent drones has simplified aerial filming and provided smarter assistant tools for users to capture a film-look footage. Existing methods of autonomous aerial filming either specify predefined camera movements for a drone to capture a footage, or employ heuristic approaches for camera motion planning. However, both predefined movements and heuristically planned motions are hardly able to provide cinematic footages for various dynamic scenarios. In this paper, we propose a data-driven learning-based approach, which can imitate a professional cameraman's intention for capturing a film-look aerial footage of a single subject in real-time. We model the decision-making process of the cameraman with two steps: 1) we train a network to predict the future image composition and camera position, and 2) our system then generates control commands to achieve the desired shot framing. At the system level, we deploy our algorithm on the limited resources of a drone and demonstrate the feasibility of running automatic filming onboard in real-time. Our experiments show how our data-driven planning approach achieves film-look footages and successfully mimics the work of a professional cameraman. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
ICRA | 5 |
| 2019 | Exposing GAN-synthesized Faces Using Landmark LocationsabstractGenerative adversary networks (GANs) have recently led to highly realistic image synthesis results. In this work, we describe a new method to expose GAN-synthesized images using the locations of the facial landmark points. Our method is based on the observations that the facial parts configuration generated by GAN models are different from those of the real faces, due to the lack of global constraints. We perform experiments demonstrating this phenomenon, and show that an SVM classifier trained using the locations of facial landmark points is sufficient to achieve good classification performance for GAN-synthesized faces. Xin Yang 0008, Yuezun Li, Honggang Qi, Siwei Lyu |
IH&MMSec | 1 |
| 2019 | Automated Pulmonary Embolism Detection from CTPA Images Using an End-to-End Convolutional Neural Network
Yi Lin 0009, Jianchao Su, Jingen Liu, Kwang-Ting Cheng, Xin Yang 0008 |
MICCAI (4) | 7 |
| 2019 | Visual-Inertial State Estimation with Pre-integration Correction for Robust Mobile Augmented RealityabstractMobile devices equipped with a monocular camera and an inertial measurement unit (IMU) are ideal platforms for augmented reality (AR) applications. However, nontrivial noises in low-cost IMUs, which are usually equipped in consumer-level mobile devices, could lead to large errors in pose estimation and in turn significantly degrade the user experience in mobile AR apps. In this study, we propose a novel monocular visual-inertial state estimation approach for robust and accurate pose estimation even for low-cost IMUs. The core of our method is an IMU pre-integration correction approach which effectively reduces the negative impact of IMU noises using the visual constraints in a sliding window and the kinematic constraint. We seamlessly integrate the IMU pre-integration correction module into a tightly-coupled,sliding-window based optimization framework for state estimation. Experimental results on public dataset EUROC demonstrate the superiority of our method to the state-of-the-art VINS-Mono in terms of smaller absolute trajectory errors (ATE) and relative pose errors (RPE). We further apply our method to real AR applications on two types of consumer-level mobile devices equipped with low-cost IMUs, i.e. an off-the-shelf smartphone and an AR glass. Experimental results demonstrate that our method can facilitate robust AR with little drifts on the two devices. Zikang Yuan, Dongfu Zhu, Jinhui Tang 0001, Chunyuan Liao, Xin Yang 0008 |
ACM Multimedia | 6 |
| 2019 | Fast and accurate visual odometry from a monocular camera
Xin Yang 0008, Tangli Xue, Hongcheng Luo, Jiabin Guo |
Frontiers Comput. Sci. | 1 |
| 2019 | Reactive obstacle avoidance of monocular quadrotors with online adapted depth prediction network
Xin Yang 0008, Hongcheng Luo, Yuhao Wu 0010, Chunyuan Liao, Kwang-Ting Cheng |
Neurocomputing | 1 |
| 2019 | A Three-Stage Deep Learning Model for Accurate Retinal Vessel SegmentationabstractAutomatic retinal vessel segmentation is a fundamental step in the diagnosis of eye-related diseases, in which both thick vessels and thin vessels are important features for symptom detection. All existing deep learning models attempt to segment both types of vessels simultaneously by using a unified pixel-wise loss that treats all vessel pixels with equal importance. Due to the highly imbalanced ratio between thick vessels and thin vessels (namely the majority of vessel pixels belong to thick vessels), the pixel-wise loss would be dominantly guided by thick vessels and relatively little influence comes from thin vessels, often leading to low segmentation accuracy for thin vessels. To address the imbalance problem, in this paper, we explore to segment thick vessels and thin vessels separately by proposing a three-stage deep learning model. The vessel segmentation task is divided into three stages, namely thick vessel segmentation, thin vessel segmentation, and vessel fusion. As better discriminative features could be learned for separate segmentation of thick vessels and thin vessels, this process minimizes the negative influence caused by their highly imbalanced ratio. The final vessel fusion stage refines the results by further identifying nonvessel pixels and improving the overall vessel thickness consistency. The experiments on public datasets DRIVE, STARE, and CHASE_DB1 clearly demonstrate that the proposed three-stage deep learning model outperforms the current state-of-the-art vessel segmentation methods. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Real-Time Dense Monocular SLAM With Online Adapted Depth Prediction NetworkabstractConsiderable advances have been achieved in estimating the depth map from a single image via convolutional neural networks (CNNs) during the past few years. Combining depth prediction from CNNs with conventional monocular simultaneous localization and mapping (SLAM) is promising for accurate and dense monocular reconstruction, in particular addressing the two long-standing challenges in conventional monocular SLAM: low map completeness and scale ambiguity. However, depth estimated by pretrained CNNs usually fails to achieve sufficient accuracy for environments of different types from the training data, which are common for certain applications such as obstacle avoidance of drones in unknown scenes. Additionally, inaccurate depth prediction of CNN could yield large tracking errors in monocular SLAM. In this paper, we present a real-time dense monocular SLAM system, which effectively fuses direct monocular SLAM with an online-adapted depth prediction network for achieving accurate depth prediction of scenes of different types from the training data and providing absolute scale information for tracking and mapping. Specifically, on one hand, tracking pose (i.e., translation and rotation) from direct SLAM is used for selecting a small set of highly effective and reliable training images, which acts as ground truth for tuning the depth prediction network on-the-fly toward better generalization ability for scenes of different types. A stage-wise Stochastic Gradient Descent algorithm with a selective update strategy is introduced for efficient convergence of the tuning process. On the other hand, the dense map produced by the adapted network is applied to address scale ambiguity of direct monocular SLAM which in turn improves the accuracy of both tracking and overall reconstruction. The system with assistance of both CPUs and GPUs, can achieve real-time performance with progressively improved reconstruction accuracy. Experimental results on public datasets and live application to obstacle avoidance of drones demonstrate that our method outperforms the state-of-the-art methods with greater map completeness and accuracy, and a smaller tracking error. Hongcheng Luo, Yuhao Wu 0010, Chunyuan Liao, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Multim. | 5 |
| 2019 | Bayesian DeNet: Monocular Depth Prediction and Frame-Wise Fusion With Synchronized UncertaintyabstractUsing deep convolutional neural networks (CNN) to predict the depth from a single image has received considerable attention in recent years due to its impressive performance. However, existing methods process each single image independently without leveraging the multiview information of video sequences in practical scenarios. Properly taking into account multiview information in video sequences beyond individual frames could offer considerable benefits in terms of depth prediction accuracy and robustness. In addition, a meaningful measure of prediction uncertainty is essential for decision making, which is not provided in existing methods. This paper presents a novel video-based depth prediction system based on a monocular camera, named Bayesian DeNet. Specifically, Bayesian DeNet consists of a 59-layer CNN that can concurrently output a depth map and an uncertainty map for each video frame. Each pixel in an uncertainty map indicates the error variance of the corresponding depth estimate. Depth estimates and uncertainties of previous frames are propagated to the current frame based on the tracked camera pose, yielding multiple depth/uncertainty hypotheses for the current frame which are then fused in a Bayesian inference framework for greater accuracy and robustness. Extensive exper-iments on three public datasets demonstrate that our Bayesian DeNet outperforms the state-of-the-art methods for monocular depth prediction. A demo video and code are publicly available.1 Xin Yang 0008, Hongcheng Luo, Chunyuan Liao, Kwang-Ting Cheng |
IEEE Trans. Multim. | 1 |
| 2018 | StitchAD-GAN for Synthesizing Apparent Diffusion Coefficient Images of Clinically Significant Prostate Cancer
Zhiwei Wang 0002, Yi Lin 0009, Chunyuan Liao, Kwang-Ting Cheng, Xin Yang 0008 |
BMVC | 5 |
| 2018 | Bi-Real Net: Enhancing the Performance of 1-Bit CNNs with Improved Representational Capability and Advanced Training Algorithm
Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang 0008, Wei Liu 0005, Kwang-Ting Cheng |
ECCV (15) | 4 |
| 2018 | ACT: An Autonomous Drone Cinematography System for Action ScenesabstractDrones are enabling new forms of cinematography. Aerial filming via drones in action scenes is difficult because it requires users to understand the dynamic scenarios and operate the drone and camera simultaneously. Existing systems allow the user to manually specify the shots and guide the drone to capture footage, while none of them employ aesthetic objectives to automate aerial filming in action scenes. Meanwhile, these drone cinematography systems depend on the external motion capture systems to perceive the human action, which is limited to the indoor environment. In this paper, we propose an Autonomous CinemaTography system “ACT” on the drone platform to address the above the challenges. To our knowledge, this is the first drone camera system which can autonomously capture cinematic shots of action scenes based on limb movements in both indoor and outdoor environments. Our system includes the following novelties. First, we propose an efficient method to extract 3D skeleton points via a stereo camera. Second, we design a real-time dynamical camera planning strategy that fulfills the aesthetic objectives for filming and respects the physical limits of a drone. At the system level, we integrate cameras and GPUs into the limited space of a drone and demonstrate the feasibility of running the entire cinematography system onboard in real-time. Experimental results in both simulation and real-world scenarios demonstrate that our cinematography system “ACT” can capture more expressive video footage of human action than that of a state-of-the-art drone camera system. Chong Huang 0005, Fei Gao 0011, Jie Pan 0004, Weihao Qiu, Peng Chen 0008, Xin Yang 0008, Shaojie Shen, Kwang-Ting Cheng |
ICRA | 7 |
| 2018 | Through-the-Lens Drone FilmingabstractAerial filming in action scenes using a drone is difficult for inexperienced flyers because manipulating a remote controller and meeting the desired image composition are two independent, while concurrent, tasks. Existing systems attempt to utilize wearable GPS-based or infrared-based sensors to track the human movement and to assist in capturing footage. However, these sensors work only in either indoor (infrared-based) or outdoor environments (GPS-based), but not both. In this paper, we introduce a novel drone filming system which integrates monocular 3D human pose estimation and localization into a drone platform to remove the constraints imposed by wearable-sensor-based solutions. Meanwhile, given the estimated position, we propose a novel drone control system, called “through-the-lens drone filming”, to allow a cameraman to conveniently control the drone by manipulating a 3D model in the preview, which closes the gap between the flight control and the viewpoint design. Our system includes two key enabling techniques: 1) subject localization based on visual-inertial fusion, and 2) through-the-lens camera planning. This is the first drone camera system which allows users to capture human actions by manipulating the camera in a virtual environment. From the drone hardware, we integrate a gimbal camera and two GPUs into the limited space of a drone and demonstrate the feasibility of running the entire system onboard with insignificant delays, which are sufficient for filming in our real-time application. Experimental results, in both simulation and real-world scenarios, demonstrate that our techniques can greatly ease camera control and capture better videos. Chong Huang 0005, Yan Kong, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 5 |
| 2018 | A Deep Model with Shape-Preserving Loss for Gland Instance Segmentation
Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
MICCAI (2) | 2 |
| 2018 | Monocular Camera Based Real-Time Dense Mapping Using Generative Adversarial NetworkabstractMonocular simultaneous localization and mapping (SLAM) is a key enabling technique for many computer vision and robotics applications. However, existing methods either can obtain only sparse or semi-dense maps in highly-textured image areas or fail to achieve a satisfactory reconstruction accuracy. In this paper, we present a new method based on a generative adversarial network,named DM-GAN, for real-time dense mapping based on a monocular camera. Specifcally, our depth generator network takes a semidense map obtained from motion stereo matching as a guidance to supervise dense depth prediction of a single RGB image. The depth generator is trained based on a combination of two loss functions, i.e. an adversarial loss for enforcing the generated depth maps to reside on the manifold of the true depth maps and a pixel-wise mean square error (MSE) for ensuring the correct absolute depth values. Extensive experiments on three public datasets demonstrate that our DM-GAN signifcantly outperforms the state-of-the-art methods in terms of greater reconstruction accuracy and higher depth completeness. Xin Yang 0008, Zhiwei Wang 0002, Qiaozhe Zhang, Wenyu Liu 0001, Chunyuan Liao, Kwang-Ting Cheng |
ACM Multimedia | 1 |
| 2018 | Neural Compatibility Modeling with Attentive Knowledge DistillationabstractRecently, the booming fashion sector and its huge potential benefits have attracted tremendous attention from many research communities. In particular, increasing research efforts have been dedicated to the complementary clothing matching as matching clothes to make a suitable outfit has become a daily headache for many people, especially those who do not have the sense of aesthetics. Thanks to the remarkable success of neural networks in various applications such as the image classification and speech recognition, the researchers are enabled to adopt the data-driven learning methods to analyze fashion items. Nevertheless, existing studies overlook the rich valuable knowledge (rules) accumulated in fashion domain, especially the rules regarding clothing matching. Towards this end, in this work, we shed light on the complementary clothing matching by integrating the advanced deep neural networks and the rich fashion domain knowledge. Considering that the rules can be fuzzy and different rules may have different confidence levels to different samples, we present a neural compatibility modeling scheme with attentive knowledge distillation based on the teacher-student network scheme. Extensive experiments on the real-world dataset show the superiority of our model over several state-of-the-art methods. Based upon the comparisons, we observe certain fashion insights that can add value to the fashion matching study. As a byproduct, we released the codes, and involved parameters to benefit other researchers. Xuemeng Song, Fuli Feng, Xianjing Han, Xin Yang 0008, Wei Liu 0005, Liqiang Nie |
SIGIR | 4 |
| 2018 | Robust and real-time pose tracking for augmented reality on mobile devices
Xin Yang 0008, Jiabin Guo, Tangli Xue, Kwang-Ting Cheng |
Multim. Tools Appl. | 1 |
| 2018 | Automated Detection of Clinically Significant Prostate Cancer in mp-MRI Images Based on an End-to-End Deep Neural NetworkabstractAutomated methods for detecting clinically significant (CS) prostate cancer (PCa) in multi-parameter magnetic resonance images (mp-MRI) are of high demand. Existing methods typically employ several separate steps, each of which is optimized individually without considering the error tolerance of other steps. As a result, they could either involve unnecessary computational cost or suffer from errors accumulated over steps. In this paper, we present an automated CS PCa detection system, where all steps are optimized jointly in an end-to-end trainable deep neural network. The proposed neural network consists of concatenated subnets: 1) a novel tissue deformation network (TDN) for automated prostate detection and multimodal registration and 2) a dual-path convolutional neural network (CNN) for CS PCa detection. Three types of loss functions, i.e., classification loss, inconsistency loss, and overlap loss, are employed for optimizing all parameters of the proposed TDN and CNN. In the training phase, the two nets mutually affect each other and effectively guide registration and extraction of representative CS PCa-relevant features to achieve results with sufficient accuracy. The entire network is trained in a weakly supervised manner by providing only image-level annotations (i.e., presence/absence of PCa) without exact priors of lesions' locations. Compared with most existing systems which require supervised labels, e.g., manual delineation of PCa lesions, it is much more convenient for clinical usage. Comprehensive evaluation based on fivefold cross validation using 360 patient data demonstrates that our system achieves a high accuracy for CS PCa detection, i.e., a sensitivity of 0.6374 and 0.8978 at 0.1 and 1 false positives per normal/benign patient. Zhiwei Wang 0002, Chaoyue Liu 0002, Danpeng Cheng, Liang Wang 0052, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 5 |
| 2018 | A Skeletal Similarity Metric for Quality Evaluation of Retinal Vessel SegmentationabstractThe most commonly used evaluation metrics for quality assessment of retinal vessel segmentation are sensitivity, specificity, and accuracy, which are based on pixel-to-pixel matching. However, due to the inter-observer problem that vessels annotated by different observers vary in both thickness and location, pixel-to-pixel matching is too restrictive to fairly evaluate the results of vessel segmentation. In this paper, the proposed skeletal similarity metric is constructed by comparing the skeleton maps generated from the reference and the source vessel segmentation maps. To address the inter-observer problem, instead of using a pixel-to-pixel matching strategy, each skeleton segment in the reference skeleton map is adaptively assigned with a searching range whose radius is determined based on its vessel thickness. Pixels in the source skeleton map located within the searching range are then selected for similarity calculation. The skeletal similarity consists of a curve similarity, which measures the structural similarity between the reference and the source skeleton maps and a thickness similarity, which measures the thickness consistency between the reference and the source vessel segmentation maps. In contrast to other metrics that provide a global score for the overall performance, we modify the definitions of true positive, false negative, true negative, and false positive based on the skeletal similarity, based on which sensitivity, specificity, accuracy, and other objective measurements can be constructed. More importantly, the skeletal similarity metric has better potential to be used as a pixelwise loss function for training deep learning models for retinal vessel segmentation. Through comparison of a set of examples, we demonstrate that the redefined metrics based on the skeletal similarity are more effective for quality evaluation, especially with greater tolerance to the inter-observer problem. Zengqiang Yan, Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Joint Classification Loss and Histogram Loss for Sketch-Based Image Retrieval
Yongluan Yan, Xinggang Wang, Xin Yang 0008, Xiang Bai, Wenyu Liu 0001 |
ICIG (1) | 3 |
| 2017 | REDBEE: A visual-inertial drone system for real-time moving object detectionabstractAerial surveillance and monitoring demand both real-time and robust motion detection from a moving camera. Most existing techniques for drones involve sending a video data streams back to a ground station with a high-end desktop computer or server. These methods share one major drawback: data transmission is subjected to considerable delay and possible corruption. Onboard computation can not only overcome the data corruption problem but also increase the range of motion. Unfortunately, due to limited weight-bearing capacity, equipping drones with computing hardware of high processing capability is not feasible. Therefore, developing a motion detection system with real-time performance and high accuracy for drones with limited computing power is highly desirable. In this paper, we propose a visual-inertial drone system for real-time motion detection, namely REDBEE, that helps overcome challenges in shooting scenes with strong parallax and dynamic background. REDBEE, which can run on the state-of-the-art commercial low-power application processor (e.g. Snapdragon Flight board used for our prototype drone), achieves real-time performance with high detection accuracy. The REDBEE system overcomes obstacles in shooting scenes with strong parallax through an inertial-aided dual-plane homography estimation; it solves the issues in shooting scenes with dynamic background by distinguishing the moving targets through a probabilistic model based on spatial, temporal, and entropy consistency. The experiments are presented which demonstrate that our system obtains greater accuracy when detecting moving targets in outdoor environments than the state-of-the-art real-time onboard detection systems. Chong Huang 0005, Peng Chen 0008, Xin Yang 0008, Kwang-Ting Cheng |
IROS | 3 |
| 2017 | Joint Detection and Diagnosis of Prostate Cancer in Multi-parametric MRI Based on Multimodal Convolutional Neural Networks
Xin Yang 0008, Zhiwei Wang 0002, Chaoyue Liu 0002, Hung Le Minh, Kwang-Ting Cheng, Liang Wang 0052 |
MICCAI (3) | 1 |
| 2017 | Real-Time Dense Monocular SLAM for Augmented RealityabstractSimultaneous localization and mapping (SLAM) via a monocular camera is a key enabling technique for many augmented reality (AR) applications. In this work, we present a monocular SLAM system which can provide real-time dense mapping even for challenging poorly-textured regions based on the piecewise planarity approximation. Specifically, our system consists of three modules. First, a tracking module based on the direct method [3] continuously estimates camera poses with respect to the scene. Second, a semi-dense mapping module takes the estimated camera pose as input and calculates depths of highly-textured pixels based on pixel matching and triangulation. Third, dense mapping module approximates textureless regions identified by a homogeneous-color region detector using piecewise plane models. The 3D piecewise planes are reconstructed via the proposed multi-plane segmentation and multi-plane fusion algorithms. Live experiments in a real AR demo with a hand-held camera demonstrate the effectiveness and efficiency of our method in practical scenario. Hongcheng Luo, Tangli Xue, Xin Yang 0008 |
ACM Multimedia | 3 |
| 2017 | DeepCADx: Automated Prostate Cancer Detection and Diagnosis in mp-MRI based on Multimodal Convolutional Neural NetworksabstractIn this paper, we present DeepCADx, a computer-aided prostate detection and diagnosis (CADx) system powered by a novel deep convolutional neural networks (CNNs). Specifically, the developed DeepCADx system processes multi-parametric magnetic resonance imaging (mp-MRI) sequences in three major steps: 1) pre-processing which registers images from different modalities and detect prostates, 2) multimodal CNNs which jointly identifies images containing prostate cancers (PCa) and generate cancer response maps (CRM) with each pixel indicating the probability to be cancerous, and 3) post-processing which localize lesion in CRMs and assess the aggressiveness (i.e. Gleason score) of each localized lesion using multimodal CNN features and a 5-class SVM classifier. Zhiwei Wang 0002, Chaoyue Liu 0002, Xiang Bai, Xin Yang 0008 |
ACM Multimedia | 4 |
| 2017 | Real-time Monocular Dense Mapping for Augmented RealityabstractMonocular simultaneous localization and mapping (SLAM) is a key enabling technique for many augmented reality (AR) applications. However, conventional methods for monocular SLAM can obtain only sparse or semi-dense maps in highly-textured image areas. Poorly-textured regions which widely exist in indoor and man-made urban environments can be hardly reconstructed, impeding interactions between virtual objects and real scenes in AR apps. In this paper,we present a novel method for real-time monocular dense mapping based on the piecewise planarity assumption for poorly textured regions. Specifically, a semi-dense map for highly-textured regions is first calculated by pixel matching and triangulation [6, 7]. Large textureless regions extracted by Maximally Stable Color Regions (MSCR) [11], which is a homogeneous-color region detector, are approximated using piecewise planar models which are estimated by the corresponding semi-dense 3D points and the proposed multi-plane segmentation algorithm. Plane models associated with the same 3D area across multiple overlapping views are linked and fused to ensure a consistent and accurate 3D reconstruction. Experimental results on two public datasets [15, 23] demonstrate that our method is 2.3X~2.9X faster than the state-of-the-art method DPPTAM [2], and meanwhile achieves better reconstruction accuracy and completeness. We also apply our method to a real AR application and live experiments with a hand-held camera demonstrate the effectiveness and efficiency of our method in practical scenario. Tangli Xue, Hongcheng Luo, Danpeng Cheng, Zikang Yuan, Xin Yang 0008 |
ACM Multimedia | 5 |
| 2017 | Joint Face Detection and Initialization for Face Alignment
Zhiwei Wang 0002, Xin Yang 0008 |
MMM (1) | 2 |
| 2017 | V-Head: Face Detection and Alignment for Facial Augmented Reality Applications
Zhiwei Wang 0002, Xin Yang 0008 |
MMM (2) | 2 |
| 2017 | Co-trained convolutional neural networks for automated detection of prostate cancer in multi-parametric MRI
Xin Yang 0008, Chaoyue Liu 0002, Zhiwei Wang 0002, Hung Le Minh, Liang Wang 0052, Kwang-Ting Cheng |
Medical Image Anal. | 1 |
| 2016 | Location-Aware Image Classification
Xinggang Wang, Xin Yang 0008, Wenyu Liu 0001, Chen Duan, Longin Jan Latecki |
MMM (1) | 2 |
| 2016 | OGB: A Distinctive and Efficient Feature for Mobile Augmented Reality
Xin Yang 0008, Xinggang Wang, Kwang-Ting Cheng |
MMM (1) | 1 |
| 2016 | Accurate and efficient pulse measurement from facial videos on smartphonesabstractNon-contact measurement of cardiac pulse signals has attracted high interests due to its convenience and cost effectiveness. However, extracting pulse signals on mobile handheld devices (e.g. smartphones) based on face videos captured by mobile cameras usually suffers from low measurement accuracy due to misalignment errors in face tracking and inevitable illumination changes in a mobile scenario, and low efficiency due to a handheld's limited computing power. We propose two techniques to address these limitations: 1) an accurate and efficient face tracking method based on an Active Shape Model (ASM) and the LDB (Local Difference Binary) feature description; 2) an adaptive temporal filtering method which can detect, and in turn denoise, sharp intensity changes in the source trace. Experimental results demonstrate that the proposed solution can achieve a speedup of 6.2X and is robust to noises in common mobile scenarios. Chong Huang 0005, Xin Yang 0008, Kwang-Ting Cheng |
WACV | 2 |
| 2016 | Renal compartment segmentation in DCE-MRI images
Xin Yang 0008, Hung Le Minh, Kwang-Ting Cheng, Kyung Hyun Sung, Wenyu Liu 0001 |
Medical Image Anal. | 1 |
| 2015 | Fusion of Vision and Inertial Sensing for Accurate and Efficient Pose Tracking on SmartphonesabstractThis paper aims at accurate and efficient pose tracking of planar targets on modern smartphones. Existing methods, relying on either visual features or motion sensing based on built-in inertial sensors, are either too computationally expensive to achieve realtime performance on a smartphone, or too noisy to achieve sufficient tracking accuracy. In this paper we present a hybrid tracking method which can achieve real-time performance with high accuracy. Based on the same framework of a state-of-the-art visual feature tracking algorithm [5] which ensures accurate and reliable pose tracking, the proposed hybrid method significantly reduces its computational cost with the assistance of a phone's built-in inertial sensors. However, noises in inertial sensors and abrupt errors in feature tracking due to severe motion blurs could result in instability of the hybrid tracking system. To address this problem, we propose to employ an adaptive Kalman filter with abrupt error detection to robustly fuse the inertial and feature tracking results. We evaluated the proposed method on a dataset consisting of 16 video clips with synchronized inertial sensing data. Experimental results demonstrated our method's superior performance and accuracy on smartphones compared to a state-of-the-art vision tracking method [5]. The dataset will be made publicly available with the publication of this paper. Xin Yang 0008, Xun Si, Tangli Xue, Kwang-Ting Cheng |
ISMAR | 1 |
| 2015 | Automatic Segmentation of Renal Compartments in DCE-MRI Images
Xin Yang 0008, Hung Le Minh, Kwang-Ting Cheng, Kyung Hyun Sung, Wenyu Liu 0001 |
MICCAI (1) | 1 |
| 2015 | Vision-Inertial Hybrid Tracking for Robust and Efficient Augmented Reality on SmartphonesabstractThis paper aims at robust and efficient pose tracking for augmented reality on modern smartphones. Existing methods, relying on either vision analysis or motion sensing, are either too computationally expensive to achieve real-time performance on a smartphone, or too noisy to achieve sufficient robustness. This paper presents a hybrid tracking system which can achieve real-time performance with high robustness. Our system utilizes an efficient featureless method based on pixel-based registration to track the object pose on every frame. The featureless tracking result is revised from time to time by a feature-based method to reduce tracking errors. Both featureless and feature-based tracking results are sensitive to large motion blurs. To improve the robustness, an adaptive Kamlan filter is proposed to fuse the visual tracking results with the inertial tracking results computed form phone's built-in sensors. Our hybrid method is evaluated on a dataset consisting of 16 video clips with synchronized inertial sensing data. Experimental results demonstrated the superior performance of our method to state-of-the-art visual tracking methods [5, 12] on smartphones. The dataset will be made publicly available with the publication of this paper. Xin Yang 0008, Xun Si, Tangli Xue, Liheng Zhang, Kwang-Ting Cheng |
ACM Multimedia | 1 |
| 2014 | Accurate Vessel Segmentation with Progressive Contrast Enhancement and Canny Refinement
Xin Yang 0008, Kwang-Ting Cheng, Aichi Chien |
ACCV (3) | 1 |
| 2014 | Geodesic Active Contours with Adaptive Configuration for Cerebral Vessel and Aneurysm SegmentationabstractActive contour is a popular technique for vascular segmentation. However, existing active contour segmentation methods require users to set values for various parameters, which requires insights to the method's mathematical formulation. Manual tuning of these parameters to optimize segmentation results is laborious for clinicians who often lack in-depth knowledge of the segmentation algorithms. Moreover, a global parameter setting applied to all voxels of an input image can hardly achieve optimized results due to vessels' high appearance variability caused by the contrast agent in homogeneity and noises. In this paper, we present a method which adaptively configures parameters for Geodesic Active Contours (GAC). The proposed method leverages shape filtering to produce a parameter image, each voxel of which is used to set parameters of GAC for the corresponding voxel of an input image. An iterative process is further developed to improve the accuracy of the shape-based parameter image. An evaluation study over 8 clinical datasets demonstrates that our method achieves greater segmentation accuracy than two popular active contour methods with manually optimized parameters. Xin Yang 0008, Kwang-Ting Cheng, Aichi Chien |
ICPR | 1 |
| 2014 | libLDB: a library for extracting ultrafast and distinctive binary feature descriptionabstractThis paper gives an overview of libLDB -- a C++ library for extracting an ultrafast and distinctive binary feature LDB (Local Difference Binary) from an image patch. LDB directly computes a binary string using simple intensity and gradient difference tests on pairwise grid cells within the patch. Relying on integral images, the average intensity and gradients of each grid cell can be obtained by only 4~8 add/subtract operations, yielding an ultrafast runtime. A multiple gridding strategy is applied to capture the distinct patterns of the patch at different spatial granularities, leading to a high distinctiveness of LDB. LDB is very suitable for vision apps which require real-time performance, especially for apps running on mobile handheld devices, such as real-time mobile object recognition and tracking, markerless mobile augmented reality, mobile panorama stitching. This software is available under the GNU General Public License (GPL) v3. Xin Yang 0008, Chong Huang 0005, Kwang-Ting Cheng |
ACM Multimedia | 1 |
| 2014 | Local Difference Binary for Ultrafast and Distinctive Feature DescriptionabstractThe efficiency and quality of a feature descriptor are critical to the user experience of many computer vision applications. However, the existing descriptors are either too computationally expensive to achieve real-time performance, or not sufficiently distinctive to identify correct matches from a large database with various transformations. In this paper, we propose a highly efficient and distinctive binary descriptor, called local difference binary (LDB). LDB directly computes a binary string for an image patch using simple intensity and gradient difference tests on pairwise grid cells within the patch. A multiple-gridding strategy and a salient bit-selection method are applied to capture the distinct patterns of the patch at different spatial granularities. Experimental results demonstrate that compared to the existing state-of-the-art binary descriptors, primarily designed for speed, LDB has similar construction efficiency, while achieving a greater accuracy and faster speed for mobile object recognition and tracking tasks. Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Learning Optimized Local Difference Binaries for Scalable Augmented Reality on Mobile DevicesabstractThe efficiency, robustness and distinctiveness of a feature descriptor are critical to the user experience and scalability of a mobile augmented reality (AR) system. However, existing descriptors are either too computationally expensive to achieve real-time performance on a mobile device such as a smartphone or tablet, or not sufficiently robust and distinctive to identify correct matches from a large database. As a result, current mobile AR systems still only have limited capabilities, which greatly restrict their deployment in practice. In this paper, we propose a highly efficient, robust and distinctive binary descriptor, called Learning-based Local Difference Binary (LLDB). LLDB directly computes a binary string for an image patch using simple intensity and gradient difference tests on pairwise grid cells within the patch. To select an optimized set of grid cell pairs, we densely sample grid cells from an image patch and then leverage a modified AdaBoost algorithm to automatically extract a small set of critical ones with the goal of maximizing the Hamming distance between mismatches while minimizing it between matches. Experimental results demonstrate that LLDB is extremely fast to compute and to match against a large database due to its high robustness and distinctiveness. Compared to the state-of-the-art binary descriptors, primarily designed for speed, LLDB has similar efficiency for descriptor construction, while achieving a greater accuracy and faster matching speed when matching over a large database with 2.3M descriptors on mobile devices. Xin Yang 0008, Kwang-Ting Cheng |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2012 | LDB: An ultra-fast feature for scalable Augmented Reality on mobile devicesabstractThe efficiency, robustness and distinctiveness of a feature descriptor are critical to the user experience and scalability of a mobile Augmented Reality (AR) system. However, existing descriptors are either too compute-expensive to achieve real-time performance on a mobile device such as a smartphone or tablet, or not sufficiently robust and distinctive to identify correct matches from a large database. As a result, current mobile AR systems still only have limited capabilities, which greatly restrict their deployment in practice. In this paper, we propose a highly efficient, robust and distinctive binary descriptor, called Local Difference Binary (LDB). LDB directly computes a binary string for an image patch using simple intensity and gradient difference tests on pairwise grid cells within the patch. A multiple gridding strategy is applied to capture the distinct patterns of the patch at different spatial granularities. Experimental results demonstrate that LDB is extremely fast to compute and to match against a large database due to its high robustness and distinctiveness. Comparing to the state-of-the-art binary descriptor BRIEF, primarily designed for speed, LDB has similar computational efficiency, while achieves a greater accuracy and 5x faster matching speed when matching over a large database with 1.7M+ descriptors. Xin Yang 0008, Kwang-Ting Cheng |
ISMAR | 1 |
| 2012 | Accelerating SURF detector on mobile devicesabstractRunning a SURF (Speeded Up Robust Features) detector on mobile devices remains too slow to support emerging applications such as mobile augmented reality. Porting it without adapting the algorithm to account for mobile platform limitations could result in significant runtime degradation. In this paper, we identify two mismatches between the SURF algorithm and the mobile hardware that cause substantial slow-down of the point detection process: 1) mismatch between the data access pattern and the small cache size, and 2) mismatch between the huge amount of branches and high pipeline hazard penalty. To address the mismatches, we propose two techniques: tiled SURF and gradient moment based orientation assignment. Tiled SURF improves data locality and greatly reduces memory traffic. A method for determining the optimal tile sizes, named content-aware tiling, is designed to minimize runtime and maximize detection accuracy. To avoid the penalties caused by pipeline hazards, we replace the original orientation operator with branching-free gradient moment computations. The proposed techniques are tested on three mobile platforms. Comparing to the original SURF, the accelerated SURF achieves a 6x~8x speedup without sacrificing recognition accuracy. Meanwhile, it achieves 59%~80% reductions in the runtime ratio of the detector running on mobile platforms compared with on x86-based PCs. Xin Yang 0008, Kwang-Ting Cheng |
ACM Multimedia | 1 |
| 2012 | MixPad: augmenting interactive paper with mice & keyboards for cross-media and fine-grained interaction with documentsabstractExisting interactive paper systems suffer from the disparate input devices for paper and computers. The finger-pen-only input on paper causes frequent devices switching (e.g. pen vs. mouse) during cross-media interactions, and may have issues of occlusion and precision. We propose MixPad, a novel interactive paper system, which allows users to exploit mice and keyboards to digitally manipulate fine-grained document content on paper, such as copying an arbitrary image region to a computer and clicking on a word for web search. With the combined input channels, MixPad enables richer digital functions on paper and facilitates bimanual operations cross different media. A preliminary user study shows positive feedback on this interaction technique. Xin Yang 0008, Chunyuan Liao, Qiong Liu 0003 |
ACM Multimedia | 1 |
| 2011 | Large-scale EMM identification based on geometry-constrained visual word correspondence votingabstractWe present a large-scale Embedded Media Marker (EMM) identification system which allows users to retrieve relevant dynamic media associated with a static paper document via camera-phones. The user supplies a query image by capturing an EMM-signified patch of a paper document through a camera phone. The system recognizes the query and in turn retrieves and plays the corresponding media on the phone. Accurate image matching is crucial for positive user experience in this application. To address the challenges posed by large datasets and variation in camera-phone-captured query images, we introduce a novel image matching scheme based on geometrically consistent correspondences. A hierarchical scheme, combined with two constraining methods, is designed to detect geometric constrained correspondences between images. A spatial neighborhood search approach is further proposed to address challenging cases of query images with a large translational shift. Experimental results on a 200k+ dataset show that our solution achieves high accuracy with low memory and time complexity and outperforms the baseline bag-of-words approach. © 2011 ACM. Xin Yang 0008, Qiong Liu 0003, Chunyuan Liao, Kwang-Ting Cheng, Andreas Girgensohn |
ICMR | 1 |
| 2009 | MyFinder: near-duplicate detection for large image collectionsabstractThe explosive growth of multimedia data poses serious challenges to data storage, management and search. Efficient near-duplicate detection is one of the required technologies for various applications. In this paper, we introduce MyFinder, an image near-duplicate detection system for large image collections. MyFinder consists of three major components: 1) a local-feature-based image representation utilizing the proposed LDP (Local-Difference-Pattern) feature, 2) the Locality-Sensitive-Hashing (LSH) as the core indexing structure to assure the most frequent data access occurred in the main memory, and 3) multi-step verification for queries to best exclude false positives and to increase the precision. Xin Yang 0008, Qiang Zhu 0006, Kwang-Ting Cheng |
ACM Multimedia | 1 |