Lei Ke

dblp:26/5225 · DBLP profile ↗
← Back
41ranked-venue papers
19as first author
26since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 29 · 11 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 first-author · 16 since 2021Computer networks · 4 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 first-authorTheory of computation · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Robust Multi-Object 4D Generation for In-the-wild Videos
abstract
We address the challenge of generating dynamic 4D scenes from monocular multi-object videos with heavy occlusions and introduce GenMOJO, a novel approach that integrates rendering-based deformable 3D Gaussian optimization with generative priors for view synthesis. While existing view-synthesis models excel at novel view generation for isolated objects, they struggle with full scenes due to their complexity and data demands. To overcome this, GenMOJO decomposes scenes into individual objects, optimizing a differentiable set of deformable Gaussians per object while capturing 2D occlusions from a 3D perspective through joint Gaussian splatting. Joint splatting ensures occlusion-aware rendering losses in observed frames, while explicit object decomposition allows the usage of object-centric diffusion models for object completion in unobserved viewpoints. To reconcile the differences between object-centric priors and the global frame-centric coordinate system of the video, GenMOJO employs differentiable transformations to unify the rendering and generative constraints within a single framework. The result is a model capable of generating 4D objects across space and time while producing 2D and 3D point tracks from monocular videos. To rigorously evaluate the quality of scene generation and the accuracy of the motion under multi-object occlusions, we introduce MOSE-PTS, a subset of the challenging MOSE benchmark, which we annotated with high-quality 2D point tracks. Quantitative evaluations and perceptual human studies confirm that GenMOJO generates more realistic novel views of scenes and produces more accurate point tracks compared to existing approaches. Project page: https://genmojo.github.io/.
Wen-Hsuan Chu, Lei Ke, Jianmeng Liu, Mingxiao Huo, Pavel Tokmakov, Katerina Fragkiadaki
CVPR2
2025 Video Depth without Video Models
abstract
Video depth estimation lifts monocular video clips to 3D by inferring dense depth at every frame. Recent advances in single-image depth estimation, brought about by the rise of large foundation models and the use of synthetic training data, have fueled a renewed interest in video depth. However, naively applying a single-image depth estimator to every frame of a video disregards temporal continuity, which not only leads to flickering but may also break when camera motion causes sudden changes in depth range. An obvious and principled solution would be to build on top of video foundation models, but these come with their own limitations; including expensive training and inference, imperfect 3D consistency, and stitching routines for the fixed-length (short) outputs. We take a step back and demonstrate how to turn a single-image latent diffusion model (LDM) into a state-of-the-art video depth estimator. Our model, which we call RollingDepth, has two main ingredients: (i) a multi-frame depth estimator that is derived from a single-image LDM and maps very short video snippets (typically frame triplets) to depth snippets. (ii) a robust, optimization-based registration algorithm that optimally assembles depth snippets sampled at various different frame rates back into a consistent video. RollingDepth is able to efficiently handle long videos with hundreds of frames and delivers more accurate depth videos than both dedicated video depth estimators and high-performing single-frame models. Project page: rollingdepth.github.io.
Bingxin Ke, Dominik Narnhofer, Lei Ke, Torben Peters, Katerina Fragkiadaki, Anton Obukhov, Konrad Schindler
CVPR4
2025 ProReflow: Progressive Reflow with Decomposed Velocity
abstract
Diffusion models have achieved significant progress in both image and video generation while still suffering from huge computation costs. As an effective solution, rectified flow aims to rectify the diffusion process of diffusion models into a straight line for few-step and even one-step generation. However, in this paper, we suggest that the original training pipeline of reflow is not optimal and introduce two techniques to improve it. Firstly, we introduce progressive reflow, which progressively reflows the diffusion models in local timesteps until the whole diffusion progresses, reducing the difficulty of flow matching. Second, we introduce aligned v-prediction, which highlights the importance of direction matching in flow matching over magnitude matching. Experimental results on SDv1.5 and SDXL demonstrate the effectiveness of our method, for example, conducting on SDv1.5 achieves an FID of 10.70 on MSCOCO2014 validation set with only 4 sampling steps, close to our teacher model (32 DDIM steps, FID = 10.05). Our codes will be released at Github.
Lei Ke, Haohang Xu, Xuefei Ning, Yu Li 0022, Haoling Li, Dongsheng Jiang, Yujiu Yang 0001, Linfeng Zhang 0001
CVPR1
2025 EntitySAM: Segment Everything in Video
abstract
Automatically tracking and segmenting every video entity remains a significant challenge. Despite rapid advancements in video segmentation, even state-of-the-art models like SAM 2 struggle to consistently track all entities across a video—a task we refer to as Video Entity Segmentation. We propose EntitySAM, a framework for zero-shot video entity segmentation. EntitySAM extends SAM 2 by removing the need for explicit prompts, allowing automatic discovery and tracking of all entities, including those appearing in later frames. We incorporate query-based entity discovery and association into SAM 2, inspired by transformer-based object detectors. Specifically, we introduce an entity decoder to facilitate inter-object communication and an automatic prompt generator using learnable object queries. Additionally, we add a semantic encoder to enhance SAM 2’s semantic awareness, improving segmentation quality. Trained on image-level mask annotations without category information from the COCO dataset, EntitySAM demonstrates strong generalization on four zero-shot video segmentation tasks: Video Entity, Panoptic, Instance, and Semantic Segmentation. Results on six popular benchmarks show that EntitySAM outperforms previous unified video segmentation methods and strong baselines, setting new standards for zero-shot video segmentation. Our code and models are at github.com/ymq2017/entitysam.
Mingqiao Ye, Seoung Wug Oh, Lei Ke, Joon-Young Lee
CVPR3
2025 Multi-View 3D Point Tracking
abstract
We introduce the first data-driven multi-view 3D point tracker, designed to track arbitrary points in dynamic scenes using multiple camera views. Unlike existing monocular trackers, which struggle with depth ambiguities and occlusion, or prior multi-camera methods that require over 20 cameras and tedious per-sequence optimization, our feed-forward model directly predicts 3D correspondences using a practical number of cameras (e.g., four), enabling robust and accurate online tracking. Given known camera poses and either sensor-based or estimated multi-view depth, our tracker fuses multi-view features into a unified point cloud and applies k-nearest-neighbors correlation alongside a transformer-based update to reliably estimate long-range 3D correspondences, even under occlusion. We train on 5K synthetic multi-view Kubric sequences and evaluate on two real-world benchmarks: Panoptic Studio and DexYCB, achieving median trajectory errors of 3.1 cm and 2.0 cm, respectively. Our method generalizes well to diverse camera setups of 1-8 views with varying vantage points and video lengths of 24-150 frames. By releasing our tracker alongside training and evaluation datasets, we aim to set a new standard for multi-view 3D tracking research and provide a practical tool for real-world applications. Project page available at https://ethz-vlg.github.io/mvtracker.
Frano Rajic, Haofei Xu, Marko Mihajlovic, Siyuan Li 0008, Irem Demir, Emircan Gündogdu, Lei Ke, Sergey Prokudin, Marc Pollefeys, Siyu Tang 0001
ICCV7
2025 Stable Segment Anything Model
abstract
The Segment Anything Model (SAM) achieves remarkable promptable segmentation given high-quality prompts which, however, often require good skills to specify. To make SAM robust to casual prompts, this paper presents the first comprehensive analysis on SAM’s segmentation stability across a diverse spectrum of prompt qualities, notably imprecise bounding boxes and insufficient points. Our key finding reveals that given such low-quality prompts, SAM’s mask decoder tends to activate image features that are biased towards the background or confined to specific object parts. To mitigate this issue, our key idea consists of calibrating solely SAM’s mask attention by adjusting the sampling locations and amplitudes of image features, while the original SAM model architecture and weights remain unchanged. Consequently, our deformable sampling plugin (DSP) enables SAM to adaptively shift attention to the prompted target regions in a data-driven manner. During inference, dynamic routing plugin (DRP) is proposed that toggles SAM between the deformable and regular grid sampling modes, conditioned on the input prompt quality. Thus, our solution, termed Stable-SAM, offers several advantages: 1) improved SAM’s segmentation stability across a wide range of prompt qualities, while 2) retaining SAM’s powerful promptable segmentation efficiency and generality, with 3) minimal learnable parameters (0.08 M) and fast adaptation. Extensive experiments validate the effectiveness and advantages of our approach, underscoring Stable-SAM as a more robust solution for segmenting anything. Codes are at https://github.com/fanq15/Stable-SAM.
Xin Tao 0001, Lei Ke, Mingqiao Ye, Di Zhang 0026, Pengfei Wan 0001, Yu-Wing Tai, Chi-Keung Tang
ICLR3
2025 M^3PC: Test-time Model Predictive Control using Pretrained Masked Trajectory Model
abstract
Recent work in Offline Reinforcement Learning (RL) has shown that a unified transformer trained under a masked auto-encoding objective can effectively capture the relationships between different modalities (e.g., states, actions, rewards) within given trajectory datasets. However, this information has not been fully exploited during the inference phase, where the agent needs to generate an optimal policy instead of just reconstructing masked components from unmasked. Given that a pretrained trajectory model can act as both a Policy Model and a World Model with appropriate mask patterns, we propose using Model Predictive Control (MPC) at test time to leverage the model's own predictive capacity to guide its action selection. Empirical results on D4RL and RoboMimic show that our inference-phase MPC significantly improves the decision-making performance of a pretrained trajectory model without any additional parameter training. Furthermore, our framework can be adapted to Offline to Online (O2O) RL and Goal Reaching RL, resulting in more substantial performance gains when an additional online interaction budget is given, and better generalization capabilities when different task targets are specified. Code is available: \href{https://github.com/wkh923/m3pc}{\texttt{https://github.com/wkh923/m3pc}}.
Kehan Wen, Yutong Hu 0001, Yao Mu 0001, Lei Ke
ICLR4
2025 TAPIP3D: Tracking Any Point in Persistent 3D Geometry
abstract
We introduce TAPIP3D, a novel approach for long-term 3D point tracking in monocular RGB and RGB-D videos. TAPIP3D represents videos as camera-stabilized spatio-temporal feature clouds, leveraging depth and camera motion information to lift 2D video features into a 3D world space where camera movement is effectively canceled out. Within this stabilized 3D representation, TAPIP3D iteratively refines multi-frame motion estimates, enabling robust point tracking over long time horizons. To handle the irregular structure of 3D point distributions, we propose a 3D Neighborhood-to-Neighborhood (N2N) attention mechanism—a 3D-aware contextualization strategy that builds informative, spatially coherent feature neighborhoods to support precise trajectory estimation. Our 3D-centric formulation significantly improves performance over existing 3D point tracking methods and even surpasses state-of-the-art 2D pixel trackers in accuracy when reliable depth is available. The model supports inference in both camera-centric (unstabilized) and world-centric (stabilized) coordinates, with experiments showing that compensating for camera motion leads to substantial gains in tracking robustness. By replacing the conventional 2D square correlation windows used in prior 2D and 3D trackers with a spatially grounded 3D attention mechanism, TAPIP3D achieves strong and consistent results across multiple 3D point tracking benchmarks. Our code and trained checkpoints will be public.
Lei Ke, Adam W. Harley, Katerina Fragkiadaki
NeurIPS2
2025 Segment Anything Meets Point Tracking
abstract
Foundation models have marked a significant stride to-ward addressing generalization challenges in deep learning. While the Segment Anything Model (SAM) has established a strong foothold in image segmentation, existing video segmentation methods still require extensive mask labeling for fine-tuning, or face performance drops on unseen data domains otherwise. In this paper, we show how foundation models for image segmentation make a step toward enhancing domain generalizability in video segmentation. We discover that, combined with long-term point tracking, image segmentation models yield state-of-the-art results in zero-shot video segmentation across multiple benchmarks. Surprisingly, point trackers exhibit generalization to domains beyond their synthetic pre-training sequences, which we attribute to the trackers' ability to harness the rich local information in the vicinity of each tracked point. Thus, we introduce SAM-PT, an innovative method for point-centric video segmentation, leveraging the capabilities of SAM alongside long-term point tracking. SAM-PT extends SAM's capability to tracking and segmenting anything in dynamic videos. Unlike traditional video segmentation methods that focus on object-centric mask propagation, our approach uniquely exploits point propagation to utilize local structure information independent of object semantics. The effectiveness of point-based tracking is underscored by direct evaluation on the zero-shot open-world UVO benchmark. Our experiments on popular video object segmentation and multi-object segmentation tracking benchmarks, including DAVIS, YouTube-VOS, and BDD100K, suggest that a point-based segmentation tracker yields better zero-shot performance and efficient interactions. We release our code at https://github.com/SysCV/sam-pt.
Frano Rajic, Lei Ke, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001
WACV2
2025 BiVM: Accurate Binarized Neural Network for Efficient Video Matting
abstract
Deep neural networks for real-time video matting suffer significant computational limitations on edge devices, hindering their adoption in widespread applications such as online conferences and short-form video production. Binarization emerges as one of the most common neural network compression approaches, substantially curbing computational and memory requirements through compact 1-bit parameters and efficient bitwise operations. However, the empirical observation reveals that accuracy and efficiency limitations exist in the binarized video matting network due to its degenerated encoder and redundant decoder. Following a theoretical analysis based on the information bottleneck principle, the limitations are mainly caused by the degradation of prediction-relevant information in the intermediate features and the redundant computation in prediction-irrelevant areas. We presentBiVM, an accurate and resource-efficientBinarized neural network forVideoMatting, where architecture and optimization allow real-time video matting to proceed on edge hardware. First, we present a series of binarized computation structures with elastic shortcuts and evolvable topologies, enabling the constructed encoder backbone to extract high-quality representation from input videos for accurate prediction. Second, we sparse the intermediate feature of the binarized decoder by masking homogeneous parts, allowing the decoder to focus on representation with diverse details while alleviating the computation burden for efficient inference. Furthermore, we construct a localized binarization-aware mimicking framework with the information-guided strategy, prompting matting-related representation in full-precision counterparts to be accurately and fully utilized. Comprehensive experiments show that the proposed BiVM surpasses alternative binarized video matting networks, including state-of-the-art (SOTA) binarization methods, by a substantial margin. For example, BiVM surpasses 16.67 MAD compared to SOTA binarization on the VM dataset. Notably, our approach can even perform comparably to the full-precision counterpart in terms of visual quality. Moreover, our BiVM achieves significant savings of 14.3× and 21.6× in computation and storage costs, respectively. We also evaluate BiVM on ARM CPU hardware, underscoring its potential for deployment in resource-constrained scenarios.
Haotong Qin, Xianglong Liu 0001, Xudong Ma, Lei Ke, Yulun Zhang 0001, Jie Luo 0004, Michele Magno
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Matching Anything by Segmenting Anything
abstract
The robust association of the same objects across video frames in complex scenes is crucial for many applications, especially multiple object tracking (MOT). Current methods predominantly rely on labeled domain-specific video datasets, which limits the cross-domain generalization of learned similarity embeddings. We propose MASA, a novel method for robust instance association learning, capable of matching any objects within videos across diverse domains without tracking labels. Leveraging the rich object segmentation from the Segment Anything Model (SAM), MASA learns instance-level correspondence through exhaustive data transformations. We treat the SAM outputs as dense object region proposals and learn to match those regions from a vast image collection. We further design a universal MASA adapter which can work in tandem with foundational segmentation or detection models and enable them to track any detected objects. Those combinations present strong zero-shot tracking ability in complex domains. Extensive tests on multiple challenging MOT and MOTS benchmarks indicate that the proposed method, using only unlabeled static images, achieves even better performance than state-of-the-art methods trained with fully annotated in-domain video sequences, in zero-shot association. Our code is available at github.com/siyuanliii/masa.
Siyuan Li 0008, Lei Ke, Martin Danelljan, Luigi Piccinelli, Mattia Segù, Luc Van Gool, Fisher Yu 0001
CVPR2
2024 SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
Siyuan Li 0008, Lei Ke, Yung-Hsu Yang, Luigi Piccinelli, Mattia Segù, Martin Danelljan, Luc Van Gool
ECCV (27)2
2024 Gaussian Grouping: Segment and Edit Anything in 3D Scenes
Mingqiao Ye, Martin Danelljan, Fisher Yu 0001, Lei Ke
ECCV (29)4
2024 Detecting Atomicity Violations for Interrupt-driven Programs via Systematic Scheduling and Prefix-directed Feedback
abstract
Interrupt-driven programs are widely used in safety-critical fields like aerospace and embedded systems. However, the unpredictable interleaving of Interrupt Service Routines (ISRs) can lead to concurrency bugs, particularly atomicity violations when ISRs preempt atomic sequences of instructions. To address this, we propose a dynamic approach for detecting atomicity violations in interrupt-driven programs. Extensive experiments demonstrate that our method is more precise and efficient than related approaches.
Ruixue Li, Bin Yu 0008, Xu Lu 0003, Lei Ke, Zixuan Yuan, Cong Tian 0001, Yansong Dong
ASE4
2024 DreamScene4D: Dynamic Multi-Object Scene Generation from Monocular Videos
abstract
View-predictive generative models provide strong priors for lifting object-centric images and videos into 3D and 4D through rendering and score distillation objectives. A question then remains: what about lifting complete multi-object dynamic scenes? There are two challenges in this direction: First, rendering error gradients are often insufficient to recover fast object motion, and second, view predictive generative models work much better for objects than whole scenes, so, score distillation objectives cannot currently be applied at the scene level directly. We present DreamScene4D, the first approach to generate 3D dynamic scenes of multiple objects from monocular videos via 360-degree novel view synthesis. Our key insight is a "decompose-recompose" approach that factorizes the video scene into the background and object tracks, while also factorizing object motion into 3 components: object-centric deformation, object-to-world-frame transformation, and camera motion. Such decomposition permits rendering error gradients and object view-predictive models to recover object 3D completions and deformations while bounding box tracks guide the large object movements in the scene. We show extensive results on challenging DAVIS, Kubric, and self-captured videos with quantitative comparisons and a user preference study. Besides 4D scene generation, DreamScene4D obtains accurate 2D persistent point track by projecting the inferred 3D trajectories to 2D. We will release our code and hope our work will stimulate more research on fine-grained 4D understanding from videos.
Wen-Hsuan Chu, Lei Ke, Katerina Fragkiadaki
NeurIPS2
2023 Mask-Free Video Instance Segmentation
abstract
The recent advancement in Video Instance Segmentation (VIS) has largely been driven by the use of deeper and increasingly data-hungry transformer-based models. However, video masks are tedious and expensive to annotate, limiting the scale and diversity of existing VIS datasets. In this work, we aim to remove the mask-annotation requirement. We propose MaskFreeVIS, achieving highly competitive VIS performance, while only using bounding box annotations for the object state. We leverage the rich temporal mask consistency constraints in videos by introducing the Temporal KNN-patch Loss (TK-Loss), providing strong mask supervision without any labels. Our TK-Loss finds one-to-many matches across frames, through an efficient patch-matching step followed by a K-nearest neighbor selection. A consistency loss is then enforced on the found matches. Our mask-free objective is simple to implement, has no trainable parameters, is computationally efficient, yet outperforms baselines employing, e.g., state-of-the-art optical flow to enforce temporal mask consistency. We validate MaskFreeVIS on the YouTube-VIS 2019/2021, OVIS and BDD100K MOTS benchmarks. The results clearly demonstrate the efficacy of our method by drastically narrowing the gap between fully and weakly-supervised VIS performance. Our code and trained models are available at http://vis.xyz/pub/maskfreevis.
Lei Ke, Martin Danelljan, Henghui Ding, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
CVPR1
2023 OVTrack: Open-Vocabulary Multiple Object Tracking
abstract
The ability to recognize, localize and track dynamic objects in a scene is fundamental to many real-world applications, such as self-driving and robotic systems. Yet, traditional multiple object tracking (MOT) benchmarks rely only on a few object categories that hardly represent the multitude of possible objects that are encountered in the real world. This leaves contemporary MOT methods limited to a small set of pre-defined object categories. In this paper, we address this limitation by tackling a novel task, open-vocabulary MOT, that aims to evaluate tracking beyond pre-defined training categories. We further develop OVTrack, an open-vocabulary tracker that is capable of tracking arbitrary object classes. Its design is based on two key ingredients: First, leveraging vision-language models for both classification and association via knowledge distillation; second, a data hallucination strategy for robust appearance feature learning from denoising diffusion probabilistic models. The result is an extremely data-efficient open-vocabulary tracker that sets a new state-of-the-art on the large-scale, large-vocabulary TAO benchmark, while being trained solely on static images.
Siyuan Li 0008, Tobias Fischer 0004, Lei Ke, Henghui Ding, Martin Danelljan, Fisher Yu 0001
CVPR3
2023 Cascade-DETR: Delving into High-Quality Universal Object Detection
abstract
Object localization in general environments is a fundamental part of vision systems. While dominating on the COCO benchmark, recent Transformer-based detection methods are not competitive in diverse domains. Moreover, these methods still struggle to very accurately estimate the object bounding boxes in complex environments.We introduce Cascade-DETR for high-quality universal object detection. We jointly tackle the generalization to diverse domains and localization accuracy by proposing the Cascade Attention layer, which explicitly integrates object-centric information into the detection decoder by limiting the attention to the previous box prediction. To further enhance accuracy, we also revisit the scoring of queries. Instead of relying on classification scores, we predict the expected IoU of the query, leading to substantially more well-calibrated confidences. Lastly, we introduce a universal object detection benchmark, UDB10, that contains 10 datasets from diverse domains. While also advancing the state-of-the-art on COCO, Cascade-DETR substantially improves DETR-based detectors on all datasets in UDB10, even by over 10 mAP in some cases. The improvements under stringent quality requirements are even more pronounced. Our code and pretrained models are at https://github.com/SysCV/cascade-detr.
Mingqiao Ye, Lei Ke, Siyuan Li 0008, Yu-Wing Tai, Chi-Keung Tang, Martin Danelljan, Fisher Yu 0001
ICCV2
2023 Segment Anything in High Quality
abstract
The recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs. We show the efficacy of HQ-SAM in a suite of 10 diverse segmentation datasets across different downstream tasks, where 8 out of them are evaluated in a zero-shot transfer protocol. Our code and pretrained models are at https://github.com/SysCV/SAM-HQ.
Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 0001, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
NeurIPS1
2023 BiMatting: Efficient Video Matting via Binarization
abstract
Real-time video matting on edge devices faces significant computational resource constraints, limiting the widespread use of video matting in applications such as online conferences and short-form video production. Binarization is a powerful compression approach that greatly reduces computation and memory consumption by using 1-bit parameters and bitwise operations. However, binarization of the video matting model is not a straightforward process, and our empirical analysis has revealed two primary bottlenecks: severe representation degradation of the encoder and massive redundant computations of the decoder. To address these issues, we propose BiMatting, an accurate and efficient video matting model using binarization. Specifically, we construct shrinkable and dense topologies of the binarized encoder block to enhance the extracted representation. We sparsify the binarized units to reduce the low-information decoding computation. Through extensive experiments, we demonstrate that BiMatting outperforms other binarized video matting models, including state-of-the-art (SOTA) binarization methods, by a significant margin. Our approach even performs comparably to the full-precision counterpart in visual quality. Furthermore, BiMatting achieves remarkable savings of 12.4$\times$ and 21.6$\times$ in computation and storage, respectively, showcasing its potential and advantages in real-world resource-constrained scenarios. Our code and models are released at https://github.com/htqin/BiMatting .
Haotong Qin, Lei Ke, Xudong Ma, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Xianglong Liu 0001, Fisher Yu 0001
NeurIPS2
2023 Occlusion-Aware Instance Segmentation Via BiLayer Network Architectures
abstract
Segmenting highly-overlapping image objects is challenging, because there is typically no distinction between real object contours and occlusion boundaries on images. Unlike previous instance segmentation methods, we model image formation as a composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top layer detects occluding objects (occluders) and the bottom layer infers partially occluded instances (occludees). The explicit modeling of occlusion relationship with bilayer structure naturally decouples the boundaries of both the occluding and occluded instances, and considers the interaction between them during mask regression. We investigate the efficacy of bilayer structure using two popular convolutional network designs, namely, Fully Convolutional Network (FCN) and Graph Convolutional Network (GCN). Further, we formulate bilayer decoupling using the vision transformer (ViT), by representing instances in the image as separate learnable occluder and occludee queries. Large and consistent improvements using one/two-stage and query-based object detectors with various backbones and network layer choices validate the generalization ability of bilayer decoupling, as shown by extensive experiments on image instance segmentation benchmarks (COCO, KINS, COCOA) and video instance segmentation benchmarks (YTVIS, OVIS, BDD100 K MOTS), especially for heavy occlusion cases.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Mask Transfiner for High-Quality Instance Segmentation
abstract
Two-stage and query-based instance segmentation methods have achieved remarkable results. However, their segmented masks are still very coarse. In this paper, we present Mask Transfiner for high-quality and efficient instance segmentation. Instead of operating on regular dense tensors, our Mask Transfiner decomposes and represents the image regions as a quadtree. Our transformer-based approach only processes detected error-prone tree nodes and self-corrects their errors in parallel. While these sparse pixels only constitute a small proportion of the total number, they are critical to the final mask quality. This allows Mask Transfiner to predict highly accurate instance masks, at a low computational cost. Extensive experiments demonstrate that Mask Transfiner outperforms current instance segmentation methods on three popular benchmarks, significantly improving both two-stage and query-based frameworks by a large margin of +3.0 mask AP on COCO and BDD100K, and +6.6 boundary AP on Cityscapes. Our code and trained models are available at https://github.com/SysCV/transfiner.
Lei Ke, Martin Danelljan, Xia Li 0005, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
CVPR1
2022 Video Mask Transfiner for High-Quality Video Instance Segmentation
Lei Ke, Henghui Ding, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
ECCV (28)1
2021 Deep Occlusion-Aware Instance Segmentation With Overlapping BiLayers
abstract
Segmenting highly-overlapping objects is challenging, because typically no distinction is made between real object contours and occlusion boundaries. Unlike previous two-stage instance segmentation methods, we model image formation as composition of two overlapping layers, and propose Bilayer Convolutional Network (BCNet), where the top GCN layer detects the occluding objects (occluder) and the bottom GCN layer infers partially occluded instance (occludee). The explicit modeling of occlusion relationship with bilayer structure naturally decouples the boundaries of both the occluding and occluded instances, and considers the interaction between them during mask regression. We validate the efficacy of bilayer decoupling on both one-stage and two-stage object detectors with different backbones and network layer choices. Despite its simplicity, extensive experiments on COCO and KINS show that our occlusion-aware BCNet achieves large and consistent performance gain especially for heavy occlusion cases. Code is available at https://github.com/lkeab/BCNet.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
CVPR1
2021 Occlusion-Aware Video Object Inpainting
abstract
Conventional video inpainting is neither object-oriented nor occlusion-aware, making it liable to obvious artifacts when large occluded object regions are inpainted. This paper presents occlusion-aware video object inpainting, which recovers both the complete shape and appearance for occluded objects in videos given their visible mask segmentation.To facilitate this new research, we construct the first large-scale video object inpainting benchmark YouTube-VOI to provide realistic occlusion scenarios with both occluded and visible object masks available. Our technical contribution VOIN jointly performs video object shape completion and occluded texture generation. In particular, the shape completion module models long-range object coherence while the flow completion module recovers accurate flow with sharp motion boundary, for propagating temporally-consistent texture to the same moving object across frames. For more realistic results, VOIN is optimized using both T-PatchGAN and a new spatio-temporal attention-based multi-class discriminator.Finally, we compare VOIN and strong baselines on YouTube-VOI. Experimental results clearly demonstrate the efficacy of our method including inpainting complex and dynamic objects. VOIN degrades gracefully with inaccurate input visible mask.
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
ICCV1
2021 Prototypical Cross-Attention Networks for Multiple Object Tracking and Segmentation
abstract
Multiple object tracking and segmentation requires detecting, tracking, and segmenting objects belonging to a set of given classes. Most approaches only exploit the temporal dimension to address the association problem, while relying on single frame predictions for the segmentation mask itself. We propose Prototypical Cross-Attention Network (PCAN), capable of leveraging rich spatio-temporal information for online multiple object tracking and segmentation. PCAN first distills a space-time memory into a set of prototypes and then employs cross-attention to retrieve rich information from the past frames. To segment each object, PCAN adopts a prototypical appearance module to learn a set of contrastive foreground and background prototypes, which are then propagated over time. Extensive experiments demonstrate that PCAN outperforms current video instance tracking and segmentation competition winners on both Youtube-VIS and BDD100K datasets, and shows efficacy to both one-stage and two-stage segmentation frameworks. Code and video resources are available at http://vis.xyz/pub/pcan.
Lei Ke, Xia Li 0005, Martin Danelljan, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001
NeurIPS1
2020 Cascaded Deep Monocular 3D Human Pose Estimation With Evolutionary Training Data
abstract
End-to-end deep representation learning has achieved remarkable accuracy for monocular 3D human pose estimation, yet these models may fail for unseen poses with limited and fixed training data. This paper proposes a novel data augmentation method that: (1) is scalable for synthesizing massive amount of training data (over 8 million valid 3D human poses with corresponding 2D projections) for training 2D-to-3D networks, (2) can effectively reduce dataset bias. Our method evolves a limited dataset to synthesize unseen 3D human skeletons based on a hierarchical human representation and heuristics inspired by prior knowledge. Extensive experiments show that our approach not only achieves state-of-the-art accuracy on the largest public benchmark, but also generalizes significantly better to unseen and rare poses. Relevant files and tools are available at the project website.
Shichao Li 0002, Lei Ke, Kevin Pratama, Yu-Wing Tai, Chi-Keung Tang, Kwang-Ting Cheng
CVPR2
2020 Commonality-Parsing Network Across Shape and Appearance for Partially Supervised Instance Segmentation
Lei Ke, Wenjie Pei, Chi-Keung Tang, Yu-Wing Tai
ECCV (8)2
2020 GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-Aware Supervision
Lei Ke, Shichao Li 0002, Yanan Sun 0005, Yu-Wing Tai, Chi-Keung Tang
ECCV (15)1
2019 Memory-Attended Recurrent Network for Video Captioning
abstract
Typical techniques for video captioning follow the encoder-decoder framework, which can only focus on one source video being processed. A potential disadvantage of such design is that it cannot capture the multiple visual context information of a word appearing in more than one relevant videos in training data. To tackle this limitation, we propose the Memory-Attended Recurrent Network (MARN) for video captioning, in which a memory structure is designed to explore the full-spectrum correspondence between a word and its various similar visual contexts across videos in training data. Thus, our model is able to achieve a more comprehensive understanding for each word and yield higher captioning quality. Furthermore, the built memory structure enables our method to model the compatibility between adjacent words explicitly instead of asking the model to learn implicitly, as most existing models do. Extensive validation on two real-word datasets demonstrates that our MARN consistently outperforms state-of-the-art methods.
Wenjie Pei, Xiangrong Wang 0002, Lei Ke, Xiaoyong Shen, Yu-Wing Tai
CVPR4
2019 Reflective Decoding Network for Image Captioning
abstract
State-of-the-art image captioning methods mostly focus on improving visual features, less attention has been paid to utilizing the inherent properties of language to boost captioning performance. In this paper, we show that vocabulary coherence between words and syntactic paradigm of sentences are also important to generate high-quality image caption. Following the conventional encoder-decoder framework, we propose the Reflective Decoding Network (RDN) for image captioning, which enhances both the long-sequence dependency and position perception of words in a caption decoder. Our model learns to collaboratively attend on both visual and textual features and meanwhile perceive each word's relative position in the sentence to maximize the information delivered in the generated caption. We evaluate the effectiveness of our RDN on the COCO image captioning datasets and achieve superior performance over the previous methods. Further experiments reveal that our approach is particularly advantageous for hard cases with complex scenes to describe by captions.
Lei Ke, Wenjie Pei, Ruiyu Li, Xiaoyong Shen, Yu-Wing Tai
ICCV1
2014 Coupling analysis of transcutaneous energy transfer coils with planar sandwich structure for a novel artificial anal sphincter
abstract
This paper presents a set of analytical expressions used to determine the coupling coefficient between primary and secondary Litz-wire planar coils used in a transcutaneous energy transfer (TET) system. A TET system has been designed to power a novel elastic scaling artificial anal sphincter system (ES-AASS) for treating severe fecal incontinence (FI), a condition that would benefit from an optimized TET. Expressions that describe the geometrical dimension dependence of self- and mutual inductances of planar coils on a ferrite substrate are provided. The effects of ferrite substrate conductivity, relative permeability, and geometrical dimensions are also considered. To verify these expressions, mutual coupling between planar coils is computed by 3D finite element analysis (FEA), and the proposed expressions show good agreement with numerical results. Different types of planar coils are fabricated with or without ferrite substrate. Measured results for each of the cases are compared with theoretical predictions and FEA solutions. The theoretical results and FEA results are in good agreement with the experimental data.
Lei Ke, Guozheng Yan, Zhiwu Wang, Dasheng Liu
J. Zhejiang Univ. Sci. C1
2012 Degrees of Freedom Region for an Interference Network With General Message Demands
abstract
We consider a single-hop interference network withKtransmitters andJreceivers, all havingMantennas. Each transmitter emits an independent message and each receiver requests an arbitrary subset of the messages. This generalizes the well-knownK-userM-antenna interference channel, where each message is requested by a unique receiver. For our setup, we derive the degrees of freedom (DoF) region. The achievability scheme generalizes the interference alignment schemes proposed by Cadambe and Jafar. In particular, we achieve general points in the DoF region by using multiple base vectors and aligning all interferers at a given receiver to the interferer with the largest DoF. As a byproduct, we obtain the DoF region for the original interference channel. We also discuss extensions of our approach where the same region can be achieved by considering a reduced set of interference alignment constraints, thus reducing the time-expansion duration needed. The DoF region for the considered system depends only on a subset of receivers whose demands meet certain characteristics. The geometric shape of the DoF region is also discussed.
Lei Ke, Aditya Ramamoorthy, Zhengdao Wang, Huarui Yin
IEEE Trans. Inf. Theory1
2012 Degrees of Freedom Regions of Two-User MIMO Z and Full Interference Channels: The Benefit of Reconfigurable Antennas
abstract
We study the degrees of freedom (DoF) regions of two-user multiple-input multiple-output Z and full interference channels in this paper. We assume that the receivers always have perfect channel state information. We first derive the DoF region of Z interference channel with channel state information at transmitter (CSIT). For full interference channel without CSIT, the DoF region has been fully characterized recently and it is shown that the previously known outer bound is not achievable. In this paper, we investigate the no-CSIT case further by assuming that one transmitter has the ability of antenna mode switching. We obtain the DoF region as a function of the number of available antenna modes and reveal the incremental gain in DoF that each additional antenna mode can bring. It is shown that, in certain cases, the reconfigurable antennas can increase the DoF. In these cases, the DoF region is maximized when the number of modes is at least equal to the number of receive antennas at the corresponding receiver, in which case the previous outer bound is achieved. In all cases, we propose systematic constructions of the beamforming and nulling matrices for achieving the DoF region. The constructions bear an interesting space-frequency coding interpretation.
Lei Ke, Zhengdao Wang
IEEE Trans. Inf. Theory1
2011 Degrees of freedom region for an interference network with general message demands
abstract
We consider a single hop interference network with K transmitters, each with an independent message and J receivers, all having the same number (M) of antennas. Each receiver requests an arbitrary subset of the messages. This generalizes the well-known K user M antenna interference channel, where each message is requested by a unique receiver. For this setup, we derive the exact degrees of freedom (DoF) region. Our achievability scheme generalizes the interference alignment scheme proposed by Cadambe and Jafar '08. In particular, we achieve general points in the DoF region by using multiple base vectors and aligning the interference at each receiver to its largest (in the DoF sense) interferer. As a byproduct of our analysis, we recover the DoF region for the original interference channel.
Lei Ke, Aditya Ramamoorthy, Zhengdao Wang, Huarui Yin
ISIT1
2011 Interference alignment and Degrees of Freedom region of cellular sigma channel
abstract
We investigate the Degrees of Freedom (DoF) Region of a cellular network, where the cells can have overlapping areas. Within an overlapping area, the mobile users can access multiple base stations. We consider a case where there are two base stations both equipped with multiple antennas. The mobile stations are all equipped with single antenna and each mobile station can belong to either a single cell or both cells. We completely characterize the DoF region for the uplink channel assuming that global channel state information is available at the transmitters. The achievability scheme is based on interference alignment at the base stations.
Huarui Yin, Lei Ke, Zhengdao Wang
ISIT2
2010 On the Degrees of Freedom Regions of Two-User MIMO Z and Full Interference Channels with Reconfigurable Antennas
abstract
We study the degrees of freedom (DoF) regions of two-user multiple-input multiple-output (MIMO) Z and full interference channels in this paper. We assume that the receivers always have perfect channel state information. We derive the DoF region of Z interference channel with channel state information at transmitter (CSIT). For full interference channel without CSIT, the DoF region has been obtained in previous work except for a special case M112, N2), where Miand Niare the number of transmit and receive antennas of user i, respectively. We show that for this case the DoF regions of the Z and full interference channels are the same. We establish the achievability based on the assumption of transmitter antenna mode switching. A systematic way of constructing the DoF-achieving nulling and beamforming matrices is presented in this paper.
Lei Ke, Zhengdao Wang
GLOBECOM1
2010 Monobit digital receivers: design, performance, and application to impulse radio
abstract
Digital receivers for future high-rate high-bandwidth communication systems will require large sampling rate. This is especially true for ultra-wideband (UWB) communications with impulse radio (IR) modulation. Due to high complexity and large power consumption, multibit high-rate analog-to-digital converter (ADC) is difficult to implement. Monobit receiver has been previously proposed to relax the need for high-rate ADC. In this paper, we derive optimal digital processing architecture for receivers based on monobit ADC with a certain over-sampling rate and the corresponding theoretically achievable performance. A practically appealing suboptimal iterative receiver is also proposed. Iterative decision-directed weight estimation, and small sample removal are distinctive features of the proposed detector. Numerical simulations show that compared with full resolution matched filter based receiver, the proposed low complexity monobit receiver incurs only 2 dB signal to noise ratio (SNR) loss in additive white Gaussian noise (AWGN) and 3.5dB SNR loss in standard UWB fading channels.
Huarui Yin, Zhengdao Wang, Lei Ke
IEEE Trans. Commun.3
2008 Capacity analysis of mimo ad hoc network with stream number selection
abstract
We study the problem of stream number selection in an ad hoc multiple-input multiple-output (MIMO) network, where there are a number of co-existent MIMO transmitter-receiver pairs. We assume that there is no channel state information (CSI) at the transmitters, and the receivers use single-user detection. It is shown that the link capacities at high signal-to-noise ratio (SNR) is mainly limited by the number of degrees of freedom available at the receiver side. In general the number of streams per transmitter should not be more than Nr/L when the interference level is high, where Nris the number of receive antennas and L is the number of simultaneous links. When the interference level is low, the number of streams should be set equal to the number of transmit antennas.
Lei Ke, Zhengdao Wang
ICASSP1
2008 Finite-Resolution Digital Receiver Design for Impulse Radio Ultra-Wideband Communication
abstract
Receiver design for impulse radio (IR) based ultra-wideband (UWB) communication is a challenge, because full-resolution digital receiver is difficult to implement under today's technology due to high sampling rate required. Some trade-offs can be made to the digital receiver, such as limiting amplitude resolution to only 1 bit, which results in a previously considered so-termed mono-bit receiver. In this paper, we consider the design of finite-resolution digital UWB receiver. We derive the optimal post-quantization processing, and analyze the achievable bit-error rate (BER) performance using an approximation of the log-likelihood ratio. We evaluate the effect of quantization threshold. The optimal threshold for two-bit quantization is obtained. Our work discloses the incremental gain that each sampling bit could bring and the results provide guidelines for designing IR UWB digital receivers.receivers.
Lei Ke, Zhengdao Wang, Huarui Yin, Weilin Gong
ICC1
2008 Finite-resolution digital receiver design for impulse radio ultra-wideband communication
abstract
Receiver design for impulse radio based ultrawideband (UWB) communication is a challenge. High sampling rate high resolution digital receiver is usually difficult to implement. Some tradeoffs can be made on the digital receiver, such as limiting amplitude resolution to only one bit, which results in a previously considered monobit receiver. In this paper, we consider the design of finite-resolution digital UWB receivers. We derive the optimal post-quantization processing, and analyze the achievable bit-error rate performance using an approximation of the log-likelihood ratio. Optimal thresholds for 4- and 3-level quantization are obtained. Training-based receiver template estimation is presented. Our work discloses the incremental gain that additional quantization levels offer and the results provide useful guidelines for designing impulse radio UWB digital receivers.
Lei Ke, Huarui Yin, Weilin Gong, Zhengdao Wang
IEEE Trans. Wirel. Commun.1