Kunming Luo

dblp:213/4872 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0002-5070-7392ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 17 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 13 since 2021
YearPublicationVenuePosition
2026 CTR3D: Cross-View Token Reduction for Dense Multi-View Generation
abstract
Recent multi-view diffusion (MVD) methods have utilized the generative capabilities of 2D image diffusion models to produce multi-view images from a single-view input. However, existing approaches often depend on dense crossview attention layers, which hinder scalability and fidelity due to their high computational costs. In this paper, we propose CTR3D, a novel method that incorporates token reduction in multi-view attention layers to efficiently generate dense, high-resolution multi-view images without restricting the camera viewpoints of the generated views. Our approach is designed into three key steps: redundancy removal, attention interaction, and token recovery. These steps leverage lightweight, projection-based techniques for multi-view token reduction and recovery, significantly improving the computational efficiency of MVD. By reducing the number of tokens in attention layers while preserving multi-view consistency, our model achieves state-of-the-art performance in novel view synthesis and 3D reconstruction while keeping efficiency for generation of dense high-resolution images and normals. Experimental results demonstrate that our method surpasses existing approaches, providing a more efficient and effective solution for multi-view generation. https://github.com/HKUST-SAIL/CTR3D
Kunming Luo, Hongyu Yan, Yuan Liu 0025, Manyuan Zhang, Wenping Wang 0001, Ping Tan 0002
3DV1
2026 Learning Efficient Meshflow and Optical Flow From Event Cameras
abstract
In this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we review the state-of-the-art in event-based flow estimation, highlighting two key areas for further research: i) the lack of meshflow-specific event datasets and methods, and ii) the underexplored challenge of event data density. First, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280 × 720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (30×faster) of our EEMFlow model compared to the recent state-of-the-art flow method. As an extension, we expand HREM into HREM+, a multi-density event dataset contributing to a thorough study of the robustness of existing methods across data with varying densities, and propose an Adaptive Density Module (ADM) to adjust the density of input event data to a more optimal range, enhancing the model's generalization ability. We empirically demonstrate that ADM helps to significantly improve the performance of EEMFlow and EEMFlow+ by 8% and 10%, respectively.
Xinglong Luo, Ao Luo, Kunming Luo, Zhengning Wang, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Ctrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout Constraints
abstract
Text-driven 3D indoor scene generation is useful for gaming, film industry, and AR/VR applications. However, existing methods cannot faithfully capture the scene layout based on text descriptions, nor do they allow flexible editing of individual objects in the room. To address these problems, we present Ctrl-Room, which can generate convincing 3D rooms with designer-style layouts and high-fidelity textures from just a text prompt. Our key insight is to separate the modeling of layouts and appearance. Our proposed method consists of two stages: a Layout Generation Stage and an Appearance Generation Stage. The Layout Generation Stage trains a text-conditional diffusion model to learn the layout distribution with our holistic scene code parameterization. Next, the Appearance Generation Stage employs a fine-tuned ControlNet to produce a vivid panoramic image of the room guided by the 3D scene layout, then further upgrades to a panoramic NeRF model. Benefiting from the scene code parameterization, we can easily edit the generated room model through our mask-guided editing module, without expensive edit-specific training. Extensive experiments on the Structured3D dataset demonstrate that our method outperforms existing methods in producing more reasonable, view-consistent, and editable 3D rooms from text prompts.
Chuan Fang, Kunming Luo, Xiaotao Hu, Rakesh Shrestha, Ping Tan 0002
3DV3
2025 Gaussianavatar-Editor: Photorealistic Animatable Gaussian Head Avatar Editor
abstract
We introduce GaussianAvatar-Editor, an innovative framework for text-driven editing of animatable Gaussian head avatars that can be fully controlled in expression, pose, and viewpoint. Unlike static 3D Gaussian editing, editing animatable 4D Gaussian avatars presents challenges related to motion occlusion and spatial-temporal inconsistency. To address these issues, we propose the Weighted Alpha Blending Equation (WABE). This function enhances the blending weight of visible Gaussians while suppressing the influence on non-visible Gaussians, effectively handling motion occlusion during editing. Furthermore, to improve editing quality and ensure 4D consistency, we incorporate conditional adversarial learning into the editing process. This strategy helps to refine the edited results and maintain consistency throughout the animation. By integrating these methods, our GaussianAvatar-Editor
Xiangyue Liu 0001, Kunming Luo, Heng Li 0009, Yuan Liu 0025, Li Yi 0001, Ping Tan 0002
3DV2
2025 SymmCompletion: High-Fidelity and High-Consistency Point Cloud Completion with Symmetry Guidance
abstract
Point cloud completion aims to recover a complete point shape from a partial point cloud. Although existing methods can form satisfactory point clouds in global completeness, they often lose the original geometry details and face the problem of geometric inconsistency between existing point clouds and reconstructed missing parts. To tackle this problem, we introduce SymmCompletion, a highly effective completion method based on symmetry guidance. Our method comprises two primary components: a Local Symmetry Transformation Network (LSTNet) and a Symmetry-Guidance Transformer (SGFormer). First, LSTNet efficiently estimates point-wise local symmetry transformation to transform key geometries of partial inputs into missing regions, thereby generating geometry-align partial-missing pairs and initial point clouds. Second, SGFormer leverages the geometric features of partial-missing pairs as the explicit symmetric guidance that can constrain the refinement process for initial point clouds. As a result, SGFormer can exploit provided priors to form high-fidelity and geometry-consistency final point clouds. Qualitative and quantitative evaluations on several benchmark datasets demonstrate that our method outperforms state-of-the-art completion networks.
Hongyu Yan, Kunming Luo
AAAI3
2024 GenN2N: Generative NeRF2NeRF Translation
abstract
We present GenN2N, a unified NeRF-to-NeRF translation framework for various NeRF translation tasks such as text-driven NeRF editing, colorization, super-resolution, in-painting, etc. Unlike previous methods designed for individual translation tasks with task-specific schemes, GenN2N achieves all these NeRF editing tasks by employing a plug-and-play image-to-image translator to perform editing in the 2D domain and lifting 2D edits into the 3D NeRF space. Since the 3D consistency of 2D edits may not be assured, we propose to model the distribution of the underlying 3D edits through a generative model that can cover all possible edited NeRFs. To model the distribution of 3D edited NeRFs from 2D edited images, we carefully design a VAE-GAN that encodes images while decoding NeRFs. The latent space is trained to align with a Gaussian distribution and the NeRFs are supervised through an adversarial loss on its renderings. To ensure the latent code does not depend on 2D viewpoints but truly reflects the 3D edits, we also regularize the latent code through a contrastive learning scheme. Extensive experiments on various editing tasks show GenN2N, as a universal framework, performs as well or better than task-specific specialists while possessing flexible generative power. More results on our project page: https://xiangyueliu.github.io/GenN2N/.
Xiangyue Liu 0001, Kunming Luo, Ping Tan 0002, Li Yi 0001
CVPR3
2024 PointRegGPT: Boosting 3D Point Cloud Registration Using Generative Point-Cloud Pairs for Training
Suyi Chen, Hao Xu 0018, Haipeng Li 0001, Kunming Luo, Guanghui Liu 0001, Chi-Wing Fu, Ping Tan 0002, Shuaicheng Liu
ECCV (51)4
2024 Diffusion Posterior Proximal Sampling for Image Restoration
abstract
Diffusion models have demonstrated remarkable efficacy in generating high-quality samples. Existing diffusion-based image restoration algorithms exploit pre-trained diffusion models to leverage data priors, yet they still preserve elements inherited from the unconditional generation paradigm. These strategies initiate the denoising process with pure white noise and incorporate random noise at each generative step, leading to over-smoothed results. In this paper, we present a refined paradigm for diffusion-based image restoration. Specifically, we opt for a sample consistent with the measurement identity at each generative step, exploiting the sampling selection as an avenue for output stability and enhancement. The number of candidate samples used for selection is adaptively determined based on the signal-to-noise ratio of the timestep. Additionally, we start the restoration process with an initialization combined with the measurement signal, providing supplementary information to better align the generative process. Extensive experimental results and analyses validate that our proposed method significantly enhances image restoration performance while consuming negligible additional computational resources.
Hongjie Wu, Linchao He, Mingqin Zhang, Dongdong Chen 0004, Kunming Luo, Mengting Luo, Jizhe Zhou 0001, Hu Chen 0002, Jiancheng Lv 0001
ACM Multimedia5
2024 GyroFlow+: Gyroscope-Guided Unsupervised Deep Homography and Optical Flow Learning
Haipeng Li 0001, Kunming Luo, Bing Zeng 0001, Shuaicheng Liu
Int. J. Comput. Vis.2
2024 GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng 0002, Kunming Luo, Shuaicheng Liu, Zun Li 0001, Ye Xiang, Lifang Wu, Bing Zeng 0001, Chang Wen Chen
Pattern Recognit. Lett.2
2023 Learning Optical Flow from Event Camera with Rendered Dataset
abstract
We study the problem of estimating optical flow from event cameras. One important issue is how to build a high-quality event-flow dataset with accurate event values and flow labels. Previous datasets are created by either capturing real scenes by event cameras or synthesizing from images with pasted foreground objects. The former case can produce real event values but with calculated flow labels, which are sparse and inaccurate. The latter case can generate dense flow labels but the interpolated events are prone to errors. In this work, we propose to render a physically correct event-flow dataset using computer graphics models. In particular, we first create indoor and outdoor 3D scenes by Blender with rich scene content variations. Second, diverse camera motions are included for the virtual capturing, producing images and accurate flow labels. Third, we render high-framerate videos between images for accurate events. The rendered dataset can adjust the density of events, based on which we further introduce an adaptive density module (ADM). Experiments show that our proposed dataset can facilitate event-flow learning, whereas previous approaches when trained on our dataset can improve their performances constantly by a relatively large margin. In addition, event-flow pipelines when equipped with our ADM can further improve performances. Our code is available at https://github.com/boomluo02/ADMFlow.
Xinglong Luo, Kunming Luo, Ao Luo, Zhengning Wang, Ping Tan 0002, Shuaicheng Liu
ICCV2
2023 AccFlow: Backward Accumulation for Long-Range Optical Flow
abstract
Recent deep learning-based optical flow estimators have exhibited impressive performance in generating local flows between consecutive frames. However, the estimation of long-range flows between distant frames, particularly under complex object deformation and large motion occlusion, remains a challenging task. One promising solution is to accumulate local flows explicitly or implicitly to obtain the desired long-range flow. Nevertheless, the accumulation errors and flow misalignment can hinder the effectiveness of this approach. This paper proposes a novel recurrent framework called AccFlow, which recursively backward accumulates local flows using a deformable module called as AccPlus. In addition, an adaptive blending module is designed along with AccPlus to alleviate the occlusion effect by backward accumulation and rectify the accumulation error. Notably, we demonstrate the superiority of backward accumulation over conventional forward accumulation, which to the best of our knowledge has not been explicitly established before. To train and evaluate the proposed AccFlow, we have constructed a large-scale high-quality dataset named CVO, which provides ground-truth optical flow labels between adjacent and distant frames. Extensive experiments validate the effectiveness of AccFlow in handling long-range optical flow estimation. Codes are available at https://github.com/mulns/AccFlow.
Guangyang Wu, Xiaohong Liu 0001, Kunming Luo, Qingqing Zheng, Shuaicheng Liu, Xinyang Jiang, Guangtao Zhai, Wenyi Wang 0005
ICCV3
2023 The Elliptic Energy Loss for Rotated Object Detection in Aerial Images
abstract
Rotated object detection is a promising yet challenging task in computer vision. Existing algorithms mainly train the rotated detector by the Ln-norm loss, which is inconsistent with the evaluation metric of Intersection over Union (IoU). However, the concise and efficient solution using the loss based on IoU between oriented boxes is hindered by its non-differentiability. In this paper, we propose Elliptic Energy Loss, a differentiable loss based on curve energy, to fit with the evaluation metric IoU. Specifically, given a pair of predicted and ground truth boxes, we first convert them to curve representations using the elliptic transformation. Then, the curve energy is calculated to measure the similarity between the predicted and ground truth curves. Finally, the curve energy is used as regression loss to optimize rotated detectors. We conduct experiments with different detectors on DOTA and HRSC2016 datasets, which demonstrate that the performance is significantly improved by our proposed loss. The code is available at https://github.com/zhangc-uestc/EEL.
Kunming Luo, Fanman Meng, Qingbo Wu 0001
ICIP2
2023 Content-Aware Unsupervised Deep Homography Estimation and its Extensions
abstract
Homography estimation is a basic image alignment method in many applications. It is usually done by extracting and matching sparse feature points, which are error-prone in low-light and low-texture images. On the other hand, previous deep homography approaches use either synthetic images for supervised learning or aerial images for unsupervised learning, both ignoring the importance of handling depth disparities and moving objects in real-world applications. To overcome these problems, in this work, we propose an unsupervised deep homography method with a new architecture design. In the spirit of the RANSAC procedure in traditional methods, we specifically learn an outlier mask to only select reliable regions for homography estimation. We calculate loss with respect to our learned deep features instead of directly comparing image content as did previously. To achieve the unsupervised training, we also formulate a novel triplet loss customized for our network. We verify our method by conducting comprehensive comparisons on a new dataset that covers a wide range of scenes with varying degrees of difficulties for the task. Experimental results reveal that our method outperforms the state-of-the-art, including deep solutions and feature-based solutions.
Shuaicheng Liu, Nianjin Ye, Chuan Wang 0001, Jirong Zhang, Lanpeng Jia, Kunming Luo, Jue Wang 0001, Jian Sun 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Learning Optical Flow with Adaptive Graph Reasoning
abstract
Estimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explicitly reason over the given scene for achieving a holistic motion understanding. In this work, taking a fresh perspective, we introduce a novel graph-based approach, called adaptive graph reasoning for optical flow (AGFlow), to emphasize the value of scene/context information in optical flow. Our key idea is to decouple the context reasoning from the matching procedure, and exploit scene information to effectively assist motion estimation by learning to reason over the adaptive graph. The proposed AGFlow can effectively exploit the context information and incorporate it within the matching procedure, producing more robust and accurate results. On both Sintel clean and final passes, our AGFlow achieves the best accuracy with EPE of 1.43 and 2.47 pixels, outperforming state-of-the-art approaches by 11.2% and 13.6%, respectively. Code is publicly available at https://github.com/megvii-research/AGFlow.
Ao Luo, Fan Yang 0054, Kunming Luo, Xin Li 0079, Haoqiang Fan, Shuaicheng Liu
AAAI3
2022 RealFlow: EM-Based Realistic Optical Flow Dataset Generation from Videos
Yunhui Han, Kunming Luo, Ao Luo, Jiangyu Liu, Haoqiang Fan, Guiming Luo, Shuaicheng Liu
ECCV (19)2
2022 ASFlow: Unsupervised Optical Flow Learning With Adaptive Pyramid Sampling
abstract
We present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propose a Content-Aware Pooling (CAP) module, which promotes local feature gathering by avoiding cross region pooling, so that the learned features become more representative. In the pyramid upsampling, we propose an Adaptive Flow Upsampling (AFU) module, where cross edge interpolation can be avoided, producing sharp motion boundaries. Equipped with these two modules, our method achieves the best performance for unsupervised optical flow estimation on multiple leading benchmarks, including MPI-Sintel, KITTI 2012 and KITTI 2015. Particularly, we achieve EPE=1.5 on KITTI 2012 and F1=9.67% KITTI 2015, which outperform the previous state-of-the-art methods by 16.7% and 13.1%, respectively.
Shuaicheng Liu, Kunming Luo, Ao Luo, Chuan Wang 0001, Fanman Meng, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 UPFlow: Upsampling Pyramid for Unsupervised Optical Flow Learning
abstract
We present an unsupervised learning approach for optical flow estimation by improving the upsampling and learning of pyramid network. We design a self-guided upsample module to tackle the interpolation blur problem caused by bilinear upsampling between pyramid levels. Moreover, we propose a pyramid distillation loss to add supervision for intermediate levels via distilling the finest flow as pseudo labels. By integrating these two components together, our method achieves the best performance for unsupervised optical flow learning on multiple leading benchmarks, including MPI-SIntel, KITTI 2012 and KITTI 2015. In particular, we achieve EPE=1.4 on KITTI 2012 and F1=9.38% on KITTI 2015, which outperform the previous state-of-the-art methods by 22.2% and 15.7%, respectively.
Kunming Luo, Chuan Wang 0001, Shuaicheng Liu, Haoqiang Fan, Jue Wang 0001, Jian Sun 0001
CVPR1
2021 GyroFlow: Gyroscope-Guided Unsupervised Optical Flow Learning
abstract
Existing optical flow methods are erroneous in challenging scenes, such as fog, rain, and night because the basic optical flow assumptions such as brightness and gradient constancy are broken. To address this problem, we present an unsupervised learning approach that fuses gyroscope into optical flow learning. Specifically, we first convert gyroscope readings into motion fields named gyro field. Second, we design a self-guided fusion module to fuse the background motion extracted from the gyro field with the optical flow and guide the network to focus on motion details. To the best of our knowledge, this is the first deep learning-based framework that fuses gyroscope data and image content for optical flow learning. To validate our method, we propose a new dataset that covers regular and challenging scenes. Experiments show that our method outperforms the state-of-art methods in both regular and challenging scenes. Code and dataset are available at https://github.com/megvii-research/GyroFlow.
Haipeng Li 0001, Kunming Luo, Shuaicheng Liu
ICCV2
2021 OIFlow: Occlusion-Inpainting Optical Flow Estimation by Unsupervised Learning
abstract
Occlusion is an inevitable and critical problem in unsupervised optical flow learning. Existing methods either treat occlusions equally as non-occluded regions or simply remove them to avoid incorrectness. However, the occlusion regions can provide effective information for optical flow learning. In this paper, we present OIFlow, an occlusion-inpainting framework to make full use of occlusion regions. Specifically, a new appearance-flow network is proposed to inpaint occluded flows based on the image content. Moreover, a boundary dilated warp is proposed to deal with occlusions caused by displacement beyond the image border. We conduct experiments on multiple leading flow benchmark datasets such as Flying Chairs, KITTI and MPI-Sintel, which demonstrate that the performance is significantly improved by our proposed occlusion handling framework.
Shuaicheng Liu, Kunming Luo, Nianjin Ye, Chuan Wang 0001, Jue Wang 0001, Bing Zeng 0001
IEEE Trans. Image Process.2
2021 Semi-Supervised Pixel-Level Scene Text Segmentation by Mutually Guided Network
abstract
In this paper we present a new data-driven method for pixel-level scene text segmentation from a single natural image. Although scene text detection, i.e. producing a text region mask, has been well studied in the past decade, pixel-level text segmentation is still an open problem due to the lack of massive pixel-level labeled data for supervised training. To tackle this issue, we incorporate text region mask as an auxiliary data into this task, considering acquiring large-scale of labeled text region mask is commonly less expensive and time-consuming. To be specific, we propose a mutually guided network which produces a polygon-level mask in one branch and a pixel-level text mask in the other. The two branches' outputs serve as guidance for each other and the whole network is trained via a semi-supervised learning strategy. Extensive experiments are conducted to demonstrate the effectiveness of our mutually guided network, and experimental results show our network outperforms the state-of-the-art in pixel-level scene text segmentation. We also demonstrate the mask produced by our network could improve the text recognition performance besides the trivial image editing application.
Chuan Wang 0001, Shan Zhao 0010, Li Zhu 0003, Kunming Luo, Yanwen Guo 0001, Jue Wang 0001, Shuaicheng Liu
IEEE Trans. Image Process.4
2020 Weakly Supervised Semantic Segmentation by a Class-Level Multiple Group Cosegmentation and Foreground Fusion Strategy
abstract
Weakly supervised semantic segmentation uses image-level labels to extract object regions. The existing methods focus on efficiently training CNN-based segmentation networks using the image-level labels. In contrast to the existing methods, this paper proposes a new fusion-based method, which first segments the foregrounds of each image by multiple group cosegmentation and then generates the semantic segmentation by combining the foregrounds. Specifically, a new CNN-based multiple group cosegmentation network is first proposed to segment foregrounds employing two cues, the discriminative cue and the local-to-global cue. Then, the fusion method is proposed to simply perform semantic segmentation based on the multiple group cosegmentation results. Experiments on the PASCAL VOC 2012 and MS COCO 2017 datasets demonstrate the effectiveness of the proposed method with mIoU values that are obviously larger than those of the existing methods.
Fanman Meng, Kunming Luo, Hongliang Li 0001, Qingbo Wu 0001, Xiaolong Xu 0004
IEEE Trans. Circuits Syst. Video Technol.2
2018 Weakly Supervised Semantic Segmentation by Multiple Group Cosegmentation
abstract
Weakly supervised semantic segmentation aims at segmenting images by image-level labels. The existing methods try to train an end-to-end CNN network, which needs to handle multiple classes that is difficult. In addition, the existing methods are sensitive to the image-level cues such as discriminative regions and the pseudo-annotations. To avoid these drawbacks, this paper proposes a new strategy, which first obtains the foregrounds of each class by multiple group cosegmentation, and then combines the results to form the semantic segmentation. In our method, three new aspects are considered. (1) we solve semantic segmentation by each class that is easy to handle. (2) we extract discriminative regions more globally by context analysis. (3) we learn local-to-global segmentation network to segment the object from local discriminative priors. A new CNN network for multiple group cosegmentation is proposed. Two subnetworks such as global context based discriminative region extraction network and local-to-global segmentation network are designed. A simple combination method based on the discriminative map is proposed to finally obtain the semantic segmentation results. We verify the proposed method on Pascal VOC dataset. The experimental results show that the proposed method can obtain mIOU value 0.563 and 0.603 (without CRF post-processing) on the validation and test dataset that outperforms many existing weakly supervised semantic segmentation methods.
Kunming Luo, Fanman Meng, Qingbo Wu 0001, Hongliang Li 0001
VCIP1