Luming Liang

dblp:46/6624 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-1127-2568ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021
YearPublicationVenuePosition
2026 ProCrop: Learning Aesthetic Image Cropping from Professional Compositions
abstract
Image cropping is crucial for enhancing the visual appeal and narrative impact of photographs, yet existing rule-based and data-driven approaches often lack diversity or require annotated training data. We introduce ProCrop, a retrieval-based method that leverages professional photography to guide cropping decisions. By fusing features from professional photographs with those of the query image, ProCrop learns from professional compositions, significantly boosting performance. Additionally, we present a large-scale dataset of 242K weakly-annotated images, generated by out-painting professional images and iteratively refining diverse crop proposals. This composition-aware dataset generation offers diverse high-quality crop proposals guided by aesthetic principles and becomes the largest publicly available dataset for image cropping. Extensive experiments show that ProCrop significantly outperforms existing methods in both supervised and weakly-supervised settings. Notably, when trained on the new dataset, our ProCrop surpasses previous weakly-supervised methods and even matches fully supervised approaches.
Tianyu Ding, Jiachen Jiang, Ilya Zharkov, Vishal M. Patel, Luming Liang
AAAI7
2025 DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
abstract
Despite the success of distillation in large language models (LLMs), most prior work applies identical loss functions to both teacher- and student-generated data. These strategies overlook the synergy between loss formulations and data types, leading to a suboptimal performance boost in student models. To address this, we propose DistiLLM-2, a contrastive approach that simultaneously increases the likelihood of teacher responses and decreases that of student responses by harnessing this synergy. Our extensive experiments show that DistiLLM-2 not only builds high-performing student models across a wide range of tasks, including instruction-following and code generation, but also supports diverse applications, such as preference alignment and vision-language extensions. These findings highlight the potential of a contrastive approach to enhance the efficacy of LLM distillation by effectively aligning teacher and student models across varied data types.
Jongwoo Ko, Sungnyun Kim, Tianyu Ding, Luming Liang, Ilya Zharkov, Se-Young Yun
ICML5
2025 Kernel Prediction Network for Offset Domain Common Image Gather Flattening and Correction
abstract
In prestack Kirchhoff depth migration, the quality of the migration profile is determined by how well common image gathers (CIGs) are flattened and corrected. The migration velocity analysis (MVA) method is proposed to flatten the offset domain CIGs (ODCIGs) by updating the migration velocities. However, conventional MVA such as the residual curvature analysis (RCA) method is typically challenging to complex structures such as lateral velocity variations or high-dip reflectors. In addition, there are structural artifacts in ODCIGs due to the multipath ray problem even if the migration velocity is accurate. To address the above problems, we developed a kernel prediction network (KPN) for ODCIG flattening and correction. Compared with conventional neural networks, the primary advantage of the KPN is that its outputs consist of a series of predicted kernels instead of pixel vectors or matrices. These predicted kernels are capable of processing the input ODCIGs pixel by pixel and slice by slice. The KPN is built by an encoder-decoder architecture, and we modified the loss function of the KPN and introduced an extra parameter associated with the migration offset to ensure that the network is more effective for the ODCIG problem.The training samples of the KPN are acquired by a random extraction algorithm, and the corresponding labels are calculated by a convolution method.Image enhancements are also applied in training samples to improve the generalization capability of the KPN. We demonstrate the effectiveness of the KPN method by comparing it with the RCA method in both synthetic and field data examples. The results show that the KPN method can flatten the events in ODCIGs, correct the depth of the improperly migrated reflectors, remove unfocused artifacts simultaneously, and further yield high-quality migration profiles in different geological examples.
Xinming Wu, Peimin Zhu, Luming Liang, Hao Zhang 0116, Zhiying Liao
IEEE Trans. Geosci. Remote. Sens.4
2024 DREAM: Diffusion Rectification and Estimation-Adaptive Models
abstract
We present DREAM, a novel training framework representing Diffusion Rectification and Estimation-Adaptive Models, requiring minimal code changes (just three lines) yet significantly enhancing the alignment of training with sampling in diffusion models. DREAM features two components: diffusion rectification, which adjusts training to reflect the sampling process, and estimation adaptation, which balances perception against distortion. When applied to image super-resolution (SR), DREAM adeptly navigates the tradeoff between minimizing distortion and preserving high image quality. Experiments demonstrate DREAM's superiority over standard diffusion-based SR methods, showing a 2 to 3× faster training convergence and a 10 to 20× reduction in sampling steps to achieve comparable results. We hope DREAM will inspire a rethinking of diffusion model training paradigms. Our source code is available at link.
Jinxin Zhou, Tianyu Ding, Jiachen Jiang, Ilya Zharkov, Zhihui Zhu, Luming Liang
CVPR7
2024 CaesarNeRF: Calibrated Semantic Representation for Few-Shot Generalizable Neural Rendering
Haidong Zhu, Tianyu Ding, Ilya Zharkov, Ramakant Nevatia, Luming Liang
ECCV (6)6
2024 Motion Graph Unleashed: A Novel Approach to Video Prediction
abstract
We introduce motion graph, a novel approach to address the video prediction problem, i.e., predicting future video frames from limited past data. The motion graph transforms patches of video frames into interconnected graph nodes, to comprehensively describe the spatial-temporal relationships among them. This representation overcomes the limitations of existing motion representations such as image differences, optical flow, and motion matrix that either fall short in capturing complex motion patterns or suffer from excessive memory consumption. We further present a video prediction pipeline empowered by motion graph, exhibiting substantial performance improvements and cost reductions. Extensive experiments on various datasets, including UCF Sports, KITTI and Cityscapes, highlight the strong representative ability of motion graph. Especially on UCF Sports, our method matches and outperforms the SOTA methods with a significant reduction in model size by 78% and a substantial decrease in GPU memory utilization by 47%.
Yiqi Zhong, Luming Liang, Bohan Tang, Ilya Zharkov, Ulrich Neumann
NeurIPS2
2023 MMVP: Motion-Matrix-based Video Prediction
abstract
A central challenge of video prediction lies where the system has to reason the objects’ future motions from image frames while simultaneously maintaining the consistency of their appearances across frames. This work introduces an end-to-end trainable two-stream video prediction framework, Motion-Matrix-based Video Prediction (MMVP), to tackle this challenge. Unlike previous methods that usually handle motion prediction and appearance maintenance within the same set of modules, MMVP decouples motion and appearance information by constructing appearance-agnostic motion matrices. The motion matrices represent the temporal similarity of each and every pair of feature patches in the input frames, and are the sole input of the motion prediction module in MMVP. This design improves video prediction in both accuracy and efficiency, and reduces the model size. Results of extensive experiments demonstrate that MMVP outperforms state-of-the-art systems on public data sets by non-negligible large margins (≈ 1 db in PSNR, UCF Sports) in significantly smaller model sizes (84% the size or smaller). Please refer to this $link$ for the official code and the datasets used in this paper.
Yiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich Neumann
ICCV2
2023 OTOv2: Automatic, Generic, User-Friendly
Luming Liang, Tianyu Ding, Zhihui Zhu, Ilya Zharkov
ICLR2
2023 CF-YOLO: Cross Fusion YOLO for Object Detection in Adverse Weather With a High-Quality Real Snow Dataset
abstract
Snow is one of the toughest adverse weather conditions for object detection (OD). Currently, not only there is a lack of snowy OD datasets to train cutting-edge detectors, but also these detectors have difficulties of learning latent information beneficial for detection in snow. To alleviate the two above problems, we first establish a real-world snowy OD dataset, named RSOD. Besides, we develop an unsupervised training strategy with a distinctive activation function, called$Peak Act$, to quantitatively evaluate the effect of snow on each object. Peak Act helps grade the images in RSOD into four-difficulty levels. To our knowledge, RSOD is the first quantitatively evaluated and graded real-world snowy OD dataset. Then, we propose a novel Cross Fusion (CF) block to construct a lightweight OD network based on YOLOv5s (called CF-YOLO). CF is a plug-and-play feature aggregation module, which integrates the advantages of Feature Pyramid Network and Path Aggregation Network in a simpler yet more flexible form. Both RSOD and CF lead our CF-YOLO to possess an optimization ability for OD in real-world snow. That is, CF-YOLO can handle unfavorable detection problems of vagueness, distortion and covering of snow. Experiments show that our CF-YOLO achieves better detection results on RSOD, compared to SOTAs. The code and dataset are available athttps://github.com/qqding77/CF-YOLO-and-RSOD.
Qiqi Ding, Peng Li 0064, Xuefeng Yan 0001, Ding Shi, Luming Liang, Weiming Wang 0002, Haoran Xie 0001, Jonathan Li 0001, Mingqiang Wei
IEEE Trans. Intell. Transp. Syst.5
2022 RSTT: Real-time Spatial Temporal Transformer for Space-Time Video Super-Resolution
abstract
Space-time video super-resolution (STVSR) is the task of interpolating videos with both Low Frame Rate (LFR) and Low Resolution (LR) to produce High-Frame-Rate (HFR) and also High-Resolution (HR) counterparts. The existing methods based on Convolutional Neural Network (CNN) succeed in achieving visually satisfied results while suffer from slow inference speed due to their heavy architec-tures. We propose to resolve this issue by using a spatial-temporal transformer that naturally incorporates the spa-tial and temporal super resolution modules into a single model. Unlike CNN-based methods, we do not explic-itly use separated building blocks for temporal interpolations and spatial super-resolutions; instead, we only use a single end-to-end transformer architecture. Specifically, a reusable dictionary is built by encoders based on the in-put LFR and LR frames, which is then utilized in the de-coder part to synthesize the HFR and HR frames. compared with the state-of-the-art TMNet [54], our network is 60% smaller (4.5M vs 12.3M parameters) and 80% faster (26.2fps vs 14.3fps on 720 x 576 frames) without sacri-ficing much performance. The source code is available at https://github.com/llmpass/RSTT.
Zhicheng Geng, Luming Liang, Tianyu Ding, Ilya Zharkov
CVPR2
2022 Accurate structure from motion using consistent cluster merging
Luming Liang, Jianquan Ouyang 0001
Multim. Tools Appl.2
2022 LOUD: Local Orthogonalization-Constrained Unsupervised Deep-Learning Denoiser
abstract
Random noise attenuation of seismic data is a fundamental problem in seismic data processing. It is not only an important problem itself but also is a crucial step for the subsequent tasks, e.g., migration and inversion. We propose a local orthogonalization constrained unsupervised deep learning denoiser (LOUD) to suppress seismic random noise based on a new loss function that specifically adapts to seismic data. Through unsupervised learning, we eliminate the common need in supervised learning approaches of collecting or generating sizable clean and noisy image pairs, which is challenging and expensive, especially for seismic data. We utilize a deep convolutional autoencoder to reconstruct the clean seismic image and leverage the local signal-and-noise orthogonalization as a constraint to guarantee that the removed noise component is orthogonal to the recovered signal. Experimental results on both synthetic and field datasets exhibit the effectiveness of our proposed method over traditional denoising methods.
Zhicheng Geng, Yangkang Chen, Sergey Fomel, Luming Liang
IEEE Trans. Geosci. Remote. Sens.4
2021 CDFI: Compression-Driven Network Design for Frame Interpolation
abstract
DNN-based frame interpolation—that generates the intermediate frames given two consecutive frames—typically relies on heavy model architectures with a huge number of features, preventing them from being deployed on systems with limited resources, e.g., mobile devices. We propose a compression-driven network design for frame interpolation (CDFI), that leverages model pruning through sparsity-inducing optimization to significantly reduce the model size while achieving superior performance. Concretely, we first compress the recently proposed AdaCoF model and show that a 10× compressed AdaCoF performs similarly as its original counterpart; then we further improve this compressed model by introducing a multi-resolution warping module, which boosts visual consistencies with multi-level details. As a consequence, we achieve a significant performance gain with only a quarter in size compared with the original AdaCoF. Moreover, our model performs favorably against other state-of-the-arts in a broad range of datasets. Finally, the proposed compression-driven framework is generic and can be easily transferred to other DNN-based frame interpolation algorithm. Our source code is available at https://github.com/tding1/CDFI.
Tianyu Ding, Luming Liang, Zhihui Zhu, Ilya Zharkov
CVPR2
2021 Only Train Once: A One-Shot Neural Network Training And Pruning Framework
abstract
Structured pruning is a commonly used technique in deploying deep neural networks (DNNs) onto resource-constrained devices. However, the existing pruning methods are usually heuristic, task-specified, and require an extra fine-tuning procedure. To overcome these limitations, we propose a framework that compresses DNNs into slimmer architectures with competitive performances and significant FLOPs reductions by Only-Train-Once (OTO). OTO contains two key steps: (i) we partition the parameters of DNNs into zero-invariant groups, enabling us to prune zero groups without affecting the output; and (ii) to promote zero groups, we then formulate a structured-sparsity optimization problem, and propose a novel optimization algorithm, Half-Space Stochastic Projected Gradient (HSPG), to solve it, which outperforms the standard proximal methods on group sparsity exploration, and maintains comparable convergence. To demonstrate the effectiveness of OTO, we train and compress full models simultaneously from scratch without fine-tuning for inference speedup and parameter reduction, and achieve state-of-the-art results on VGG16 for CIFAR10, ResNet50 for CIFAR10 and Bert for SQuAD and competitive result on ResNet50 for ImageNet. The source code is available at https://github.com/tianyic/onlytrainonce.
Bo Ji 0003, Tianyu Ding, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Xiao Tu
NeurIPS7
2021 Convolutional neural network with median layers for denoising salt-and-pepper contaminations
Luming Liang, Lionel Gueguen, Mingqiang Wei, Xinming Wu, Harry Qin
Neurocomputing1
2020 Detail-recovery Image Deraining via Context Aggregation Networks
abstract
This paper looks at this intriguing question: are single images with their details lost during deraining, reversible to their artifact-free status? We propose an end-to-end detail-recovery image deraining network (termed a DRDNet) to solve the problem. Unlike existing image deraining approaches that attempt to meet the conflicting goal of simultaneously deraining and preserving details in a unified framework, we propose to view rain removal and detail recovery as two seperate tasks, so that each part could specialize rather than trade-off between two conflicting goals. Specifically, we introduce two parallel sub-networks with a comprehensive loss function which synergize to derain and recover the lost details caused by deraining. For complete rain removal, we present a rain residual network with the squeeze-and-excitation (SE) operation to remove rain streaks from the rainy images. For detail recovery, we construct a specialized detail repair network consisting of welldesigned blocks, named structure detail context aggregation block (SDCAB), to encourage the lost details to return for eliminating image degradations. Moreover, the detail recovery branch of our proposed detail repair framework is detachable and can be incorporated into existing deraining methods to boost their performances. DRD-Net has been validated on several well-known benchmark datasets in terms of deraining robustness and detail accuracy. Comparisons show clear visual and numerical improvements of our method over the state-of-the-arts.
Mingqiang Wei, Jun Wang 0039, Yidan Feng, Luming Liang, Haoran Xie 0001, Fu Lee Wang, Meng Wang 0001
CVPR5
2020 Accurate 3D motion tracking by combining image alignment and feature matching
Luming Liang, Jianquan Ouyang 0001
Multim. Tools Appl.2
2019 FaultNet3D: Predicting Fault Probabilities, Strikes, and Dips With a Single Convolutional Neural Network
abstract
We simultaneously estimate fault probabilities, strikes, and dips directly from a seismic image by using a single convolutional neural network (CNN). In this method, we assume a local 3-D fault is a plane defined by a single combination of strike and dip angles. We assume the fault strikes and dips, respectively, are in the ranges of [0°, 360°] and [64°, 85°], which are divided into 577 classes corresponding to the situation of no fault and 576 different combinations of strikes and dips. We construct a 7-layer CNN to classify the fault strike and dip in a local seismic cube and obtain the classification probability at the same time. With the fault probability, strike and dip estimated at some seismic pixel, we further compute a fault cube (centered at the pixel) with fault features elongated along the fault plane. By sliding the classification window within a full seismic image, we are able to obtain a lot of overlapping fault cubes which are stacked to compute three full images of enhanced and continuous fault probabilities, strikes, and dips. To train the CNN model, we propose an effective and efficient workflow to automatically create 900 000 synthetic seismic cubes and the corresponding fault class labels. Although trained with only synthetic data sets, our CNN model can be applied to accurately estimate fault probabilities, strikes, and dips within field seismic images that are acquired at totally different surveys. With the estimated three fault images, we further construct fault cells that are represented as small 3-D squares, each square is colored by fault probability and oriented by fault strike and dip. We recursively link the fault cells by following the fault strikes and dips to finally construct fault skins, which are simple linked data structures to represent fault surfaces.
Xinming Wu, Yunzhi Shi, Sergey Fomel, Luming Liang, Qie Zhang, Anar Z. Yusifov
IEEE Trans. Geosci. Remote. Sens.4
2018 Corrections to "Image Interpolation by Blending Kernels"
abstract
Presents corrections to the paper, “Image interpolation by blending kernels,” (Liang, L.), IEEE Signal Process. Lett., vol. 15, pp. 805–808, 2008.
Luming Liang, Zeze Zhang
IEEE Signal Process. Lett.1
2017 Tensor Voting Guided Mesh Denoising
abstract
Mesh denoising is imperative for improving imperfect surfaces acquired by scanning devices. The main challenge is to faithfully retain geometric features and avoid introducing additional artifacts when removing noise. Unlike the existing mesh denoising techniques that focus only on either the first-order features or high-order differential properties, our approach exploits the synergy when facet normals and quadric surfaces are integrated to recover a piecewise smooth surface. In specific, we vote on surface normal tensors from robust statistics to guide the creation of consistent subneighborhoods subsequently used by moving least squares (MLS). This voting naturally leads to a conceptually simple way that gives a unified mesh-denoising framework for not only handling noise but also enabling the recovering of surfaces with both sharp and small-scale features. The effectiveness of our framework stems from: 1) the multiscale tensor voting that avoids the influence from noise; 2) the effective energy minimization strategy to searching the consistent subneighborhoods; and 3) the piecewise MLS that fully prevents the side effects from different subneighborhoods during surface fitting. Our framework is direct, practical, and easy to understand. Comparisons with the state-of-the-art methods demonstrate its outstanding performance on feature preservation and artifact suppression.
Mingqiang Wei, Luming Liang, Wai-Man Pang, Jun Wang 0039, Huisi Wu
IEEE Trans Autom. Sci. Eng.2
2016 3D Pose Tracking With Multitemplate Warping and SIFT Correspondences
abstract
Template warping is a popular technique in vision-based 3D motion tracking and 3D pose estimation due to its flexibility of being applicable to monocular video sequences. However, the method suffers from two major limitations that hamper its successful use in practice. First, it requires the camera to be calibrated prior to applying the method. Second, it may fail to provide good results if the inter-frame displacements are too large. To overcome the first problem, we propose to estimate the unknown focal length of the camera from several initial frames by an iterative optimization process. To alleviate the second problem, we propose a tracking method based on combining complementary information provided by dense optical flow and tracked scale-invariant feature transform (SIFT) features. While optical flow is good for small displacements and provides accurate local information, tracked SIFT features are better at handling larger displacements or global transformations. To combine these two pieces of complementary information, we introduce a forgetting factor to bootstrap the 3D pose estimates provided by SIFT features, and refine the final results using optical flow. Experiments are performed on three public databases, i.e., the Biwi Head Pose dataset, the BU dataset, and the McGill Faces datasets. The results illustrate that the proposed solution provides more accurate results than baseline methods that rely solely on either template warping or SIFT features. In addition, the approach can be applied in a larger variety of scenarios, due to circumventing the need for camera calibration, thus providing a more flexible solution to the problem than existing methods.
Luming Liang, Wenzhang Liang, Hassan Foroosh
IEEE Trans. Circuits Syst. Video Technol.2
2016 Spin Contour
abstract
Spin image is a powerful shape descriptor, useful in a point set or surface registration. However, the usage of spin images is hampered by issues such as sensitivity to noise and sampling rate and time-consuming matching process. We propose a novel spin-image-based local surface descriptor named spin contour to alleviate these problems. This descriptor is not an image but a 2-D point set. Comparisons show that the spin contour is robust to noise and sampling differences. The matching time is also improved over spin images.
Luming Liang, Mingqiang Wei, Andrzej Szymczak, Wai-Man Pang, Meng Wang 0001
IEEE Trans. Multim.1
2015 Geodesic spin contour for partial near-isometric matching
Luming Liang, Andrzej Szymczak, Mingqiang Wei
Comput. Graph.1
2009 Curvature normal vector driven interpolatory subdivision
abstract
We present an intrinsically nonlinear interpolatory subdivision scheme with geometric information and some free parameters via discrete curvatures normal vector. Our scheme can produce fair G1-continuous curves, which can avoid the potential pitfalls and unacceptable cases appeared in the four-point subdivision scheme. Furthermore, with the proper parameter choice, the proposed scheme is convexity-preserving, and reproduces the conic curve. Finally, the experimental results show our scheme is effective.
Huanxi Zhao, Xia Qiu, Luming Liang, Beiji Zou 0001
Shape Modeling International3
2008 A note on the paper "Normal based subdivision scheme for curve design" by Xunnian Yang
Luming Liang, Huanxi Zhao, Beiji Zou 0001
Comput. Aided Geom. Des.1
2008 Image Interpolation by Blending Kernels
abstract
A new convolution-based image interpolation method is presented, whose kernel function is designed via blending some well-known kernels. The new kernel is a better approximation to sinc function both in the space domain and the frequency domain. Comparative experiments with several polynomial spline type algorithms indicate that our approach exhibits a significant improvement in image quality.
Luming Liang
IEEE Signal Process. Lett.1