EDBT 2026 Demo / reviewers in the wild / expert
Shao-Ping Lu
dblp:17/9088
· DBLP profile ↗
38ranked-venue papers
7as first author
17since 2021 · last 2025
0000-0002-8492-0925ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 34 · 6 first-author · 16 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Effective Wavelet Neural Network for Ultra-High-Definition Image Deraining
Zhuoran Zheng, Shao-Ping Lu, Yulu Yang |
CGI (3) | 3 |
| 2024 | An Aesthetic-Guided Multimodal Framework for Video SummarizationabstractVideo summarization aims to generate a concise yet informative version of a lengthy video for efficient viewing. Generally, humans can discern important shots using audiovisual information and aesthetic elements of the video. However, existing methods often rely solely on unimodal information or ignore aesthetic elements, resulting in sub-optimal summaries. To tackle these issues, we present an aesthetic-guided multimodal framework (AMF) for video summarization. Specifically, we design a visual-aesthetics encoder to extract diverse aesthetic elements, and jointly integrate with visual content to create a comprehensive feature representation. Furthermore, we construct a multimodal fusion module to leverage the complementary properties among modalities, focusing on the most important content for the desired summary. Extensive experiments on four popular datasets show that the proposed method outperforms various compared methods in commonly used evaluation metrics. Jiehang Xie, Xuanbai Chen, Shao-Ping Lu |
ICME | 3 |
| 2024 | Video summarization via knowledge-aware multimodal deep networks
Jiehang Xie, Xuanbai Chen, Sicheng Zhao, Shao-Ping Lu |
Knowl. Based Syst. | 4 |
| 2024 | A transferability-aware covariance alignment network for image steganalysis
Jiao Liu 0003, Shao-Ping Lu, Yulu Yang |
Multim. Tools Appl. | 2 |
| 2023 | Self-Supervised Implicit 3D Reconstruction via RGB-D ScansabstractRecently, 3D reconstruction methods based on the neural radiance fields have demonstrated remarkable generative performance. However, these methods frequently tend to be resource hungry and are challenging to regulate large low-textured regions in typical indoor scenes. In this work, we analyze and integrate inherent semantic geometry cues for self-supervised 3D reconstruction training via a unified framework of volume rendering and signed distance implicit representations. In contrast to previous neural implicit methods, we simultaneously incorporate the pixel-aligned features and image patches for multi-view consistency, thereby enabling us to depict a large indoor scene from challenging scenarios with rich visual details and large smooth backgrounds. Extensive experiments and comparisons demonstrate that our proposed method has achieved state-of-the-art results by a large margin in various tasks (e.g. actual surface reconstruction, novel view synthesis, and learning a universal scheme in occlusion or distorted regions). Jiao Liu 0003, Shao-Ping Lu, Bo Ren 0003 |
ICME | 3 |
| 2023 | Multimodal-Based and Aesthetic-Guided Narrative Video SummarizationabstractNarrative videos usually illustrate the main content through multiple narrative information such as audios, video frames and subtitles. Existing video summarization approaches rarely consider the multiple dimensional narrative inputs, or ignore the impact of shots artistic assembly when directly applied to narrative videos. This paper introduces a multimodal-based and aesthetic-guided narrative video summarization method. Our method leverages multimodal information including visual content, subtitles and audio information through our specified key shots selection, subtitle summarization, and highlight extraction components. Furthermore, under the guidance of cinematographic aesthetic, we design a novel shots assembly module to ensure the shot content completeness and then assemble the selected shots into a desired summary. Besides, our method also provides the flexible specification for shots selection, to achieve which it automatically selects semantically related shots according to the user-designed text. By conducting a large number of quantitative experimental evaluations and user studies, we demonstrate that our method effectively preserves important narrative information of the original video, and it is capable of rapidly producing high-quality and aesthetic-guided narrative video summaries. Jiehang Xie, Xuanbai Chen, Tianyi Zhang 0013, Shao-Ping Lu, Pablo César, Yulu Yang |
IEEE Trans. Multim. | 5 |
| 2022 | Multi-Mode Interactive Image SegmentationabstractLarge-scale pixel-level annotations are scarce for current data-hungry medical image analysis models. For the fast acquisition of annotations, an economical and efficient interactive medical image segmentation method is urgently needed. However, current techniques usually fail in many cases, as their interaction styles cannot work on various inherent ambiguities of medical images, such as irregular shapes and fuzzy boundaries. To address this problem, we propose a multi-mode interactive segmentation framework for medical images, where diverse interaction modes can be chosen and allowed to cooperate with each other. In our framework, users can encircle the target regions with various initial interaction modes according to the structural complexity. Then, based on the initial segmentation, users can jointly utilize the region and boundary interactions to refine the mislabeled regions caused by different ambiguities. We evaluate our framework on extensive medical images, including X-ray, CT, MRI, ultrasound, endoscopy, and photo. Sufficient experimental results and user study show that our framework is a reliable choice for image annotation in various real scenes. Zheng Lin 0005, Zhao Zhang 0018, Linghao Han, Shao-Ping Lu |
ACM Multimedia | 4 |
| 2022 | A Knowledge Augmented and Multimodal-Based Framework for Video SummarizationabstractVideo summarization aims to generate a compact version of a lengthy video that retains its primary content. In general, humans are gifted with producing a high-quality video summary, because they acquire crucial content through multiple dimensional information and own abundant background knowledge about the original video. However, existing methods rarely consider multichannel information and ignore the impact of external knowledge, resulting in the limited quality of the generated summaries. This paper proposes a knowledge augmented and multimodal-based video summarization method, termed KAMV, to address the problem above. Specifically, we design a knowledge encoder with a hybrid method consisting of generation and retrieval, to capture descriptive content and latent connections between events and entities based on the external knowledge base, which can provide rich implicit knowledge for better comprehending the video viewed. Furthermore, for the sake of exploring the interactions among visual, audio, implicit knowledge and emphasizing the content that is most relevant to the desired summary, we present a fusion module under the supervision of these multimodal information. By conducting extensive experiments on four public datasets, the results demonstrate the superior performance yielded by the proposed KAMV compared to the state-of-the-art video summarization approaches. Jiehang Xie, Xuanbai Chen, Shao-Ping Lu, Yulu Yang |
ACM Multimedia | 3 |
| 2022 | Towards natural object-based image recoloringabstractExisting color editing algorithms enable users to edit the colors in an image according to their own aesthetics. Unlike artists who have an accurate grasp of color, ordinary users are inexperienced in color selection and matching, and allowing non-professional users to edit colors arbitrarily may lead to unrealistic editing results. To address this issue, we introduce a palette-based approach for realistic object-level image recoloring. Our data-driven approach consists of an offline learning part that learns the color distributions for different objects in the real world, and an online recoloring part that first recognizes the object category, and then recommends appropriate realistic candidate colors learned in the offline step for that category. We also provide an intuitive user interface for efficient color manipulation. After color selection, image matting is performed to ensure smoothness of the object boundary. Comprehensive evaluation on various color editing examples demonstrates that our approach outperforms existing state-of-the-art color editing algorithms. Mengyao Cui 0001, Zhe Zhu, Yulu Yang, Shao-Ping Lu |
Comput. Vis. Media | 4 |
| 2021 | News Content Completion with Location-Aware Image Selection
Zhengkun Zhang, Jun Wang 0023, Adam Jatowt, Zhe Sun 0009, Shao-Ping Lu, Zhenglu Yang |
AAAI | 5 |
| 2021 | DOTS: Decoupling Operation and Topology in Differentiable Architecture SearchabstractDifferentiable Architecture Search (DARTS) has attracted extensive attention due to its efficiency in searching for cell structures. DARTS mainly focuses on the operation search and derives the cell topology from the operation weights. However, the operation weights can not indicate the importance of cell topology and result in poor topology rating correctness. To tackle this, we propose to Decouple the Operation and Topology Search (DOTS), which decouples the topology representation from operation weights and makes an explicit topology search. DOTS is achieved by introducing a topology search space that contains combinations of candidate edges. The proposed search space directly reflects the search objective and can be easily extended to support a flexible number of edges in the searched cell. Existing gradient-based NAS methods can be incorporated into DOTS for further improvement by the topology search. Considering that some operations (e.g., Skip-Connection) can affect the topology, we propose a group operation search scheme to preserve topology-related operations for a better topology search. The experiments on CI-FAR10/100 and ImageNet demonstrate that DOTS is an effective solution for differentiable NAS. The code is released at https://github.com/guyuchao/DOTS. Yuchao Gu, Yun Liu 0011, Yi Yang 0001, Yu-Huan Wu, Shao-Ping Lu, Ming-Ming Cheng |
CVPR | 6 |
| 2021 | Large-Capacity Image Steganography Based on Invertible Neural NetworksabstractMany attempts have been made to hide information in images, where one main challenge is how to increase the payload capacity without the container image being detected as containing a message. In this paper, we propose a large-capacity Invertible Steganography Network (ISN) for image steganography. We take steganography and the recovery of hidden images as a pair of inverse problems on image domain transformation, and then introduce the forward and backward propagation operations of a single invertible network to leverage the image embedding and extracting problems. Sharing all parameters of our single ISN architecture enables us to efficiently generate both the container image and the revealed hidden image(s) with high quality. Moreover, in our architecture the capacity of image steganography is significantly improved by naturally increasing the number of channels of the hidden image branch. Comprehensive experiments demonstrate that with this significant improvement of the steganography payload capacity, our ISN achieves state-of-the-art in both visual and quantitative comparisons. Shao-Ping Lu, Paul L. Rosin |
CVPR | 1 |
| 2021 | iNAS: Integral NAS for Device-Aware Salient Object DetectionabstractExisting salient object detection (SOD) models usually focus on either backbone feature extractors or saliency heads, ignoring their relations. A powerful backbone could still achieve sub-optimal performance with a weak saliency head and vice versa. Moreover, the balance between model performance and inference latency poses a great challenge to model design, especially when considering different deployment scenarios. Considering all components in an integral neural architecture search (iNAS) space, we propose a flexible device-aware search scheme that only trains the SOD model once and quickly finds high-performance but low-latency models on multiple devices. An evolution search with latency-group sampling (LGS) is proposed to explore the entire latency area of our enlarged search space. Models searched by iNAS achieve similar performance with SOTA methods but reduce the 3.8×, 3.3×, 2.6×, 1.9× latency on Huawei Nova6 SE, Intel Core CPU, the Jetson Nano, and Nvidia Titan Xp. The code is released at https://mmcheng.net/inas/. Yuchao Gu, Shanghua Gao, Shao-Ping Lu, Ming-Ming Cheng |
ICCV | 5 |
| 2021 | Deep Symmetric Network for Underexposed Image Enhancement with Recurrent Attentional LearningabstractUnderexposed image enhancement is of importance in many research domains. In this paper, we take this problem as image feature transformation between the underexposed image and its paired enhanced version, and we propose a deep symmetric network for the issue. Our symmetric network adapts invertible neural networks (INN) for bidirectional feature learning between images, and to ensure the mutual propagation invertible we specifically construct two pairs of encoder-decoder with the same pretrained parameters. This invertible mechanism with bidirectional feature transformations enable us to both avoid colour bias and recover the content effectively for image enhancement. In addition, we propose a new recurrent residual-attention module (RRAM), where the recurrent learning network is designed to gradually perform the desired colour adjustments. Ablation experiments are executed to show the role of each component of our new architecture. We conduct a large number of experiments on two datasets to demonstrate that our method achieves the state-of-the-art effect in underexposed image enhancement. Code is available at https://www.shaopinglu.net/proj-iccv21/ImageEnhancement.html. Shao-Ping Lu, Tao Chen 0015, Zhenglu Yang, Ariel Shamir |
ICCV | 2 |
| 2021 | Low-Rank Constrained Super-Resolution for Mixed-Resolution Multiview VideoabstractMultiview video allows for simultaneously presenting dynamic imaging from multiple viewpoints, enabling a broad range of immersive applications. This paper proposes a novel super-resolution (SR) approach to mixed-resolution (MR) multiview video, whereby the low-resolution (LR) videos produced by MR camera setups are up-sampled based on the neighboring HR videos. Our solution analyzes the statistical correlation of different resolutions between multiple views, and introduces a low-rank prior based SR optimization framework using local linear embedding and weighted nuclear norm minimization. The target HR patch is reconstructed by learning texture details from the neighboring HR camera views using local linear embedding. A low-rank constrained patch optimization solution is introduced to effectively restrain visual artifacts and the ADMM framework is used to solve the resulting optimization problem. Comprehensive experiments including objective and subjective test metrics demonstrate that the proposed method outperforms the state-of-the-art SR methods for MR multiview video. Shao-Ping Lu, Senmao Li, Gauthier Lafruit, Ming-Ming Cheng, Adrian Munteanu 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Bilateral Attention Network for RGB-D Salient Object DetectionabstractRGB-D salient object detection (SOD) aims to segment the most attractive objects in a pair of cross-modal RGB and depth images. Currently, most existing RGB-D SOD methods focus on the foreground region when utilizing the depth images. However, the background also provides important information in traditional SOD methods for promising performance. To better explore salient information in both foreground and background regions, this paper proposes a Bilateral Attention Network (BiANet) for the RGB-D SOD task. Specifically, we introduce a Bilateral Attention Module (BAM) with a complementary attention mechanism: foreground-first (FF) attention and background-first (BF) attention. The FF attention focuses on the foreground region with a gradual refinement style, while the BF one recovers potentially useful salient information in the background region. Benefited from the proposed BAM module, our BiANet can capture more meaningful foreground and background cues, and shift more attention to refining the uncertain details between foreground and background regions. Additionally, we extend our BAM by leveraging the multi-scale techniques for better SOD performance. Extensive experiments on six benchmark datasets demonstrate that our BiANet outperforms other state-of-the-art RGB-D SOD methods in terms of objective metrics and subjective visual comparison. Our BiANet can run up to 80 fps on 224×224 RGB-D images, with an NVIDIA GeForce RTX 2080Ti GPU. Comprehensive ablation studies also validate our contributions. Zhao Zhang 0018, Zheng Lin 0005, Jun Xu 0019, Wenda Jin, Shao-Ping Lu, Deng-Ping Fan |
IEEE Trans. Image Process. | 5 |
| 2021 | Aesthetic-guided outward image croppingabstractImage cropping is a commonly used post-processing operation for adjusting the scene composition of an input photography, therefore improving its aesthetics. Existing automatic image cropping methods are all bounded by the image border, thus have very limited freedom for aesthetics improvement if the original scene composition is far from ideal, e.g. the main object is too close to the image border. In this paper, we propose a novel, aesthetic-guided outward image cropping method. It can go beyond the image border to create a desirable composition that is unachievable using previous cropping methods. Our method first evaluates the input image to determine how much the content of the image should be extrapolated by a field of view (FOV) evaluation model. We then synthesize the image content in the extrapolated region, and seek an optimal aesthetic crop within the expanded FOV, by jointly considering the aesthetics of the cropped view, and the local image quality of the extrapolated image content. Experimental results show that our method can generate more visually pleasing image composition in cases that are difficult for previous image cropping tools due to the border constraint, and can also automatically degrade to an inward method when high quality image extrapolation is infeasible. Feng-Heng Li, Hao-Zhi Huang 0001, Yong Zhang 0034, Shao-Ping Lu, Jue Wang 0001 |
ACM Trans. Graph. | 5 |
| 2020 | Pyramid Constrained Self-Attention Network for Fast Video Salient Object DetectionabstractSpatiotemporal information is essential for video salient object detection (VSOD) due to the highly attractive object motion for human's attention. Previous VSOD methods usually use Long Short-Term Memory (LSTM) or 3D ConvNet (C3D), which can only encode motion information through step-by-step propagation in the temporal domain. Recently, the non-local mechanism is proposed to capture long-range dependencies directly. However, it is not straightforward to apply the non-local mechanism into VSOD, because i) it fails to capture motion cues and tends to learn motion-independent global contexts; ii) its computation and memory costs are prohibitive for video dense prediction tasks such as VSOD. To address the above problems, we design a Constrained Self-Attention (CSA) operation to capture motion cues, based on the prior that objects always move in a continuous trajectory. We group a set of CSA operations in Pyramid structures (PCSA) to capture objects at various scales and speeds. Extensive experimental results demonstrate that our method outperforms previous state-of-the-art methods in both accuracy and speed (110 FPS on a single Titan Xp) on five challenge datasets. Our code is available at https://github.com/guyuchao/PyramidCSA. Yuchao Gu, Ziqin Wang, Yun Liu 0011, Ming-Ming Cheng, Shao-Ping Lu |
AAAI | 6 |
| 2020 | Interactive Image Segmentation With First Click AttentionabstractIn the task of interactive image segmentation, users initially click one point to segment the main body of the target object and then provide more points on mislabeled regions iteratively for a precise segmentation. Existing methods treat all interaction points indiscriminately, ignoring the difference between the first click and the remaining ones. In this paper, we demonstrate the critical role of the first click about providing the location and main body information of the target object. A deep framework, named First Click Attention Network (FCA-Net), is proposed to make better use of the first click. In this network, the interactive segmentation result can be much improved with the following benefits: focus invariance, location guidance, and error-tolerant ability. We then put forward a click-based loss function and a structural integrity strategy for better segmentation effect. The visualized segmentation results and sufficient experiments on five datasets demonstrate the importance of the first click and the superiority of our FCA-Net. Zheng Lin 0005, Zhao Zhang 0018, Lin-Zhuo Chen, Ming-Ming Cheng, Shao-Ping Lu |
CVPR | 5 |
| 2020 | 3D computational modeling and perceptual analysis of kinetic depth effectsabstractHumans have the ability to perceive kinetic depth effects , i.e., to perceived 3D shapes from 2D projections of rotating 3D objects. This process is based on a variety of visual cues such as lighting and shading effects. However, when such cues are weak or missing, perception can become faulty, as demonstrated by the famous silhouette illusion example of the spinning dancer . Inspired by this, we establish objective and subjective evaluation models of rotated 3D objects by taking their projected 2D images as input. We investigate five different cues: ambient luminance, shading, rotation speed, perspective, and color difference between the objects and background. In the objective evaluation model, we first apply 3D reconstruction algorithms to obtain an objective reconstruction quality metric, and then use quadratic stepwise regression analysis to determine weights of depth cues to represent the reconstruction quality. In the subjective evaluation model, we use a comprehensive user study to reveal correlations with reaction time and accuracy, rotation speed, and perspective. The two evaluation models are generally consistent, and potentially of benefit to inter-disciplinary research into visual perception and 3D reconstruction. Mengyao Cui 0001, Shao-Ping Lu, Miao Wang 0004, Yongliang Yang 0002, Yukun Lai, Paul L. Rosin |
Comput. Vis. Media | 2 |
| 2020 | Structure-Preserving Neural Style TransferabstractState-of-the-art neural style transfer methods have demonstrated amazing results by training feed-forward convolutional neural networks or using an iterative optimization strategy. The image representation used in these methods, which contains two components: style representation and content representation, is typically based on high-level features extracted from pretrained classification networks. Because the classification networks are originally designed for object recognition, the extracted features often focus on the central object and neglect other details. As a result, the style textures tend to scatter over the stylized outputs and disrupt the content structures. To address this issue, we present a novel image stylization method that involves an additional structure representation. Our structure representation, which considers two factors: i) the global structure represented by the depth map and ii) the local structure details represented by the image edges, effectively reflects the spatial distribution of all the components in an image as well as the structure of dominant objects respectively. Experimental results demonstrate that our method achieves an impressive visual effectiveness, which is particularly significant when processing images sensitive to structure distortion, e.g. images containing multiple objects potentially at different depths, or dominant objects with clear structures. Ming-Ming Cheng, Xiao-Chang Liu, Shao-Ping Lu, Yukun Lai, Paul L. Rosin |
IEEE Trans. Image Process. | 4 |
| 2019 | Geometry-Aware ICP for Scene Reconstruction from RGB-D Camera
Bo Ren 0003, Jiacheng Wu 0001, Ya-Lei Lv, Ming-Ming Cheng, Shao-Ping Lu |
J. Comput. Sci. Technol. | 5 |
| 2019 | Consistent video projection on curved displays
Yangxintong Lyu, Shao-Ping Lu, Quentin Bolsee, Adrian Munteanu 0001 |
Signal Process. Image Commun. | 2 |
| 2019 | Dictionary Learning-Based, Directional, and Optimized Prediction for Lenslet Image CodingabstractIn this paper, a novel approach to encode lenslet (LL) images is proposed. The method departs from traditional block-based coding structures and employs a hexagonal-shaped pixel cluster, called macro-pixel, as an elementary coding unit. A novel prediction mode based on dictionary learning is proposed, whereby macro-pixels are represented by a sparse linear combination of atoms from a generic dictionary. Additionally, an optimized linear prediction mode and a directional prediction mode specifically designed for macro-pixels are proposed. Rate-distortion optimization is utilized to select the best intra prediction mode for each macro-pixel. Experimental results on the light field image data set show that the proposed coding system outperforms HEVC and the state-of-the-art in LL image coding with an average peak signal to noise ratio gain of 3.33 and 1.41 dB, respectively, and with rate savings of 67.13% and 34.30%, respectively. Rui Zhong 0005, Ionut Schiopu, Bruno Cornelis, Shao-Ping Lu, Junsong Yuan 0001, Adrian Munteanu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Deep Online Video Stabilization With Multi-Grid Warping Transformation LearningabstractVideo stabilization techniques are essential for most hand-held captured videos due to high-frequency shakes. Several 2D, 2.5D and 3D-based stabilization techniques have been presented previously, but to our knowledge, no solutions based on deep neural networks had been proposed to date. The main reason for this omission is shortage in training data as well as the challenge of modeling the problem using neural networks. In this paper, we present a video stabilization technique using a convolutional neural network. Previous works usually propose an offline algorithm that smoothes a holistic camera path based on feature matching. Instead, we focus on low-latency, real-time camera path smoothing, that does not explicitly represent the camera path, and does not use future frames. Our neural network model, called StabNet, learns a set of mesh-grid transformations progressively for each input frame from the previous set of stabalized camera frames, and creates stable corresponding latent camera paths implicitly. To train the network, we collect a dataset of synchronized steady and unsteady video pairs via a specially designed hand-held hardware. Experimental results show that our proposed online method performs comparatively to traditional offline video stabilization methods without using future frames, while running about 10× faster. More importantly, our proposed StabNet is able to handle low-quality videos such as night-scene videos, watermarked videos, blurry videos and noisy videos, where existing methods fail in feature extraction or matching. Miao Wang 0004, Guo-Ye Yang, Jin-Kun Lin, Song-Hai Zhang, Ariel Shamir, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | Synthesis of Shaking Video Using Motion Capture Data and Dynamic 3D Scene ModelingabstractImportant video processing methods such as video stabilization and deblurring often do not have ground-truth data available. This poses a great challenge in the development and parameter tunning of such methods. Synthetic shaken video is very useful to generate well-defined ground-truth datasets. Existing shaking video synthesis methods simulate shaky camera motion by performing 2D view warping using only a single 2D video, which does not always correspond to realistic 3D motions. In this paper, we introduce a novel shaking video synthesis approach. The proposed framework constructs the camera motion trajectory by making use of human motion information that is captured in the real-world. Moreover, we render the shaken video from man-made dynamic 3D scenes with detailed camera pose information. Our novel approach provides both accurate 2D visual content and camera motion trajectory in the 3D scene, which allows for evaluating the visual distortion as well as the offsets of the recovered camera trajectory. The proposed synthesis method of shaking video will benefit and ease future research on 3D-aware video stabilization. Shao-Ping Lu, Beerend Ceulemans, Miao Wang 0004, Adrian Munteanu 0001 |
ICIP | 1 |
| 2018 | Structured Skip List: A Compact Data Structure for 3D ReconstructionabstractThe model produced by 3D reconstruction algorithm is usually represented by voxels. The management of these voxels is usually divided into two categories: ordered and unordered methods. The ordered method holds too many empty voxels to maintain data order which leads to a low storage efficiency. On the contrary, the unordered method keeps massive index data to only store nonempty voxels. In this paper, we design a new data management method for real-time indoor 3D reconstruction, called Structured Skip List (SSL). The SSL can be treated as a semi-ordered method, because the advantages of both the ordered and unordered methods are taken into account: 1) it only holds nonempty voxels similar to the unordered method; 2) the structured information is introduced to reduce the storage space of index data. By these designs, the SSL has a better performance on storage efficiency. To handle the data collision in voxel allocation, a hash allocation list (HAL) is proposed. The length of each Skip List is kept balanced by fusing the IMU (Inertial Measurement Unit) information for a high operation efficiency. The storage efficiency analysis of different data management methods is shown in this paper. What's more, exhaustive investigation is carried out on several datasets with these methods. The experimental result demonstrates that our design can achieve a high storage efficiency with little time loss compared to the state-of-the-art methods. Shijie Li 0006, Ming-Ming Cheng, Yun Liu 0011, Shao-Ping Lu, Victor Adrian Prisacariu |
IROS | 4 |
| 2018 | Historical Context-based Style Classification of Painting Images via Label Distribution LearningabstractAnalyzing and categorizing the style of visual art images, especially paintings, is gaining popularity owing to its importance in understanding and appreciating the art. The evolution of painting style is both continuous, in a sense that new styles may inherit, develop or even mutate from their predecessors and multi-modal because of various issues such as the visual appearance, the birthplace, the origin time and the art movement. Motivated by this peculiarity, we introduce a novel knowledge distilling strategy to assist visual feature learning in the convolutional neural network for painting style classification. More specifically, a multi-factor distribution is employed as soft-labels to distill complementary information with visual input, which extracts from different historical context via label distribution learning. The proposed method is well-encapsulated in a multi-task learning framework which allows end-to-end training. We demonstrate the superiority of the proposed method over the state-of-the-art approaches on Painting91, OilPainting, and Pandora datasets. Jufeng Yang, Liyi Chen 0003, Le Zhang 0001, Xiaoxiao Sun 0002, Dongyu She, Shao-Ping Lu, Ming-Ming Cheng |
ACM Multimedia | 6 |
| 2018 | Hyper-Lapse From Multiple Spatially-Overlapping VideosabstractHyper-lapse video with high speed-up rate is an efficient way to overview long videos, such as a human activity in first-person view. Existing hyper-lapse video creation methods produce a fast-forward video effect using only one video source. In this paper, we present a novel hyper-lapse video creation approach based on multiple spatially-overlapping videos. We assume the videos share a common view or location, and find transition points where jumps from one video to another may occur. We represent the collection of videos using a hyper-lapse transition graph; the edges between nodes represent possible hyper-lapse frame transitions. To create a hyper-lapse video, a shortest path search is performed on this digraph to optimize frame sampling and assembly simultaneously. Finally, we render the hyper-lapse results using video stabilization and appearance smoothing techniques on the selected frames. Our technique can synthesize novel virtual hyper-lapse routes, which may not exist originally. We show various application results on both indoor and outdoor video collections with static scenes, moving objects, and crowds. Miao Wang 0004, Jun-Bang Liang, Song-Hai Zhang, Shao-Ping Lu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | BiggerSelfie: Selfie Video Expansion With Hand-Held CameraabstractSelfie photography from the hand-held camera is becoming a popular media type. Although being convenient and flexible, it suffers from low camera motion stability, small field of view, and limited background content. These limitations can annoy users, especially, when touring a place of interest and taking selfie videos. In this paper, we present a novel method to create what we call a BiggerSelfie that deals with these shortcomings. Using a video of the environment that has partial content overlap with the selfie video, we stitch plausible frames selected from the environment video to the original selfie frames and stabilize the composed video content with a portrait-preserving constraint. Using the proposed method, one can easily obtain a stable selfie video with expanded background content by merely capturing some background shots. We show various results and several evaluations to demonstrate the applicability of our method. Miao Wang 0004, Ariel Shamir, Guo-Ye Yang, Jin-Kun Lin, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Robust Multiview Synthesis for Wide-Baseline Camera ArraysabstractIn many advanced multimedia systems, multiview content can offer more immersion compared to classical stereoscopy. The feeling of immersiveness is increased substantially by offering motion-parallax, as well as stereopsis. This drives both the so-called free-navigation and super-multiview technologies. However, it is currently still challenging to acquire, store, process, and transmit this type of content. This paper presents a novel multiview-interpolation framework for wide-baseline camera arrays. The proposed method comprises several novel components, including point cloud-based filtering, improved de-ghosting, multireference color blending, and depth-aware MRF-based disocclusion in painting. The method offers robustness against depth errors caused by quantization and smoothing across object boundaries. Furthermore, the available input color and depth are maximally exploited while preventing propagation of unreliable information to virtual viewpoints. The experimental results show that the proposed method outperforms the state-of-the-art View Synthesis Reference Software (VSRS 4.1) both in objective terms as well as subjectively, based on a visual assessment on a high-end light-field three-dimensional display. Beerend Ceulemans, Shao-Ping Lu, Gauthier Lafruit, Adrian Munteanu 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Color correction for large-baseline multiview videoabstractColor misalignment correction is an important, yet unsolved problem , especially for multiview video captured by large disparity camera setups. In this paper, we introduce a robust large-baseline color correction method that preserves the original manifold structure of the input video. The manifold structure is extracted by locally linear embedding (LLE), aimed at linearly representing each pixel based on its neighbors, assuming that they are all clustered in a high-dimensional feature space. Besides the proposed manifold structure preservation constraint, the proposed method enforces spatio-temporal color consistencies and gradient preservation. The multiview color correction solution is obtained by solving a global optimization problem . Thorough objective and subjective experimental results demonstrate that our proposed approach significantly and systematically outperforms the state-of-the-art color correction methods on large-baseline multiview video data . Siqi Ye, Shao-Ping Lu, Adrian Munteanu 0001 |
Signal Process. Image Commun. | 2 |
| 2017 | Wavelet-Based L∞ Semi-regular Mesh CodingabstractPolygonal meshes are popular three-dimensional virtual representations employed in a wide range of applications. Users have very high expectations with respect to the accuracy of these virtual representations, fueling a steady increase in the processing power and performance of graphics processing hardware. This accuracy is closely related to how detailed the virtual representations are. The more detailed these representations become, the higher the amount of data that will need to be displayed, stored, or transmitted. Efficient compression techniques are of critical importance in this context. State-ofthe-art compression performance of semi-regular mesh coding systems has been achieved through the use of subdivision-based wavelet coding techniques. However, the vast majority of these codecs are optimized with respect to the L2distortion metric, i.e., the average error. This makes them unsuitable for applications where each input signal sample has a certain significance. To alleviate this problem, we propose to optimize the mesh codec with respect to the L∞metric, which allows for the control of the local reconstruction error. This paper proposes novel data-dependent formulations for the L∞distortion. The proposed L∞estimators are incorporated in a state-of-the-art wavelet-based semi-regular mesh codec. The resulting coding system offers scalability in L∞sense. The experiments demonstrate the advantages of L∞coding in providing a tight control on the local reconstruction error. Furthermore, the proposed data-dependent L∞approaches significantly improve estimation accuracy, reducing the classical low-rate gap between the estimated and actual L∞distortion observed for previous L∞estimators. Ruxandra-Marina Florea, Adrian Munteanu 0001, Shao-Ping Lu, Peter Schelkens |
IEEE Trans. Multim. | 3 |
| 2016 | Efficient MRF-based disocclusion inpainting in multiview videoabstractView synthesis using depth image-based rendering generates virtual viewpoints of a 3D scene based on texture and depth information from a set of available cameras. One of the core components in view synthesis is image inpainting which performs the reconstruction of areas that were occluded in the available cameras but are visible from the virtual viewpoint. Inpainting methods based on Markov random fields (MRFs) have been shown to be very effective in inpainting large areas in images. In this paper, we propose a novel MRF-based in-painting method for multiview video. The proposed method steers the MRF optimization towards completion from background to foreground and exploits the available depth information in order to avoid bleeding artifacts. The proposed approach allows for efficiently filling-in large disocclusion areas and greatly accelerates execution compared to traditional MRF-based inpainting techniques. The experimental results show that view synthesis based on the proposed inpainting method systematically improves performance over the state-of-the-art in multiview view synthesis. Average PSNR gains up to 1.88 dB compared to the MPEG View Synthesis Reference software were observed. Beerend Ceulemans, Shao-Ping Lu, Gauthier Lafruit, Peter Schelkens, Adrian Munteanu 0001 |
ICME | 2 |
| 2015 | Color retargeting: Interactive time-varying color image composition from time-lapse sequencesabstractIn this paper, we present an interactive static image composition approach, namely color retargeting , to flexibly represent time-varying color editing effect based on time-lapse video sequences. Instead of performing precise image matting or blending techniques, our approach treats the color composition as a pixel-level resampling problem. In order to both satisfy the user’s editing requirements and avoid visual artifacts, we construct a globally optimized interpolation field. This field defines from which input video frames the output pixels should be resampled. Our proposed resampling solution ensures that (i) the global color transition in the output image is as smooth as possible, (ii) the desired colors/objects specified by the user from different video frames are well preserved, and (iii) additional local color transition directions in the image space assigned by the user are also satisfied. Various examples have been shown to demonstrate that our efficient solution enables the user to easily create time-varying color image composition results. Shao-Ping Lu, Guillaume Dauphin, Gauthier Lafruit, Adrian Munteanu 0001 |
Comput. Vis. Media | 1 |
| 2015 | Spatio-Temporally Consistent Color and Structure Optimization for Multiview Video Color CorrectionabstractWhen compared to conventional 2-D video, multiview video can significantly enhance the visual 3-D experience in 3-D applications by offering horizontal parallax. However, when processing images originating from different views, it is common that the colors between the different cameras are not well- calibrated . To solve this problem, a novel energy function -based color correction method for multiview camera setups is proposed to enforce that colors are as close as possible to those in the reference image but also that the overall structural information is well-preserved. The proposed system introduces a spatio-temporal correspondence matching method to ensure that each pixel in the input image gets bijectively mapped to a reference pixel. By combining this mapping with the original structural information, we construct a global optimization algorithm in a Laplacian matrix formulation and solve it using a sparse matrix solver. We further introduce a novel forward-reverse objective evaluation model to overcome the problem of lack of ground truth in this field. The visual comparisons are shown to outperform state-of-the-art multiview color correction methods, while the objective evaluation reports PSNR gains of up to 1.34 dB and SSIM gains of up to 3.2%, respectively. Shao-Ping Lu, Beerend Ceulemans, Adrian Munteanu 0001, Peter Schelkens |
IEEE Trans. Multim. | 1 |
| 2013 | Timeline Editing of Objects in VideoabstractWe present a video editing technique based on changing the timelines of individual objects in video, which leaves them in their original places but puts them at different times. This allows the production of object-level slow motion effects, fast motion effects, or even time reversal. This is more flexible than simply applying such effects to whole frames, as new relationships between objects can be created. As we restrict object interactions to the same spatial locations as in the original video, our approach can produce highquality results using only coarse matting of video objects. Coarse matting can be done efficiently using automatic video object segmentation, avoiding tedious manual matting. To design the output, the user interactively indicates the desired new life spans of objects, and may also change the overall running time of the video. Our method rearranges the timelines of objects in the video whilst applying appropriate object interaction constraints. We demonstrate that, while this editing technique is somewhat restrictive, it still allows many interesting results. Shao-Ping Lu, Song-Hai Zhang, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2011 | Saliency-Based Fidelity Adaptation Preprocessing for Video Coding
Shao-Ping Lu, Song-Hai Zhang |
J. Comput. Sci. Technol. | 1 |