EDBT 2026 Demo / reviewers in the wild / expert
Wenbin Li 0002
dblp:27/1736-2
· DBLP profile ↗
28ranked-venue papers
9as first author
8since 2021 · last 2026
0000-0002-5593-2599ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 5 first-author · 4 since 2021Systems, architecture and hardware · 7 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ProjectiveShading: Inserting 3D Objects into Indoor Images with Complex ShadowsabstractAbstract Realistically inserting virtual 3D objects into real‐world images requires perceptually coherent shadowing of the object and background scene. Achieving this in single‐view indoor scenes with sunlight is challenging due to complex, partially visible occluders and indirect lighting. Environment maps alone cannot produce realistic shadows on virtual objects, and any representation (scene parameters) used for rendering must be practically estimable. We introduce ProjectiveShading, the first automatic method for inverse‐ and re‐rendering that handles bi‐directional shadow interactions for realistic object composition. Our key innovation is the sunlight map, a 2D image encoding direct sunlight and arbitrary occlusions. It is generated from single‐view estimations using off‐the‐shelf models and is compatible with standard rendering engines. We also propose algorithms to estimate sunlight direction and to blend virtual and real shadows while preserving background textures. Experiments on synthetic and in‐the‐wild images show our method outperforms previous approaches. Jundan Luo, Nanxuan Zhao, Lu Wang 0007, Wenbin Li 0002, Christian Richardt |
Comput. Graph. Forum | 5 |
| 2024 | Camera Chameleon - The Creative Impact of Tracked Tangible Interfaces for Virtual Film Pre-ProductionabstractShows such as The Mandalorian have used Virtual Production (VP) to great creative advantage. This combination, of tracked cameras and virtual environments, gives film crews the flexibility to rapidly iterate ideas and discover the best telling of their story within the time constraints of movie production. We explore if the same ideas can translate to pre-production, specifically if similar advantages can be found during the storyboarding phase. Developing a novel virtual production interface named Camera Chameleon, we compare user interaction against traditional keyboard/mouse peripherals when creating storyboards. Similar expert-rated quality is measured on storyboard output, however, users report increased enjoyment and complete shots more quickly. Novice users explore more when using our approach, demonstrating its creative potential. Will Kerr, Crescent Jicol, Tom S. F. Haines, Wenbin Li 0002 |
ICME | 4 |
| 2024 | Reinforcement Learning of Dolly-In Filming Using a Ground-Based RobotabstractFree-roaming dollies enhance filmmaking with dynamic movement, but challenges in automated camera control remain unresolved. Our study advances this field by applying Reinforcement Learning (RL) to automate dolly-in shots using free-roaming ground-based filming robots, overcoming traditional control hurdles. We demonstrate the effectiveness of combined control for precise film tasks by comparing it to independent control strategies. Our robust RL pipeline surpasses traditional Proportional-Derivative controller performance in simulation and proves its efficacy in real-world tests on a modified ROSBot 2.0 platform equipped with a camera turret. This validates our approach’s practicality and sets the stage for further research in complex filming scenarios, contributing significantly to the fusion of technology with cinematic creativity. This work presents a leap forward in the field and opens new avenues for research and development, effectively bridging the gap between technological advancement and creative filmmaking. Philip Lorimer, Jack Saunders, Alan Hunter 0001, Wenbin Li 0002 |
IROS | 4 |
| 2024 | Identifying Optimal Launch Sites of High-Altitude Latex-Balloons using Bayesian Optimisation for the Task of Station-KeepingabstractStation-keeping tasks for high-altitude balloons show promise in areas such as ecological surveys, atmospheric analysis, and communication relays. However, identifying the optimal time and position to launch a latex high-altitude balloon is still a challenging and multifaceted problem. For example, tasks such as forest fire tracking place geometric constraints on the launch location of the balloon. Furthermore, identifying the most optimal location also heavily depends on atmospheric conditions. We first illustrate how reinforcement learning-based controllers, frequently used for station-keeping tasks, can exploit the environment. This exploitation can degrade performance on unseen weather patterns and affect station-keeping performance when identifying an optimal launch configuration. Valuing all states equally in the region, the agent exploits the region’s geometry by flying near the edge, leading to risky behaviours. We propose a modification which compensates for this exploitation and finds this leads to, on average, higher steps within the target region on unseen data. Then, we illustrate how Bayesian Optimisation (BO) can identify the optimal launch location to perform station-keeping tasks, maximising the return from a given rollout. We show BO can find this launch location in fewer steps compared to other optimisation methods. Results indicate that, surprisingly, the most optimal location to launch from is not commonly within the target region. Please find further information about our project at https://sites.google.com/view/bo-lauch-balloon/. Jack Saunders, Sajad Saeedi G., Adam Hartshorne, Binbin Xu 0001, Özgür Simsek, Alan Hunter 0001, Wenbin Li 0002 |
IROS | 7 |
| 2024 | CRefNet: Learning Consistent Reflectance Estimation With a Decoder-Sharing TransformerabstractWe present CRefNet, a hybrid transformer-convolutional deep neural network for consistent reflectance estimation in intrinsic image decomposition. Estimating consistent reflectance is particularly challenging when the same material appears differently due to changes in illumination. Our method achieves enhanced global reflectance consistency via a novel transformer module that converts image features to reflectance features. At the same time, this module also exploits long-range data interactions. We introduce reflectance reconstruction as a novel auxiliary task that shares a common decoder with the reflectance estimation task, and which substantially improves the quality of reconstructed reflectance maps. Finally, we improve local reflectance consistency via a new rectified gradient filter that effectively suppresses small variations in predictions without any overhead at inference time. Our experiments show that our contributions enable CRefNet to predict highly consistent reflectance maps and to outperform the state of the art by 10% WHDR. Jundan Luo, Nanxuan Zhao, Wenbin Li 0002, Christian Richardt |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Parallel Reinforcement Learning Simulation for Visual Quadrotor NavigationabstractReinforcement learning (RL) is an agent-based approach for teaching robots to navigate within the physical world. Gathering data for RL is known to be a laborious task, and real-world experiments can be risky. Simulators facilitate the collection of training data in a quicker and more cost-effective manner. However, RL frequently requires a significant number of simulation steps for an agent to become skilful at simple tasks. This is a prevalent issue within the field of RL-based visual quadrotor navigation where state dimensions are typically very large and dynamic models are complex. Furthermore, rendering images and obtaining physical properties of the agent can be computationally expensive. To solve this, we present a simulation framework, built on AirSim, which provides efficient parallel training. Building on this framework, Ape-X is modified to incorporate parallel training of AirSim environments to make use of numerous networked computers. Through experiments we were able to achieve a reduction in training time from 3.9 hours to 11 minutes, for a toy problem, using the aforementioned framework and a total of 74 agents and two networked computers. Further details including a github repo and videos about our project, PRL4AirSim, can be found at https://sites.google.com/view/prl4airsim/home Jack Saunders, Sajad Saeedi G., Wenbin Li 0002 |
ICRA | 3 |
| 2023 | Resource-Constrained Station-Keeping for Latex Balloons Using Reinforcement LearningabstractHigh altitude balloons have proved useful for ecological aerial surveys, atmospheric monitoring, and communication relays. However, due to weight and power constraints, there is a need to investigate alternate modes of propulsion to navigate in the stratosphere. Very recently, reinforcement learning has been proposed as a control scheme to maintain balloons in the region of a fixed location, facilitated through diverse opposing wind-fields at different altitudes. Although air-pump based station keeping has been explored, there is no research on the control problem for venting and ballasting actuated balloons, which is commonly used as a low-cost alternative. We show how reinforcement learning can be used for this type of balloon. Specifically, we use the soft actor-critic algorithm, which on average is able to station-keep within 50 km for on average 25% of the flight, consistent with state-of-the-art. Furthermore, we show that the proposed controller effectively minimises the consumption of resources, thereby supporting long duration flights. We frame the controller as a continuous control reinforcement learning problem, which allows for a more diverse range of trajectories, as opposed to current state-of-the-art work, which uses discrete action spaces. Furthermore, through continuous control, we can make use of larger ascent rates which are not possible using air-pumps. The desired ascent-rate is decoupled into desired altitude and time-factor to provide a more transparent policy, compared to low-level control commands used in previous works. Finally, by applying the equations of motion, we establish appropriate thresholds for venting and ballasting to prevent the agent from exploiting the environment. More specifically, we ensure actions are physically feasible by enforcing constraints on venting and ballasting. Jack Saunders, Loïc Prenevost, Özgür Simsek, Alan Hunter 0001, Wenbin Li 0002 |
IROS | 5 |
| 2021 | Fast, High-Quality Hierarchical Depth-Map Super-ResolutionabstractThe low spatial resolution of acquired depth maps is a major drawback of most RGBD sensors. However, there are many scenarios in which fast acquisition of high-resolution and high-quality depth maps would be desirable. One approach to achieve higher quality depth maps is through super-resolution. However, edge preservation is challenging, and artifacts such as depth confusion and blurring are easily introduced near boundaries. In view of this, we propose a method for fast, high-quality hierarchical depth-map super-resolution (HDS). In our method, a high-resolution RGB image is degraded layer by layer to guide the bilateral filtering of the depth map. To improve the upsampled depth map quality, we construct a feature-based bilateral filter (FBF) for the interpolation, by using the extracted RGB shallow and multi-layer features. To accelerate the process, we perform filtering only near depth boundaries and through matrix operations. We also propose an extension of our HDS model to a Classification-based Hierarchical Depth-map Super-resolution (C-HDS) model, where a context-aware trilateral filter reduces the contributions of unreliable neighbors to the current missing depth location. Experimental results show that the proposed method is significantly faster than existing methods for generating high-resolution depth maps, while also significantly improving depth quality compared to the current state-of-the-art approaches, especially for large-scale 16x super-resolution. Yiguo Qiao, Licheng Jiao, Wenbin Li 0002, Christian Richardt, Darren Cosker |
ACM Multimedia | 3 |
| 2020 | RGBD-Dog: Predicting Canine Pose from RGBD SensorsabstractThe automatic extraction of animal 3D pose from images without markers is of interest in a range of scientific fields. Most work to date predicts animal pose from RGB images, based on 2D labelling of joint positions. However, due to the difficult nature of obtaining training data, no ground truth dataset of 3D animal motion is available to quantitatively evaluate these approaches. In addition, a lack of 3D animal pose data also makes it difficult to train 3D pose-prediction methods in a similar manner to the popular field of body-pose prediction. In our work, we focus on the problem of 3D canine pose estimation from RGBD images, recording a diverse range of dog breeds with several Microsoft Kinect v2s, simultaneously obtaining the 3D ground truth skeleton via a motion capture system. We generate a dataset of synthetic RGBD images from this data. A stacked hourglass network is trained to predict 3D joint locations, which is then constrained using prior models of shape and pose. We evaluate our model on both synthetic and real RGBD images and compare our results to previously published work fitting canine models to images. Finally, despite our training set consisting only of dog data, visual inspection implies that our network can produce good predictions for images of other quadrupeds - e.g. horses or cats - when their pose is similar to that contained in our training set. Sinead Kearney, Wenbin Li 0002, Martin Parsons, Kwang In Kim, Darren Cosker |
CVPR | 2 |
| 2020 | High-quality depth up-sampling via a supervised classification guided MRF model
Yiguo Qiao, Licheng Jiao, Xu Tang 0004, Wenbin Li 0002, Darren Cosker |
Pattern Recognit. Lett. | 4 |
| 2019 | Characterizing Visual Localization and Mapping DatasetsabstractBenchmarking mapping and motion estimation algorithms is established practice in robotics and computer vision. As the diversity of datasets increases, in terms of the trajectories, models, and scenes, it becomes a challenge to select datasets for a given benchmarking purpose. Inspired by the Wasserstein distance, this paper addresses this concern by developing novel metrics to evaluate trajectories and the environments without relying on any SLAM or motion estimation algorithm. The metrics, which so far have been missing in the research community, can be applied to the plethora of datasets that exist. Additionally, to improve the robotics SLAM benchmarking, the paper presents a new dataset for visual localization and mapping algorithms. A broad range of real-world trajectories is used in very high-quality scenes and a rendering framework to create a set of synthetic datasets with ground-truth trajectory and dense map which are representative of key SLAM applications such as virtual reality (VR), micro aerial vehicle (MAV) flight, and ground robotics. Sajad Saeedi G., Eduardo D. C. Carvalho, Wenbin Li 0002, Dimos Tzoumanikas, Stefan Leutenegger, Paul H. J. Kelly, Andrew J. Davison |
ICRA | 3 |
| 2019 | MID-Fusion: Octree-based Object-Level Multi-Instance Dynamic SLAMabstractWe propose a new multi-instance dynamic RGB-D SLAM system using an object-level octree-based volumetric representation. It can provide robust camera tracking in dynamic environments and at the same time, continuously estimate geometric, semantic, and motion properties for arbitrary objects in the scene. For each incoming frame, we perform instance segmentation to detect objects and refine mask boundaries using geometric and motion information. Meanwhile, we estimate the pose of each existing moving object using an object-oriented tracking method and robustly track the camera pose against the static scene. Based on the estimated camera pose and object poses, we associate segmented masks with existing models and incrementally fuse corresponding colour, depth, semantic, and foreground object probabilities into each object model. In contrast to existing approaches, our system is the first system to generate an object-level dynamic volumetric map from a single RGB-D camera, which can be used directly for robotic tasks. Our method can run at 2-3 Hz on a CPU, excluding the instance segmentation part. We demonstrate its effectiveness by quantitatively and qualitatively testing it on both synthetic and real-world sequences. Binbin Xu 0001, Wenbin Li 0002, Dimos Tzoumanikas, Michael Bloesch, Andrew J. Davison, Stefan Leutenegger |
ICRA | 2 |
| 2019 | Anisotropic Surface Remeshing without Obtuse AnglesabstractAbstract We present a novel anisotropic surface remeshing method that can efficiently eliminate obtuse angles. Unlike previous work that can only suppress obtuse angles with expensive resampling and Lloyd‐type iterations, our method relies on a simple yet efficient connectivity and geometry refinement, which can not only remove all the obtuse angles, but also preserves the original mesh connectivity as much as possible. Our method can be directly used as a post‐processing step for anisotropic meshes generated from existing algorithms to improve mesh quality. We evaluate our method by testing on a variety of meshes with different geometry and topology, and comparing with representative prior work. The results demonstrate the effectiveness and efficiency of our approach. Qun-Ce Xu, Dong-Ming Yan 0001, Wenbin Li 0002, Yongliang Yang 0002 |
Comput. Graph. Forum | 3 |
| 2018 | Motion Estimation and Segmentation of Natural Phenomena
Da Chen 0003, Wenbin Li 0002, Peter Hall 0001 |
BMVC | 2 |
| 2018 | InteriorNet: Mega-scale Multi-sensor Photo-realistic Indoor Scenes Dataset
Wenbin Li 0002, Sajad Saeedi G., John McCormac, Ronald Clark, Dimos Tzoumanikas, Yuzhong Huang, Rui Tang 0015, Stefan Leutenegger |
BMVC | 1 |
| 2018 | Learning system in real-time machine vision
Wenbin Li 0002, Zhihan Lyu, Darren Cosker |
Neurocomputing | 1 |
| 2018 | Learn to model blurry motion via directional similarity and filtering
Wenbin Li 0002, Da Chen 0003, Zhihan Lyu, Yan Yan 0002, Darren Cosker |
Pattern Recognit. | 1 |
| 2017 | Dense RGB-D-inertial SLAM with map deformationsabstractWhile dense visual SLAM methods are capable of estimating dense reconstructions of the environment, they suffer from a lack of robustness in their tracking step, especially when the optimisation is poorly initialised. Sparse visual SLAM systems have attained high levels of accuracy and robustness through the inclusion of inertial measurements in a tightly-coupled fusion. Inspired by this performance, we propose the first tightly-coupled dense RGB-D-inertial SLAM system. Our system has real-time capability while running on a GPU. It jointly optimises for the camera pose, velocity, IMU biases and gravity direction while building up a globally consistent, fully dense surfel-based 3D reconstruction of the environment. Through a series of experiments on both synthetic and real world datasets, we show that our dense visual-inertial SLAM system is more robust to fast motions and periods of low texture and low geometric variation than a related RGB-D-only SLAM system. Tristan Laidlow, Michael Bloesch, Wenbin Li 0002, Stefan Leutenegger |
IROS | 3 |
| 2017 | Guest Editorial: Advances in Big Data Methods for Image ProcessingabstractNowadays, big data analysts in scientific, government, industrial, and commercial domains are confronted with the challenges associated with rapidly growing volumes of data that are collected in multiple applications, such as biochemical and genetics research, fundamental physical experiments and astronomical observations, social networks, and consumer behaviour studies. In these applications, large amounts of raw data can be applied for decision making and action planning, yet their volume and increasingly complex structure limit the applicability of multiple well-known approaches that are widely utilised with small datasets, including principal component analysis, singular value decomposition, and spectral analysis. Visual data in particular usually includes a large amount of image data, massive unstructured pixels and multispectral images. Visual data has always been 'Big Data', which includes surveillance, multimedia, YouTube/Flickr, and medical imaging by image analysis techniques – for example feature extraction, dimensionality reduction, parallel/GPU processing, and tracking. New paradigms, techniques, and algorithms are required to address the issue of big data. Several approaches have been put forward for representing and processing large datasets with complex structures. Multi-dimensional data, described by multiple parameters, can be expressed and analysed using multi-way arrays, which have been applied in image processing, biomedical signal processing, telecommunications and sensor array processing as well as other domains. The main goal of large image data processing is to extract the characteristics of the image itself, including semantics, quality, relevancy, and other physical senses. Animal biometrics based recognition systems are gradually gaining more proliferation due to their diversity of application and uses. The recognition system is applied for representation, recognition of generic visual features, and classification of different species based on their phenotype appearances, the morphological image pattern, and biometric characteristics. The muzzle point image pattern is a primary animal biometric characteristic for the recognition of individual cattle. It is similar to the identification of minutiae points in human fingerprints. This study presents an automatic recognition algorithm of muzzle point image pattern of cattle for the identification of individual cattle, verification of false insurance claims, registration, and traceability process. In 'Muzzle point pattern based techniques for individual cattle identification', the proposed recognition algorithm uses the texture feature descriptors, such as speeded up robust features and local binary pattern for the extraction of features from the muzzle point images at different smoothed levels of Gaussian pyramid. The feature descriptors acquired at each Gaussian smoothed level are combined using fusion weighted sum-rule method. With a muzzle point image pattern database of 500 cattle, the proposed algorithm yields the desired level of 93.87% identification accuracy. The comparative analysis of experimental results for proposed work and appearance based face recognition algorithms has been done at each level. With the rapid development of computer science, problems with digital products piracy and copyright disputes become more serious; therefore, it is an urgent task to find solutions for these problems. In 'Digital image watermarking method based on DCT and fractal encoding', the authors' develop a digital watermarking algorithm based on a fractal encoding method and the discrete cosine transform (DCT). The proposed method combines fractal encoding method and DCT method for double encryptions to improve traditional DCT method. The image is encoded by fractal encoding as the first encryption, and then encoded parameters are used in DCT method as the second encryption. First, the fractal encoding method is adopted to encode a private image with private scales. Encoding parameters are applied as digital watermarking. Then, digital watermarking is added to the original image to reversibly use DCT, which means the authors can extract the private image from the carrier image with private encoding scales. Finally, attacking experiments are carried out on the carrier image by using several attacking methods. Experimental results show that the presented method has higher performance characteristics such as robustness and peak signal to noise ratio than classical methods. A 3D model watermarking method robust to geometric attacks is proposed in '3D model watermarking algorithm robust to geometric attacks'. The vertices of the model are classified into three groups. The vertices of the low-resolution group are used to establish an invariant space in which to resist geometric attacks. In the medium-resolution group, appropriate vertices in which to embed the watermarking information are selected. The selection of vertices that contain information about the watermark is based on the area of the local set of the vertex, the curvature of which determines the embedding strength of the watermark. The vertices of the high-resolution group are reserved for resisting simplification and smoothing attacks. The choice of the embedding position and the embedding strength can provide a suitable trade-off between good transparency and maximum robustness of the proposed method. The simulation results show that, compared to existing state-of-the-art methods, the proposed method is robust against attacks such as noise, smoothing, simplification, cropping, rotation, translation, and scaling while ensuring high visual quality of the watermarked model. A 3D model watermarking method robust to geometric attacks is proposed. The vertices of the model are classified into three groups. The vertices of the low-resolution group are used to establish an invariant space in which to resist geometric attacks. In the medium-resolution group, appropriate vertices in which to embed the watermarking information are selected. The selection of vertices that contain information about the watermark is based on the area of the local set of the vertex, the curvature of which determines the embedding strength of the watermark. The vertices of the high-resolution group are reserved for resisting simplification and smoothing attacks. The choice of the embedding position and the embedding strength can provide a suitable trade-off between good transparency and maximum robustness of the proposed method. The simulation results show that, compared to existing state-of-the-art methods, the proposed method is robust against attacks such as noise, smoothing, simplification, cropping, rotation, translation, and scaling while ensuring high visual quality of the watermarked model. Pedestrian detection has become one of the hottest topics in intelligent traffic systems because of its potential applications in driver assistance and automatic driving. In 'Fast pedestrian detection and dynamic tracking for intelligent vehicles within V2V cooperative environment', a fast pedestrian detection and dynamic tracking method within vehicle-to-vehicle (V2V) cooperative environment is proposed. A dynamic tracking-by-detection framework for real-time pedestrian detection is developed. First, cascade classifiers, based on selected Haar-like features, are trained to detect pedestrians. Then, CamShift algorithm combined with extended Kalman filtering is used in pedestrian dynamic tracking. Finally, with the crowdsourcing detected information, a smartphone-based V2V cooperative warning system is developed to share useful detection results within blind spots. The experiment results show that the proposed method has a real-time and accurate performance, which can provide a reference for road traffic safety monitoring technology. Due to storage conditions and material's non-planar shape, geometric distortion of the two-dimensional content is widely present in scanned document images. Effective geometric restoration of these distorted document images considerably increases character recognition rate in large-scale digitisation. For large-scale digitisation of historical books, geometric restoration solutions expect to be accurate, generic, robust, unsupervised and reversible. However, most methods in the literature concentrate on improving restoration accuracy for specific distortion effect, but not their applicability in large-scale digitisation. 'Effective geometric restoration of distorted historical document for large-scale digitisation' proposes an effective mesh based geometric restoration system (GRLSD) for large-scale distorted historical document digitisation. In this system, an automatic mesh generation based dewarping tool is proposed to geometrically model and correct arbitrary warping historical documents. An XML-based mesh recorder is proposed to record the mesh of distortion information for reversible use. A graphic user interface (GUI) toolkit is designed to visually display and manually manipulate the mesh for improving geometric restoration accuracy. Experimental results show that the proposed automatic dewarping approach efficiently corrects arbitrarily warped historical documents, with an improved performance over several state-of-the-art geometric restoration methods. By using XML mesh recorder and GUI toolkit, the GRLSD system greatly aids users to flexibly monitor and correct ambiguous points of mesh for the prevention of damaging historical document images without distortions in large-scale digitalisation. The ultimate receiver of image and video is human visual system (HVS). An important problem in the domain of image and video processing is how to establish visual information representation model meeting the HVS perception property. In 'Structured entropy of primitive: big databased stereoscopic image quality assessment', authors give theory analysis and experiment results to prove that l_1 norm-based entropy of primitive (EoP) is superior to the l_0 norm-based EoP for the monocular cue in image quality assessment. By developing the concept of mutual information of primitive (MIP) as the binocular cue, an l_1 EoP-based stereoscopic image quality assessment metric is proposed. With EoP as monocular cue and MIP as binocular cue, the relative entropy between the original stereoscopic image and the distorted one is explored to predict the quality score with support vector regression. To avoid destroying images' structured information, the structured EoP (SEoP) is further explored to measure the stereoscopic image information. Extensive experimental results demonstrate that the stereoscopic image quality assessment algorithm with SEoP as monocular cue and MIP as binocular cue outperforms many state-of-the-art ones. Zhihan Lyu, Wenbin Li 0002 |
IET Image Process. | 2 |
| 2017 | Video interpolation using optical flow and Laplacian smoothness
Wenbin Li 0002, Darren Cosker |
Neurocomputing | 1 |
| 2017 | Blur robust optical flow using motion channel
Wenbin Li 0002, Jee Hang Lee, Gang Ren 0001, Darren Cosker |
Neurocomputing | 1 |
| 2017 | Virtual reality geographical interactive scene semantics research for immersive geography learning
Zhihan Lyu, Xiaoming Li 0009, Wenbin Li 0002 |
Neurocomputing | 3 |
| 2016 | Dense Motion Estimation for Smoke
Da Chen 0003, Wenbin Li 0002, Peter Hall 0001 |
ACCV (4) | 2 |
| 2016 | Roto++: accelerating professional rotoscoping using shape manifoldsabstractRotoscoping (cutting out different characters/objects/layers in raw video footage) is a ubiquitous task in modern post-production and represents a significant investment in person-hours. In this work, we study the particular task of professional rotoscoping for high-end, live action movies and propose a new framework that works with roto-artists to accelerate the workflow and improve their productivity. Working with the existing keyframing paradigm, our first contribution is the development of a shape model that is updated as artists add successive keyframes. This model is used to improve the output of traditional interpolation and tracking techniques, reducing the number of keyframes that need to be specified by the artist. Our second contribution is to use the same shape model to provide a new interactive tool that allows an artist to reduce the time spent editing each keyframe. The more keyframes that are edited, the better the interactive tool becomes, accelerating the process and making the artist more efficient without compromising their control. Finally, we also provide a new, professionally rotoscoped dataset that enables truly representative, real-world evaluation of rotoscoping methods. We used this dataset to perform a number of experiments, including an expert study with professional roto-artists, to show, quantitatively, the advantages of our approach. Wenbin Li 0002, Fabio Viola, Jonathan Starck, Gabriel J. Brostow, Neill D. F. Campbell |
ACM Trans. Graph. | 1 |
| 2015 | Multi-view Reconstruction of Highly Specular Surfaces in Uncontrolled EnvironmentsabstractReconstructing the surface of highly specular objects is a challenging task. The shapes of diffuse and rough specular objects can be captured in an uncontrolled setting using consumer equipment. In contrast, highly specular objects have previously deterred capture in uncontrolled environments and have only been reconstructed using tailor-made hardware. We propose a method to reconstruct such objects in uncontrolled environments using only commodity hardware. As input, our method expects multi-view photographs of the specular object, its silhouettes and an environment map of its surroundings. We compare the reflected colors in the photographs with the ones in the environment to form probability distributions over the surface normals. As the effect of inter-reflections cannot be ignored for highly specular objects, we explicitly model them when forming the probability distributions. We recover the shape of the object in an iterative process where we alternate between estimating normals and updating the shape of the object to better explain these normals. We run experiments on both synthetic and real-world data, that show our method is robust and produces accurate reconstructions with as few as 25 input photographs. Clément Godard, Peter Hedman, Wenbin Li 0002, Gabriel J. Brostow |
3DV | 3 |
| 2014 | Robust optical flow estimation for continuous blurred scenes using RGB-motion imaging and directional filteringabstractOptical flow estimation is a difficult task given real-world video footage with camera and object blur. In this paper, we combine a 3D pose&position tracker with an RGB sensor allowing us to capture video footage together with 3D camera motion. We show that the additional camera motion information can be embedded into a hybrid optical flow framework by interleaving an iterative blind deconvolution and warping based minimization scheme. Such a hybrid framework significantly improves the accuracy of optical flow estimation in scenes with strong blur. Our approach yields improved overall performance against three state-of-the-art baseline methods applied to our proposed ground truth sequences, as well as in several other real-world sequences captured by our novel imaging system. Wenbin Li 0002, Jee Hang Lee, Gang Ren 0001, Darren Cosker |
WACV | 1 |
| 2013 | Optical Flow Estimation Using Laplacian Mesh EnergyabstractIn this paper we present a novel non-rigid optical flow algorithm for dense image correspondence and non-rigid registration. The algorithm uses a unique Laplacian Mesh Energy term to encourage local smoothness whilst simul-taneously preserving non-rigid deformation. Laplacian de-formation approaches have become popular in graphics re-search as they enable mesh deformations to preserve local surface shape. In this work we propose a novel Laplacian Mesh Energy formula to ensure such sensible local defor-mations between image pairs. We express this wholly with-in the optical flow optimization, and show its application in a novel coarse-to-fine pyramidal approach. Our algorith-m achieves the state-of-the-art performance in all trials on the Garg et al. dataset, and top tier performance on the Middlebury evaluation. 1. Wenbin Li 0002, Darren Cosker, Matthew Brown 0001, Rui Tang 0015 |
CVPR | 1 |
| 2012 | An Anchor Patch Based Optimization Framework for Reducing Optical Flow Drift in Long Image Sequences
Wenbin Li 0002, Darren Cosker, Matthew Brown 0001 |
ACCV (3) | 1 |