Mude Lin

dblp:176/5573 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorArtificial intelligence and machine learning · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
3D vision · 73% Reinforcement learning · 15% Video understanding and tracking · 10%
Computer graphics and multimedia
2 papers
Image and video processing · 67% Computational photography and imaging · 33%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision › stereo vision
stereo matching
0.822020
Visually Imbalanced Stereo Matching · CVPR 2020
Single View Stereo Matching · CVPR 2018
Computational photography and imaging
high dynamic range imaging
0.412020
Learning a Reinforced Agent for Flexible Exposure Bracketing Selection · CVPR 2020
Image and video processing
image restoration
0.412020
Visually Imbalanced Stereo Matching · CVPR 2020
Image and video processing › image fusion
multi-exposure image fusion
0.412020
Learning a Reinforced Agent for Flexible Exposure Bracketing Selection · CVPR 2020
Computer vision › 3D vision
depth estimation
0.312018
Single View Stereo Matching · CVPR 2018
Computer vision › 3D vision › depth estimation
monocular depth estimation
0.312018
Single View Stereo Matching · CVPR 2018
Computer vision › 3D vision
novel view synthesis
0.312018
Single View Stereo Matching · CVPR 2018
Computer vision › 3D vision
3d human pose estimation
0.312017
Recurrent 3D Pose Sequence Machines · CVPR 2017
Computer vision › Video understanding and tracking
temporal modeling
0.312017
Recurrent 3D Pose Sequence Machines · CVPR 2017
Computer vision › 3D vision
geometric constraints
0.112018
Single View Stereo Matching · CVPR 2018
Computer vision › Face, body and person analysis › human pose estimation
2d human pose estimation
0.112017
Recurrent 3D Pose Sequence Machines · CVPR 2017

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.9joint restoration and reconstruction · 0.9deep neural network · 0.9view synthesis · 0.3end-to-end training · 0.3recurrent neural network · 0.3multi-stage sequential refinement · 0.3feature adaptation · 0.3
YearPublicationVenuePosition
2020 Visually Imbalanced Stereo Matching
abstract
Understanding of human vision system (HVS) has inspired many computer vision algorithms. Stereo matching, which borrows the idea from human stereopsis, has been extensively studied in the existing literature. However, scant attention has been drawn on a typical scenario where binocular inputs are qualitatively different (e.g., high-res master camera and low-res slave camera in a dual-lens module). Recent advances in human optometry reveal the capability of the human visual system to maintain coarse stereopsis under such visually imbalanced conditions. Bionically aroused, it is natural to question that: do stereo machines share the same capability? In this paper, we carry out a systematic comparison to investigate the effect of various imbalanced conditions on current popular stereo matching algorithms. We show that resembling the human visual system, those algorithms can handle limited degrees of monocular downgrading but also prone to collapses beyond a certain threshold. To avoid such collapse, we propose a solution to recover the stereopsis by a joint guided-view-restoration and stereo-reconstruction framework. We show the superiority of our framework on KITTI dataset and its extension on real-world applications.
Yicun Liu, Jimmy S. J. Ren, Jiawei Zhang 0002, Mude Lin
CVPR5
2020 Learning a Reinforced Agent for Flexible Exposure Bracketing Selection
abstract
Automatically selecting exposure bracketing (images exposed differently) is important to obtain a high dynamic range image by using multi-exposure fusion. Unlike previous methods that have many restrictions such as requiring camera response function, sensor noise model, and a stream of preview images with different exposures (not accessible in some scenarios e.g. mobile applications), we propose a novel deep neural network to automatically select exposure bracketing, named EBSNet, which is sufficiently flexible without having the above restrictions. EBSNet is formulated as a reinforced agent that is trained by maximizing rewards provided by a multi-exposure fusion network (MEFNet). By utilizing the illumination and semantic information extracted from just a single auto-exposure preview image, EBSNet enables to select an optimal exposure bracketing for multi-exposure fusion. EBSNet and MEFNet can be jointly trained to produce favorable results against recent state-of-the-art approaches. To facilitate future research, we provide a new benchmark dataset for multi-exposure selection and fusion.
Zhouxia Wang, Jiawei Zhang 0002, Mude Lin, Ping Luo 0002, Jimmy S. J. Ren
CVPR3
2018 Single View Stereo Matching
abstract
Previous monocular depth estimation methods take a single view and directly regress the expected results. Though recent advances are made by applying geometrically inspired loss functions during training, the inference procedure does not explicitly impose any geometrical constraint. Therefore these models purely rely on the quality of data and the effectiveness of learning to generalize. This either leads to suboptimal results or the demand of huge amount of expensive ground truth labelled data to generate reasonable results. In this paper, we show for the first time that the monocular depth estimation problem can be reformulated as two sub-problems, a view synthesis procedure followed by stereo matching, with two intriguing properties, namely i) geometrical constraints can be explicitly imposed during inference; ii) demand on labelled depth data can be greatly alleviated. We show that the whole pipeline can still be trained in an end-to-end fashion and this new formulation plays a critical role in advancing the performance. The resulting model outperforms all the previous monocular depth estimation methods as well as the stereo block matching method in the challenging KITTI dataset by only using a small number of real training data. The model also generalizes well to other monocular depth estimation benchmarks. We also discuss the implications and the advantages of solving monocular depth estimation using stereo methods.
Jimmy S. J. Ren, Mude Lin, Jiahao Pang, Wenxiu Sun, Hongsheng Li 0001, Liang Lin 0004
CVPR3
2017 Recurrent 3D Pose Sequence Machines
abstract
3D Human articulated pose recovery from monocular image sequences is very challenging due to the diverse appearances, viewpoints, occlusions, and also the human 3D pose is inherently ambiguous from the monocular imagery. It is thus critical to exploit rich spatial and temporal long-range dependencies among body joints for accurate 3D pose sequence prediction. Existing approaches usually manually design some elaborate prior terms and human body kinematic constraints for capturing structures, which are often insufficient to exploit all intrinsic structures and not scalable for all scenarios. In contrast, this paper presents a Recurrent 3D Pose Sequence Machine(RPSM) to automatically learn the image-dependent structural constraint and sequence-dependent temporal context by using a multi-stage sequential refinement. At each stage, our RPSM is composed of three modules to predict the 3D pose sequences based on the previously learned 2D pose representations and 3D poses: (i) a 2D pose module extracting the image-dependent pose representations, (ii) a 3D pose recurrent module regressing 3D poses and (iii) a feature adaption module serving as a bridge between module (i) and (ii) to enable the representation transformation from 2D to 3D domain. These three modules are then assembled into a sequential prediction framework to refine the predicted poses with multiple recurrent stages. Extensive evaluations on the Human3.6M dataset and HumanEva-I dataset show that our RPSM outperforms all state-of-the-art approaches for 3D pose estimation.
Mude Lin, Liang Lin 0004, Xiaodan Liang, Keze Wang
CVPR1
2016 Character proposal network for robust text extraction
abstract
Maximally stable extremal regions (MSER), which is a popular method to generate character proposals/candidates, has shown superior performance in scene text detection. However, the pixel-level operation limits its capability for handling some challenging cases (e.g., multiple connected characters, separated parts of one character and non-uniform illumination). To better tackle these cases, we design a character proposal network (CPN) by taking advantage of the high capacity and fast computing of fully convolutional network (FCN). Specifically, the network simultaneously predicts character-ness scores and refines the corresponding locations. The character-ness scores can be used for proposal ranking to reject non-character proposals and the refining process aims to obtain the more accurate locations. Furthermore, considering the situation that different characters have different aspect ratios, we propose a multi-template strategy, designing a refiner for each aspect ratio. The extensive experiments indicate our method achieves recall rates of 93.88%, 93.60% and 96.46% on ICDAR 2013, SVT and Chinese 2k datasets respectively using less than 1000 proposals, demonstrating promising performance of our character proposal network.
Shuye Zhang, Mude Lin, Tianshui Chen, Liang Lin 0004
ICASSP2