Peng Sun 0011

dblp:88/619-11 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-8187-9736ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
17 papers
Segmentation and scene understanding · 25% Reinforcement learning · 24% Video understanding and tracking · 20%
Theoretical computer science
3 papers
Mathematical optimization · 88% Coding theory · 8% Algorithms and data structures · 4%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 87% Computational science and engineering · 13%

Topics — the 30 heaviest of 45, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › object tracking
active target tracking
2.152021
AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Towards Distraction-Robust Active Visual Tracking · ICML 2021
End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.842021
AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Towards Distraction-Robust Active Visual Tracking · ICML 2021
Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI · ICML 2019
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
1.532022
Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022
Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation · Int. J. Comput. Vis. 2021
Graph-Guided Architecture Search for Real-Time Semantic Segmentation · CVPR 2020
Computer vision › Video understanding and tracking
object tracking
1.332021
AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Towards Distraction-Robust Active Visual Tracking · ICML 2021
End-to-end Active Object Tracking via Reinforcement Learning · ICML 2018
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
RGB-D salient object detection
1.122022
Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022
Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion · CVPR 2021
Computer vision › Segmentation and scene understanding › semantic segmentation › efficient semantic segmentation
real-time semantic segmentation
0.922021
Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation · Int. J. Comput. Vis. 2021
Graph-Guided Architecture Search for Real-Time Semantic Segmentation · CVPR 2020
Computer vision › Segmentation and scene understanding
semantic segmentation
0.922021
Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation · Int. J. Comput. Vis. 2021
Graph-Guided Architecture Search for Real-Time Semantic Segmentation · CVPR 2020
Machine learning › Reinforcement learning
deep reinforcement learning
0.822020
End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2020
End-to-end Active Object Tracking via Reinforcement Learning · ICML 2018
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › task-specific architecture search
multimodal fusion architecture search
0.612022
Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022
Computer vision › Segmentation and scene understanding
saliency detection
0.612022
Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022
Machine learning › Reinforcement learning › robust reinforcement learning
adversarial reinforcement learning
0.512021
AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Machine learning › Efficient and distributed learning › adaptive computation
anytime prediction
0.512021
Anytime Recognition with Routing Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Computer vision › Image recognition and object detection
image classification
0.512021
Anytime Recognition with Routing Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection
0.512021
Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion · CVPR 2021
Computer vision › Segmentation and scene understanding
scene parsing
0.512021
Anytime Recognition with Routing Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2021
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.532014
A Convergence Rate Analysis for LogitBoost, MART and Their Variant · ICML 2014
Saving Evaluation Time for the Decision Function in Boosting: Representation and Reordering Base Learner · ICML (3) 2013
AOSO-LogitBoost: Adaptive One-Vs-One LogitBoost for Multi-Class Problem · ICML 2012
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer
0.522020
End-to-end Active Object Tracking via Reinforcement Learning · ICML 2018
End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning
0.412019
Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI · ICML 2019
Machine learning › Generative modeling › image generation
conditional image generation
0.312018
Tagging Like Humans: Diverse and Distinct Image Annotation · CVPR 2018
Machine learning › Generative modeling
generative adversarial network
0.312018
Tagging Like Humans: Diverse and Distinct Image Annotation · CVPR 2018
Machine learning › Reinforcement learning
imitation learning
0.312018
Exponentially Weighted Imitation Learning for Batched Historical Data · NeurIPS 2018
Machine learning › Reinforcement learning
offline reinforcement learning
0.312018
Exponentially Weighted Imitation Learning for Batched Historical Data · NeurIPS 2018
Multimedia analysis and retrieval
image annotation
0.312018
Tagging Like Humans: Diverse and Distinct Image Annotation · CVPR 2018
Mathematical optimization › continuous optimization › convex optimization › proximal methods
alternating direction method of multipliers
0.312018
An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018
Mathematical optimization › continuous optimization
convex optimization
0.312018
An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018
Mathematical optimization › continuous optimization › convex optimization
operator splitting
0.312018
An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018
Mathematical optimization
primal-dual method
0.312018
An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018
Mathematical optimization › continuous optimization › convex optimization
proximal methods
0.312018
An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018
Medical and health informatics › medical imaging
cardiac imaging
0.312017
Comprehensive Modeling and Visualization of Cardiac Anatomy and Physiology from CT Imaging and Computer Simulations · IEEE Trans. Vis. Comput. Graph. 2017
Medical and health informatics
computer-aided diagnosis
0.312017
Comprehensive Modeling and Visualization of Cardiac Anatomy and Physiology from CT Imaging and Computer Simulations · IEEE Trans. Vis. Comput. Graph. 2017

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 1.4environment augmentation · 1.3neural architecture search · 1.0asymmetric dueling · 0.9reward shaping · 0.8ConvNet-LSTM · 0.8determinantal point process · 0.7GAN · 0.7multimodal fusion · 0.6hemodynamic simulation · 0.6geometric modeling · 0.6depth-sensitive attention · 0.6CT imaging · 0.6multimodal feature fusion · 0.5cross-modal teacher-student learning · 0.5variable metric · 0.3over-relaxation · 0.3hybrid proximal extragradient · 0.3
YearPublicationVenuePosition
2024 Stereo Matching Method with Integrated Geometric Encoding for Disparity Refinement
abstract
Neural network-based stereo matching algorithmms have made significant progress in fields such as robot navigation and autonomous driving. These application scenarios where fast and accurate obtaining the disparity of stereo images is critical for real-time stereo matching decisions. However, current stereo matching algorithms face the challenge of balancing real-time and accuracy while maintaining high accuracy. In this paper, we propose a disparity update strategy based on geometric encoding (MDStereo), which uses global and non-local geometric feature information to update disparity for highly accurate and quick matching. The proposed MDStereo constructs a group of geometric encoding volumes to encode the local information of the image; Next, a new GEDU method for disparity updating is proposed, which retrieves the correlation of high-resolution cost volumes in the form of sampling, and then fuses the geometric encoding information to iteratively update the disparity. Compared to RAFT-Stereo which retrieves correlations from all cascade cost volumes, our GEDU not only provides rich information but also has a more concise architecture. Furthermore, to speed up the inference of the algorithm, we improve the 3D stacked hourglass network, which effectively increases the receptive field and reduces the computational complexity. Our MDStereo has validated its effectiveness and accuracy on several benchmarks, achieving an EPE (end point error) of 0.58 pixels, a 3-pixel error of 2.58%, and a runtime of 43ms on the Scene Flow dataset. At the time of writing, MDStereo outperformed the published real-time methods at the popular KITTI 2012 and KITTI 2015. Compared with existing iteratively updating disparity methods (e.g., RAFT-Stereo), our method reduces the memory consumption by 54% and greatly improves the inference speed.
Shujia Ye, Ligang Cao, Chun Yuan 0003, Qianghua Li, Peng Sun 0011
IJCNN7
2024 The Fittest Wins: A Multistage Framework Achieving New SOTA in ViZDoom Competition
abstract
This article offers an integrated solution for first-person shooter (FPS) games to train agents with adaptive strategies. Solving such complex decision tasks requires generalization ability and adaptive strategies. We develop a framework using a novel adaptive strategic control algorithm combined with advanced techniques, such as the hindsight experience replay, multiagent reinforcement learning, and league training. The approach adopts a multistage learning scheme, consisting of learning a goal-conditioned navigation policy, then transferring to learn sophisticated shooting skills by playing against a league of players, and finally learning adaptive strategies. Our agent achieves the SOTA result in pastViZDoomAI Competitions, surpassing previous top-ranked agents (never seen during training) by a large margin. We provide comprehensive analysis and experiments to elaborate the effect of each component in affecting the agent performance and demonstrate that the proposed and adopted techniques are essential to achieve superior performance inViZDoomCompetition and potentially valuable for general end-to-end FPS games.
Shuxing Li, Honghua Dong, Yu Yang 0016, Chun Yuan 0003, Peng Sun 0011, Lei Han 0001
IEEE Trans. Games6
2022 Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search
Peng Sun 0011, Wenhu Zhang, Congli Song, Xi Li 0001
Int. J. Comput. Vis.1
2021 Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion
abstract
RGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e., RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RGB feature modeling scheme using the depth-wise geometric prior of salient objects. In principle, the feature modeling scheme is carried out in a depth-sensitive attention module, which leads to the RGB feature enhancement as well as the background distraction reduction by capturing the depth geometry prior. More-over, to perform effective multi-modal feature fusion, we further present an automatic architecture search approach for RGB-D SOD, which does well in finding out a feasible architecture from our specially designed multi-modal multi-scale search space. Extensive experiments on seven standard benchmarks demonstrate the effectiveness of the proposed approach against the state-of-the-art.
Peng Sun 0011, Wenhu Zhang, Xi Li 0001
CVPR1
2021 Towards Distraction-Robust Active Visual Tracking
abstract
In active visual tracking, it is notoriously difficult when distracting objects appear, as distractors often mislead the tracker by occluding the target or bringing a confusing appearance. To address this issue, we propose a mixed cooperative-competitive multi-agent game, where a target and multiple distractors form a collaborative team to play against a tracker and make it fail to follow. Through learning in our game, diverse distracting behaviors of the distractors naturally emerge, thereby exposing the tracker’s weakness, which helps enhance the distraction-robustness of the tracker. For effective learning, we then present a bunch of practical methods, including a reward function for distractors, a cross-modal teacher-student learning strategy, and a recurrent attention mechanism for the tracker. The experimental results show that our tracker performs desired distraction-robust active visual tracking and can be well generalized to unseen environments. We also show that the multi-agent game can be used to adversarially test the robustness of trackers.
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001
ICML2
2021 Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation
Peng Sun 0011, Jiaxiang Wu 0001, Peiwen Lin, Junzhou Huang, Xi Li 0001
Int. J. Comput. Vis.1
2021 Anytime Recognition with Routing Convolutional Networks
abstract
Achieving an automatic trade-off between accuracy and efficiency for a single deep neural network is highly desired in time-sensitive computer vision applications. To achieve anytime prediction, existing methods only embed fixed exits to neural networks and make the predictions with the fixed exits for all the samples (refer to the "latest-all" strategy). However, it is observed that the latest exit within a time budget does not always provide a more accurate prediction than the earlier exits for testing samples of various difficulties, making the "latest-all" strategy a sub-optimal solution. Motivated by this, we propose to improve the anytime prediction accuracy by allowing each sample to adaptively select its own optimal exit within a specific time budget. Specifically, we propose a new Routing Convolutional Network (RCN). For any given time budget, it adaptively selects the optimal layer as exit for a specific testing sample. To learn an optimal policy for sample routing, a Q-network is embedded into the RCN at each exit, considering both potential information gain and time-cost. To further boost the anytime prediction accuracy, the exits and the Q-networks are optimized alternately to mutually boost each other under the cost-sensitive environment. Apart from applying to whole image classification, RCN can also be adapted to dense prediction tasks, e.g., scene parsing, to achieve the pixel-level anytime prediction. Extensive experimental results on CIFAR-10, CIFAR-100, and ImageNet classification benchmarks, and Cityscapes scene parsing benchmark demonstrate the efficacy of the proposed RCN for anytime recognition.
Zequn Jie, Peng Sun 0011, Xi Li 0001, Jiashi Feng, Wei Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking
abstract
Visual Active Tracking (VAT) aims at following a target object by autonomously controlling the motion system of a tracker given visual observations. To learn a robust tracker for VAT, in this article, we propose a novel adversarial reinforcement learning (RL) method which adopts an Asymmetric Dueling mechanism, referred to as AD-VAT. In the mechanism, the tracker and target, viewed as two learnable agents, are opponents and can mutually enhance each other during the dueling/competition: i.e., the tracker intends to lockup the target, while the target tries to escape from the tracker. The dueling is asymmetric in that the target is additionally fed with the tracker's observation and action, and learns to predict the tracker's reward as an auxiliary task. Such an asymmetric dueling mechanism produces a stronger target, which in turn induces a more robust tracker. To improve the performance of the tracker in the case of challenging scenarios such as obstacles, we employ more advanced environment augmentation technique and two-stage training strategies, termed as AD-VAT+. For a better understanding of the asymmetric dueling mechanism, we also analyze the target's behaviors as the training proceeds and visualize the latent space of the tracker. The experimental results, in both 2D and 3D environments, demonstrate that the proposed method leads to a faster convergence in training and yields more robust tracking behaviors in different testing scenarios. The potential of the active tracker is also shown in real-world videos.
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2020 Graph-Guided Architecture Search for Real-Time Semantic Segmentation
abstract
Designing a lightweight semantic segmentation network often requires researchers to find a trade-off between performance and speed, which is always empirical due to the limited interpretability of neural networks. In order to release researchers from these tedious mechanical trials, we propose a Graph-guided Architecture Search (GAS) pipeline to automatically search real-time semantic segmentation networks. Unlike previous works that use a simplified search space and stack a repeatable cell to form a network, we introduce a novel search mechanism with a new search space where a lightweight model can be effectively explored through the cell-level diversity and latency oriented constraint. Specifically, to produce the cell-level diversity, the cell-sharing constraint is eliminated through the cell-independent manner. Then a graph convolution network (GCN) is seamlessly integrated as a communication mechanism between cells. Finally, a latency-oriented constraint is endowed into the search process to balance the speed and performance. Extensive experiments on Cityscapes and CamVid datasets demonstrate that GAS achieves the new state-of-the-art trade-off between accuracy and speed. In particular, on Cityscapes dataset, GAS achieves the new best performance of 73.5% mIoU with the speed of 108.4 FPS on Titan Xp.
Peiwen Lin, Peng Sun 0011, Sirui Xie, Xi Li 0001, Jianping Shi
CVPR2
2020 End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning
abstract
We study active object tracking, where a tracker takes visual observations (i.e., frame sequences) as input and produces the corresponding camera control signals as output (e.g., move forward, turn left, etc.). Conventional methods tackle tracking and camera control tasks separately, and the resulting system is difficult to tune jointly. These methods also require significant human efforts for image labeling and expensive trial-and-error system tuning in the real world. To address these issues, we propose, in this paper, an end-to-end solution via deep reinforcement learning. A ConvNet-LSTM function approximator is adopted for the direct frame-to-action prediction. We further propose an environment augmentation technique and a customized reward function, which are crucial for successful training. The tracker trained in simulators (ViZDoom and Unreal Engine) demonstrates good generalization behaviors in the case of unseen object moving paths, unseen object appearances, unseen backgrounds, and distracting objects. The system is robust and can restore tracking after occasional lost of the target being tracked. We also find that the tracking ability, obtained solely from simulators, can potentially transfer to real-world scenarios. We demonstrate successful examples of such transfer, via experiments over the VOT dataset and the deployment of a real-world robot using the proposed active tracker trained in simulation.
Wenhan Luo, Peng Sun 0011, Fangwei Zhong, Wei Liu 0005, Tong Zhang 0001, Yizhou Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Generative adversarial exploration for reinforcement learning
abstract
Exploration is crucial for training the optimal reinforcement learning (RL) policy, where the key is to discriminate whether a state visiting is novel. Most previous work focuses on designing heuristic rules or distance metrics to check whether a state is novel without considering such a discrimination process that can be learned. In this paper, we propose a novel method called generative adversarial exploration (GAEX) to encourage exploration in RL via introducing an intrinsic reward output from a generative adversarial network, where the generator provides fake samples of states that help discriminator identify those less frequently visited states. Thus the agent is encouraged to visit those states which the discriminator is less confident to judge as visited. GAEX is easy to implement and of high training efficiency. In our experiments, we apply GAEX into DQN and the DQN-GAEX algorithm achieves convincing performance on challenging exploration problems, including the game Venture, Montezuma's Revenge and Super Mario Bros, without further fine-tuning on complicate learning algorithms. To our knowledge, this is the first work to employ GAN in RL exploration problems.
Weijun Hong, Menghui Zhu, Minghuan Liu, Weinan Zhang 0001, Ming Zhou 0006, Yong Yu 0001, Peng Sun 0011
DAI7
2019 AD-VAT: An Asymmetric Dueling mechanism for learning Visual Active Tracking
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001
ICLR (Poster)2
2019 Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI
abstract
We consider the problem of multi-agent reinforcement learning (MARL) in video game AI, where the agents are located in a spatial grid-world environment and the number of agents varies both within and across episodes. The challenge is to flexibly control an arbitrary number of agents while achieving effective collaboration. Existing MARL methods usually suffer from the trade-off between these two considerations. To address the issue, we propose a novel architecture that learns a spatial joint representation of all the agents and outputs grid-wise actions. Each agent will be controlled independently by taking the action from the grid it occupies. By viewing the state information as a grid feature map, we employ a convolutional encoder-decoder as the policy network. This architecture naturally promotes agent communication because of the large receptive field provided by the stacked convolutional layers. Moreover, the spatially shared convolutional parameters enable fast parallel exploration that the experiences discovered by one agent can be immediately transferred to others. The proposed method can be conveniently integrated with general reinforcement learning algorithms, e.g., PPO and Q-learning. We demonstrate the effectiveness of the proposed method in extensive challenging multi-agent tasks in StarCraft II.
Lei Han 0001, Peng Sun 0011, Yali Du 0001, Jiechao Xiong, Qing Wang 0015, Xinghai Sun, Han Liu 0001, Tong Zhang 0001
ICML2
2018 Tagging Like Humans: Diverse and Distinct Image Annotation
abstract
In this work we propose a new automatic image annotation model, dubbed diverse and distinct image annotation (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet distinct and diverse tags. In D2IA, we generate a relevant and distinct tag subset, in which the tags are relevant to the image contents and semantically distinct to each other, using sequential sampling from a determinantal point process (DPP) model. Multiple such tag subsets that cover diverse semantic aspects or diverse semantic levels of the image contents are generated by randomly perturbing the DPP sampling process. We leverage a generative adversarial network (GAN) model to train D2IA. Extensive experiments including quantitative and qualitative comparisons, as well as human subject studies, on two benchmark datasets demonstrate that the proposed model can produce more diverse and distinct tags than the state-of-the-arts.
Baoyuan Wu, Peng Sun 0011, Wei Liu 0005, Bernard Ghanem, Siwei Lyu
CVPR3
2018 End-to-end Active Object Tracking via Reinforcement Learning
abstract
We study active object tracking, where a tracker takes as input the visual observation (i.e. frame sequence) and produces the camera control signal (e.g., move forward, turn left, etc). Conventional methods tackle the tracking and the camera control separately, which is challenging to tune jointly. It also incurs many human efforts for labeling and many expensive trial-and-errors in real-world. To address these issues, we propose, in this paper, an end-to-end solution via deep reinforcement learning, where a ConvNet-LSTM function approximator is adopted for the direct frame-to-action prediction. We further propose an environment augmentation technique and a customized reward function, which are crucial for a successful training. The tracker trained in simulators (ViZDoom, Unreal Engine) shows good generalization in the case of unseen object moving path, unseen object appearance, unseen background, and distracting object. It can restore tracking when occasionally losing the target. With the experiments over the VOT dataset, we also find that the tracking ability, obtained solely from simulators, can potentially transfer to real-world scenarios.
Wenhan Luo, Peng Sun 0011, Fangwei Zhong, Wei Liu 0005, Tong Zhang 0001, Yizhou Wang 0001
ICML2
2018 An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method
abstract
We propose a novel algorithmic framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-gradient (VMOR-HPE) method with a global convergence guarantee for the maximal monotone operator inclusion problem. Its iteration complexities and local linear convergence rate are provided, which theoretically demonstrate that a large over-relaxed step-size contributes to accelerating the proposed VMOR-HPE as a byproduct. Specifically, we find that a large class of primal and primal-dual operator splitting algorithms are all special cases of VMOR-HPE. Hence, the proposed framework offers a new insight into these operator splitting algorithms. In addition, we apply VMOR-HPE to the Karush-Kuhn-Tucker (KKT) generalized equation of linear equality constrained multi-block composite convex optimization, yielding a new algorithm, namely nonsymmetric Proximal Alternating Direction Method of Multipliers with a preconditioned Extra-gradient step in which the preconditioned metric is generated by a blockwise Barzilai-Borwein line search technique (PADMM-EBB). We also establish iteration complexities of PADMM-EBB in terms of the KKT residual. Finally, we apply PADMM-EBB to handle the nonnegative dual graph regularized low-rank representation problem. Promising results on synthetic and real datasets corroborate the efficacy of PADMM-EBB.
Li Shen 0005, Peng Sun 0011, Wei Liu 0005, Tong Zhang 0001
ICML2
2018 Exponentially Weighted Imitation Learning for Batched Historical Data
abstract
We consider deep policy learning with only batched historical trajectories. The main challenge of this problem is that the learner no longer has a simulator or ``environment oracle'' as in most reinforcement learning settings. To solve this problem, we propose a monotonic advantage reweighted imitation learning strategy that is applicable to problems with complex nonlinear function approximation and works well with hybrid (discrete and continuous) action space. The method does not rely on the knowledge of the behavior policy, thus can be used to learn from data generated by an unknown policy. Under mild conditions, our algorithm, though surprisingly simple, has a policy improvement bound and outperforms most competing methods empirically. Thorough numerical results are also provided to demonstrate the efficacy of the proposed methodology.
Qing Wang 0015, Jiechao Xiong, Lei Han 0001, Peng Sun 0011, Han Liu 0001, Tong Zhang 0001
NeurIPS4
2017 Comprehensive Modeling and Visualization of Cardiac Anatomy and Physiology from CT Imaging and Computer Simulations
abstract
In clinical cardiology, both anatomy and physiology are needed to diagnose cardiac pathologies. CT imaging and computer simulations provide valuable and complementary data for this purpose. However, it remains challenging to gain useful information from the large amount of high-dimensional diverse data. The current tools are not adequately integrated to visualize anatomic and physiologic data from a complete yet focused perspective. We introduce a new computer-aided diagnosis framework, which allows for comprehensive modeling and visualization of cardiac anatomy and physiology from CT imaging data and computer simulations, with a primary focus on ischemic heart disease. The following visual information is presented: (1) Anatomy from CT imaging: geometric modeling and visualization of cardiac anatomy, including four heart chambers, left and right ventricular outflow tracts, and coronary arteries; (2) Function from CT imaging: motion modeling, strain calculation, and visualization of four heart chambers; (3) Physiology from CT imaging: quantification and visualization of myocardial perfusion and contextual integration with coronary artery anatomy; (4) Physiology from computer simulation: computation and visualization of hemodynamics (e.g., coronary blood velocity, pressure, shear stress, and fluid forces on the vessel wall). Substantially, feedback from cardiologists have confirmed the practical utility of integrating these features for the purpose of computer-aided diagnosis of ischemic heart disease.
Guanglei Xiong, Peng Sun 0011, Haoyin Zhou, Seongmin Ha, Briain o Hartaigh, Quynh A. Truong, James K. Min
IEEE Trans. Vis. Comput. Graph.2
2014 A Convergence Rate Analysis for LogitBoost, MART and Their Variant
abstract
LogitBoost, MART and their variant can be viewed as additive tree regression using logistic loss and boosting style optimization. We analyze their convergence rates based on a new weak learnability formulation. We show that it has O(\frac1T) rate when using gradient descent only, while a linear rate is achieved when using Newton descent. Moreover, introducing Newton descent when growing the trees, as LogitBoost does, leads to a faster linear rate. Empirical results on UCI datasets support our analysis.
Peng Sun 0011, Tong Zhang 0001, Jie Zhou 0001
ICML1
2014 An improved multiclass LogitBoost using adaptive-one-vs-one
Peng Sun 0011, Mark D. Reid, Jie Zhou 0001
Mach. Learn.1
2013 Saving Evaluation Time for the Decision Function in Boosting: Representation and Reordering Base Learner
abstract
For a well trained Boosting classifier, we are interested in how to save the testing time, i.e., to make the decision without evaluating all the base learners. To address this problem, in previous work the base learners are sequentially calculated and early stopping is allowed if the decision function has been confident enough to output its value. In such a chain structure, the order of base learners is critical: better order can lead to less evaluation time. In this paper, we present a novel method for ordering. We base our discussion on the data structure representing Boosting’s decision function. Viewing the decision function a boolean expression, we propose a Binary Valued Tree for its representation. As a secondary contribution, such a representation unifies the work by previous researchers and helps devise new representation. Also, its connection to Binary Decision Diagram(BDD) is discussed.
Peng Sun 0011, Jie Zhou 0001
ICML (3)1
2012 The Convexity and Design of Composite Multiclass Losses
Mark D. Reid, Robert C. Williamson, Peng Sun 0011
ICML3
2012 AOSO-LogitBoost: Adaptive One-Vs-One LogitBoost for Multi-Class Problem
Peng Sun 0011, Mark D. Reid, Jie Zhou 0001
ICML1