EDBT 2026 Demo / reviewers in the wild / expert
Peng Sun 0011
dblp:88/619-11
· DBLP profile ↗
23ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0002-8187-9736ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 7 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
17 papers |
Segmentation and scene understanding · 25% Reinforcement learning · 24% Video understanding and tracking · 20% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 88% Coding theory · 8% Algorithms and data structures · 4% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 87% Computational science and engineering · 13% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › object tracking
active target tracking |
2.1 | 5 | 2021 | AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021 Towards Distraction-Robust Active Visual Tracking · ICML 2021 End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
1.8 | 4 | 2021 | AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021 Towards Distraction-Robust Active Visual Tracking · ICML 2021 Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI · ICML 2019 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
1.5 | 3 | 2022 | Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022 Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation · Int. J. Comput. Vis. 2021 Graph-Guided Architecture Search for Real-Time Semantic Segmentation · CVPR 2020 |
Computer vision › Video understanding and tracking
object tracking |
1.3 | 3 | 2021 | AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021 Towards Distraction-Robust Active Visual Tracking · ICML 2021 End-to-end Active Object Tracking via Reinforcement Learning · ICML 2018 |
Computer vision › Segmentation and scene understanding › saliency detection › salient object detection
RGB-D salient object detection |
1.1 | 2 | 2022 | Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022 Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion · CVPR 2021 |
Computer vision › Segmentation and scene understanding › semantic segmentation › efficient semantic segmentation
real-time semantic segmentation |
0.9 | 2 | 2021 | Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation · Int. J. Comput. Vis. 2021 Graph-Guided Architecture Search for Real-Time Semantic Segmentation · CVPR 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.9 | 2 | 2021 | Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation · Int. J. Comput. Vis. 2021 Graph-Guided Architecture Search for Real-Time Semantic Segmentation · CVPR 2020 |
Machine learning › Reinforcement learning
deep reinforcement learning |
0.8 | 2 | 2020 | End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2020 End-to-end Active Object Tracking via Reinforcement Learning · ICML 2018 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search › task-specific architecture search
multimodal fusion architecture search |
0.6 | 1 | 2022 | Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022 |
Computer vision › Segmentation and scene understanding
saliency detection |
0.6 | 1 | 2022 | Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search · Int. J. Comput. Vis. 2022 |
Machine learning › Reinforcement learning › robust reinforcement learning
adversarial reinforcement learning |
0.5 | 1 | 2021 | AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Machine learning › Efficient and distributed learning › adaptive computation
anytime prediction |
0.5 | 1 | 2021 | Anytime Recognition with Routing Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | Anytime Recognition with Routing Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Computer vision › Segmentation and scene understanding › saliency detection
salient object detection |
0.5 | 1 | 2021 | Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal Fusion · CVPR 2021 |
Computer vision › Segmentation and scene understanding
scene parsing |
0.5 | 1 | 2021 | Anytime Recognition with Routing Convolutional Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2021 |
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting |
0.5 | 3 | 2014 | A Convergence Rate Analysis for LogitBoost, MART and Their Variant · ICML 2014 Saving Evaluation Time for the Decision Function in Boosting: Representation and Reordering Base Learner · ICML (3) 2013 AOSO-LogitBoost: Adaptive One-Vs-One LogitBoost for Multi-Class Problem · ICML 2012 |
Machine learning › Transfer learning and domain adaptation
sim-to-real transfer |
0.5 | 2 | 2020 | End-to-end Active Object Tracking via Reinforcement Learning · ICML 2018 End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2020 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning
cooperative multi-agent reinforcement learning |
0.4 | 1 | 2019 | Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AI · ICML 2019 |
Machine learning › Generative modeling › image generation
conditional image generation |
0.3 | 1 | 2018 | Tagging Like Humans: Diverse and Distinct Image Annotation · CVPR 2018 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2018 | Tagging Like Humans: Diverse and Distinct Image Annotation · CVPR 2018 |
Machine learning › Reinforcement learning
imitation learning |
0.3 | 1 | 2018 | Exponentially Weighted Imitation Learning for Batched Historical Data · NeurIPS 2018 |
Machine learning › Reinforcement learning
offline reinforcement learning |
0.3 | 1 | 2018 | Exponentially Weighted Imitation Learning for Batched Historical Data · NeurIPS 2018 |
Multimedia analysis and retrieval
image annotation |
0.3 | 1 | 2018 | Tagging Like Humans: Diverse and Distinct Image Annotation · CVPR 2018 |
Mathematical optimization › continuous optimization › convex optimization › proximal methods
alternating direction method of multipliers |
0.3 | 1 | 2018 | An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018 |
Mathematical optimization › continuous optimization
convex optimization |
0.3 | 1 | 2018 | An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018 |
Mathematical optimization › continuous optimization › convex optimization
operator splitting |
0.3 | 1 | 2018 | An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018 |
Mathematical optimization
primal-dual method |
0.3 | 1 | 2018 | An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018 |
Mathematical optimization › continuous optimization › convex optimization
proximal methods |
0.3 | 1 | 2018 | An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient Method · ICML 2018 |
Medical and health informatics › medical imaging
cardiac imaging |
0.3 | 1 | 2017 | Comprehensive Modeling and Visualization of Cardiac Anatomy and Physiology from CT Imaging and Computer Simulations · IEEE Trans. Vis. Comput. Graph. 2017 |
Medical and health informatics
computer-aided diagnosis |
0.3 | 1 | 2017 | Comprehensive Modeling and Visualization of Cardiac Anatomy and Physiology from CT Imaging and Computer Simulations · IEEE Trans. Vis. Comput. Graph. 2017 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.4environment augmentation · 1.3neural architecture search · 1.0asymmetric dueling · 0.9reward shaping · 0.8ConvNet-LSTM · 0.8determinantal point process · 0.7GAN · 0.7multimodal fusion · 0.6hemodynamic simulation · 0.6geometric modeling · 0.6depth-sensitive attention · 0.6CT imaging · 0.6multimodal feature fusion · 0.5cross-modal teacher-student learning · 0.5variable metric · 0.3over-relaxation · 0.3hybrid proximal extragradient · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Stereo Matching Method with Integrated Geometric Encoding for Disparity RefinementabstractNeural network-based stereo matching algorithmms have made significant progress in fields such as robot navigation and autonomous driving. These application scenarios where fast and accurate obtaining the disparity of stereo images is critical for real-time stereo matching decisions. However, current stereo matching algorithms face the challenge of balancing real-time and accuracy while maintaining high accuracy. In this paper, we propose a disparity update strategy based on geometric encoding (MDStereo), which uses global and non-local geometric feature information to update disparity for highly accurate and quick matching. The proposed MDStereo constructs a group of geometric encoding volumes to encode the local information of the image; Next, a new GEDU method for disparity updating is proposed, which retrieves the correlation of high-resolution cost volumes in the form of sampling, and then fuses the geometric encoding information to iteratively update the disparity. Compared to RAFT-Stereo which retrieves correlations from all cascade cost volumes, our GEDU not only provides rich information but also has a more concise architecture. Furthermore, to speed up the inference of the algorithm, we improve the 3D stacked hourglass network, which effectively increases the receptive field and reduces the computational complexity. Our MDStereo has validated its effectiveness and accuracy on several benchmarks, achieving an EPE (end point error) of 0.58 pixels, a 3-pixel error of 2.58%, and a runtime of 43ms on the Scene Flow dataset. At the time of writing, MDStereo outperformed the published real-time methods at the popular KITTI 2012 and KITTI 2015. Compared with existing iteratively updating disparity methods (e.g., RAFT-Stereo), our method reduces the memory consumption by 54% and greatly improves the inference speed. Shujia Ye, Ligang Cao, Chun Yuan 0003, Qianghua Li, Peng Sun 0011 |
IJCNN | 7 |
| 2024 | The Fittest Wins: A Multistage Framework Achieving New SOTA in ViZDoom CompetitionabstractThis article offers an integrated solution for first-person shooter (FPS) games to train agents with adaptive strategies. Solving such complex decision tasks requires generalization ability and adaptive strategies. We develop a framework using a novel adaptive strategic control algorithm combined with advanced techniques, such as the hindsight experience replay, multiagent reinforcement learning, and league training. The approach adopts a multistage learning scheme, consisting of learning a goal-conditioned navigation policy, then transferring to learn sophisticated shooting skills by playing against a league of players, and finally learning adaptive strategies. Our agent achieves the SOTA result in pastViZDoomAI Competitions, surpassing previous top-ranked agents (never seen during training) by a large margin. We provide comprehensive analysis and experiments to elaborate the effect of each component in affecting the agent performance and demonstrate that the proposed and adopted techniques are essential to achieve superior performance inViZDoomCompetition and potentially valuable for general end-to-end FPS games. Shuxing Li, Honghua Dong, Yu Yang 0016, Chun Yuan 0003, Peng Sun 0011, Lei Han 0001 |
IEEE Trans. Games | 6 |
| 2022 | Learnable Depth-Sensitive Attention for Deep RGB-D Saliency Detection with Multi-modal Fusion Architecture Search
Peng Sun 0011, Wenhu Zhang, Congli Song, Xi Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Deep RGB-D Saliency Detection With Depth-Sensitive Attention and Automatic Multi-Modal FusionabstractRGB-D salient object detection (SOD) is usually formulated as a problem of classification or regression over two modalities, i.e., RGB and depth. Hence, effective RGB-D feature modeling and multi-modal feature fusion both play a vital role in RGB-D SOD. In this paper, we propose a depth-sensitive RGB feature modeling scheme using the depth-wise geometric prior of salient objects. In principle, the feature modeling scheme is carried out in a depth-sensitive attention module, which leads to the RGB feature enhancement as well as the background distraction reduction by capturing the depth geometry prior. More-over, to perform effective multi-modal feature fusion, we further present an automatic architecture search approach for RGB-D SOD, which does well in finding out a feasible architecture from our specially designed multi-modal multi-scale search space. Extensive experiments on seven standard benchmarks demonstrate the effectiveness of the proposed approach against the state-of-the-art. Peng Sun 0011, Wenhu Zhang, Xi Li 0001 |
CVPR | 1 |
| 2021 | Towards Distraction-Robust Active Visual TrackingabstractIn active visual tracking, it is notoriously difficult when distracting objects appear, as distractors often mislead the tracker by occluding the target or bringing a confusing appearance. To address this issue, we propose a mixed cooperative-competitive multi-agent game, where a target and multiple distractors form a collaborative team to play against a tracker and make it fail to follow. Through learning in our game, diverse distracting behaviors of the distractors naturally emerge, thereby exposing the tracker’s weakness, which helps enhance the distraction-robustness of the tracker. For effective learning, we then present a bunch of practical methods, including a reward function for distractors, a cross-modal teacher-student learning strategy, and a recurrent attention mechanism for the tracker. The experimental results show that our tracker performs desired distraction-robust active visual tracking and can be well generalized to unseen environments. We also show that the multi-agent game can be used to adversarially test the robustness of trackers. Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001 |
ICML | 2 |
| 2021 | Real-Time Semantic Segmentation via Auto Depth, Downsampling Joint Decision and Feature Aggregation
Peng Sun 0011, Jiaxiang Wu 0001, Peiwen Lin, Junzhou Huang, Xi Li 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | Anytime Recognition with Routing Convolutional NetworksabstractAchieving an automatic trade-off between accuracy and efficiency for a single deep neural network is highly desired in time-sensitive computer vision applications. To achieve anytime prediction, existing methods only embed fixed exits to neural networks and make the predictions with the fixed exits for all the samples (refer to the "latest-all" strategy). However, it is observed that the latest exit within a time budget does not always provide a more accurate prediction than the earlier exits for testing samples of various difficulties, making the "latest-all" strategy a sub-optimal solution. Motivated by this, we propose to improve the anytime prediction accuracy by allowing each sample to adaptively select its own optimal exit within a specific time budget. Specifically, we propose a new Routing Convolutional Network (RCN). For any given time budget, it adaptively selects the optimal layer as exit for a specific testing sample. To learn an optimal policy for sample routing, a Q-network is embedded into the RCN at each exit, considering both potential information gain and time-cost. To further boost the anytime prediction accuracy, the exits and the Q-networks are optimized alternately to mutually boost each other under the cost-sensitive environment. Apart from applying to whole image classification, RCN can also be adapted to dense prediction tasks, e.g., scene parsing, to achieve the pixel-level anytime prediction. Extensive experimental results on CIFAR-10, CIFAR-100, and ImageNet classification benchmarks, and Cityscapes scene parsing benchmark demonstrate the efficacy of the proposed RCN for anytime recognition. Zequn Jie, Peng Sun 0011, Xi Li 0001, Jiashi Feng, Wei Liu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | AD-VAT+: An Asymmetric Dueling Mechanism for Learning and Understanding Visual Active TrackingabstractVisual Active Tracking (VAT) aims at following a target object by autonomously controlling the motion system of a tracker given visual observations. To learn a robust tracker for VAT, in this article, we propose a novel adversarial reinforcement learning (RL) method which adopts an Asymmetric Dueling mechanism, referred to as AD-VAT. In the mechanism, the tracker and target, viewed as two learnable agents, are opponents and can mutually enhance each other during the dueling/competition: i.e., the tracker intends to lockup the target, while the target tries to escape from the tracker. The dueling is asymmetric in that the target is additionally fed with the tracker's observation and action, and learns to predict the tracker's reward as an auxiliary task. Such an asymmetric dueling mechanism produces a stronger target, which in turn induces a more robust tracker. To improve the performance of the tracker in the case of challenging scenarios such as obstacles, we employ more advanced environment augmentation technique and two-stage training strategies, termed as AD-VAT+. For a better understanding of the asymmetric dueling mechanism, we also analyze the target's behaviors as the training proceeds and visualize the latent space of the tracker. The experimental results, in both 2D and 3D environments, demonstrate that the proposed method leads to a faster convergence in training and yields more robust tracking behaviors in different testing scenarios. The potential of the active tracker is also shown in real-world videos. Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Graph-Guided Architecture Search for Real-Time Semantic SegmentationabstractDesigning a lightweight semantic segmentation network often requires researchers to find a trade-off between performance and speed, which is always empirical due to the limited interpretability of neural networks. In order to release researchers from these tedious mechanical trials, we propose a Graph-guided Architecture Search (GAS) pipeline to automatically search real-time semantic segmentation networks. Unlike previous works that use a simplified search space and stack a repeatable cell to form a network, we introduce a novel search mechanism with a new search space where a lightweight model can be effectively explored through the cell-level diversity and latency oriented constraint. Specifically, to produce the cell-level diversity, the cell-sharing constraint is eliminated through the cell-independent manner. Then a graph convolution network (GCN) is seamlessly integrated as a communication mechanism between cells. Finally, a latency-oriented constraint is endowed into the search process to balance the speed and performance. Extensive experiments on Cityscapes and CamVid datasets demonstrate that GAS achieves the new state-of-the-art trade-off between accuracy and speed. In particular, on Cityscapes dataset, GAS achieves the new best performance of 73.5% mIoU with the speed of 108.4 FPS on Titan Xp. Peiwen Lin, Peng Sun 0011, Sirui Xie, Xi Li 0001, Jianping Shi |
CVPR | 2 |
| 2020 | End-to-End Active Object Tracking and Its Real-World Deployment via Reinforcement LearningabstractWe study active object tracking, where a tracker takes visual observations (i.e., frame sequences) as input and produces the corresponding camera control signals as output (e.g., move forward, turn left, etc.). Conventional methods tackle tracking and camera control tasks separately, and the resulting system is difficult to tune jointly. These methods also require significant human efforts for image labeling and expensive trial-and-error system tuning in the real world. To address these issues, we propose, in this paper, an end-to-end solution via deep reinforcement learning. A ConvNet-LSTM function approximator is adopted for the direct frame-to-action prediction. We further propose an environment augmentation technique and a customized reward function, which are crucial for successful training. The tracker trained in simulators (ViZDoom and Unreal Engine) demonstrates good generalization behaviors in the case of unseen object moving paths, unseen object appearances, unseen backgrounds, and distracting objects. The system is robust and can restore tracking after occasional lost of the target being tracked. We also find that the tracking ability, obtained solely from simulators, can potentially transfer to real-world scenarios. We demonstrate successful examples of such transfer, via experiments over the VOT dataset and the deployment of a real-world robot using the proposed active tracker trained in simulation. Wenhan Luo, Peng Sun 0011, Fangwei Zhong, Wei Liu 0005, Tong Zhang 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Generative adversarial exploration for reinforcement learningabstractExploration is crucial for training the optimal reinforcement learning (RL) policy, where the key is to discriminate whether a state visiting is novel. Most previous work focuses on designing heuristic rules or distance metrics to check whether a state is novel without considering such a discrimination process that can be learned. In this paper, we propose a novel method called generative adversarial exploration (GAEX) to encourage exploration in RL via introducing an intrinsic reward output from a generative adversarial network, where the generator provides fake samples of states that help discriminator identify those less frequently visited states. Thus the agent is encouraged to visit those states which the discriminator is less confident to judge as visited. GAEX is easy to implement and of high training efficiency. In our experiments, we apply GAEX into DQN and the DQN-GAEX algorithm achieves convincing performance on challenging exploration problems, including the game Venture, Montezuma's Revenge and Super Mario Bros, without further fine-tuning on complicate learning algorithms. To our knowledge, this is the first work to employ GAN in RL exploration problems. Weijun Hong, Menghui Zhu, Minghuan Liu, Weinan Zhang 0001, Ming Zhou 0006, Yong Yu 0001, Peng Sun 0011 |
DAI | 7 |
| 2019 | AD-VAT: An Asymmetric Dueling mechanism for learning Visual Active Tracking
Fangwei Zhong, Peng Sun 0011, Wenhan Luo, Tingyun Yan, Yizhou Wang 0001 |
ICLR (Poster) | 2 |
| 2019 | Grid-Wise Control for Multi-Agent Reinforcement Learning in Video Game AIabstractWe consider the problem of multi-agent reinforcement learning (MARL) in video game AI, where the agents are located in a spatial grid-world environment and the number of agents varies both within and across episodes. The challenge is to flexibly control an arbitrary number of agents while achieving effective collaboration. Existing MARL methods usually suffer from the trade-off between these two considerations. To address the issue, we propose a novel architecture that learns a spatial joint representation of all the agents and outputs grid-wise actions. Each agent will be controlled independently by taking the action from the grid it occupies. By viewing the state information as a grid feature map, we employ a convolutional encoder-decoder as the policy network. This architecture naturally promotes agent communication because of the large receptive field provided by the stacked convolutional layers. Moreover, the spatially shared convolutional parameters enable fast parallel exploration that the experiences discovered by one agent can be immediately transferred to others. The proposed method can be conveniently integrated with general reinforcement learning algorithms, e.g., PPO and Q-learning. We demonstrate the effectiveness of the proposed method in extensive challenging multi-agent tasks in StarCraft II. Lei Han 0001, Peng Sun 0011, Yali Du 0001, Jiechao Xiong, Qing Wang 0015, Xinghai Sun, Han Liu 0001, Tong Zhang 0001 |
ICML | 2 |
| 2018 | Tagging Like Humans: Diverse and Distinct Image AnnotationabstractIn this work we propose a new automatic image annotation model, dubbed diverse and distinct image annotation (D2IA). The generative model D2IA is inspired by the ensemble of human annotations, which create semantically relevant, yet distinct and diverse tags. In D2IA, we generate a relevant and distinct tag subset, in which the tags are relevant to the image contents and semantically distinct to each other, using sequential sampling from a determinantal point process (DPP) model. Multiple such tag subsets that cover diverse semantic aspects or diverse semantic levels of the image contents are generated by randomly perturbing the DPP sampling process. We leverage a generative adversarial network (GAN) model to train D2IA. Extensive experiments including quantitative and qualitative comparisons, as well as human subject studies, on two benchmark datasets demonstrate that the proposed model can produce more diverse and distinct tags than the state-of-the-arts. Baoyuan Wu, Peng Sun 0011, Wei Liu 0005, Bernard Ghanem, Siwei Lyu |
CVPR | 3 |
| 2018 | End-to-end Active Object Tracking via Reinforcement LearningabstractWe study active object tracking, where a tracker takes as input the visual observation (i.e. frame sequence) and produces the camera control signal (e.g., move forward, turn left, etc). Conventional methods tackle the tracking and the camera control separately, which is challenging to tune jointly. It also incurs many human efforts for labeling and many expensive trial-and-errors in real-world. To address these issues, we propose, in this paper, an end-to-end solution via deep reinforcement learning, where a ConvNet-LSTM function approximator is adopted for the direct frame-to-action prediction. We further propose an environment augmentation technique and a customized reward function, which are crucial for a successful training. The tracker trained in simulators (ViZDoom, Unreal Engine) shows good generalization in the case of unseen object moving path, unseen object appearance, unseen background, and distracting object. It can restore tracking when occasionally losing the target. With the experiments over the VOT dataset, we also find that the tracking ability, obtained solely from simulators, can potentially transfer to real-world scenarios. Wenhan Luo, Peng Sun 0011, Fangwei Zhong, Wei Liu 0005, Tong Zhang 0001, Yizhou Wang 0001 |
ICML | 2 |
| 2018 | An Algorithmic Framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-Gradient MethodabstractWe propose a novel algorithmic framework of Variable Metric Over-Relaxed Hybrid Proximal Extra-gradient (VMOR-HPE) method with a global convergence guarantee for the maximal monotone operator inclusion problem. Its iteration complexities and local linear convergence rate are provided, which theoretically demonstrate that a large over-relaxed step-size contributes to accelerating the proposed VMOR-HPE as a byproduct. Specifically, we find that a large class of primal and primal-dual operator splitting algorithms are all special cases of VMOR-HPE. Hence, the proposed framework offers a new insight into these operator splitting algorithms. In addition, we apply VMOR-HPE to the Karush-Kuhn-Tucker (KKT) generalized equation of linear equality constrained multi-block composite convex optimization, yielding a new algorithm, namely nonsymmetric Proximal Alternating Direction Method of Multipliers with a preconditioned Extra-gradient step in which the preconditioned metric is generated by a blockwise Barzilai-Borwein line search technique (PADMM-EBB). We also establish iteration complexities of PADMM-EBB in terms of the KKT residual. Finally, we apply PADMM-EBB to handle the nonnegative dual graph regularized low-rank representation problem. Promising results on synthetic and real datasets corroborate the efficacy of PADMM-EBB. Li Shen 0005, Peng Sun 0011, Wei Liu 0005, Tong Zhang 0001 |
ICML | 2 |
| 2018 | Exponentially Weighted Imitation Learning for Batched Historical DataabstractWe consider deep policy learning with only batched historical trajectories. The main challenge of this problem is that the learner no longer has a simulator or ``environment oracle'' as in most reinforcement learning settings. To solve this problem, we propose a monotonic advantage reweighted imitation learning strategy that is applicable to problems with complex nonlinear function approximation and works well with hybrid (discrete and continuous) action space. The method does not rely on the knowledge of the behavior policy, thus can be used to learn from data generated by an unknown policy. Under mild conditions, our algorithm, though surprisingly simple, has a policy improvement bound and outperforms most competing methods empirically. Thorough numerical results are also provided to demonstrate the efficacy of the proposed methodology. Qing Wang 0015, Jiechao Xiong, Lei Han 0001, Peng Sun 0011, Han Liu 0001, Tong Zhang 0001 |
NeurIPS | 4 |
| 2017 | Comprehensive Modeling and Visualization of Cardiac Anatomy and Physiology from CT Imaging and Computer SimulationsabstractIn clinical cardiology, both anatomy and physiology are needed to diagnose cardiac pathologies. CT imaging and computer simulations provide valuable and complementary data for this purpose. However, it remains challenging to gain useful information from the large amount of high-dimensional diverse data. The current tools are not adequately integrated to visualize anatomic and physiologic data from a complete yet focused perspective. We introduce a new computer-aided diagnosis framework, which allows for comprehensive modeling and visualization of cardiac anatomy and physiology from CT imaging data and computer simulations, with a primary focus on ischemic heart disease. The following visual information is presented: (1) Anatomy from CT imaging: geometric modeling and visualization of cardiac anatomy, including four heart chambers, left and right ventricular outflow tracts, and coronary arteries; (2) Function from CT imaging: motion modeling, strain calculation, and visualization of four heart chambers; (3) Physiology from CT imaging: quantification and visualization of myocardial perfusion and contextual integration with coronary artery anatomy; (4) Physiology from computer simulation: computation and visualization of hemodynamics (e.g., coronary blood velocity, pressure, shear stress, and fluid forces on the vessel wall). Substantially, feedback from cardiologists have confirmed the practical utility of integrating these features for the purpose of computer-aided diagnosis of ischemic heart disease. Guanglei Xiong, Peng Sun 0011, Haoyin Zhou, Seongmin Ha, Briain o Hartaigh, Quynh A. Truong, James K. Min |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2014 | A Convergence Rate Analysis for LogitBoost, MART and Their VariantabstractLogitBoost, MART and their variant can be viewed as additive tree regression using logistic loss and boosting style optimization. We analyze their convergence rates based on a new weak learnability formulation. We show that it has O(\frac1T) rate when using gradient descent only, while a linear rate is achieved when using Newton descent. Moreover, introducing Newton descent when growing the trees, as LogitBoost does, leads to a faster linear rate. Empirical results on UCI datasets support our analysis. Peng Sun 0011, Tong Zhang 0001, Jie Zhou 0001 |
ICML | 1 |
| 2014 | An improved multiclass LogitBoost using adaptive-one-vs-one
Peng Sun 0011, Mark D. Reid, Jie Zhou 0001 |
Mach. Learn. | 1 |
| 2013 | Saving Evaluation Time for the Decision Function in Boosting: Representation and Reordering Base LearnerabstractFor a well trained Boosting classifier, we are interested in how to save the testing time, i.e., to make the decision without evaluating all the base learners. To address this problem, in previous work the base learners are sequentially calculated and early stopping is allowed if the decision function has been confident enough to output its value. In such a chain structure, the order of base learners is critical: better order can lead to less evaluation time. In this paper, we present a novel method for ordering. We base our discussion on the data structure representing Boosting’s decision function. Viewing the decision function a boolean expression, we propose a Binary Valued Tree for its representation. As a secondary contribution, such a representation unifies the work by previous researchers and helps devise new representation. Also, its connection to Binary Decision Diagram(BDD) is discussed. Peng Sun 0011, Jie Zhou 0001 |
ICML (3) | 1 |
| 2012 | The Convexity and Design of Composite Multiclass Losses
Mark D. Reid, Robert C. Williamson, Peng Sun 0011 |
ICML | 3 |
| 2012 | AOSO-LogitBoost: Adaptive One-Vs-One LogitBoost for Multi-Class Problem
Peng Sun 0011, Mark D. Reid, Jie Zhou 0001 |
ICML | 1 |