Haonan Yu

dblp:97/8693 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 10 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
18 papers
Reinforcement learning · 39% Trustworthy machine learning · 17% Vision and language · 12%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 52% Image and video processing · 48%

Topics — the 30 heaviest of 42, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
interpretability
1.012026
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › interpretability › post-hoc explanation
model-agnostic explanation
1.012026
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models · ACL (1) 2026
Machine learning › Trustworthy machine learning › interpretability
post-hoc explanation
1.012026
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models · ACL (1) 2026
Computer vision › Vision and language
grounded language learning
0.832018
Interactive Grounded Language Acquisition and Generalization in a 2D World · ICLR (Poster) 2018
Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game · ACL (1) 2018
Grounded Language Learning from Video Described with Sentences · ACL (1) 2013
Machine learning › Efficient and distributed learning › model compression › sparse training
lottery ticket hypothesis
0.822020
Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP · ICLR 2020
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers · NeurIPS 2019
Machine learning › Representation and self-supervised learning › representation learning
object-centric representation learning
0.812024
VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAE · ICLR 2024
Machine learning › Reinforcement learning › offline reinforcement learning
offline-to-online reinforcement learning
0.712023
Policy Expansion for Bridging Offline-to-Online Reinforcement Learning · ICLR 2023
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.612022
Towards Safe Reinforcement Learning with a Safety Editor Policy · NeurIPS 2022
Machine learning › Reinforcement learning
exploration
0.612022
Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning
safe reinforcement learning
0.612022
Towards Safe Reinforcement Learning with a Safety Editor Policy · NeurIPS 2022
Computer vision › Vision and language
video captioning
0.532016
Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks · CVPR 2016
Learning to Describe Video with Weak Supervision by Exploiting Negative Sentential Information · AAAI 2015
Grounded Language Learning from Video Described with Sentences · ACL (1) 2013
Machine learning › Reinforcement learning › action space design
action repetition
0.512021
TAAC: Temporally Abstract Actor-Critic for Continuous Control · NeurIPS 2021
Machine learning › Reinforcement learning
actor-critic methods
0.512021
TAAC: Temporally Abstract Actor-Critic for Continuous Control · NeurIPS 2021
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.512021
Hierarchical Reinforcement Learning by Discovering Intrinsic Options · ICLR 2021
Machine learning › Reinforcement learning
off-policy reinforcement learning
0.512021
TAAC: Temporally Abstract Actor-Critic for Continuous Control · NeurIPS 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning
option discovery
0.512021
Hierarchical Reinforcement Learning by Discovering Intrinsic Options · ICLR 2021
Machine learning › Reinforcement learning › hierarchical reinforcement learning
temporal abstraction
0.512021
TAAC: Temporally Abstract Actor-Critic for Continuous Control · NeurIPS 2021
Machine learning › Efficient and distributed learning
model compression
0.412019
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers · NeurIPS 2019
Visual content generation and editing › 3d content generation
3d generative modeling
0.412019
Order-Aware Generative Modeling Using the 3D-Craft Dataset · ICCV 2019
Natural language and speech › Language models and text generation › language acquisition
interactive language learning
0.312018
Interactive Grounded Language Acquisition and Generalization in a 2D World · ICLR (Poster) 2018
Computer vision › Vision and language › visual grounding
language grounding
0.312018
Interactive Grounded Language Acquisition and Generalization in a 2D World · ICLR (Poster) 2018
Computer vision › Image recognition and object detection
visual concept learning
0.312018
Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game · ACL (1) 2018
Natural language and speech › Language models and text generation › prompting
prompt engineering
0.312026
Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models · ACL (1) 2026
Computer vision › Video understanding and tracking › video analytics › video object analysis › object-centric video understanding
video object discovery
0.312017
Sentence Directed Video Object Codiscovery · Int. J. Comput. Vis. 2017
Machine learning › Deep learning architectures and training › recurrent neural network › deep recurrent network
hierarchical recurrent neural network
0.212016
Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks · CVPR 2016
Computer vision › Vision and language › video captioning
video paragraph captioning
0.212016
Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks · CVPR 2016
Machine learning › Learning paradigms
weakly supervised learning
0.212015
Learning to Describe Video with Weak Supervision by Exploiting Negative Sentential Information · AAAI 2015
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan generation
generative planning
0.212022
Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning · ICLR 2022
Machine learning › Reinforcement learning
model-free reinforcement learning
0.212022
Towards Safe Reinforcement Learning with a Safety Editor Policy · NeurIPS 2022
Computer vision › Video understanding and tracking › activity recognition
human activity recognition
0.212013
Recognize Human Activities from Partially Observed Videos · CVPR 2013

Methods — techniques the papers use, named apart from their topics

imitation learning · 1.1screen-and-apply verification · 1.0proxy model · 1.0u-net attention · 0.8transformer decoder · 0.8sequential VAE · 0.8policy expansion · 0.7temporal coordination · 0.6hinge loss · 0.6generative planning · 0.6VoxelCNN · 0.4saliency detection · 0.1image segmentation · 0.1sketch-like and envelope-like saliency maps · 0.1classification · 0.1
YearPublicationVenuePosition
2026 Revitalizing Black-Box Interpretability: Actionable Interpretability for LLMs via Proxy Models
abstract
Post-hoc explanations provide transparency and are essential for guiding model optimization, such as prompt engineering and data sanitation.However, applying model-agnostic techniques to Large Language Models (LLMs) is hindered by prohibitive computational costs, rendering these tools dormant for real-world applications.To revitalize model-agnostic interpretability, we propose a budget-friendly proxy framework that leverages efficient models to approximate the decision boundaries of expensive LLMs.We introduce a screen-and-apply mechanism to statistically verify local alignment before deployment.Our empirical evaluation confirms that proxy explanations achieve over 90% fidelity with only 9.5% of the oracle's cost.Building on this foundation, we demonstrate the actionable utility of our framework in prompt compression and poisoned example removal.Results show that reliable proxy explanations effectively guide optimization, transforming interpretability from a passive observation tool into a scalable primitive for LLM development.Additionally, we open-source code and datasets to facilitate future research 1 .
Junhao Liu 0002, Haonan Yu, Xin Zhang 0035
ACL (1)2
2024 VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAE
abstract
Unsupervised video object learning seeks to decompose video scenes into structural object representations without any supervision from depth, optical flow, or segmentation. We present VONet, an innovative approach that is inspired by MONet. While utilizing a U-Net architecture, VONet employs an efficient and effective parallel attention inference process, generating attention masks for all slots simultaneously. Additionally, to enhance the temporal consistency of each mask across consecutive video frames, VONet develops an object-wise sequential VAE framework. The integration of these innovative encoder-side techniques, in conjunction with an expressive transformer-based decoder, establishes VONet as the leading unsupervised method for object learning across five MOVI datasets, encompassing videos of diverse complexities. Code is available at https://github.com/hnyu/vonet.
Haonan Yu, Wei Xu 0017
ICLR1
2023 Policy Expansion for Bridging Offline-to-Online Reinforcement Learning
Haichao Zhang 0001, Wei Xu 0017, Haonan Yu
ICLR3
2022 Generative Planning for Temporally Coordinated Exploration in Reinforcement Learning
Haichao Zhang 0001, Wei Xu 0017, Haonan Yu
ICLR3
2022 Towards Safe Reinforcement Learning with a Safety Editor Policy
abstract
We consider the safe reinforcement learning (RL) problem of maximizing utility with extremely low constraint violation rates. Assuming no prior knowledge or pre-training of the environment safety model given a task, an agent has to learn, via exploration, which states and actions are safe. A popular approach in this line of research is to combine a model-free RL algorithm with the Lagrangian method to adjust the weight of the constraint reward relative to the utility reward dynamically. It relies on a single policy to handle the conflict between utility and constraint rewards, which is often challenging. We present SEditor, a two-policy approach that learns a safety editor policy transforming potentially unsafe actions proposed by a utility maximizer policy into safe ones. The safety editor is trained to maximize the constraint reward while minimizing a hinge loss of the utility state-action values before and after an action is edited. SEditor extends existing safety layer designs that assume simplified safety models, to general safe RL scenarios where the safety model can in theory be arbitrarily complex. As a first-order method, it is easy to implement and efficient for both inference and training. On 12 Safety Gym tasks and 2 safe racing tasks, SEditor obtains much a higher overall safety-weighted-utility (SWU) score than the baselines, and demonstrates outstanding utility performance with constraint violation rates as low as once per 2k time steps, even in obstacle-dense environments. On some tasks, this low violation rate is up to 200 times lower than that of an unconstrained RL method with similar utility performance. Code is available at https://github.com/hnyu/seditor.
Haonan Yu, Wei Xu 0017, Haichao Zhang 0001
NeurIPS1
2021 Hierarchical Reinforcement Learning by Discovering Intrinsic Options
Jesse Zhang, Haonan Yu, Wei Xu 0017
ICLR2
2021 TAAC: Temporally Abstract Actor-Critic for Continuous Control
abstract
We present temporally abstract actor-critic (TAAC), a simple but effective off-policy RL algorithm that incorporates closed-loop temporal abstraction into the actor-critic framework. TAAC adds a second-stage binary policy to choose between the previous action and a new action output by an actor. Crucially, its "act-or-repeat" decision hinges on the actually sampled action instead of the expected behavior of the actor. This post-acting switching scheme let the overall policy make more informed decisions. TAAC has two important features: a) persistent exploration, and b) a new compare-through Q operator for multi-step TD backup, specially tailored to the action repetition scenario. We demonstrate TAAC's advantages over several strong baselines across 14 continuous control tasks. Our surprising finding reveals that while achieving top performance, TAAC is able to "mine" a significant number of repeated actions with the trained policy even on continuous tasks whose problem structures on the surface seem to repel action repetition. This suggests that aside from encouraging persistent exploration, action repetition can find its place in a good policy behavior. Code is available at https://github.com/hnyu/taac.
Haonan Yu, Wei Xu 0017, Haichao Zhang 0001
NeurIPS1
2021 Interactive Hepatic Parenchymal Transection Simulation with Haptic Feedback
abstract
Liver resection involves surgical removal of a portion of the liver. It is used to treat liver tumors and liver injuries. The complexity and high-risk nature of this surgery prevents novice doctors from practicing it on real patients. Virtual surgery simulation was developed to simulate surgical procedures to enable medical professionals to be trained without requiring a patient, a cadaver, or an animal. Therefore, there is a strong need for the development of a liver resection surgery simulation system. We propose a real-time simulation system that provides realistic visual and tactile feedback for hepatic parenchymal transection. The tetrahedron structure and cluster-based shape matching are used for physical model construction, topology update of a three-dimensional liver model soft deformation simulation, and haptic rendering acceleration. During the liver parenchyma separation simulation, a tetrahedral mesh is used for surface triangle subdivision and surface generation of the surgical wound. The shape-matching cluster is separated via component detection on an undirected graph constructed using the tetrahedral mesh. In our system, cluster-based shape matching is implemented on a GPU, whereas haptic rendering and topology updates are implemented on a CPU. Experimental results show that haptic rendering can be performed at a high frequency (> 900 Hz), whereas mesh skinning and graphics rendering can be performed at 45 fps. The topology update can be executed at an interactive rate (> 10 Hz) on a single CPU thread. We propose an interactive hepatic parenchymal transection simulation method based on a tetrahedral structure. The tetrahedral mesh simultaneously supports physical model construction, topology update, and haptic rendering acceleration.
Haonan Yu, Aimin Hao
Virtual Real. Intell. Hardw.2
2020 Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP
Haonan Yu, Sergey Edunov, Yuandong Tian, Ari S. Morcos
ICLR1
2020 QTMS: A quadratic time complexity topology-aware process mapping method for large-scale parallel applications on shared HPC system
Baicheng Yan, Limin Xiao 0002, Guangjun Qin, Bin Dong 0004, Haonan Yu
Parallel Comput.6
2019 Order-Aware Generative Modeling Using the 3D-Craft Dataset
abstract
Research on 2D and 3D generative models typically focuses on the final artifact being created, e.g., an image or a 3D structure. Unlike 2D image generation, the generation of 3D objects in the real world is commonly constrained by the process and order in which the object is constructed. For instance, gravity needs to be taken into account when building a block tower. In this paper, we explore the prediction of ordered actions to construct 3D objects. Instead of predicting actions based on physical constraints, we propose learning through observing human actions. To enable large-scale data collection, we use the Minecraft1 environment. We introduce 3D-Craft, a new dataset of 2,500 Minecraft houses each built by human players sequentially from scratch. To learn from these human action sequences, we propose an order-aware 3D generative model called VoxelCNN. In contrast to other 3D generative models which either have no explicit order (e.g. holistic generation with 3DGAN [35]), or follow a simple heuristic order (e.g. raster-scan), VoxelCNN is trained to imitate human building order with spatial awareness. We also transferred the order to other dataset such as ShapeNet[10]. The 3D-Craft dataset, models, and benchmark system will be made publicly available, which may inspire new directions for future research exploration. https://github.com/facebookresearch/VoxelCNN.
Zhuoyuan Chen, Kavya Srinet, Charles R. Qi, Haoqi Fan 0001, Jerry Ma, C. Lawrence Zitnick, Demi Guo, Tong Xiao 0003, Saining Xie, Xinlei Chen, Arthur Szlam, Shubham Tulsiani, Haonan Yu, Jonathan Gray
ICCV13
2019 One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
abstract
The success of lottery ticket initializations (Frankle and Carbin, 2019) suggests that small, sparsified networks can be trained so long as the network is initialized appropriately. Unfortunately, finding these "winning ticket'' initializations is computationally expensive. One potential solution is to reuse the same winning tickets across a variety of datasets and optimizers. However, the generality of winning ticket initializations remains unclear. Here, we attempt to answer this question by generating winning tickets for one training configuration (optimizer and dataset) and evaluating their performance on another configuration. Perhaps surprisingly, we found that, within the natural images domain, winning ticket initializations generalized across a variety of datasets, including Fashion MNIST, SVHN, CIFAR-10/100, ImageNet, and Places365, often achieving performance close to that of winning tickets generated on the same dataset. Moreover, winning tickets generated using larger datasets consistently transferred better than those generated using smaller datasets. We also found that winning ticket initializations generalize across optimizers with high performance. These results suggest that winning ticket initializations generated by sufficiently large datasets contain inductive biases generic to neural networks more broadly which improve training across many settings and provide hope for the development of better initialization methods.
Ari S. Morcos, Haonan Yu, Michela Paganini, Yuandong Tian
NeurIPS2
2018 Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game
abstract
Building intelligent agents that can communicate with and learn from humans in natural language is of great value.Supervised language learning is limited by the ability of capturing mainly the statistics of training data, and is hardly adaptive to new scenarios or flexible for acquiring new knowledge without inefficient retraining or catastrophic forgetting.We highlight the perspective that conversational interaction serves as a natural interface both for language learning and for novel knowledge acquisition and propose a joint imitation and reinforcement approach for grounded language learning through an interactive conversational game.The agent trained with this approach is able to actively acquire information by asking questions about novel objects and use the justlearned knowledge in subsequent conversations in a one-shot fashion.Results compared with other methods verified the effectiveness of the proposed approach.
Haichao Zhang 0001, Haonan Yu, Wei Xu 0017
ACL (1)2
2018 Interactive Grounded Language Acquisition and Generalization in a 2D World
Haonan Yu, Haichao Zhang 0001, Wei Xu 0017
ICLR (Poster)1
2018 Driving Under the Influence (of Language)
abstract
We present a unified framework which supports grounding natural-language semantics in robotic driving. This framework supports acquisition (learning grounded meanings of nouns and prepositions from human sentential annotation of robotic driving paths), generation (using such acquired meanings to generate sentential description of new robotic driving paths), and comprehension (using such acquired meanings to support automated driving to accomplish navigational goals specified in natural language). We evaluate the performance of these three tasks by having independent human judges rate the semantic fidelity of the sentences associated with paths. Overall, machine performance is 74.9%, while the performance of human annotators is 83.8%.
Daniel Paul Barrett, Scott Alan Bronikowski, Haonan Yu, Jeffrey Mark Siskind
IEEE Trans. Neural Networks Learn. Syst.3
2017 Sentence Directed Video Object Codiscovery
abstract
Video object codiscovery can leverage the weak semantic constraint implied by sentences that describe the video content. Our codiscovery method, like other object codetection techniques, does not employ any pretrained object models or detectors. Unlike most prior work that focuses on codetecting large objects which are usually salient both in size and appearance, our method can discover small or medium sized objects as well as ones that may be occluded for part of the video. More importantly, our method can codiscover multiple object instances of different classes within a single video clip. Although the semantic information employed is usually simple and weak, it can greatly boost performance by constraining the hypothesized object locations. Experiments show promising results on three datasets: an average IoU score of 0.423 on a new dataset with 15 object classes, an average IoU score of 0.373 on a subset of CAD-120 with 5 object classes, and an average IoU score of 0.358 on a subset of MPII-Cooking with 7 object classes. Our result on this subset of MPII-Cooking improves upon those of the previous state-of-the-art methods by significant margins.
Haonan Yu, Jeffrey Mark Siskind
Int. J. Comput. Vis.1
2016 Video Paragraph Captioning Using Hierarchical Recurrent Neural Networks
abstract
We present an approach that exploits hierarchical Recurrent Neural Networks (RNNs) to tackle the video captioning problem, i.e., generating one or multiple sentences to describe a realistic video. Our hierarchical framework contains a sentence generator and a paragraph generator. The sentence generator produces one simple short sentence that describes a specific short video interval. It exploits both temporal-and spatial-attention mechanisms to selectively focus on visual elements during generation. The paragraph generator captures the inter-sentence dependency by taking as input the sentential embedding produced by the sentence generator, combining it with the paragraph history, and outputting the new initial state for the sentence generator. We evaluate our approach on two large-scale benchmark datasets: YouTubeClips and TACoS-MultiLevel. The experiments demonstrate that our approach significantly outperforms the current state-of-the-art methods with BLEU@4 scores 0.499 and 0.305 respectively.
Haonan Yu, Jiang Wang 0001, Zhiheng Huang, Yi Yang 0007, Wei Xu 0017
CVPR1
2016 Collecting and annotating the large continuous action dataset
abstract
We make available to the community a new dataset to support action recognition research. This dataset is different from prior datasets in several key ways. It is significantly larger. It contains streaming video with long segments containing multiple action occurrences that often overlap in space and/or time. All actions were filmed in the same collection of backgrounds so that background gives little clue as to action class. We had five humans to replicate the annotation of temporal extent of action occurrences labeled with their classes and measured a surprisingly low level of intercoder agreement. Baseline experiments show that recent state-of-the-art methods perform poorly on this dataset. This suggests that this will be a challenging dataset to foster advances in action recognition research. This manuscript serves to describe the novel content and characteristics of the LCA dataset, present the design decisions made when filming the dataset, document the novel methods employed to annotate the dataset, and present the results of our baseline experiments.
Daniel Paul Barrett, Haonan Yu, Jeffrey Mark Siskind
Mach. Vis. Appl.3
2015 Learning to Describe Video with Weak Supervision by Exploiting Negative Sentential Information
abstract
Most previous work on video description trains individualparts of speech independently. It is more appealing from a linguistic point of view, for word models for all parts of speech to be learned simultaneously from whole sentences, a hypothesis suggested by some linguists for child language acquisition. In this paper, we learn to describe video by discriminatively training positive sentential labels against negative ones in a weakly supervised fashion: the meaning representations (i.e., HMMs) of individual words in these labels are learned from whole sentences without any correspondence annotation of what those words denote in the video. Textual descriptions are then generated for new video using trained word models.
Haonan Yu, Jeffrey Mark Siskind
AAAI1
2015 A Compositional Framework for Grounding Language Inference, Generation, and Acquisition in Video
abstract
We present an approach to simultaneously reasoning about a video clip and an entire natural-language sentence. The compositional nature of language is exploited to construct models which represent the meanings of entire sentences composed out of the meanings of the words in those sentences mediated by a grammar that encodes the predicate-argument relations. We demonstrate that these models faithfully represent the meanings of sentences and are sensitive to how the roles played by participants (nouns), their characteristics (adjectives), the actions performed (verbs), the manner of such actions (adverbs), and changing spatial relations between participants (prepositions) affect the meaning of a sentence and how it is grounded in video. We exploit this methodology in three ways. In the first, a video clip along with a sentence are taken as input and the participants in the event described by the sentence are highlighted, even when the clip depicts multiple similar simultaneous events. In the second, a video clip is taken as input without a sentence and a sentence is generated that describes an event in that clip. In the third, a corpus of video clips is paired with sentences which describe some of the events in those clips and the meanings of the words in those sentences are learned. We learn these meanings without needing to specify which attribute of the video clips each word in a given sentence refers to. The learned meaning representations are shown to be intelligible to humans.
Haonan Yu, N. Siddharth 0001, Andrei Barbu, Jeffrey Mark Siskind
J. Artif. Intell. Res.1
2013 Grounded Language Learning from Video Described with Sentences
Haonan Yu, Jeffrey Mark Siskind
ACL (1)1
2013 Recognize Human Activities from Partially Observed Videos
abstract
Recognizing human activities in partially observed videos is a challenging problem and has many practical applications. When the unobserved subsequence is at the end of the video, the problem is reduced to activity prediction from unfinished activity streaming, which has been studied by many researchers. However, in the general case, an unobserved subsequence may occur at any time by yielding a temporal gap in the video. In this paper, we propose a new method that can recognize human activities from partially observed videos in the general case. Specifically, we formulate the problem into a probabilistic framework: 1) dividing each activity into multiple ordered temporal segments, 2) using spatiotemporal features of the training video samples in each segment as bases and applying sparse coding (SC) to derive the activity likelihood of the test video sample at each segment, and 3) finally combining the likelihood at each segment to achieve a global posterior for the activities. We further extend the proposed method to include more bases that correspond to a mixture of segments with different temporal lengths (MSSC), which can better represent the activities with large intra-class variations. We evaluate the proposed methods (SC and MSSC) on various real videos. We also evaluate the proposed methods on two special cases: 1) activity prediction where the unobserved subsequence is at the end of the video, and 2) human activity recognition on fully observed videos. Experimental results show that the proposed methods outperform existing state-of-the-art comparison methods.
Yu Cao 0003, Daniel Paul Barrett, Andrei Barbu, N. Siddharth 0001, Haonan Yu, Aaron Michaux, Yuewei Lin, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002
CVPR5
2011 Salient region detection and segmentation for general object recognition and image understanding
Tiejun Huang 0001, Yonghong Tian 0001, Jia Li 0003, Haonan Yu
Sci. China Inf. Sci.4
2010 Automatic interesting object extraction from images using complementary saliency maps
abstract
Automatic interesting object extraction is widely used in many image applications. Among various extraction approaches, saliency-based ones usually have a better performance since they well accord with human visual perception. However, nearly all existing saliency-based approaches suffer the integrity problem, namely, the extracted result is either a small part of the object (referred to as sketch-like) or a large region that contains some redundant part of the background (referred to as envelope-like). In this paper, we propose a novel object extraction approach by integrating two kinds of "complementary" saliency maps (i.e., sketch-like and envelope-like maps). In our approach, the extraction process is decomposed into two sub-processes, one used to extract a high-precision result based on the sketch-like map, and the other used to extract a high-recall result based on the envelope-like map. Then a classification step is used to extract an exact object based on the two results. By transferring the complex extraction task to an easier classification problem, our approach can effectively break down the integrity problem. Experimental results show that the proposed approach outperforms six state-of-art saliency-based methods remarkably in automatic object extraction, and is even comparable to some interactive approaches.
Haonan Yu, Jia Li 0003, Yonghong Tian 0001, Tiejun Huang 0001
ACM Multimedia1