EDBT 2026 Demo / reviewers in the wild / expert
Ming Jiang 0019
dblp:02/10377-19
· DBLP profile ↗
36ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0001-6439-5476ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 6 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Defying Distractions in Multimodal Tasks: A Novel Benchmark for Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) with "multimodal distractibility," where plausible but irrelevant visual or textual inputs cause significant drops in reasoning consistency and lead to unreliable outputs. This paper introduces a comprehensive framework to systematically diagnose, evaluate, and mitigate this critical challenge. We present three core components: the large-scale IR-VQA benchmark to surface these vulnerabilities across four paradigms; novel diagnostic metrics, Positive Consistency (PC) and Negative Consistency (NC), which move beyond standard accuracy to rigorously measure a model's reasoning stability; and the Relevance-Gated Multimodal Routing (RGMR) mechanism, a novel, lightweight module that proactively and dynamically filters distractions at inference time. Our experiments reveal that state-of-the-art models exhibit significant drops in consistency on IR-VQA. We demonstrate that finetuning on IR-VQA and deploying RGMR substantially improve model robustness where standard prompting fails. Our comprehensive analysis of model behaviors under different types of distractions and the underlying reasoning failures provides a clear path forward for developing more reliable multimodal systems. Jinhui Yang, Ming Jiang 0019, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Explainable Saliency: Articulating Reasoning with Contextual PrioritizationabstractDeep saliency models, which predict what parts of an image capture our attention, are often like black boxes. This limits their use, especially in areas where understanding why a model makes a decision is crucial. Our research tackles this challenge by developing an explainable saliency (XSal) model that not only identifies what is important in an image, but also explains its choices in a way that makes sense to humans. We achieve this by using vision-language models to reason about images and by focusing the model’s attention on the most crucial information using a contextual prioritization mechanism. Unlike prior approaches that rely on fixation descriptions or soft-attention based semantic aggregation, our method directly models the reasoning steps involved in saliency prediction, generating selectively prioritized explanations clarify why specific regions are prioritized. Comprehensive evaluations demonstrate the effectiveness of our model in generating high-quality saliency maps and coherent, contextually relevant explanations. This research is a step towards more transparent and trustworthy AI systems that can help us understand and navigate the world around us. Ming Jiang 0019, Qi Zhao 0001 |
CVPR | 2 |
| 2024 | Beyond Average: Individualized Visual Scanpath PredictionabstractUnderstanding how attention varies across individuals has significant scientific and societal impacts. However, existing visual scanpath models treat attention uniformly, neglecting individual differences. To bridge this gap, this paper focuses on individualized scanpath prediction (ISP), a new attention modeling task that aims to accurately predict how different individuals shift their attention in diverse visual tasks. It proposes an ISP method featuring three novel technical components: (1) an observer encoder to characterize and integrate an observer's unique attention traits, (2) an observer-centric feature integration approach that holistically combines visual features, task guidance, and observer-specific characteristics, and (3) an adaptive fixation prioritization mechanism that refines scanpath predictions by dynamically prioritizing semantic feature maps based on individual observers' attention traits. These novel components allow scanpath models to effectively address the attention variations across different observers. Our method is generally applicable to different datasets, model architectures, and visual tasks, offering a comprehensive tool for transforming general scanpath models into individualized ones. Comprehensive evaluations using value-based and ranking-based metrics verify the method's effectiveness and generalizability. Ming Jiang 0019, Qi Zhao 0001 |
CVPR | 2 |
| 2024 | GazeXplain: Learning to Predict Natural Language Explanations of Visual Scanpaths
Ming Jiang 0019, Qi Zhao 0001 |
ECCV (8) | 2 |
| 2024 | Learning Chain of Counterfactual Thought for Bias-Robust Vision-Language Reasoning
Ming Jiang 0019, Qi Zhao 0001 |
ECCV (8) | 2 |
| 2024 | GRACE: Graph-Based Contextual Debiasing for Fair Visual Question Answering
Ming Jiang 0019, Qi Zhao 0001 |
ECCV (17) | 2 |
| 2024 | Every Problem, Every Step, All in Focus: Learning to Solve Vision-Language Problems With Integrated AttentionabstractIntegrating information from vision and language modalities has sparked interesting applications in the fields of computer vision and natural language processing. Existing methods, though promising in tasks like image captioning and visual question answering, face challenges in understanding real-life issues and offering step-by-step solutions. In particular, they typically limit their scope to solutions with a sequential structure, thus ignoring complex inter-step dependencies. To bridge this gap, we propose a graph-based approach to vision-language problem solving. It leverages a novel integrated attention mechanism that jointly considers the importance of features within each step as well as across multiple steps. Together with a graph neural network method, this attention mechanism can be progressively learned to predict sequential and non-sequential solution graphs depending on the characterization of the problem-solving process. To tightly couple attention with the problem-solving procedure, we further design new learning objectives with attention metrics that quantify this integrated attention, which better aligns visual and language information within steps, and more accurately captures information flow between steps. Experimental results on VisualHow, a comprehensive dataset of varying solution structures, show significant improvements in predicting steps and dependencies, demonstrating the effectiveness of our approach in tackling various vision-language problems. Jinhui Yang, Shi Chen 0001, Louis Wang, Ming Jiang 0019, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | What Do Deep Saliency Models Learn about Visual Attention?abstractIn recent years, deep saliency models have made significant progress in predicting human visual attention. However, the mechanisms behind their success remain largely unexplained due to the opaque nature of deep neural networks. In this paper, we present a novel analytic framework that sheds light on the implicit features learned by saliency models and provides principled interpretation and quantification of their contributions to saliency prediction. Our approach decomposes these implicit features into interpretable bases that are explicitly aligned with semantic attributes and reformulates saliency prediction as a weighted combination of probability maps connecting the bases and saliency. By applying our framework, we conduct extensive analyses from various perspectives, including the positive and negative weights of semantics, the impact of training data and architectural designs, the progressive influences of fine-tuning, and common error patterns of state-of-the-art deep saliency models. Additionally, we demonstrate the effectiveness of our framework by exploring visual attention characteristics in various application scenarios, such as the atypical attention of people with autism spectrum disorder, attention to emotion-eliciting stimuli, and attention evolution over time. Our code is publicly available at \url{https://github.com/szzexpoi/saliency_analysis}. Shi Chen 0001, Ming Jiang 0019, Qi Zhao 0001 |
NeurIPS | 2 |
| 2023 | Emotional Attention: From Eye Tracking to Computational ModelingabstractAttending selectively to emotion-eliciting stimuli is intrinsic to human vision. In this research, we investigate how emotion-elicitation features of images relate to human selective attention. We create the EMOtional attention dataset (EMOd). It is a set of diverse emotion-eliciting images, each with (1) eye-tracking data from 16 subjects, (2) image context labels at both object- and scene-level. Based on analyses of human perceptions of EMOd, we report an emotion prioritization effect: emotion-eliciting content draws stronger and earlier human attention than neutral content, but this advantage diminishes dramatically after initial fixation. We find that human attention is more focused on awe eliciting and aesthetic vehicle and animal scenes in EMOd. Aiming to model the above human attention behavior computationally, we design a deep neural network (CASNet II), which includes a channel weighting subnetwork that prioritizes emotion-eliciting objects, and an Atrous Spatial Pyramid Pooling (ASPP) structure that learns the relative importance of image regions at multiple scales. Visualizations and quantitative analyses demonstrate the model's ability to simulate human attention behavior, especially on emotion-eliciting content. Shaojing Fan, Zhiqi Shen 0002, Ming Jiang 0019, Bryan L. Koenig, Mohan Kankanhalli, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | VisualHow: Multimodal Problem SolvingabstractRecent progress in the interdisciplinary studies of computer vision (CV) and natural language processing (NLP) has enabled the development of intelligent systems that can describe what they see and answer questions accordingly. However, despite showing usefulness in performing these vision-language tasks, existing methods still struggle in understanding real-life problems (i.e., how to do something) and suggesting step-by-step guidance to solve them. With an overarching goal of developing intelligent systems to assist humans in various daily activities, we propose VisualHow, a free-form and open-ended research that focuses on understanding a real-life problem and deriving its solution by incorporating key components across multiple modalities. We develop a new dataset with 20,028 real-life problems and 102,933 steps that constitute their solutions, where each step consists of both a visual illustration and a textual description that guide the problem solving. To establish better understanding of problems and solutions, we also provide annotations of multimodal attention that localizes important components across modalities and solution graphs that encapsulate different steps in structured representations. These data and annotations enable a family of new vision-language tasks that solve real-life problems. Through extensive experiments with representative models, we demonstrate their effectiveness on training and testing models for the new tasks, and there is significant scope for improvement by learning effective attention mechanisms. Our dataset and models are available at https://github.com/formidify/VisualHow. Jinhui Yang, Ming Jiang 0019, Shi Chen 0001, Louis Wang, Qi Zhao 0001 |
CVPR | 3 |
| 2022 | Query and Attention Augmentation for Knowledge-Based Explainable ReasoningabstractExplainable visual question answering (VQA) models have been developed with neural modules and query-based knowledge incorporation to answer knowledge-requiring questions. Yet, most reasoning methods cannot effectively generate queries or incorporate external knowledge during the reasoning process, which may lead to suboptimal results. To bridge this research gap, we present Query and Attention Augmentation, a general approach that augments neural module networks to jointly reason about visual and external knowledge. To take both knowledge sources into account during reasoning, it parses the input question into a functional program with queries augmented through a novel reinforcement learning method, and jointly directs augmented attention to visual and external knowledge based on intermediate reasoning results. With extensive experiments on multiple VQA datasets, our method demonstrates significant performance, explainability, and generalizability over state-of-the-art models in answering questions requiring different extents of knowledge. Our source code is available at https://github.com/SuperJohnZhang/QAA. Ming Jiang 0019, Qi Zhao 0001 |
CVPR | 2 |
| 2022 | New Datasets and Models for Contextual Reasoning in Visual Dialog
Ming Jiang 0019, Qi Zhao 0001 |
ECCV (36) | 2 |
| 2022 | Attention in Reasoning: Dataset, Analysis, and ModelingabstractWhile attention has been an increasingly popular component in deep neural networks to both interpret and boost the performance of models, little work has examined how attention progresses to accomplish a task and whether it is reasonable. In this work, we propose an Attention with Reasoning capability (AiR) framework that uses attention to understand and improve the process leading to task outcomes. We first define an evaluation metric based on a sequence of atomic reasoning operations, enabling a quantitative measurement of attention that considers the reasoning process. We then collect human eye-tracking and answer correctness data, and analyze various machine and human attention mechanisms on their reasoning capability and how they impact task performance. To improve the attention and reasoning ability of visual question answering models, we propose to supervise the learning of attention progressively along the reasoning process and to differentiate the correct and incorrect attention patterns. We demonstrate the effectiveness of the proposed framework in analyzing and modeling attention with better reasoning capability and task performance. The code and data are available at https://github.com/szzexpoi/AiR. Shi Chen 0001, Ming Jiang 0019, Jinhui Yang, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Predicting Human Scanpaths in Visual Question AnsweringabstractAttention has been an important mechanism for both humans and computer vision systems. While state-of-the-art models to predict attention focus on estimating a static probabilistic saliency map with free-viewing behavior, real-life scenarios are filled with tasks of varying types and complexities, and visual exploration is a temporal process that contributes to task performance. To bridge the gap, we conduct a first study to understand and predict the temporal sequences of eye fixations (a.k.a. scanpaths) during performing general tasks, and examine how scanpaths affect task performance. We present a new deep reinforcement learning method to predict scanpaths leading to different performances in visual question answering. Conditioned on a task guidance map, the proposed model learns question-specific attention patterns to generate scanpaths. It addresses the exposure bias in scanpath prediction with self-critical sequence training and designs a Consistency-Divergence loss to generate distinguishable scanpaths between correct and incorrect answers. The proposed model not only accurately predicts the spatio-temporal patterns of human behavior in visual question answering, such as fixation position, duration, and order, but also generalizes to free-viewing and visual search tasks, achieving human-level performance in all tasks and significantly outperforming the state of the art. Ming Jiang 0019, Qi Zhao 0001 |
CVPR | 2 |
| 2021 | Explicit Knowledge Incorporation for Visual ReasoningabstractExisting explainable and explicit visual reasoning methods only perform reasoning based on visual evidence but do not take into account knowledge beyond what is in the visual scene. To addresses the knowledge gap between visual reasoning methods and the semantic complexity of real-world images, we present the first explicit visual reasoning method that incorporates external knowledge and models high-order relational attention for improved generalizability and explainability. Specifically, we propose a knowledge incorporation network that explicitly creates and includes new graph nodes for entities and predicates from external knowledge bases to enrich the semantics of the scene graph used in explicit reasoning. We then create a novel Graph-Relate module to perform high-order relational attention on the enriched scene graph. By explicitly introducing structured external knowledge and high-order relational attention, our method demonstrates significant generalizability and explainability over the state-of-the-art visual reasoning approaches on the GQA and VQAv2 datasets. Ming Jiang 0019, Qi Zhao 0001 |
CVPR | 2 |
| 2021 | Leveraging Human Attention in Novel Object CaptioningabstractImage captioning models depend on training with paired image-text corpora, which poses various challenges in describing images containing novel objects absent from the training data. While previous novel object captioning methods rely on external image taggers or object detectors to describe novel objects, we present the Attention-based Novel Object Captioner (ANOC) that complements novel object captioners with human attention features that characterize generally important information independent of tasks. It introduces a gating mechanism that adaptively incorporates human attention with self-learned machine attention, with a Constrained Self-Critical Sequence Training method to address the exposure bias while maintaining constraints of novel object descriptions. Extensive experiments conducted on the nocaps and Held-Out COCO datasets demonstrate that our method considerably outperforms the state-of-the-art novel object captioners. Our source code is available at https://github.com/chenxy99/ANOC. Ming Jiang 0019, Qi Zhao 0001 |
IJCAI | 2 |
| 2021 | Self-Distillation for Few-Shot Image CaptioningabstractThe development of large-scale image-captioning datasets is expensive, while the abundance of unpaired images and text corpus can potentially help reduce the efforts of manual annotation. In this paper, we study the few-shot image captioning problem that only requires a small amount of annotated image-caption pairs. We propose an ensemble- based self-distillation method that allows image captioning models to be trained with unpaired images and captions. The ensemble consists of multiple base models trained with different data samples in each iteration. For learning from unpaired images, we generate multiple pseudo captions with the ensemble and allocate different weights according to their confidence levels. For learning from unpaired captions, we propose a simple yet effective pseudo feature generation method based on Gradient Descent. The pseudo captions and pseudo features from the ensemble are used to train the base models in future iterations. The proposed method is general over different image captioning models and datasets. Our experiments demonstrate significant performance improvements and meaningful captions generated with only 1% of paired training data. Source code is available at https://github.com/chenxy99/SD-FSIC. Ming Jiang 0019, Qi Zhao 0001 |
WACV | 2 |
| 2021 | Saliency Prediction with External KnowledgeabstractThe last decades have seen great progress in saliency prediction, with the success of deep neural networks that are able to encode high-level semantics. Yet, while humans have the innate capability in leveraging their knowledge to decide where to look (e.g. people pay more attention to familiar faces such as celebrities), saliency prediction models have only been trained with large eye-tracking datasets. This work proposes to bridge this gap by explicitly incorporating external knowledge for saliency models as humans do. We develop networks that learn to highlight regions by incorporating prior knowledge of semantic relationships, be it general or domain-specific, depending on the task of interest. At the core of the method is a new Graph Semantic Saliency Network (GraSSNet) that constructs a graph that encodes semantic relationships learned from external knowledge. A Spatial Graph Attention Network is then developed to update saliency features based on the learned graph. Experiments show that the proposed model learns to predict saliency from the external knowledge and outperforms the state-of-the-art on four saliency benchmarks. Ming Jiang 0019, Qi Zhao 0001 |
WACV | 2 |
| 2020 | Fantastic Answers and Where to Find Them: Immersive Question-Directed Visual AttentionabstractWhile most visual attention studies focus on bottom-up attention with restricted field-of-view, real-life situations are filled with embodied vision tasks. The role of attention is more significant in the latter due to the information overload, and attention to the most important regions is critical to the success of tasks. The effects of visual attention on task performance in this context have also been widely ignored. This research addresses a number of challenges to bridge this research gap, on both the data and model aspects. Specifically, we introduce the first dataset of top-down attention in immersive scenes. The Immersive Question-directed Visual Attention (IQVA) dataset features visual attention and corresponding task performance (i.e., answer correctness). It consists of 975 questions and answers collected from people viewing 360° videos in a head-mounted display. Analyses of the data demonstrate a significant correlation between people's task performance and their eye movements, suggesting the role of attention in task performance. With that, a neural network is developed to encode the differences of correct and incorrect attention and jointly predict the two. The proposed attention model for the first time takes into account answer correctness, whose outputs naturally distinguish important regions from distractions. This study with new data and features may enable new tasks that leverage attention and answer correctness, and inspire new research that reveals the process behind decision making in performing various tasks. Ming Jiang 0019, Shi Chen 0001, Jinhui Yang, Qi Zhao 0001 |
CVPR | 1 |
| 2020 | AiR: Attention with Reasoning Capability
Shi Chen 0001, Ming Jiang 0019, Jinhui Yang, Qi Zhao 0001 |
ECCV (1) | 2 |
| 2020 | A structure-guided approach to the prediction of natural image saliency
Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
Neurocomputing | 2 |
| 2019 | CapVis: Toward Better Understanding of Visual-Verbal Saliency ConsistencyabstractWhen looking at an image, humans shift their attention toward interesting regions, making sequences of eye fixations. When describing an image, they also come up with simple sentences that highlight the key elements in the scene. What is the correlation between where people look and what they describe in an image? To investigate this problem intuitively, we develop a visual analytics system, CapVis, to look into visual attention and image captioning, two types of subjective annotations that are relatively task-free and natural. Using these annotations, we propose a word-weighting scheme to extract visual and verbal saliency ranks to compare against each other. In our approach, a number of low-level and semantic-level features relevant to visual-verbal saliency consistency are proposed and visualized for a better understanding of image content. Our method also shows the different ways that a human and a computational model look at and describe images, which provides reliable information for a captioning model. Experiment also shows that the visualized feature can be integrated into a computational model to effectively predict the consistency between the two modalities on an image dataset with both types of annotations. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2018 | Emotional Attention: A Study of Image Sentiment and Visual AttentionabstractImage sentiment influences visual perception. Emotion-eliciting stimuli such as happy faces and poisonous snakes are generally prioritized in human attention. However, little research has evaluated the interrelationships of image sentiment and visual saliency. In this paper, we present the first study to focus on the relation between emotional properties of an image and visual attention. We first create the EMOtional attention dataset (EMOd). It is a diverse set of emotion-eliciting images, and each image has (1) eye-tracking data collected from 16 subjects, (2) intensive image context labels including object contour, object sentiment, object semantic category, and high-level perceptual attributes such as image aesthetics and elicited emotions. We perform extensive analyses on EMOd to identify how image sentiment relates to human attention. We discover an emotion prioritization effect: for our images, emotion-eliciting content attracts human attention strongly, but such advantage diminishes dramatically after initial fixation. Aiming to model the human emotion prioritization computationally, we design a deep neural network for saliency prediction, which includes a novel subnetwork that learns the spatial and semantic context of the image scene. The proposed network outperforms the state-of-the-art on three benchmark datasets, by effectively capturing the relative importance of human attention within an image. The code, models, and dataset are available online at https://nus-sesame.top/emotionalattention/. Shaojing Fan, Zhiqi Shen 0002, Ming Jiang 0019, Bryan L. Koenig, Mohan Kankanhalli, Qi Zhao 0001 |
CVPR | 3 |
| 2018 | Image Visual Realism: From Human Perception to Machine ComputationabstractVisual realism is defined as the extent to which an image appears to people as a photo rather than computer generated. Assessing visual realism is important in applications like computer graphics rendering and photo retouching. However, current realism evaluation approaches use either labor-intensive human judgments or automated algorithms largely dependent on comparing renderings to reference images. We develop a reference-free computational framework for visual realism prediction to overcome these constraints. First, we construct a benchmark dataset of 2,520 images with comprehensive human annotated attributes. From statistical modeling on this data, we identify image attributes most relevant for visual realism. We propose both empirically-based (guided by our statistical modeling of human data) and deep convolutional neural network models to predict visual realism of images. Our framework has the following advantages: (1) it creates an interpretable and concise empirical model that characterizes human perception of visual realism; (2) it links computational features to latent factors of human image perception. Shaojing Fan, Tian-Tsong Ng, Bryan L. Koenig, Jonathan S. Herberg, Ming Jiang 0019, Zhiqi Shen 0002, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2017 | Learning Visual Attention to Identify People with Autism Spectrum DisorderabstractThis paper presents a novel method for quantitative and objective diagnoses of Autism Spectrum Disorder (ASD) using eye tracking and deep neural networks. ASD is prevalent, with 1.5% of people in the US. The lack of clinical resources for early diagnoses has been a long-lasting issue. This work differentiates itself with three unique features: first, the proposed approach is data-driven and free of assumptions, important for new discoveries in understanding ASD as well as other neurodevelopmental disorders. Second, we concentrate our analyses on the differences in eye movement patterns between healthy people and those with ASD. An image selection method based on Fisher scores allows feature learning with the most discriminative contents, leading to efficient and accurate diagnoses. Third, we leverage the recent advances in deep neural networks for both prediction and visualization. Experimental results show the superior performance of our method in terms of multiple evaluation metrics used in diagnostic tests. Ming Jiang 0019, Qi Zhao 0001 |
ICCV | 1 |
| 2017 | The Role of Visual Attention in Sentiment PredictionabstractAutomated assessment of visual sentiment has many applications, such as monitoring social media and facilitating online advertising. In current research on automated visual sentiment assessment, images are mainly input and processed as a whole. However, human attention is biased, and a focal region with high acuity can disproportionately influence visual sentiment. To investigate how attention influences visual sentiment, we conducted experiments that reveal critical insights into human perception. We discover that negative sentiments are elicited by the focal region without a notable influence of contextual information, whereas positive sentiments are influenced by both focal and contextual information. Building on these insights, we create new deep convolutional neural networks for sentiment prediction that have additional channels devoted to encoding focal information. On two benchmark datasets, the proposed models demonstrate superior performance compared with the state-of-the-art methods. Extensive visualizations and statistical analyses indicate that the focal channels are more effective on images with focal objects, especially for images that also elicit negative sentiments. Shaojing Fan, Ming Jiang 0019, Zhiqi Shen 0002, Bryan L. Koenig, Mohan Kankanhalli, Qi Zhao 0001 |
ACM Multimedia | 2 |
| 2017 | Saliency prediction with scene structural guidanceabstractPrevious works have suggested the role of scene information in directing gaze. The structure of a scene provides global contextual information that complements local object information in saliency prediction. In this study, we explore how scene envelopes such as openness, depth, and perspective affect visual attention in natural outdoor images. To facilitate this study, an eye tracking dataset is first built with 500 natural scene images and eye tracking data with 15 subjects free-viewing the images. We make observations on scene layout properties and propose a set of scene structural features relating to visual attention. We further integrate features from deep neural networks and use the set of complementary features for saliency prediction. Our features are independent of and can work together with many computational modules, and this work demonstrates the use of Multiple kernel learning (MKL) as an example to integrate the features at low- and high-levels. Experimental results demonstrate that our model outperforms existing methods and our scene structural features can improve the performance of other saliency models in outdoor scenes. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
SMC | 2 |
| 2017 | Visual-verbal consistency of image saliencyabstractWhen looking at an image, humans shift their attention towards interesting regions, making sequences of eye fixations. When describing an image, they also come up with simple sentences that highlight the key elements in the scene. What is the correlation between where people look and what they describe in an image? To investigate this problem, we look into eye fixations and image captions, two types of subjective annotations that are relatively task-free and natural. From the annotations, we extract visual and verbal saliency ranks to compare against each other. We then propose a number of low-level and semantic-level features relevant to the visualverbal consistency. Integrated into a computational model, the proposed features effectively predict the consistency between the two modalities on a large dataset with both types of annotations, namely SALICON [1]. Haoran Liang 0001, Ming Jiang 0019, Ronghua Liang, Qi Zhao 0001 |
SMC | 2 |
| 2017 | Determining child orientation from overhead video: A multiple kernel learning approachabstractOur goal is to automatically detect which direction a child is facing based on a single, simple overhead picture, and track that direction across time. Engaging in joint attention, which is the shared focus of two individuals on some object of interest, is a strong cue of typically developing children, and the lack thereof can be an indicator of autism spectrum disorder or other pervasive developmental disorder. Therefore, the goal of many psychology experiments with children is to determine when, for how long, and towards what the child looks after some bid for attention or reaction. While much research looks for the orientation of faces based on frontal or profile pictures, or non-morphable, larger objects like cars, fewer studies work in the setting of minimally-invasive overhead person gaze or orientation detection. To automatically detect the child's orientation during a human-robot interaction experiment, we mount a camera on the ceiling of a child development laboratory and analyze the video footage. We use multiple kernel learning on eight potential orientation directions to determine a child's orientation during the video recorded interaction. We also contribute the labelled dataset we used on this challenging problem. Marie D. Manner, Ming Jiang 0019, Qi Zhao 0001, Maria L. Gini, Jed T. Elison |
SMC | 2 |
| 2016 | A Paradigm for Building Generalized Models of Human Image Perception through Data FusionabstractIn many sub-fields, researchers collect datasets of human ground truth that are used to create a new algorithm. For example, in research on image perception, datasets have been collected for topics such as what makes an image aesthetic or memorable. Despite high costs for human data collection, datasets are infrequently reused beyond their own fields of interest. Moreover, the algorithms built from them are domain-specific (predict a small set of attributes) and usually unconnected to one another. In this paper, we present a paradigm for building generalized and expandable models of human image perception. First, we fuse multiple fragmented and partially-overlapping datasets through data imputation. We then create a theoretically-structured statistical model of human image perception that is fit to the fused datasets. The resulting model has many advantages. (1) It is generalized, going beyond the content of the constituent datasets, and can be easily expanded by fusing additional datasets. (2) It provides a new ontology usable as a network to expand human data in a cost-effective way. (3) It can guide the design of a generalized computational algorithm for multi-dimensional visual perception. Indeed, experimental results show that a model-based algorithm outperforms state-of-the-art methods on predicting visual sentiment, visual realism and interestingness. Our paradigm can be used in various visual tasks (e.g., video summarization). Shaojing Fan, Tian-Tsong Ng, Bryan L. Koenig, Ming Jiang 0019, Qi Zhao 0001 |
CVPR | 4 |
| 2016 | Learning to Predict Sequences of Human Visual FixationsabstractMost state-of-the-art visual attention models estimate the probability distribution of fixating the eyes in a location of the image, the so-called saliency maps. Yet, these models do not predict the temporal sequence of eye fixations, which may be valuable for better predicting the human eye fixations, as well as for understanding the role of the different cues during visual exploration. In this paper, we present a method for predicting the sequence of human eye fixations, which is learned from the recorded human eye-tracking data. We use least-squares policy iteration (LSPI) to learn a visual exploration policy that mimics the recorded eye-fixation examples. The model uses a different set of parameters for the different stages of visual exploration that capture the importance of the cues during the scanpath. In a series of experiments, we demonstrate the effectiveness of using LSPI for combining multiple cues at different stages of the scanpath. The learned parameters suggest that the low-level and high-level cues (semantics) are similarly important at the first eye fixation of the scanpath, and the contribution of high-level cues keeps increasing during the visual exploration. Results show that our approach obtains the state-of-the-art performances on two challenging data sets: 1) OSIE data set and 2) MIT data set. Ming Jiang 0019, Xavier Boix, Gemma Roig, Luc Van Gool, Qi Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Saliency Prediction with Active Semantic SegmentationabstractMing Jiang1 [email protected] Xavier Boix1,3 [email protected] Juan Xu1 [email protected] Gemma Roig2,3 [email protected] Luc Van Gool3 [email protected] Qi Zhao1 [email protected] 1 Department of Electrical and Computer Engineering National University of Singapore Singapore 2 CBMM, LCSL Massachusetts Institute of Technology Istituto Italiano di Tecnologia Cambridge, MA 3 Computer Vision Laboratory ETH Zurich Switzerland Ming Jiang 0019, Xavier Boix, Gemma Roig, Luc Van Gool, Qi Zhao 0001 |
BMVC | 1 |
| 2015 | SALICON: Saliency in ContextabstractSaliency in Context (SALICON) is an ongoing effort that aims at understanding and predicting visual attention. This paper presents a new method to collect large-scale human data during natural explorations on images. While current datasets present a rich set of images and task-specific annotations such as category labels and object segments, this work focuses on recording and logging how humans shift their attention during visual exploration. The goal is to offer new possibilities to (1) complement task-specific annotations to advance the ultimate goal in visual understanding, and (2) understand visual attention and learn saliency models, all with human attentional data at a much larger scale. We designed a mouse-contingent multi-resolutional paradigm based on neurophysiological and psychophysical studies of peripheral vision, to simulate the natural viewing behavior of humans. The new paradigm allowed using a general-purpose mouse instead of an eye tracker to record viewing behaviors, thus enabling large-scale data collection. The paradigm was validated with controlled laboratory as well as large-scale online data. We report in this paper a proof-of-concept SALICON dataset of human “free-viewing” data on 10,000 images from the Microsoft COCO (MS COCO) dataset with rich contextual information. We evaluated the use of the collected data in the context of saliency prediction, and demonstrated them a good source as ground truth for the evaluation of saliency algorithms. Ming Jiang 0019, Shengsheng Huang, Juanyong Duan, Qi Zhao 0001 |
CVPR | 1 |
| 2015 | Multi-Camera SaliencyabstractA significant body of literature on saliency modeling predicts where humans look in a single image or video. Besides the scientific goal of understanding how information is fused from multiple visual sources to identify regions of interest in a holistic manner, there are tremendous engineering applications of multi-camera saliency due to the widespread of cameras. This paper proposes a principled framework to smoothly integrate visual information from multiple views to a global scene map, and to employ a saliency algorithm incorporating high-level features to identify the most important regions by fusing visual information. The proposed method has the following key distinguishing features compared with its counterparts: (1) the proposed saliency detection is global (salient regions from one local view may not be important in a global context), (2) it does not require special ways for camera deployment or overlapping field of view, and (3) the key saliency algorithm is effective in highlighting interesting object regions though not a single detector is used. Experiments on several data sets confirm the effectiveness of the proposed principled framework. Yan Luo 0002, Ming Jiang 0019, Yongkang Wong, Qi Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Saliency in Crowd
Ming Jiang 0019, Qi Zhao 0001 |
ECCV (7) | 1 |
| 2013 | Leveraging Human Fixations in Sparse Coding: Learning a Discriminative Dictionary for Saliency PredictionabstractThis paper proposes to learn a discriminative dictionary for saliency detection. In addition to the conventional sparse coding mechanism that learns a representational dictionary of natural images for saliency prediction, this work uses supervised information from eye tracking experiments in training to enhance the discriminative power of the learned dictionary. Furthermore, we explicitly model saliency at multi-scale by formulating it as a multi-class problem, and a label consistency term is incorporated into the framework to encourage class (salient vs. non-salient) and scale consistency in the learned sparse codes. K-SVD is employed as the central computational module to efficiently obtain the optimal solution. Experiments demonstrate the superior performance of the proposed algorithm compared with the state-of-the-art in saliency prediction. Ming Jiang 0019, Mingli Song, Qi Zhao 0001 |
SMC | 1 |