VLDB 2026 Research / reviewers in the wild / expert
Xuming He 0001
dblp:03/4230 · also Xu-Ming He 0001
· DBLP profile ↗
121ranked-venue papers
6as first author
55since 2021 · last 2025
0000-0003-2150-1237ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 5 first-author · 35 since 2021Artificial intelligence and machine learning · 87 · 6 first-author · 43 since 2021Systems, architecture and hardware · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FastGrasp: Efficient Grasp Synthesis with DiffusionabstractEffectively modeling the interaction between human hands and objects is challenging due to the complex physical constraints and the requirement for high generation efficiency in applications. Prior approaches often employ computationally intensive two-stage approaches, which first generate an intermediate representation, such as contact maps, followed by an iterative optimization procedure that updates hand meshes to capture the hand-object relation. However, due to the high computation complexity during the optimization stage, such strategies often suffer from low efficiency in inference. To address this limitation, this work introduces a novel diffusion-modelbased approach that generates the grasping pose in a one-stage manner. This allows us to significantly improve generation speed and the diversity of generated hand poses. In particular, we develop a Latent Diffusion Model with an Adaptation Module for object-conditioned hand pose generation and a contact-aware loss to enforce the physical constraints between hands and objects. Extensive experiments demonstrate that our method achieves faster inference, higher diversity, and superior pose quality than state-of-the-art approaches. Code is available at https://github.com/wuxiaofei01/FastGrasp. Caoji Li, Yuexin Ma, Yujiao Shi 0002, Xuming He 0001 |
3DV | 6 |
| 2025 | Relation-aware Hierarchical Prompt for Open-vocabulary Scene Graph GenerationabstractOpen-vocabulary Scene Graph Generation (OV-SGG) overcomes the limitations of the closed-set assumption by aligning visual relationship representations with open-vocabulary textual representations. This enables the identification of novel visual relationships, making it applicable to real-world scenarios with diverse relationships. However, existing OV-SGG methods are constrained by fixed text representations, limiting diversity and accuracy in image-text alignment. To address these challenges, we propose the Relation-Aware Hierarchical Prompting (RAHP) framework, which enhances text representation by integrating subject-object and region-specific relation information. Our approach utilizes entity clustering to address the complexity of relation triplet categories, enabling the effective integration of subject-object information. Additionally, we utilize a large language model (LLM) to generate detailed region-aware prompts, capturing fine-grained visual interactions and improving alignment between visual and textual modalities. RAHP also introduces a dynamic selection mechanism within Vision-Language Models (VLMs), which adaptively selects relevant text prompts based on the visual content, reducing noise from irrelevant prompts. Extensive experiments on the Visual Genome and Open Images v6 datasets demonstrate that our framework consistently achieves state-of-the-art performance, demonstrating its effectiveness in addressing the challenges of open-vocabulary scene graph generation. Rongjie Li, Chongyu Wang, Xuming He 0001 |
AAAI | 4 |
| 2025 | LVM-MO: A Large Vision Model Pioneer on Full-Chip Mask OptimizationabstractMoving toward the post-Moore era, full-chip mask optimization (MO) has become a pivotal step for semiconductor designers and manufacturers in extending current resolution enhancement techniques. The majority of recent research efforts have focused on clip-level restoration, employing a divide-and-conquer approach to mitigate the impacts of optical proximity and process bias across entire chips. Nevertheless, when confronted with industrial full-chip mask optimization challenges, these works exhibit limited correction capabilities, struggle with generalization, and are time-inefficient. In this paper, we propose a novel full-chip mask optimization paradigm based on a massive lithography data-driven large vision model. Our approach features a foundation layout feature extractor, which is aware of the mutual influence of polygons in long-range pattern perception as well as optical physics and chemical characteristics of lithography, matters. Compared with state-of-the-art (SOTA) works, our work demonstrates significant advantages in terms of resolution fidelity, correction speed, and the ability to handle full-chip scale layouts. Xuming He 0001, Hao Geng, Jingyi Yu 0001 |
DAC | 6 |
| 2025 | LMLitho: A Large Vision Model-Driven Lithography Simulation FrameworkabstractAs IC fabrication advances toward smaller process nodes, design technology co-optimization (DTCO) has emerged as a critical enabler of chip performance advancements. Lithography simulation, vital for bridging design and manufacturing, now plays an indispensable role in designing litho-friendly layouts/masks and developing resolution enhancement techniques (RETs). While academia and industry have explored statistical techniques and machine learning models for simulators, the computing paradigm and hardware prevent these solutions from efficiently and accurately simulating the complicated optical imaging coupled with resist film imaging. In this paper, we propose a new simulation paradigm: LMLitho (large vision model-driven lithography simulator), trained on circa one hundred thousand triplets of illumination maps, masks, and resist images. The cross-attention mechanism in our simulator inherently captures diffraction patterns akin to light wave interference within mask features, while hierarchical attention layers enable the modeling of long-range diffraction effects (e.g., proximity effects). A comprehensive dataset encompassing diverse classical types of source and mask patterns, including both metal-1 and via layers, is generated to meet the requirements of training our large vision model-based simulator1. The experimental results demonstrate that our simulator achieves over 120× speedup compared to existing commercial solutions while preserving comparable high fidelity, and exhibits superior generalization to advanced process nodes. When deployed in inverse lithography technology (ILT)-guided mask optimization workflows, masks of higher quality are generated than existing solutions. Zhen Wang 0030, Hongquan He, Xuming He 0001, Qi Sun 0002, Cheng Zhuo, Bei Yu 0001, Jingyi Yu 0001, Hao Geng |
ICCAD | 4 |
| 2025 | GeoDistill: Geometry-Guided Self-Distillation for Weakly Supervised Cross-View LocalizationabstractCross-view localization, the task of estimating a camera's 3-degrees-of-freedom (3-DoF) pose by aligning ground-level images with satellite images, is crucial for large-scale outdoor applications like autonomous navigation and augmented reality. Existing methods often rely on fully supervised learning, which requires costly ground-truth pose annotations. In this work, we propose GeoDistill, a Geometry guided weakly supervised self distillation framework that uses teacher-student learning with Field-of-View (FoV)-based masking to enhance local feature learning for robust cross-view localization. In GeoDistill, the teacher model localizes a panoramic image, while the student model predicts locations from a limited FoV counterpart created by FoV-based masking. By aligning the student's predictions with those of the teacher, the student focuses on key features like lane lines and ignores textureless regions, such as roads. This results in more accurate predictions and reduced uncertainty, regardless of whether the query images are panoramas or limited FoV images. Our experiments show that GeoDistill significantly improves localization performance across different frameworks. Additionally, we introduce a novel orientation estimation network that predicts relative orientation without requiring precise planar position ground truth. GeoDistill provides a scalable and efficient solution for real-world cross-view localization challenges. Code and model can be found at https://github.com/tongshw/GeoDistill. Shaowen Tong, Zimin Xia, Alexandre Alahi, Xuming He 0001, Yujiao Shi 0002 |
ICCV | 4 |
| 2025 | LithoSim: A Large, Holistic Lithography Simulation Benchmark for AI-Driven Semiconductor ManufacturingabstractLithography orchestrates a symphony of light, mask and photochemicals to transfer the integrated circuit patterns onto the wafer. Lithography simulation serves as the critical nexus between circuit design and manufacturing, where its speed and accuracy fundamentally govern the optimization quality of downstream resolution enhancement techniques (RET). While machine learning promises to circumvent computational limitations of lithography process through data-driven or physics-informed approximations of computational lithography, existing simulators suffer from inadequate lithographic awareness due to insufficient training data capturing essential process variations and mask correction rules. We present LithoSim, the most comprehensive lithography simulation benchmark to date, featuring over $4$ million high-resolution input-output pairs with rigorous physical correspondence. The dataset systematically incorporates alterable optical source distributions, metal and via mask topologies with optical proximity correction (OPC) variants, and process windows reflecting fab-realistic variations. By integrating domain-specific metrics spanning AI performance and lithographic fidelity, LithoSim establishes a unified evaluation framework for data-driven and physics-informed computational lithography. The data (https://huggingface.co/datasets/grandiflorum/LithoSim), code (https://dw-hongquan.github.io/LithoSim), and pre-trained models (https://huggingface.co/grandiflorum/LithoSim) are released openly to support the development of hybrid ML-based and high-fidelity lithography simulation for the benefit of semiconductor manufacturing. Hongquan He, Zhen Wang 0030, Jingya Wang 0001, Xuming He 0001, Bei Yu 0001, Jingyi Yu 0001, Hao Geng |
NeurIPS | 5 |
| 2025 | GUI-Rise: Structured Reasoning and History Summarization for GUI NavigationabstractWhile Multimodal Large Language Models (MLLMs) have advanced GUI navigation agents, current approaches face limitations in cross-domain generalization and effective history utilization. We present a reasoning-enhanced framework that systematically integrates structured reasoning, action prediction, and history summarization. The structured reasoning component generates coherent Chain-of-Thought analyses combining progress estimation and decision reasoning, which inform both immediate action predictions and compact history summaries for future steps. Based on this framework, we train a GUI agent, GUI-Rise, through supervised fine-tuning on pseudo-labeled trajectories and reinforcement learning with Group Relative Policy Optimization (GRPO). This framework employs specialized rewards, including a history-aware objective, directly linking summary quality to subsequent action performance. Comprehensive evaluations on standard benchmarks demonstrate state-of-the-art results under identical training data conditions, with particularly strong performance in out-of-domain scenarios. These findings validate our framework's ability to maintain robust reasoning and generalization across diverse GUI navigation tasks. Chongyu Wang, Rongjie Li, Yingchen Yu, Xuming He 0001, Song Bai 0001 |
NeurIPS | 5 |
| 2025 | TokMan: Tokenize Manhattan Mask Optimization for Inverse LithographyabstractManhattan representations, defined by axis-aligned, orthogonal structures, are widely used in vision, robotics, and semiconductor design for their geometric regularity and algorithmic simplicity. In integrated circuit (IC) design, Manhattan geometry is key for routing, design rule checking, and lithographic manufacturability. However, as feature sizes shrink, optical system distortions lead to inconsistency between intended layout and printed wafer. Although Inverse Lithography Technology(ILT) is proposed to compensates these effects, learning-based ILT methods, while achieving high simulation fidelity, often generate curvilinear masks on continuous pixel grids, violating Manhattan constraints. Therefore, we propose TokMan, the first framework to formulate mask optimization as a discrete, structure-aware sequence modeling task. Our method leverages a Diffusion Transformer to tokenize layouts into discrete geometric primitives with polygon-wise dependencies and denoise Manhattan-aligned point sequences corrupted by optical proximity effects, while ensuring binary, manufacturable masks. Trained with self-supervised lithographic feedback through differentiable simulation and refined with ILT post-processing, TokMan achieves state-of-the-art fidelity, runtime efficiency, and strict manufacturing compliance on a large-scale dataset of IC layouts. Jingya Wang 0001, Xuming He 0001, Hao Geng, Jingyi Yu 0001 |
NeurIPS | 6 |
| 2024 | Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only TrainingabstractImage captioning aims at generating descriptive and meaningful textual descriptions of images, enabling a broad range of vision-language applications. Prior works have demonstrated that harnessing the power of Contrastive Image Language Pre-training (CLIP) offers a promising approach to achieving zero-shot captioning, eliminating the need for expensive caption annotations. However, the widely observed modality gap in the latent space of CLIP harms the performance of zero-shot captioning by breaking the alignment between paired image-text features. To address this issue, we conduct an analysis on the CLIP latent space which leads to two findings. Firstly, we observe that the CLIP's visual feature of image subregions can achieve closer proximity to the paired caption due to the inherent information loss in text descriptions. In addition, we show that the modality gap between a paired image-text can be empirically modeled as a zero-mean Gaussian distribution. Motivated by the findings, we propose a novel zero-shot image captioning framework with text-only training to reduce the modality gap. In particular, we introduce a subregion feature aggregation to leverage local region information, which produces a compact visual representation for matching text representation. Moreover, we incorporate a noise injection and CLIP reranking strategy to boost captioning performance. We also extend our framework to build a zero-shot VQA pipeline, demonstrating its generality. Through extensive experiments on common captioning and VQA datasets such as MSCOCO, Flickr30k and VQAV2, we show that our method achieves remarkable performance improvements. Code is available at https://github.com/Artanic30/MacCap. Longtian Qiu, Shan Ning, Xuming He 0001 |
AAAI | 3 |
| 2024 | DSGG: Dense Relation Transformer for an End-to-End Scene Graph GenerationabstractScene graph generation aims to capture detailed spatial and semantic relationships between objects in an image, which is challenging due to incomplete labelling, longtailed relationship categories, and relational semantic over-lap. Existing Transformer-based methods either employ distinct queries for objects and predicates or utilize holistic queries for relation triplets and hence often suffer from limited capacity in learning low-frequency relationships. In this paper, we present a new Transformer-based method, called DSGG, that views scene graph detection as a direct graph prediction problem based on a unique set of graph-aware queries. In particular, each graph-aware query encodes a compact representation of both the node and all of its relations in the graph, acquired through the utilization of a relaxed sub-graph matching during the training process. Moreover, to address the problem of relational semantic overlap, we utilize a strategy for relation distillation, aiming to efficiently learn multiple instances of semantic relationships. Extensive experiments on the VG and the PSG datasets show that our model achieves state-of-the-art results, showing a significant improvement of 3.5% and 6.7% in mR@50 and mR@100 for the scene-graph generation task and achieves an even more substantial improvement of 8.5% and 10.3% in mR@50 and mR@100 for the panoptic scene graph generation task. Code is available at https://github.com/zeeshanhayder/DSGG. Zeeshan Hayder, Xuming He 0001 |
CVPR | 2 |
| 2024 | Learning by Correction: Efficient Tuning Task for Zero-Shot Generative Vision-Language ReasoningabstractGenerative vision-language models (VLMs) have shown impressive performance in zero-shot vision-language tasks like image captioning and visual question answering. How-ever, improving their zero-shot reasoning typically requires second-stage instruction tuning, which relies heavily on human-labeled or large language model-generated annotation, incurring high labeling costs. To tackle this challenge, we introduce Image-Conditioned Caption Correction (ICCC), a novel pre-training task designed to enhance VLMs' zero-shot performance without the need for labeled task-aware data. The ICCC task compels VLMs to rectify mismatches between visual and language concepts, thereby enhancing instruction following and text generation conditioned on visual inputs. Leveraging language structure and a lightweight dependency parser, we construct data samples of ICCC taskfrom image-text datasets with low labeling and computation costs. Experimental results on BLIP-2 and InstructBLIP demonstrate significant improvements in zero-shot image-text generation-based VL tasks through ICCC instruction tuning. Rongjie Li, Yu Wu 0014, Xuming He 0001 |
CVPR | 3 |
| 2024 | From Pixels to Graphs: Open-Vocabulary Scene Graph Generation with Vision-Language ModelsabstractScene graph generation (SGG) aims to parse a visual scene into an intermediate graph representation for downstream reasoning tasks. Despite recent advancements, existing methods struggle to generate scene graphs with novel visual relation concepts. To address this challenge, we introduce a new open-vocabulary SGG framework based on sequence generation. Our framework leverages vision-language pre-trained models (VLM) by incorporating an image-to-graph generation paradigm. Specifically, we generate scene graph sequences via image-to-text generation with VLM and then construct scene graphs from these sequences. By doing so, we harness the strong capabilities of VLM for open-vocabulary SGG and seamlessly integrate explicit relational modeling for enhancing the VL tasks. Experimental results demonstrate that our design not only achieves superior performance with an open vocabulary but also enhances downstream vision-language task performance through explicit relation modeling knowledge. Rongjie Li, Songyang Zhang 0001, Dahua Lin, Kai Chen 0026, Xuming He 0001 |
CVPR | 5 |
| 2024 | LLM-HD: Layout Language Model for Hotspot Detection with GDS Semantic EncodingabstractLayout hotspot detection approaches are challenged by the time-to-market constraint and complex designs under rapid downscaling of technology nodes. Pattern matching and learning-based detectors are proposed as quick detection methods. These layout image-based detectors use images transformed from binary database files of layout like GDSII as their inputs. Italy leads to foreground information (e.g., metal polygons) loss and even distortion when shrinking the image size to fit the approach input. Moreover, plenty of irrelevant background information such as non-polygon pixels is also fed into the model, which hinders the fitting of the model and results in a waste of computational resources. In this work, for the first time, we propose a new layout hotspot detection paradigm, where hotspots are directly detected on binary database files by exploiting a hierarchical GDS semantic representation scheme and a well-designed pre-trained natural language processing (NLP) model. Compared with state-of-the-art (SOTA) works, the proposed detector achieves better results on both the ICCAD2012 metal layer benchmark and the more challenging ICCAD2020 via layer benchmark, which demonstrates the effectiveness and efficiency. Jingya Wang 0001, Xuming He 0001, Jingyi Yu 0001, Hao Geng |
DAC | 5 |
| 2024 | SPHINX: A Mixer of Weights, Visual Embeddings and Image Scales for Multi-modal Large Language Models
Renrui Zhang, Peng Gao 0007, Longtian Qiu, Han Xiao 0010, Han Qiu 0010, Wenqi Shao, Keqin Chen, Jiaming Han, Siyuan Huang 0004, Xuming He 0001, Yu Qiao 0001, Hongsheng Li 0001 |
ECCV (62) | 13 |
| 2024 | Dual-Level Adaptive Self-labeling for Novel Class Discovery in Point Cloud Segmentation
Ruijie Xu 0006, Chuyu Zhang, Hui Ren 0003, Xuming He 0001 |
ECCV (13) | 4 |
| 2024 | P2OT: Progressive Partial Optimal Transport for Deep Imbalanced ClusteringabstractDeep clustering, which learns representation and semantic clustering without labels information, poses a great challenge for deep learning-based approaches. Despite significant progress in recent years, most existing methods focus on uniformly distributed datasets, significantly limiting the practical applicability of their methods. In this paper, we first introduce a more practical problem setting named deep imbalanced clustering, where the underlying classes exhibit an imbalance distribution. To tackle this problem, we propose a novel pseudo-labeling-based learning framework. Our framework formulates pseudo-label generation as a progressive partial optimal transport problem, which progressively transports each sample to imbalanced clusters under prior distribution constraints, thus generating imbalance-aware pseudo-labels and learning from high-confident samples.
In addition, we transform the initial formulation into an unbalanced optimal transport problem with augmented constraints, which can be solved efficiently by a fast matrix scaling algorithm. Experiments on various datasets, including a human-curated long-tailed CIFAR100, challenging ImageNet-R, and large-scale subsets of fine-grained iNaturalist2018 datasets, demonstrate the superiority of our method. Chuyu Zhang, Hui Ren 0003, Xuming He 0001 |
ICLR | 3 |
| 2024 | Multi-Level Progressive Reinforcement Learning for Control Policy in Physical SimulationsabstractTraining model-free intelligent agents in complex real-world scenarios using reinforcement learning (RL) often necessitates simulation-based environments due to high physical expenses. However, when simulation takes a long time, e.g., in an unsteady 3D fluid simulation with interactions to the controllable solids, existing RL algorithms meet difficulty to accomplish training within a reasonable timeframes. In this paper, we propose a novel multi-level framework for RL to accelerate convergence as the first attempt to address this difficulty. Motivated by the idea of multi-grid solver, the control policy on a virtual agent over time can be decomposed into different frequency levels, which can be progressively learned via a set of simulations in a coarse-to-fine manner. It is expected that most RL trials are performed in coarser simulations to learn lower control frequency levels with more efficient convergence, while higher frequency levels require much less RL trials, thus significantly accelerating the learning process. To implement our idea, we designed a novel multi-level residual network with a filter module attached, where each level of the network is learned by performing RL for a given simulation resolution. The proposed framework is evaluated by conducting policy learning experiments on a virtual aerial (2D) and an underwater (3D) robot, both requiring time-consuming physical simulations. Our results demonstrate a decrease in almost half in learning time compared to a direct RL approach, while achieving similar control performance. Kefei Wu, Xuming He 0001, Yang Wang 0063, Xiaopei Liu |
ICRA | 2 |
| 2024 | RealDex: Towards Human-like Grasping for Robotic Dexterous Hand
Yaxun Yang, Youzhuo Wang, Yichen Yao 0001, Sören Schwertfeger, Sibei Yang, Wenping Wang 0001, Jingyi Yu 0001, Xuming He 0001, Yuexin Ma |
IJCAI | 11 |
| 2024 | Generalize or Detect? Towards Robust Semantic Segmentation Under Multiple Distribution ShiftsabstractIn open-world scenarios, where both novel classes and domains may exist, an ideal segmentation model should detect anomaly classes for safety and generalize to new domains. However, existing methods often struggle to distinguish between domain-level and semantic-level distribution shifts, leading to poor OOD detection or domain generalization performance. In this work, we aim to equip the model to generalize effectively to covariate-shift regions while precisely identifying semantic-shift regions. To achieve this, we design a novel generative augmentation method to produce coherent images that incorporate both anomaly (or novel) objects and various covariate shifts at both image and object levels. Furthermore, we introduce a training strategy that recalibrates uncertainty specifically for semantic shifts and enhances the feature extractor to align features associated with domain shifts. We validate the effectiveness of our method across benchmarks featuring both semantic and domain shifts. Our method achieves state-of-the-art performance across all benchmarks for both OOD detection and domain generalization. Code is available at https://github.com/gaozhitong/MultiShiftSeg. Zhitong Gao, Mathieu Salzmann, Xuming He 0001 |
NeurIPS | 4 |
| 2024 | CryoGEM: Physics-Informed Generative Cryo-Electron MicroscopyabstractIn the past decade, deep conditional generative models have revolutionized the generation of realistic images, extending their application from entertainment to scientific domains. Single-particle cryo-electron microscopy (cryo-EM) is crucial in resolving near-atomic resolution 3D structures of proteins, such as the SARS-COV-2 spike protein. To achieve high-resolution reconstruction, a comprehensive data processing pipeline has been adopted. However, its performance is still limited as it lacks high-quality annotated datasets for training. To address this, we introduce physics-informed generative cryo-electron microscopy (CryoGEM), which for the first time integrates physics-based cryo-EM simulation with a generative unpaired noise translation to generate physically correct synthetic cryo-EM datasets with realistic noises. Initially, CryoGEM simulates the cryo-EM imaging process based on a virtual specimen. To generate realistic noises, we leverage an unpaired noise translation via contrastive learning with a novel mask-guided sampling scheme. Extensive experiments show that CryoGEM is capable of generating authentic cryo-EM images. The generated dataset can be used as training data for particle picking and pose estimation models, eventually improving the reconstruction resolution. Jiakai Zhang, Qihe Chen, Wenyuan Gao, Xuming He 0001, Jingyi Yu 0001 |
NeurIPS | 5 |
| 2024 | SGTR+: End-to-End Scene Graph Generation With TransformerabstractScene Graph Generation (SGG) remains a challenging visual understanding task due to its compositional property. Most previous works adopt a bottom-up, two-stage or point-based, one-stage approach, which often suffers from high time complexity or suboptimal designs. In this paper, we propose a novel SGG method to address the aforementioned issues, formulating the task as a bipartite graph construction problem. To address the issues above, we create a transformer-based end-to-end framework to generate the entity and entity-aware predicate proposal set, and infer directed edges to form relation triplets. Moreover, we design a graph assembling module to infer the connectivity of the bipartite scene graph based on our entity-aware structure, enabling us to generate the scene graph in an end-to-end manner. Based on bipartite graph assembling paradigm, we further propose a new technical design to address the efficacy of entity-aware modeling and optimization stability of graph assembling. Equipped with the enhanced entity-aware design, our method achieves optimal performance and time-complexity. Extensive experimental results show that our design is able to achieve the state-of-the-art or comparable performance on three challenging benchmarks, surpassing most of the existing approaches and enjoying higher efficiency in inference. Rongjie Li, Songyang Zhang 0001, Xuming He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Multi-Modal Modality-Masked Diffusion Network for Brain MRI Synthesis With Random Modality MissingabstractSynthesis of unavailable imaging modalities from available ones can generate modality-specific complementary information and enable multi-modality based medical images diagnosis or treatment. Existing generative methods for medical image synthesis are usually based on cross-modal translation between acquired and missing modalities. These methods are usually dedicated to specific missing modality and perform synthesis in one shot, which cannot deal with varying number of missing modalities flexibly and construct the mapping across modalities effectively. To address the above issues, in this paper, we propose a unified Multi-modal Modality-masked Diffusion Network (M2DN), tackling multi-modal synthesis from the perspective of "progressive whole-modality inpainting", instead of "cross-modal translation". Specifically, our M2DN considers the missing modalities as random noise and takes all the modalities as a unity in each reverse diffusion step. The proposed joint synthesis scheme performs synthesis for the missing modalities and self-reconstruction for the available ones, which not only enables synthesis for arbitrary missing scenarios, but also facilitates the construction of common latent space and enhances the model representation ability. Besides, we introduce a modality-mask scheme to encode availability status of each incoming modality explicitly in a binary mask, which is adopted as condition for the diffusion model to further enhance the synthesis performance of our M2DN for arbitrary missing scenarios. We carry out experiments on two public brain MRI datasets for synthesis and downstream segmentation tasks. Experimental results demonstrate that our M2DN outperforms the state-of-the-art models significantly and shows great generalizability for arbitrary missing modalities. Kaicong Sun, Jun Xu 0019, Xuming He 0001, Dinggang Shen |
IEEE Trans. Medical Imaging | 4 |
| 2023 | CALIP: Zero-Shot Enhancement of CLIP with Parameter-Free AttentionabstractContrastive Language-Image Pre-training (CLIP) has been shown to learn visual representations with promising zero-shot performance. To further improve its downstream accuracy, existing works propose additional learnable modules upon CLIP and fine-tune them by few-shot training sets. However, the resulting extra training cost and data requirement severely hinder the efficiency for model deployment and knowledge transfer. In this paper, we introduce a free-lunch enhancement method, CALIP, to boost CLIP's zero-shot performance via a parameter-free attention module. Specifically, we guide visual and textual representations to interact with each other and explore cross-modal informative features via attention. As the pre-training has largely reduced the embedding distances between two modalities, we discard all learnable parameters in the attention and bidirectionally update the multi-modal features, enabling the whole process to be parameter-free and training-free. In this way, the images are blended with textual-aware signals and the text representations become visual-guided for better adaptive zero-shot alignment. We evaluate CALIP on various benchmarks of 14 datasets for both 2D image and 3D point cloud few-shot classification, showing consistent zero-shot performance improvement over CLIP. Based on that, we further insert a small number of linear layers in CALIP's attention module and verify our robustness under the few-shot settings, which also achieves leading performance compared to existing methods. Those extensive experiments demonstrate the superiority of our approach for efficient enhancement of CLIP. Code is available at https://github.com/ZiyuGuo99/CALIP. Renrui Zhang, Longtian Qiu, Xianzheng Ma, Xupeng Miao, Xuming He 0001, Bin Cui 0001 |
AAAI | 6 |
| 2023 | Cascade Sparse Feature Propagation Network for Interactive Segmentation
Chuyu Zhang, Hui Ren 0003, Chuanyang Hu, Yongfei Liu, Xuming He 0001 |
BMVC | 5 |
| 2023 | HOICLIP: Efficient Knowledge Transfer for HOI Detection with Vision-Language ModelsabstractHuman-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions. Recently, Contrastive Language-Image Pre-training (CLIP) has shown great potential in providing interaction prior for HOI detectors via knowledge distillation. However, such approaches often rely on large-scale training data and suffer from inferior performance under few/zero-shot scenarios. In this paper, we propose a novel HOI detection framework that efficiently extracts prior knowledge from CLIP and achieves better generalization. In detail, we first introduce a novel interaction decoder to extract informative regions in the visual feature map of CLIP via a cross-attention mechanism, which is then fused with the detection backbone by a knowledge integration block for more accurate human- object pair detection. In addition, prior knowledge in CLIP text encoder is leveraged to generate a classifier by embedding HOI descriptions. To distinguish fine-grained interactions, we build a verb classifier from training data via visual semantic arithmetic and a lightweight verb representation adapter. Furthermore, we propose a training-free enhancement to exploit global HOI predictions from CLIP. Extensive experiments demonstrate that our method outperforms the state of the art by a large margin on various settings, e.g. +4.04 mAP on HICO-Det. The source code is available in https://github.com/Artanic30/HOICLIP. Shan Ning, Longtian Qiu, Yongfei Liu, Xuming He 0001 |
CVPR | 4 |
| 2023 | Part-aware Prototypical Graph Network for One-shot Skeleton-based Action RecognitionabstractIn this paper, we study the problem of one-shot skeleton-based action recognition, which poses unique challenges in learning transferable representation from base classes to novel classes, particularly for fine-grained actions. Existing meta-learning frameworks typically rely on the body-level representations in spatial dimension, which limits the generalisation to capture subtle visual differences in the fine-grained label space. To overcome the above limitation, we propose a part-aware prototypical representation for one-shot skeleton-based action recognition. Our method captures skeleton motion patterns at two distinctive spatial levels, one for global contexts among all body joints, referred to as body level, and the other attends to local spatial regions of body parts, referred to as the part level. We also devise a class-agnostic attention mechanism to highlight important parts for each action class. Specifically, we develop a part-aware prototypical graph network consisting of three modules: a cascaded embedding module for our dual-level modelling, an attention-based part fusion module to fuse parts and generate part-aware prototypes, and a matching module to perform classification with the part-aware representations. We demonstrate the effectiveness of our method on two public skeleton-based action recognition datasets: NTU RGB+D 120 and NW-UCLA. Tailin Chen, Desen Zhou, Jian Wang 0066, Qian He 0001, Chuanyang Hu, Errui Ding, Yu Guan 0001, Xuming He 0001 |
FG | 9 |
| 2023 | Class-relation Knowledge Distillation for Novel Class DiscoveryabstractWe tackle the problem of novel class discovery, which aims to learn novel classes without supervision based on labeled data from known classes. A key challenge lies in transferring the knowledge in the known-class data to the learning of novel classes. Previous methods mainly focus on building a shared representation space for knowledge transfer and often ignore modeling class relations. To address this, we introduce a class relation representation for the novel classes based on the predicted class distribution of a model trained on known classes. Empirically, we find that such class relation becomes less informative during typical discovery training. To prevent such information loss, we propose a novel knowledge distillation framework, which utilizes our class-relation representation to regularize the learning of novel classes. In addition, to enable a flexible knowledge distillation scheme for each data point in novel classes, we develop a learnable weighting function for the regularization, which adaptively promotes knowledge transfer based on the semantic similarity between the novel and known classes. To validate the effectiveness and generalization of our method, we conduct extensive experiments on multiple benchmarks, including CIFAR100, Stanford Cars, CUB, and FGVC-Aircraft datasets. Our results demonstrate that the proposed method outperforms the previous state-of-the-art methods by a significant margin on almost all benchmarks. Code is available at here. Peiyan Gu, Chuyu Zhang, Ruijie Xu 0006, Xuming He 0001 |
ICCV | 4 |
| 2023 | Grounded Image Text Matching with Mismatched Relation ReasoningabstractThis paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image, then localize referred objects or ground the mismatched parts of the text. We provide a benchmark for evaluating vision-language (VL) models on this task, with a focus on the challenging settings of limited training data and out-of-distribution sentence lengths. Our evaluation demonstrates that pre-trained VL models often lack data efficiency and length generalization ability. To address this, we propose the Relation-sensitive Correspondence Reasoning Network (RCRN), which incorporates relation-aware reasoning via bi-directional message propagation guided by language structure. Our RCRN can he interpreted as a modular program and delivers strong performance in terms of both length generalization and data efficiency. The code and data are available on https://githuh.coin/SHTUPLUS/GITM-MR. Yu Wu 0014, Yana Wei, Haozhe Wang 0002, Yongfei Liu, Sibei Yang, Xuming He 0001 |
ICCV | 6 |
| 2023 | Human-centric Scene Understanding for 3D Large-scale ScenariosabstractHuman-centric scene understanding is significant for real-world applications, but it is extremely challenging due to the existence of diverse human poses and actions, complex human-environment interactions, severe occlusions in crowds, etc. In this paper, we present a large-scale multi-modal dataset for human-centric scene under-standing, dubbed HuCenLife, which is collected in diverse daily-life scenarios with rich and fine-grained annotations. Our HuCenLife can benefit many 3D perception tasks, such as segmentation, detection, action recognition, etc., and we also provide benchmarks for these tasks to facilitate related research. In addition, we design novel modules for LiDAR-based segmentation and action recognition, which are more applicable for large-scale human-centric scenarios and achieve state-of-the-art performance. The dataset and code can be found at https://github.com/4DVLab/HuCenLife.git. Yiteng Xu, Peishan Cong, Yichen Yao 0001, Runnan Chen, Yuenan Hou, Xinge Zhu, Xuming He 0001, Jingyi Yu 0001, Yuexin Ma |
ICCV | 7 |
| 2023 | Modeling Multimodal Aleatoric Uncertainty in Segmentation with Mixture of Stochastic Experts
Zhitong Gao, Yucong Chen, Chuyu Zhang, Xuming He 0001 |
ICLR | 4 |
| 2023 | Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning
Yongfei Liu, Desen Zhou, Tinne Tuytelaars, Xuming He 0001 |
ICLR | 5 |
| 2023 | MILD: Modeling the Instance Learning Dynamics for Learning with Noisy LabelsabstractDespite deep learning has achieved great success, it often relies on a large amount of training data with accurate labels, which are expensive and time-consuming to collect. A prominent direction to reduce the cost is to learn with noisy labels, which are ubiquitous in the real-world applications. A critical challenge for such a learning task is to reduce the effect of network memorization on the falsely-labeled data. In this work, we propose an iterative selection approach based on the Weibull mixture model, which identifies clean data by considering the overall learning dynamics of each data instance. In contrast to the previous small-loss heuristics, we leverage the observation that deep network is easy to memorize and hard to forget clean data. In particular, we measure the difficulty of memorization and forgetting for each instance via the transition times between being misclassified and being memorized in training, and integrate them into a novel metric for selection. Based on the proposed metric, we retain a subset of identified clean data and repeat the selection procedure to iteratively refine the clean subset, which is finally used for model training. To validate our method, we perform extensive experiments on synthetic noisy datasets and real-world web data, and our strategy outperforms existing noisy-label learning methods. Chuanyang Hu, Shipeng Yan, Zhitong Gao, Xuming He 0001 |
IJCAI | 4 |
| 2023 | Exploring Learning-Based Control Policy for Fish-Like Robots in Altered Background FlowsabstractThe study of motion control for the fish-like robots in complex fluid fields is of great importance in improving the performance of underwater vehicles, due to its strong maneuverability, propulsion efficiency, and deceptive visual appearance. In this article, a novel learning-based control framework is first proposed to autonomously explore efficient control policies that are capable of performing motion control tasks in non-quiescent and unknown background flows. First, we utilize a high-fidelity simulation system, named FishGym, to generate various uniform flows. Next, a DRL-based algorithm is incorporated with the FishGym to train the fish-like robot to control its motion to optimally complete a delicately designed task (Approaching Target and Stay) in both quiescent and uniform flow. Then, the obtained control policy together with an online estimator is directly applied to a Path-Following Task. The proposed framework well balances the simulation accuracy and the computational efficiency, which is of crucial importance for effective coupling with the learning algorithm. The simulation results indicate that, via the proposed learning framework, the robot successfully acquired a swimming strategy that can be used to adapt to different background flows and tasks. Furthermore, we also observe some adaptation behavior of the robot, such as rheotaxis, that is similar to the fish in nature, which gains us more insight into the mechanism underlying the adaptation behavior of fish in a complex environment. Xiaozhu Lin, Wenbin Song, Xiaopei Liu, Xuming He 0001, Yang Wang 0063 |
IROS | 4 |
| 2023 | HC-Net: Hybrid Classification Network for Automatic Periodontal Disease Diagnosis
Lanzhuju Mei, Yu Fang 0008, Zhiming Cui 0001, Nizhuan Wang 0001, Xuming He 0001, Yiqiang Zhan, Xiang Sean Zhou, Maurizio Tonetti, Dinggang Shen |
MICCAI (6) | 6 |
| 2023 | OccluBEV: Occlusion Aware Spatiotemporal Modeling for Multi-view 3D Object DetectionabstractBird's-Eye-View (BEV) based 3D visual perception, which formulates a unified space for multi-view representation, has received wide attention in autonomous driving due to its scalability for downstream tasks. However, view transform in transformer-based BEV methods is agnostic of 3D occlusion relationships, resulting in model degradation. To construct a higher-quality BEV space, this paper analyzes the mutual occlusion problems in the view transform process and proposes a new transformer-based method named OccluBEV. OccluBEV alleviates the occlusion issue via point cloud information distillation in both the image and BEV space. Specifically, in the image space, we perform depth estimation for each pixel and utilize it to guide image feature mapping. Further, since predicting depth directly from monocular image is ill-posed, ignoring stereo information such as multi-view and temporal cues, this paper introduces a voxel visibility segmentation task in 3D BEV space. The task explicitly predicts whether each voxel in the 3D BEV grid is occupied or not. In addition, to alleviate the overfitting problem in BEV feature learning under a single task, we design a multi-head learning framework which jointly models multiple strongly-correlated tasks in a unified BEV space. The effectiveness of the proposed method is fully validated on the nuScenes dataset, achieving a competetive NDS/mAP score of 57.5/47.9 on the nuScenes test leaderboard using ResNet101 backbone, which is superior to state-of-the-art camera-based solutions. Ziteng Wen, Jinshui Hu, Xuming He 0001, Fengren Wang, Shun Lou, Haibo Fan |
ACM Multimedia | 6 |
| 2023 | ATTA: Anomaly-aware Test-Time Adaptation for Out-of-Distribution Detection in SegmentationabstractRecent advancements in dense out-of-distribution (OOD) detection have primarily focused on scenarios where the training and testing datasets share a similar domain, with the assumption that no domain shift exists between them. However, in real-world situations, domain shift often exits and significantly affects the accuracy of existing out-of-distribution (OOD) detection models. In this work, we propose a dual-level OOD detection framework to handle domain shift and semantic shift jointly. The first level distinguishes whether domain shift exists in the image by leveraging global low-level features, while the second level identifies pixels with semantic shift by utilizing dense high-level feature maps. In this way, we can selectively adapt the model to unseen domains as well as enhance model's capacity in detecting novel classes. We validate the efficacy of our proposed method on several OOD segmentation benchmarks, including those with significant domain shifts and those without, observing consistent performance improvements across various baseline models. Code is available at https://github.com/gaozhitong/ATTA. Zhitong Gao, Shipeng Yan, Xuming He 0001 |
NeurIPS | 3 |
| 2022 | SGTR: End-to-end Scene Graph Generation with TransformerabstractScene Graph Generation (SGG) remains a challenging visual understanding task due to its compositional property. Most previous works adopt a bottom-up two-stage or a point-based one-stage approach, which often suffers from high time complexity or sub-optimal designs. In this work, we propose a novel SGG method to address the aforementioned issues, formulating the task as a bipartite graph construction problem. To solve the problem, we develop a transformer-based end-to-end framework that first generates the entity and predicate proposal set, followed by inferring directed edges to form the relation triplets. In particular, we develop a new entity-aware predicate representation based on a structural predicate generator that leverages the compositional property of relationships. Moreover, we design a graph assembling module to infer the connectivity of the bipartite scene graph based on our entity-aware structure, enabling us to generate the scene graph in an end-to-end manner. Extensive experimental results show that our design is able to achieve the state-of-the-art or comparable performance on two challenging benchmarks, surpassing most of the existing approaches and enjoying higher efficiency in inference. We hope our model can serve as a strong baseline for the Transformer-based scene graph generation.11Code is available: https://github.com/Scarecrow0/SGTR Rongjie Li, Songyang Zhang 0001, Xuming He 0001 |
CVPR | 3 |
| 2022 | General Incremental Learning with Domain-aware Categorical RepresentationsabstractContinual learning is an important problem for achieving human-level intelligence in real-world applications as an agent must continuously accumulate knowledge in response to streaming data/tasks. In this work, we consider a general and yet under-explored incremental learning problem in which both the class distribution and class-specific domain distribution change over time. In addition to the typical challenges in class incremental learning, this setting also faces the intra-class stability-plasticity dilemma and intra-class domain imbalance problems. To address above issues, we develop a novel domain-aware continual learning method based on the EM framework. Specifically, we introduce a flexible class representation based on the von Mises-Fisher mixture model to capture the intra-class structure, using an expansion-and- reduction strategy to dynamically increase the number of components according to the class complexity. Moreover, we design a bi-level balanced memory to cope with data imbalances within and across classes, which combines with a distillation loss to achieve better inter- and intra-class stability-plasticity trade-off. We conduct exhaustive experiments on three benchmarks: iDigits, iDomainNet and iCIFAR-20. The results show that our approach consistently outperforms previous methods by a significant margin, demonstrating its superiority. Jiangwei Xie, Shipeng Yan, Xuming He 0001 |
CVPR | 3 |
| 2022 | Learning Semantic Correspondence with Sparse Annotations
Shuaiyi Huang, Luyu Yang, Bo He 0004, Songyang Zhang 0001, Xuming He 0001, Abhinav Shrivastava |
ECCV (14) | 5 |
| 2022 | Generative Negative Text Replay for Continual Vision-Language Pretraining
Shipeng Yan, Lanqing Hong, Hang Xu 0004, Jianhua Han, Tinne Tuytelaars, Zhenguo Li, Xuming He 0001 |
ECCV (36) | 7 |
| 2022 | Robust Temporally-Coherent Strategy for Few-shot Video Instance SegmentationabstractTraditional video instance segmentation (VIS) aims to detect, segment, and track object instances from a known class set in videos. In real-world applications, however, video instance segmentation typically need to cope with novel-class instances and to fast adapt with a few labeled videos. In this work, we aim to tackle the task of few-shot video instance segmentation (FVIS), which is challenging due to large variations in object appearance and motion. We propose a robust temporally coherent strategy, termed as VTFA, based on a two-stage fine-tuning approach. VTFA enforces the instance segmentation of novel classes to be temporally smooth and reduces the classification bias between novel and base classes. The proposed Memory-aware Temporal Context Encoding Module (MTCE) in VTFA encodes the temporal context information, which contributes to the consistency in the final predictions. We also propose a loss named Instance-level Pair-wise Contrastive (IPC) Loss on both the novel and base classes to enhance the robustness of instance classification. To validate our method, we develop a YouTube-VIS-FS benchmark to compare our method with several baselines. The experimental evaluation shows that our strategy is superior or competitive to those strong baselines. Qiuyue Wang, Songyang Zhang 0001, Xuming He 0001 |
ICIP | 3 |
| 2022 | FishGym: A High-Performance Physics-based Simulation Framework for Underwater Robot LearningabstractBionic underwater robots have demonstrated their superiority in many applications. Yet, training their intelligence for a variety of tasks that mimic the behavior of underwater creatures poses a number of challenges in practice, mainly due to lack of a large amount of available training data as well as the high cost in real physical environment. Alternatively, simulation has been considered as a viable and important tool for acquiring datasets in different environments, but it mostly targeted rigid and soft body systems. There is currently dearth of work for more complex fluid systems interacting with immersed solids that can be efficiently and accurately simulated for robot training purposes. In this paper, we propose a new platform called “FishGym”, which can be used to train fish-like underwater robots. The framework consists of a robotic fish modeling module using articulated body with skinning, a GPU-based high-performance localized two-way coupled fluid-structure interaction simulation module that handles both finite and infinitely large domains, as well as a reinforcement learning module. We leveraged existing training methods with adaptations to underwater fish-like robots and obtained learned control policies for multiple benchmark tasks. The training results are demonstrated with reasonable motion trajectories, with comparisons and analyses to empirical models as well as known real fish swimming behaviors to highlight the advantages of the proposed platform. Wenji Liu, Kai Bai, Xuming He 0001, Shuran Song, Changxi Zheng, Xiaopei Liu |
ICRA | 3 |
| 2022 | ROI-Constrained Bidding via Curriculum-Guided Bayesian Reinforcement LearningabstractReal-Time Bidding (RTB) is an important mechanism in modern online advertising systems. Advertisers employ bidding strategies in RTB to optimize their advertising effects subject to various financial requirements, especially the return-on-investment (ROI) constraint. ROIs change non-monotonically during the sequential bidding process, and often induce a see-saw effect between constraint satisfaction and objective optimization. While some existing approaches show promising results in static or mildly changing ad markets, they fail to generalize to highly dynamic ad markets with ROI constraints, due to their inability to adaptively balance constraints and objectives amidst non-stationarity and partial observability. In this work, we specialize in ROI-Constrained Bidding in non-stationary markets. Based on a Partially Observable Constrained Markov Decision Process, our method exploits an indicator-augmented reward function free of extra trade-off parameters and develops a Curriculum-Guided Bayesian Reinforcement Learning (CBRL) framework to adaptively control the constraint-objective trade-off in non-stationary ad markets. Extensive experiments on a large-scale industrial dataset with two problem settings reveal that CBRL generalizes well in both in-distribution and out-of-distribution data regimes, and enjoys superior learning efficiency and stability. Haozhe Wang 0002, Panyan Fang, Xuming He 0001, Liang Wang 0001, Bo Zheng 0007 |
KDD | 5 |
| 2021 | Bipartite Graph Network With Adaptive Message Passing for Unbiased Scene Graph GenerationabstractScene graph generation is an important visual understanding task with a broad range of vision applications. Despite recent tremendous progress, it remains challenging due to the intrinsic long-tailed class distribution and large intra-class variation. To address these issues, we introduce a novel confidence-aware bipartite graph neural network with adaptive message propagation mechanism for unbiased scene graph generation. In addition, we propose an efficient bi-level data resampling strategy to alleviate the imbalanced data distribution problem in training our graph network. Our approach achieves superior or competitive performance over previous methods on several challenging datasets, including Visual Genome, Open Images V4/V6, demonstrating its effectiveness and generality. Rongjie Li, Songyang Zhang 0001, Xuming He 0001 |
CVPR | 4 |
| 2021 | Relation-aware Instance Refinement for Weakly Supervised Visual GroundingabstractVisual grounding, which aims to build a correspondence between visual objects and their language entities, plays a key role in cross-modal scene understanding. One promising and scalable strategy for learning visual grounding is to utilize weak supervision from only image-caption pairs. Previous methods typically rely on matching query phrases directly to a precomputed, fixed object candidate pool, which leads to inaccurate localization and ambiguous matching due to lack of semantic relation constraints. In our paper, we propose a novel context-aware weakly-supervised learning method that incorporates coarse-to-fine object refinement and entity relation modeling into a two-stage deep network, capable of producing more accurate object representation and matching. To effectively train our network, we introduce a self-taught regression loss for the proposal locations and a classification loss based on parsed entity relations. Extensive experiments on two public benchmarks Flickr30K Entities and ReferItGame demonstrate the efficacy of our weakly grounding framework. The results show that we outperform the previous methods by a considerable margin, achieving 59.27% top-1 accuracy in Flickr30K Entities and 37.68% in the ReferItGame dataset respectively1. Yongfei Liu, Lin Ma 0002, Xuming He 0001 |
CVPR | 4 |
| 2021 | DER: Dynamically Expandable Representation for Class Incremental LearningabstractWe address the problem of class incremental learning, which is a core step towards achieving adaptive vision intelligence. In particular, we consider the task setting of incremental learning with limited memory and aim to achieve better stability-plasticity trade-off. To this end, we propose a novel two-stage learning approach that utilizes a dynamically expandable representation for more effective incremental concept modeling. Specifically, at each incremental step, we freeze the previously learned representation and augment it with additional feature dimensions from a new learnable feature extractor. This enables us to integrate new visual concepts with retaining learned knowledge. We dynamically expand the representation according to the complexity of novel concepts by introducing a channel-level mask-based pruning strategy. Moreover, we introduce an auxiliary loss to encourage the model to learn diverse and discriminate features for novel concepts. We conduct extensive experiments on the three class incremental learning benchmarks and our method consistently outperforms other methods with a large margin.1 Shipeng Yan, Jiangwei Xie, Xuming He 0001 |
CVPR | 3 |
| 2021 | Distribution Alignment: A Unified Framework for Long-Tail Visual RecognitionabstractDespite the recent success of deep neural networks, it remains challenging to effectively model the long-tail class distribution in visual recognition tasks. To address this problem, we first investigate the performance bottleneck of the two-stage learning framework via ablative study. Motivated by our discovery, we propose a unified distribution alignment strategy for long-tail visual recognition. Specifically, we develop an adaptive calibration function that enables us to adjust the classification scores for each data point. We then introduce a generalized re-weight method in the two-stage learning to balance the class prior, which provides a flexible and unified solution to diverse scenarios in visual recognition tasks. We validate our method by extensive experiments on four tasks, including image classification, semantic segmentation, object detection, and instance segmentation. Our approach achieves the state-of-the-art results across all four recognition tasks with a simple and unified framework. Songyang Zhang 0001, Shipeng Yan, Xuming He 0001, Jian Sun 0001 |
CVPR | 4 |
| 2021 | GNeRF: GAN-based Neural Radiance Field without Posed CameraabstractWe introduce GNeRF, a framework to marry Generative Adversarial Networks (GAN) with Neural Radiance Field (NeRF) reconstruction for the complex scenarios with unknown and even randomly initialized camera poses. Recent NeRF-based advances have gained popularity for remarkable realistic novel view synthesis. However, most of them heavily rely on accurate camera poses estimation, while few recent methods can only optimize the unknown camera poses in roughly forward-facing scenes with relatively short camera trajectories and require rough camera poses initialization. Differently, our GNeRF only utilizes randomly initialized poses for complex outside-in scenarios. We propose a novel two-phases end-to-end framework. The first phase takes the use of GANs into the new realm for optimizing coarse camera poses and radiance fields jointly, while the second phase refines them with additional photometric loss. We overcome local minima using a hybrid and iterative optimization scheme. Extensive experiments on a variety of synthetic and natural scenes demonstrate the effectiveness of GNeRF. More impressively, our approach outperforms the baselines favorably in those scenes with repeated patterns or even low textures that are regarded as extremely challenging before. Quan Meng, Anpei Chen, Haimin Luo, Minye Wu, Hao Su 0001, Lan Xu 0003, Xuming He 0001, Jingyi Yu 0001 |
ICCV | 7 |
| 2021 | Learning Implicit Temporal Alignment for Few-shot Video ClassificationabstractFew-shot video classification aims to learn new video categories with only a few labeled examples, alleviating the burden of costly annotation in real-world applications. However, it is particularly challenging to learn a class-invariant spatial-temporal representation in such a setting. To address this, we propose a novel matching-based few-shot learning strategy for video sequences in this work. Our main idea is to introduce an implicit temporal alignment for a video pair, capable of estimating the similarity between them in an accurate and robust manner. Moreover, we design an effective context encoding module to incorporate spatial and feature channel context, resulting in better modeling of intra-class variations. To train our model, we develop a multi-task loss for learning video matching, leading to video features with better generalization. Extensive experimental results on two challenging benchmarks, show that our method outperforms the prior arts with a sizable margin on Something-Something-V2 and competitive results on Kinetics. Songyang Zhang 0001, Xuming He 0001 |
IJCAI | 3 |
| 2021 | Superpixel-Guided Iterative Learning from Noisy Labels for Medical Image Segmentation
Shuailin Li, Zhitong Gao, Xuming He 0001 |
MICCAI (1) | 3 |
| 2021 | Single Image 3D Object Estimation with Primitive Graph NetworksabstractReconstructing 3D object from a single image (RGB or depth) is a fundamental problem in visual scene understanding and yet remains challenging due to its ill-posed nature and complexity in real-world scenes. To address those challenges, we adopt a primitive-based representation for 3D object, and propose a two-stage graph network for primitive-based 3D object estimation, which consists of a sequential proposal module and a graph reasoning module. Given a 2D image, our proposal module first generates a sequence of 3D primitives from input image with local feature attention. Then the graph reasoning module performs joint reasoning on a primitive graph to capture the global shape context for each primitive. Such a framework is capable of taking into account rich geometry and semantic constraints during 3D structure recovery, producing 3D objects with more coherent structure even under challenging viewing conditions. We train the entire graph neural network in a stage-wise strategy and evaluate it on three benchmarks: Pix3D, ModelNet and NYU Depth V2. Extensive experiments show that our approach outperforms the previous state of the arts with a considerable margin. Qian He 0001, Desen Zhou, Xuming He 0001 |
ACM Multimedia | 4 |
| 2021 | Learning Multi-Granular Spatio-Temporal Graph Network for Skeleton-based Action RecognitionabstractThe task of skeleton-based action recognition remains a core challenge in human-centred scene understanding due to the multiple granularities and large variation in human motion. Existing approaches typically employ a single neural representation for different motion patterns, which has difficulty in capturing fine-grained action classes given limited training data. To address the aforementioned problems, we propose a novel multi-granular spatio-temporal graph network for skeleton-based action classification that jointly models the coarse- and fine-grained skeleton motion patterns. To this end, we develop a dual-head graph network consisting of two interleaved branches, which enables us to extract features at two spatio-temporal resolutions in an effective and efficient manner. Moreover, our network utilises a cross-head communication strategy to mutually enhance the representations of both heads. We conducted extensive experiments on three large-scale datasets, namely NTU RGB+D 60, NTU RGB+D 120, and Kinetics-Skeleton, and achieves the state-of-the-art performance on all the benchmarks, which validates the effectiveness of our method1. Tailin Chen, Desen Zhou, Jian Wang 0066, Yu Guan 0001, Xuming He 0001, Errui Ding |
ACM Multimedia | 6 |
| 2021 | An EM Framework for Online Incremental Learning of Semantic SegmentationabstractIncremental learning of semantic segmentation has emerged as a promising strategy for visual scene interpretation in the open-world setting. However, it remains challenging to acquire novel classes in an online fashion for the segmentation task, mainly due to its continuously-evolving semantic label space, partial pixelwise ground-truth annotations, and constrained data availability. To address this, we propose an incremental learning strategy that can fast adapt deep segmentation models without catastrophic forgetting, using a streaming input data with pixel annotations on the novel classes only. To this end, we develop a unified learning strategy based on the Expectation-Maximization (EM) framework, which integrates an iterative relabeling strategy that fills in the missing labels and a rehearsal-based incremental learning step that balances the stability-plasticity of the model. Moreover, our EM algorithm adopts an adaptive sampling method to select informative training data and a class-balancing training strategy in the incremental model updates, both improving the efficacy of model learning. We validate our approach on the PASCAL VOC 2012 and ADE20K datasets, and the results demonstrate its superior performance over the existing incremental methods. Shipeng Yan, Jiangwei Xie, Songyang Zhang 0001, Xuming He 0001 |
ACM Multimedia | 5 |
| 2021 | Dynamic Grained Encoder for Vision TransformersabstractTransformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained Encoder for vision transformers, which can adaptively assign a suitable number of queries to each spatial region. Thus it achieves a fine-grained representation in discriminative regions while keeping high efficiency. Besides, the dynamic grained encoder is compatible with most vision transformer frameworks. Without bells and whistles, our encoder allows the state-of-the-art vision transformers to reduce computational complexity by 40%-60% while maintaining comparable performance on image classification. Extensive experiments on object detection and segmentation further demonstrate the generalizability of our approach. Code is available at https://github.com/StevenGrove/vtpack. Lin Song 0002, Songyang Zhang 0001, Xuming He 0001, Hongbin Sun 0001, Jian Sun 0001, Nanning Zheng 0001 |
NeurIPS | 5 |
| 2021 | Fixed-Price Diffusion Mechanism Design
Dengji Zhao, Xuming He 0001 |
PRICAI (1) | 4 |
| 2020 | Learning Cross-Modal Context Graph for Visual GroundingabstractVisual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the resulting semantic ambiguities. Prior works typically focus on learning representations of individual phrases with limited context information. To address their limitations, this paper proposes a language-guided graph representation to capture the global context of grounding entities and their relations, and develop a cross-modal graph matching strategy for the multiple-phrase visual grounding task. In particular, we introduce a modular graph neural network to compute context-aware representations of phrases and object proposals respectively via message propagation, followed by a graph-based matching module to generate globally consistent localization of grounding phrases. We train the entire graph neural network jointly in a two-stage strategy and evaluate it on the Flickr30K Entities benchmark. Extensive experiments show that our method outperforms the prior state of the arts by a sizable margin, evidencing the efficacy of our grounding framework. Code is available at https://github.com/youngfly11/LCMCG-PyTorch. Yongfei Liu, Xiaodan Zhu 0001, Xuming He 0001 |
AAAI | 4 |
| 2020 | Part-Aware Prototype Network for Few-Shot Semantic Segmentation
Yongfei Liu, Xiangyi Zhang, Songyang Zhang 0001, Xuming He 0001 |
ECCV (9) | 4 |
| 2020 | Disentangled Representation Learning for Controllable Image Synthesis: An Information-Theoretic PerspectiveabstractIn this paper, we look into the problem of disentangled representation learning and controllable image synthesis in a deep generative model. We develop an encoder-decoder architecture for a variant of the Variational Auto-Encoder (VAE) with two latent codes z1and z2. Our framework uses z2to capture specified factors of variation while z1captures the complementary factors of variation. To this end, we analyze the learning problem from the perspective of multivariate mutual information, derive optimizable lower bounds of the conditional mutual information in the image synthesis processes and incorporate them into the training objective. We validate our method empirically on the Color MNIST dataset and the CelebA dataset by showing controllable image syntheses. Our proposed paradigm is simple yet effective and is applicable to many situations, including those where there is not an explicit factorization of features available, or where the features are non-categorical. Shichang Tang, Xu Zhou 0005, Xuming He 0001, Yi Ma 0001 |
ICPR | 3 |
| 2020 | Shape-Aware Semi-supervised 3D Semantic Segmentation for Medical Images
Shuailin Li, Chuyu Zhang, Xuming He 0001 |
MICCAI (1) | 3 |
| 2020 | LGNN: A Context-aware Line Segment DetectorabstractWe present a novel real-time line segment detection scheme called Line Graph Neural Network (LGNN). Existing approaches require a computationally expensive verification or postprocessing step. Our LGNN employs a deep convolutional neural network (DCNN) for proposing line segment directly, with a graph neural network (GNN) module for reasoning their connectivities. Specifically, LGNN exploits a new quadruplet representation for each segment where the GNN module takes the predicted candidates as vertexes and constructs a sparse graph to enforce structural context. Compared with the state-of-the-art, LGNN achieves near real-time performance without compromising accuracy. LGNN further enables time-sensitive 3D applications. When a 3D point cloud is accessible, we present a multi-modal line segment classification technique for extracting a 3D wireframe of the environment robustly and efficiently. Quan Meng, Jiakai Zhang, Qiang Hu 0003, Xuming He 0001, Jingyi Yu 0001 |
ACM Multimedia | 4 |
| 2020 | Confidence-Aware Adversarial Learning for Self-supervised Semantic Matching
Shuaiyi Huang, Qiuyue Wang, Xuming He 0001 |
PRCV (1) | 3 |
| 2020 | Learning a Layout Transfer Network for Context Aware Object DetectionabstractWe present a context aware object detection method based on a retrieve-and-transform scene layout model. Given an input image, our approach first retrieves a coarse scene layout from a codebook of typical layout templates. In order to handle large layout variations, we use a variant of the spatial transformer network to transform and refine the retrieved layout, resulting in a set of interpretable and semantically meaningful feature maps of object locations and scales. The above steps are implemented as a Layout Transfer Network which we integrate into Faster RCNN to allow for joint reasoning of object detection and scene layout estimation. Extensive experiments on three public datasets verified that our approach provides consistent performance improvements to the state-of-the-art object detection baselines on a variety of challenging tasks in the traffic surveillance and the autonomous driving domains. Tao Wang 0047, Xuming He 0001, Yuanzheng Cai, Guobao Xiao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | A Dual Attention Network with Semantic Embedding for Few-Shot LearningabstractDespite recent success of deep neural networks, it remains challenging to efficiently learn new visual concepts from limited training data. To address this problem, a prevailing strategy is to build a meta-learner that learns prior knowledge on learning from a small set of annotated data. However, most of existing meta-learning approaches rely on a global representation of images and a meta-learner with complex model structures, which are sensitive to background clutter and difficult to interpret. We propose a novel meta-learning method for few-shot classification based on two simple attention mechanisms: one is a spatial attention to localize relevant object regions and the other is a task attention to select similar training data for label prediction. We implement our method via a dual-attention network and design a semantic-aware meta-learning loss to train the meta-learner network in an end-to-end manner. We validate our model on three few-shot image classification datasets with extensive ablative study, and our approach shows competitive performances over these datasets with fewer parameters. For facilitating the future research, code and data split are available: https://github.com/tonysy/STANet-PyTorch Shipeng Yan, Songyang Zhang 0001, Xuming He 0001 |
AAAI | 3 |
| 2019 | Dynamic Context Correspondence Network for Semantic AlignmentabstractEstablishing semantic correspondence is a core problem in computer vision and remains challenging due to large intra-class variations and lack of annotated data. In this paper, we aim to incorporate global semantic context in a flexible manner to overcome the limitations of prior work that relies on local semantic representations. To this end, we first propose a context-aware semantic representation that incorporates spatial layout for robust matching against local ambiguities. We then develop a novel dynamic fusion strategy based on attention mechanism to weave the advantages of both local and context features by integrating semantic cues from multiple scales. We instantiate our strategy by designing an end-to-end learnable deep network, named as Dynamic Context Correspondence Network (DCCNet). To train the network, we adopt a multi-auxiliary task loss to improve the efficiency of our weakly-supervised learning procedure. Our approach achieves superior or competitive performance over previous methods on several challenging datasets, including PF-Pascal, PF-Willow, and TSS, demonstrating its effectiveness and generality. Shuaiyi Huang, Qiuyue Wang, Songyang Zhang 0001, Shipeng Yan, Xuming He 0001 |
ICCV | 5 |
| 2019 | Pose-Aware Multi-Level Feature Network for Human Object Interaction DetectionabstractReasoning human object interactions is a core problem in human-centric scene understanding and detecting such relations poses a unique challenge to vision systems due to large variations in human-object configurations, multiple co-occurring relation instances and subtle visual difference between relation categories. To address those challenges, we propose a multi-level relation detection strategy that utilizes human pose cues to capture global spatial configurations of relations and as an attention mechanism to dynamically zoom into relevant regions at human part level. We develop a multi-branch deep network to learn a pose-augmented relation representation at three semantic levels, incorporating interaction context, object features and detailed semantic part cues. As a result, our approach is capable of generating robust predictions on fine-grained human object interactions with interpretable outputs. Extensive experimental evaluations on public benchmarks show that our model outperforms prior methods by a considerable margin, demonstrating its efficacy in handling complex scenes. Desen Zhou, Yongfei Liu, Rongjie Li, Xuming He 0001 |
ICCV | 5 |
| 2019 | LatentGNN: Learning Efficient Non-local Relations for Visual RecognitionabstractCapturing long-range dependencies in feature representations is crucial for many visual recognition tasks. Despite recent successes of deep convolutional networks, it remains challenging to model non-local context relations between visual features. A promising strategy is to model the feature context by a fully-connected graph neural network (GNN), which augments traditional convolutional features with an estimated non-local context representation. However, most GNN-based approaches require computing a dense graph affinity matrix and hence have difficulty in scaling up to tackle complex real-world visual problems. In this work, we propose an efficient and yet flexible non-local relation representation based on a novel class of graph neural networks. Our key idea is to introduce a latent space to reduce the complexity of graph, which allows us to use a low-rank representation for the graph affinity matrix and to achieve a linear complexity in computation. Extensive experimental evaluations on three major visual recognition tasks show that our method outperforms the prior works with a large margin while maintaining a low computation cost. Songyang Zhang 0001, Xuming He 0001, Shipeng Yan |
ICML | 2 |
| 2018 | 3D Box Proposals From a Single Monocular Image of an Indoor SceneabstractModern object detection methods typically rely on bounding box proposals as input. While initially popularized in the 2D case, this idea has received increasing attention for 3D bounding boxes. Nevertheless, existing 3D box proposal techniques all assume having access to depth as input, which is unfortunately not always available in practice. In this paper, we therefore introduce an approach to generating 3D box proposals from a single monocular RGB image. To this end, we develop an integrated, fully differentiable framework that inherently predicts a depth map, extracts a 3D volumetric scene representation and generates 3D object proposals. At the core of our approach lies a novel residual, differentiable truncated signed distance function module, which, accounting for the relatively low accuracy of the predicted depth map, extracts a 3D volumetric representation of the scene. Our experiments on the standard NYUv2 dataset demonstrate that our framework lets us generate high-quality 3D box proposals and that it outperforms the two-stage technique consisting of successively performing state-of-the-art depth prediction and depth-based 3D proposal generation. Wei Zhuo 0004, Mathieu Salzmann, Xuming He 0001, Miaomiao Liu 0001 |
AAAI | 3 |
| 2018 | 3D Object Structure Recovery via Semi-supervised Learning on Videos
Qian He 0001, Desen Zhou, Xuming He 0001 |
BMVC | 3 |
| 2018 | Geometry-Aware Deep Network for Single-Image Novel View SynthesisabstractThis paper tackles the problem of novel view synthesis from a single image. In particular, we target real-world scenes with rich geometric structure, a challenging task due to the large appearance variations of such scenes and the lack of simple 3D models to represent them. Modern, learning-based approaches mostly focus on appearance to synthesize novel views and thus tend to generate predictions that are inconsistent with the underlying scene structure. By contrast, in this paper, we propose to exploit the 3D geometry of the scene to synthesize a novel view. Specifically, we approximate a real-world scene by a fixed number of planes, and learn to predict a set of homographies and their corresponding region masks to transform the input image into a novel view. To this end, we develop a new region-aware geometric transform network that performs these multiple tasks in a common framework. Our results on the outdoor KITTI and the indoor ScanNet datasets demonstrate the effectiveness of our network in generating high-quality synthetic views that respect the scene geometry, thus outperforming the state-of-the-art methods. Miaomiao Liu 0001, Xuming He 0001, Mathieu Salzmann |
CVPR | 2 |
| 2018 | SemStyle: Learning to Generate Stylised Image Captions Using Unaligned TextabstractLinguistic style is an essential part of written communication, with the power to affect both clarity and attractiveness. With recent advances in vision and language, we can start to tackle the problem of generating image captions that are both visually grounded and appropriately styled. Existing approaches either require styled training captions aligned to images or generate captions with low relevance. We develop a model that learns to generate visually relevant styled captions from a large corpus of styled text without aligned images. The core idea of this model, called SemStyle, is to separate semantics and style. One key component is a novel and concise semantic term representation generated using natural language processing techniques and frame semantics. In addition, we develop a unified language model that decodes sentences with diverse word choices and syntax for different styles. Evaluations, both automatic and manual, show captions from SemStyle preserve image semantics, are descriptive, and are style shifted. More broadly, this work provides possibilities to learn richer image descriptions from the plethora of linguistic data available on the web. Alexander Patrick Mathews, Lexing Xie, Xuming He 0001 |
CVPR | 3 |
| 2018 | One-Shot Action Localization by Learning Sequence Matching NetworkabstractLearning based temporal action localization methods require vast amounts of training data. However, such large-scale video datasets, which are expected to capture the dynamics of every action category, are not only very expensive to acquire but are also not practical simply because there exists an uncountable number of action classes. This poses a critical restriction to the current methods when the training samples are few and rare (e.g. when the target action classes are not present in the current publicly available datasets). To address this challenge, we conceptualize a new example-based action detection problem where only a few examples are provided, and the goal is to find the occurrences of these examples in an untrimmed video sequence. Towards this objective, we introduce a novel one-shot action localization method that alleviates the need for large amounts of training samples. Our solution adopts the one-shot learning technique of Matching Network and utilizes correlations to mine and localize actions of previously unseen classes. We evaluate our one-shot action localization method on the THUMOS14 and ActivityNet datasets, of which we modified the configuration to fit our one-shot problem setup. Xuming He 0001, Fatih Porikli |
CVPR | 2 |
| 2018 | Instance-Aware Detailed Action Labeling in VideosabstractWe address the problem of detailed sequence labeling of complex activities in videos, which aims to assign an action label to every frame. Previous work typically focus on predicting action class labels for each frame in a sequence without reasoning action instances. However, such category-level labeling is inefficient in encoding the global constraints at the action instance level and tends to produce inconsistent results. In this work we consider a fusion approach that exploits the synergy between action detection and sequence labeling for complex activities. To this end, we propose an instance-aware sequence labeling method that utilizes the cues from action instance detection. In particular, we design an LSTM-based fusion network that integrates framewise action labeling and action instance prediction to produce a final consistent labeling. To evaluate our method, we create a large-scale RGBD video dataset on gym activities for sequence labeling and action detection called GADD. The experimental results on GADD dataset show that our method outperforms all the state-of-the-art methods consistently in terms of labeling accuracy. Xuming He 0001, Fatih Porikli |
WACV | 2 |
| 2018 | Learning to refine depth for robust stereo estimation
Feiyang Cheng, Xuming He 0001, Hong Zhang 0018 |
Pattern Recognit. | 2 |
| 2017 | Boundary-Aware Instance SegmentationabstractWe address the problem of instance-level semantic segmentation, which aims at jointly detecting, segmenting and classifying every individual object in an image. In this context, existing methods typically propose candidate objects, usually as bounding boxes, and directly predict a binary mask within each such proposal. As a consequence, they cannot recover from errors in the object candidate generation process, such as too small or shifted boxes. In this paper, we introduce a novel object segment representation based on the distance transform of the object masks. We then design an object mask network (OMN) with a new residual-deconvolution architecture that infers such a representation and decodes it into the final binary object mask. This allows us to predict masks that go beyond the scope of the bounding boxes and are thus robust to inaccurate object candidates. We integrate our OMN into a Multitask Network Cascade framework, and learn the resulting boundary-aware instance segmentation (BAIS) network in an end-to-end manner. Our experiments on the PASCAL VOC 2012 and the Cityscapes datasets demonstrate the benefits of our approach, which outperforms the state-of-the-art in both object proposal generation and instance segmentation. Zeeshan Hayder, Xuming He 0001, Mathieu Salzmann |
CVPR | 2 |
| 2017 | Predicting Salient Face in Multiple-Face VideosabstractAlthough the recent success of convolutional neural network (CNN) advances state-of-the-art saliency prediction in static images, few work has addressed the problem of predicting attention in videos. On the other hand, we find that the attention of different subjects consistently focuses on a single face in each frame of videos involving multiple faces. Therefore, we propose in this paper a novel deep learning (DL) based method to predict salient face in multiple-face videos, which is capable of learning features and transition of salient faces across video frames. In particular, we first learn a CNN for each frame to locate salient face. Taking CNN features as input, we develop a multiple-stream long short-term memory (M-LSTM) network to predict the temporal transition of salient faces in video sequences. To evaluate our DL-based method, we build a new eye-tracking database of multiple-face videos. The experimental results show that our method outperforms the prior state-of-the-art methods in predicting visual attention on faces in multiple-face videos. Yufan Liu 0001, Songyang Zhang 0001, Mai Xu, Xuming He 0001 |
CVPR | 4 |
| 2017 | Indoor Scene Parsing with Instance Segmentation, Semantic Labeling and Support Relationship InferenceabstractOver the years, indoor scene parsing has attracted a growing interest in the computer vision community. Existing methods have typically focused on diverse subtasks of this challenging problem. In particular, while some of them aim at segmenting the image into regions, such as object or surface instances, others aim at inferring the semantic labels of given regions, or their support relationships. These different tasks are typically treated as separate ones. However, they bear strong connections: good regions should respect the semantic labels, support can only be defined for meaningful regions, support relationships strongly depend on semantics. In this paper, we therefore introduce an approach to jointly segment the instances and infer their semantic labels and support relationships from a single input image. By exploiting a hierarchical segmentation, we formulate our problem as that of jointly finding the regions in the hierarchy that correspond to instances and estimating their class labels and pairwise support relationships. We express this via a Markov Random Field, which allows us to further encode links between the different types of variables. Inference in this model can be done exactly via integer linear programming, and we learn its parameters in a structural SVM framework. Our experiments on NYUv2 demonstrate the benefits of reasoning jointly about all these subtasks of indoor scene parsing. Wei Zhuo 0004, Mathieu Salzmann, Xuming He 0001, Miaomiao Liu 0001 |
CVPR | 3 |
| 2017 | Deep Free-Form Deformation Network for Object-Mask Registration
Xuming He 0001 |
ICCV | 2 |
| 2017 | Learning deep structured network for weakly supervised change detectionabstractConventional change detection methods require a large number of images to learn background models or depend on tedious pixel-level labeling by humans. In this paper, we present a weakly supervised approach that needs only image-level labels to simultaneously detect and localize changes in a pair of images. To this end, we employ a deep neural network with DAG topology to learn patterns of change from image-level labeled training data. On top of the initial CNN activations, we define a CRF model to incorporate the local differences and context with the dense connections between individual pixels. We apply a constrained mean-field algorithm to estimate the pixel-level labels, and use the estimated labels to update the parameters of the CNN in an iterative EM framework. This enables imposing global constraints on the observed foreground probability mass function. Our evaluations on four benchmark datasets demonstrate superior detection and localization performance. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
IJCAI | 2 |
| 2017 | Learning Spatial Transforms for Refining Object Segment ProposalsabstractWe address the problem of object segment proposal generation, which is a critical step in many instance-level semantic segmentation and scene understanding pipelines. In contrast to prior works that predict binary segment masks from images, we take an alternative refinement approach to improve the quality of a given segment candidate pool. In particular, we propose an efficient deep network that learns 2D spatial transforms to warp an initial object mask towards nearby object region. We formulate this segment refinement task as a regression problem and design a novel feature pooling strategy in our deep network to predict an affine transformation for each object mask. We evaluate our method extensively on two challenging public benchmarks and apply our refinement network to three different initial segment proposal settings. Our results show sizable improvements in average recall across all the settings, achieving the state-of-the-art performances. Xuming He 0001, Fatih Porikli |
WACV | 2 |
| 2017 | Forest Change Detection in Incomplete Satellite Images With Deep Neural NetworksabstractLand cover change monitoring is an important task from the perspective of regional resource monitoring, disaster management, land development, and environmental planning. In this paper, we analyze imagery data from remote sensing satellites to detect forest cover changes over a period of 29 years (1987-2015). Since the original data are severely incomplete and contaminated with artifacts, we first devise a spatiotemporal inpainting mechanism to recover the missing surface reflectance information. The spatial filling process makes use of the available data of the nearby temporal instances followed by a sparse encoding-based reconstruction. We formulate the change detection task as a region classification problem. We build a multiresolution profile (MRP) of the target area and generate a candidate set of bounding-box proposals that enclose potential change regions. In contrast to existing methods that use handcrafted features, we automatically learn region representations using a deep neural network in a data-driven fashion. Based on these highly discriminative representations, we determine forest changes and predict their onset and offset timings by labeling the candidate set of proposals. Our approach achieves the state-of-the-art average patch classification rate of 91.6% (an improvement of ~16%) and the mean onset/offset prediction error of 4.9 months (an error reduction of five months) compared with a strong baseline. We also qualitatively analyze the detected changes in the unlabeled image regions, which demonstrate that the proposed forest change detection approach is scalable to new regions. Salman Khan 0001, Xuming He 0001, Fatih Porikli, Mohammed Bennamoun |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2017 | Stacked Learning to Search for Scene LabelingabstractSearch-based structured prediction methods have shown promising successes in both computer vision and natural language processing recently. However, most existing search-based approaches lead to a complex multi-stage learning process, which is ill-suited for scene labeling problems with a high-dimensional output space. In this paper, a stacked learning to search method is proposed to address scene labeling tasks. We design a simplified search process consisting of a sequence of ranking functions, which are learned based on a stacked learning strategy to prevent over-fitting. Our method is able to encode rich prior knowledge by incorporating a variety of local and global scene features. In addition, we estimate a labeling confidence map to further improve the search efficiency from two aspects: first, it constrains the search space more effectively by pruning out low-quality solutions based on confidence scores; second, we employ the confidence map as an additional ranking feature to improve its prediction performance and thus reduce the search steps. Our approach is evaluated on both semantic segmentation and geometric labeling tasks, including the Stanford Background, Sift Flow, Geometric Context and NYUv2 RGB-D dataset. The competitive results demonstrate that our stacked learning to search method provides an effective alternative paradigm for scene labeling. Feiyang Cheng, Xuming He 0001, Hong Zhang 0018 |
IEEE Trans. Image Process. | 2 |
| 2016 | SentiCap: Generating Image Descriptions with SentimentsabstractThe recent progress on image recognition and language modeling is making automatic description of image content a reality. However, stylized, non-factual aspects of the written description are missing from the current systems. One such style is descriptions with emotions, which is commonplace in everyday communication, and influences decision-making and interpersonal relationships. We design a system to describe an image with emotions, and present a model that automatically generates captions with positive or negative sentiments. We propose a novel switching recurrent neural network with word-level regularization, which is able to produce emotional image captions using only 2000+ training sentences containing sentiments. We evaluate the captions with different automatic and crowd-sourcing metrics. Our model compares favourably in common quality metrics for image captioning. In 84.6% of cases the generated positive captions were judged as being at least as descriptive as the factual captions. Of these positive captions 88% were confirmed by the crowd-sourced workers as having the appropriate sentiment. Alexander Patrick Mathews, Lexing Xie, Xuming He 0001 |
AAAI | 3 |
| 2016 | Object-Aware Dictionary Learning with Deep Features
Yurui Xie, Fatih Porikli, Xuming He 0001 |
ACCV (2) | 3 |
| 2016 | Learning to Generate Object Segment Proposals with Multi-modal Cues
Xuming He 0001, Fatih Porikli |
ACCV (1) | 2 |
| 2016 | Learning to Co-Generate Object Proposals with a Deep Structured NetworkabstractGenerating object proposals has become a key component of modern object detection pipelines. However, most existing methods generate the object candidates independently of each other. In this paper, we present an approach to co-generating object proposals in multiple images, thus leveraging the collective power of multiple object candidates. In particular, we introduce a deep structured network that jointly predicts the objectness scores and the bounding box locations of multiple object candidates. Our deep structured network consists of a fully-connected Conditional Random Field built on top of a set of deep Convolutional Neural Networks, which learn features to model both the individual object candidates and the similarity between multiple candidates. To train our deep structured network, we develop an end-to-end learning algorithm that, by unrolling the CRF inference procedure, lets us backpropagate the loss gradient throughout the entire structured network. We demonstrate the effectiveness of our approach on two benchmark datasets, showing significant improvement over state-of-the-art object proposal algorithms. Zeeshan Hayder, Xuming He 0001, Mathieu Salzmann |
CVPR | 2 |
| 2016 | Learning Dynamic Hierarchical Models for Anytime Scene Labeling
Buyu Liu, Xuming He 0001 |
ECCV (6) | 2 |
| 2016 | Building Scene Models by Completing and Hallucinating Depth and Semantics
Miaomiao Liu 0001, Xuming He 0001, Mathieu Salzmann |
ECCV (6) | 2 |
| 2016 | Semantic context and depth-aware object proposal generationabstractThis paper presents a context-aware object proposal generation method for stereo images. Unlike existing methods which mostly rely on image-based or depth features to generate object candidates, we propose to incorporate additional geometric and high-level semantic context information into the proposal generation. Our method starts from an initial object proposal set, and encode objectness for each proposal using three types of features , including a CNN feature, a geometric feature computed from dense depth map, and a semantic context feature from pixel-wise scene labeling. We then train an efficient random forest classifier to re-rank the initial proposals and a set of linear regressors to fine-tune the location of each proposal. Experiments on the KITTI dataset show our approach significantly improves the quality of the initial proposals and achieves the state-of-the-art performance using only a fraction of original object candidates. Xuming He 0001, Fatih Porikli, Laurent Kneip |
ICIP | 2 |
| 2016 | Learning Hough Transform with Latent Structures for Joint Object Detection and Pose Estimation
Xuming He 0001, Nick Barnes, Mingwen Wang 0001 |
MMM (2) | 2 |
| 2016 | Contour Completion Without Region SegmentationabstractContour completion plays an important role in visual perception, where the goal is to group fragmented low-level edge elements into perceptually coherent and salient contours. Most existing methods for contour completion have focused on pixelwise detection accuracy. In contrast, fewer methods have addressed the global contour closure effect, despite psychological evidences for its importance. This paper proposes a purely contour-based higher order CRF model to achieve contour closure, through local connectedness approximation. This leads to a simplified problem structure, where our higher order inference problem can be transformed into an integer linear program and be solved efficiently. Compared with the methods based on the same bottom-up edge detector, our method achieves a superior contour grouping ability (measured by Rand index), a comparable precision-recall performance, and more visually pleasing results. Our results suggest that contour closure can be effectively achieved in contour domain, in contrast to a popular view that segmentation is essential for this purpose. Yansheng Ming, Hongdong Li, Xuming He 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Separating objects and clutter in indoor scenesabstractObjects' spatial layout estimation and clutter identification are two important tasks to understand indoor scenes. We propose to solve both of these problems in a joint framework using RGBD images of indoor scenes. In contrast to recent approaches which focus on either one of these two problems, we perform ‘fine grained structure categorization’ by predicting all the major objects and simultaneously labeling the cluttered regions. A conditional random field model is proposed to incorporate a rich set of local appearance, geometric features and interactions between the scene elements. We take a structural learning approach with a loss of 3D localisation to estimate the model parameters from a large annotated RGBD dataset, and a mixed integer linear programming formulation for inference. We demonstrate that our approach is able to detect cuboids and estimate cluttered regions across many different object and scene categories in the presence of occlusion, illumination and appearance variations. Salman Khan 0001, Xuming He 0001, Mohammed Bennamoun, Ferdous Sohel, Roberto Togneri |
CVPR | 2 |
| 2015 | Multiclass semantic video segmentation with object-level active inferenceabstractWe address the problem of integrating object reasoning with supervoxel labeling in multiclass semantic video segmentation. To this end, we first propose an object-augmented dense CRF in spatio-temporal domain, which captures long-range dependency between supervoxels, and imposes consistency between object and supervoxel labels. We develop an efficient mean field inference algorithm to jointly infer the supervoxel labels, object activations and their occlusion relations for a moderate number of object hypotheses. To scale up our method, we adopt an active inference strategy to improve the efficiency, which adaptively selects object subgraphs in the object-augmented dense CRF. We formulate the problem as a Markov Decision Process, which learns an approximate optimal policy based on a reward of accuracy improvement and a set of well-designed model and input features. We evaluate our method on three publicly available multiclass video semantic segmentation datasets and demonstrate superior efficiency and accuracy. Buyu Liu, Xuming He 0001 |
CVPR | 2 |
| 2015 | Indoor scene structure analysis for single image depth estimationabstractWe tackle the problem of single image depth estimation, which, without additional knowledge, suffers from many ambiguities. Unlike previous approaches that only reason locally, we propose to exploit the global structure of the scene to estimate its depth. To this end, we introduce a hierarchical representation of the scene, which models local depth jointly with mid-level and global scene structures. We formulate single image depth estimation as inference in a graphical model whose edges let us encode the interactions within and across the different layers of our hierarchy. Our method therefore still produces detailed depth estimates, but also leverages higher-level information about the scene. We demonstrate the benefits of our approach over local depth estimation methods on standard indoor datasets. Wei Zhuo 0004, Mathieu Salzmann, Xuming He 0001, Miaomiao Liu 0001 |
CVPR | 3 |
| 2015 | Structural Kernel Learning for Large Scale Multiclass Object Co-detectionabstractExploiting contextual relationships across images has recently proven key to improve object detection. The resulting object co-detection algorithms, however, fail to exploit the correlations between multiple classes and, for scalability reasons are limited to modeling object instance similarity with relatively low-dimensional hand-crafted features. Here, we address the problem of multiclass object co-detection for large scale datasets. To this end, we formulate co-detection as the joint multiclass labeling of object candidates obtained in a class-independent manner. To exploit the correlations between objects, we build a fully-connected CRF on the candidates, which explicitly incorporates both geometric layout relations across object classes and similarity relations across multiple images. We then introduce a structural boosting algorithm that lets us exploits rich, high-dimensional deep network features to learn object similarity within our fully-connected CRF. Our experiments on PASCAL VOC 2007 and 2012 evidences the benefits of our approach over object detection with RCNN, single-image CRF methods and state-of-the-art co-detection algorithms. Zeeshan Hayder, Xuming He 0001, Mathieu Salzmann |
ICCV | 2 |
| 2015 | Multi-class Semantic Video Segmentation with Exemplar-Based Object ReasoningabstractWe tackle the problem of semantic segmentation of dynamic scene in video sequences. We propose to incorporate foreground object information into pixel labeling by jointly reasoning semantic labels of super-voxels, object instance tracks and geometric relations between objects. We take an exemplar approach to object modeling by using a small set of object annotations and exploring the temporal consistency of object motion. After generating a set of moving object hypotheses, we design a CRF framework that jointly models the super voxel and object instances. The optimal semantic labeling is inferred by the MAP estimation of the model, which is solved by a single move-making based optimization procedure. We demonstrate the effectiveness of our method on three public datasets and show that our model can achieve superior or comparable results than the state of-the-art with less object-level supervision. Buyu Liu, Xuming He 0001, Stephen Gould |
WACV | 2 |
| 2015 | Choosing Basic-Level Concept Names Using Visual and Language ContextabstractWe study basic-level categories for describing visual concepts, and empirically observe context-dependant basic level names across thousands of concepts. We propose methods for predicting basic-level names using a series of classification and ranking tasks, producing the first large scale catalogue of basic-level names for hundreds of thousands of images depicting thousands of visual concepts. We also demonstrate the usefulness of our method with a picture-to-word task, showing strong improvement over recent work by Ordonez et al, by modeling of both visual and language context. Our study suggests that a model for naming visual concepts is an important part of any automatic image/video captioning and visual story-telling system. Alexander Patrick Mathews, Lexing Xie, Xuming He 0001 |
WACV | 3 |
| 2015 | Motion Segmentation of Truncated Signed Distance Function Based Volumetric SurfacesabstractTruncated signed distance function (TSDF) based volumetric surface reconstructions of static environments can be readily acquired using recent RGB-D camera based mapping systems. If objects in the environment move then a previously obtained TSDF reconstruction is no longer current. Handling this problem requires segmenting moving objects from the reconstruction. To this end, we present a novel solution to the motion segmentation of TSDF volumes. The segmentation problem is cast as CRF-based MAP inference in the voxel space. We propose: a novel data term by solving sparse multi-body motion segmentation and computing likelihoods for each motion label in the RGB-D image space, and, a novel pairwise term based on gradients of the TSDF volume. Experimental evaluation shows that the proposed approach achieves successful segmentations on reconstructions acquired with Kinect Fusion. Unlike the existing solutions which only work if the objects move completely from their initially occupied spaces, the proposed method permits segmentation of objects when they start to move. Samunda Perera, Nick Barnes, Xuming He 0001, Shahram Izadi, Pushmeet Kohli, Ben Glocker |
WACV | 3 |
| 2015 | Winding Number Constrained Contour DetectionabstractSalient contour detection can benefit from the integration of both contour cues and region cues. However, this task is difficult due to different nature of region representations and contour representations. To solve this problem, this paper proposes an energy minimization framework based on winding number constraints. In this framework, both region cues, such as color/texture homogeneity, and contour cues, such as local contrast and continuity, are represented in a joint objective function, which has both region and contour labels. The key problem is how to design constraints that ensure the topological consistency of the two kinds of labels. Our technique is based on the topological concept of winding number. Using a fast method for winding number computation, a small number of linear constraints are derived to ensure label consistency. Our method is instantiated by ratio-based energy functions. By successfully integrating both region and contour cues, our method shows advantages over competitive methods. Our method is extended to incorporate user interaction, which leads to further improvements. Yansheng Ming, Hongdong Li, Xuming He 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Robust Face Alignment Under Occlusion via Regional Predictive Power EstimationabstractFace alignment has been well studied in recent years, however, when a face alignment model is applied on facial images with heavy partial occlusion, the performance deteriorates significantly. In this paper, instead of training an occlusion-aware model with visibility annotation, we address this issue via a model adaptation scheme that uses the result of a local regression forest (RF) voting method. In the proposed scheme, the consistency of the votes of the local RF in each of several oversegmented regions is used to determine the reliability of predicting the location of the facial landmarks. The latter is what we call regional predictive power (RPP). Subsequently, we adapt a holistic voting method (cascaded pose regression based on random ferns) by putting weights on the votes of each fern according to the RPP of the regions used in the fern tests. The proposed method shows superior performance over existing face alignment models in the most challenging data sets (COFW and 300-W). Moreover, it can also estimate with high accuracy (72.4% overlap ratio) which image areas belong to the face or nonface objects, on the heavily occluded images of the COFW data set, without explicit occlusion modeling. Heng Yang 0001, Xuming He 0001, Xuhui Jia, Ioannis Patras |
IEEE Trans. Image Process. | 2 |
| 2014 | An Exemplar-Based CRF for Multi-instance Object SegmentationabstractWe address the problem of joint detection and segmentation of multiple object instances in an image, a key step towards scene understanding. Inspired by data-driven methods, we propose an exemplar-based approach to the task of instance segmentation, in which a set of reference image/shape masks is used to find multiple objects. We design a novel CRF framework that jointly models object appearance, shape deformation, and object occlusion. To tackle the challenging MAP inference problem, we derive an alternating procedure that interleaves object segmentation and shape/appearance adaptation. We evaluate our method on two datasets with instance labels and show promising results. Xuming He 0001, Stephen Gould |
CVPR | 1 |
| 2014 | Discrete-Continuous Depth Estimation from a Single ImageabstractIn this paper, we tackle the problem of estimating the depth of a scene from a single image. This is a challenging task, since a single image on its own does not provide any depth cue. To address this, we exploit the availability of a pool of images for which the depth is known. More specifically, we formulate monocular depth estimation as a discrete-continuous optimization problem, where the continuous variables encode the depth of the superpixels in the input image, and the discrete ones represent relationships between neighboring superpixels. The solution to this discrete-continuous optimization problem is then obtained by performing inference in a graphical model using particle belief propagation. The unary potentials in this graphical model are computed by making use of the images with known depth. We demonstrate the effectiveness of our model in both the indoor and outdoor scenarios. Our experimental evaluation shows that our depth estimates are more accurate than existing methods on standard datasets. Miaomiao Liu 0001, Mathieu Salzmann, Xuming He 0001 |
CVPR | 3 |
| 2014 | Superpixel Graph Label Transfer with Learned Distance Metric
Stephen Gould, Jiecheng Zhao, Xuming He 0001, Yuhang Zhang 0001 |
ECCV (1) | 3 |
| 2014 | Object Co-detection via Efficient Inference in a Fully-Connected CRF
Zeeshan Hayder, Mathieu Salzmann, Xuming He 0001 |
ECCV (3) | 3 |
| 2014 | Joint semantic and geometric segmentation of videos with a stage modelabstractWe address the problem of geometric and semantic consistent video segmentation for outdoor scenes. With no assumption on camera movement, we jointly model the semantic-geometric class of spatio-temporal regions (supervoxels) and geometric scene layout in each frame. Our main contribution is to propose a stage scene model to efficiently capture the dependency between the semantic and geometric labels. We build a unified CRF model on supervoxel labels and stage parameters, and design an alternating inference algorithm to minimize the resulting energy function. We also extend smoothing based on hierarchical image segmentation to spatio-temporal setting and show it achieves better performance than a pairwise random field model. Our method is evaluated on the CamVid dataset and achieves state-of-the-art per-pixel as well as per-class accuracy in predicting both semantic and geometric labels. Buyu Liu, Xuming He 0001, Stephen Gould |
WACV | 2 |
| 2013 | Winding Number for Region-Boundary Consistent Salient Contour ExtractionabstractThis paper aims to extract salient closed contours froman image. For this vision task, both region segmentation cues (e.g. color/texture homogeneity) and boundary detection cues (e.g. local contrast, edge continuity and contour closure) play important and complementary roles. In this paper we show how to combine both cues in a unified framework. The main focus is given to how to maintain the consistency (compatibility) between the region cues and the boundary cues. To this ends, we introduce the use of winding number-a well-known concept in topology-as a powerful mathematical device. By this device, the region-boundary consistency is represented as aset of simple linear relationships. Our method is applied to the figure-ground segmentation problem. The experiments show clearly improved results. Yansheng Ming, Hongdong Li, Xuming He 0001 |
CVPR | 3 |
| 2013 | Learning Structured Hough Voting for Joint Object Detection and Occlusion ReasoningabstractWe propose a structured Hough voting method for detecting objects with heavy occlusion in indoor environments. First, we extend the Hough hypothesis space to include both object location and its visibility pattern, and design a new score function that accumulates votes for object detection and occlusion prediction. In addition, we explore the correlation between objects and their environment, building a depth-encoded object-context model based on RGB-D data. Particularly, we design a layered context representation and allow image patches from both objects and backgrounds voting for the object hypotheses. We demonstrate that using a data-driven 2.1D representation we can learn visual codebooks with better quality, and more interpretable detection results in terms of spatial relationship between objects and viewer. We test our algorithm on two challenging RGB-D datasets with significant occlusion and intraclass variation, and demonstrate the superior performance of our method. Tao Wang 0047, Xuming He 0001, Nick Barnes |
CVPR | 2 |
| 2013 | Symmetry detection via contour groupingabstractThis paper presents a simple but effective model for detecting the symmetric axes of bilaterally symmetric objects in unsegmented natural scene images. Our model constructs a directed graph of symmetry interaction. Every node in the graph represents a matched pair of features, and every directed edge represents the interaction between nodes. The bilateral symmetry detection problem is then formulated as finding the star subgraph with maximal weight. The star structure ensures the consistency between grouped nodes while the optimal star subgraph can be found in polynomial time. Our model makes prediction based on contour cue: each node in the graph represents a pair of edge segments. Compared with the Loy and Eklundh's method which used SIFT feature, our model can often produce better results for the images containing limited texture. This advantage is demonstrated on two natural scene image sets. Yansheng Ming, Hongdong Li, Xuming He 0001 |
ICIP | 3 |
| 2013 | Glass object segmentation by label transfer on joint depth and appearance manifoldsabstractWe address the glass object localization problem with a RGB-D camera. Our approach uses a nonparametric, data-driven label transfer scheme for local glass boundary estimation. A weighted voting scheme based on a joint feature manifold is adopted to integrate depth and appearance cues, and we learn a distance metric on the depth-encoded feature manifold. Local boundary evidence is then integrated into a MRF framework for spatially coherent glass object detection and segmentation. The efficacy of our approach is verified on a challenging RGB-D glass dataset where we obtained a clear improvement over the state-of-the-art both in terms of accuracy and speed. Tao Wang 0047, Xuming He 0001, Nick Barnes |
ICIP | 2 |
| 2013 | Picture tags and world knowledge: learning tag relations from visual semantic sourcesabstractThis paper studies the use of everyday words to describe images. The common saying has it that 'a picture is worth a thousand words', here we ask which thousand? The proliferation of tagged social multimedia data presents a challenge to understanding collective tag-use at large scale -- one can ask if patterns from photo tags help understand tag-tag relations, and how it can be leveraged to improve visual search and recognition. We propose a new method to jointly analyze three distinct visual knowledge resources: Flickr, ImageNet/WordNet, and ConceptNet. This allows us to quantify the visual relevance of both tags learn their relationships. We propose a novel network estimation algorithm, Inverse Concept Rank, to infer incomplete tag relationships. We then design an algorithm for image annotation that takes into account both image and tag features. We analyze over 5 million photos with over 20,000 visual tags. The statistics from this collection leads to good results for image tagging, relationship estimation, and generalizing to unseen tags. This is a first step in analyzing picture tags and everyday semantic knowledge. Potential other applications include generating natural language descriptions of pictures, as well as validating and supplementing knowledge databases. Lexing Xie, Xuming He 0001 |
ACM Multimedia | 2 |
| 2013 | Tracking Large-Scale Video Remix in Real-World EventsabstractContent sharing networks, such as YouTube, contain traces of both explicit online interactions (such as likes, comments, or subscriptions), as well as latent interactions (such as quoting, or remixing, parts of a video). We propose visual memes, or frequently re-posted short video segments, for detecting and monitoring such latent video interactions at scale. Visual memes are extracted by scalable detection algorithms that we develop, with high accuracy. We further augment visual memes with text, via a statistical model of latent topics. We model content interactions on YouTube with visual memes, defining several measures of influence and building predictive models for meme popularity. Experiments are carried out with over 2 million video shots from more than 40,000 videos on two prominent news events in 2009: the election in Iran and the swine flu epidemic. In these two events, a high percentage of videos contain remixed content, and it is apparent that traditional news media and citizen journalists have different roles in disseminating remixed content. We perform two quantitative evaluations for annotating visual memes and predicting their popularity. The proposed joint statistical model of visual memes and words outperforms an alternative concurrence model, with an average error of 2% for predicting meme volume and 17% for predicting meme lifespan. Lexing Xie, Apostol Natsev, Xuming He 0001, John R. Kender, Matthew L. Hill, John R. Smith |
IEEE Trans. Multim. | 3 |
| 2012 | Connected contours: A new contour completion model that respects the closure effectabstractContour Completion plays an important role in visual perception, where the goal is to group fragmented low-level edge elements into perceptually coherent and salient contours. This process is often considered as guided by some middle-level Gestalt principles. Most existing methods for contour completion have focused on utilizing rather local Gestalt laws such as good-continuity and proximity. In contrast, much fewer methods have addressed the global contour closure effect, despite that many psychological evidences have shown the usefulness of closure in perceptual grouping. This paper proposes a novel higher-order CRF model to address the contour closure effect, through local connectedness approximation. This leads to a simplified problem structure, where the higher-order inference can be formulated as an integer linear program (ILP) and solved by an efficient cutting-plane variant. Tested on the BSDS benchmark, our method achieves a comparable precision-recall performance, a superior contour grouping ability (measured by Rand index), and more visually pleasing results, compared with existing methods. Yansheng Ming, Hongdong Li, Xuming He 0001 |
CVPR | 3 |
| 2012 | Glass object localization by joint inference of boundary and depth
Tao Wang 0047, Xuming He 0001, Nick Barnes |
ICPR | 2 |
| 2011 | Efficient Image Denoising by MRF Approximation with Uniform-Sampled Multi-spanning-treeabstractTraditionally, image processing based on Markov Random Field (MRF) is often addressed on a 4-connected grid graph defined on the image. This structure is not computationally efficient. In our work, we develop a multiple-trees structure to approximate the 4-connected grid. A set of spanning trees are generated by a new algorithm: re-weighted random walk (RWRW). This structure effectively covers the original grid and guarantees uniformly distributed occurrence of each edge. Exact maximum a posterior (MAP) inference is performed on each tree structure by dynamic programming and a median filter is chosen to merge the results together. As an important application, image denoising is used to validate our method. Experimentally, our algorithm provides better performance and higher computational efficiency than traditional methods (such as Loopy Belief Propagation) on a 4-connected MRF. Hongdong Li, Xuming He 0001 |
ICIG | 3 |
| 2010 | Occlusion Boundary Detection Using Pseudo-depth
Xuming He 0001, Alan L. Yuille |
ECCV (4) | 1 |
| 2010 | A unified model of short-range and long-range motion perceptionabstractThe human vision system is able to effortlessly perceive both short-range and long-range motion patterns in complex dynamic scenes. Previous work has assumed that two different mechanisms are involved in processing these two types of motion. In this paper, we propose a hierarchical model as a unified framework for modeling both short-range and long-range motion perception. Our model consists of two key components: a data likelihood that proposes multiple motion hypotheses using nonlinear matching, and a hierarchical prior that imposes slowness and spatial smoothness constraints on the motion field at multiple scales. We tested our model on two types of stimuli, random dot kinematograms and multiple-aperture stimuli, both commonly used in human vision research. We demonstrate that the hierarchical model adequately accounts for human performance in psychophysical experiments. Xuming He 0001, Hongjing Lu, Alan L. Yuille |
NIPS | 2 |
| 2008 | Latent topic random fields: Learning using a taxonomy of labelsabstractAn important problem in image labeling concerns learning with images labeled at varying levels of specificity. We propose an approach that can incorporate images with labels drawn from a semantic hierarchy, and can also readily cope with missing labels, and roughly-specified object boundaries. We introduce a new form of latent topic model, learning a novel context representation in the joint label-and-image space by capturing co-occurring patterns within and between image features and object labels. Given a topic, the model generates the input data, as well as a topic-dependent probabilistic classifier to predict labels for image regions. We present results on two real-world datasets, demonstrating significant improvements gained by including the coarsely labeled images. Xuming He 0001, Richard S. Zemel |
CVPR | 1 |
| 2008 | Using latent Dirichlet allocation to incorporate domain knowledge for topic transition detectionabstractThis paper studies automatic detection of topic transitions for recorded presentations. This can be achieved by matching slide content with presentation transcripts directly with some similarity metrics. Such literal matching, however, misses domain-specific knowledge and is sensitive to speech recognition errors. In this paper, we incorporate relevant written materials, e.g., textbooks for lectures, which convey semantic relationships, in particular domain-specific relationships, between words. To this end, we train latent Dirichlet allocation (LDA) models on these materials and measure the similarity between slides and transcripts in the acquired hidden-topic space. This similarity is then combined with literal matchings. Experiments show that the proposed approach reduces the errors in slide transition detection by 17-41 % on manual transcripts and 27-37% on automatic transcripts. Index Terms: slides transition detection, boundary detection. 1. Xiaodan Zhu 0001, Xuming He 0001, Cosmin Munteanu, Gerald Penn |
INTERSPEECH | 2 |
| 2008 | Learning Hybrid Models for Image Annotation with Partially Labeled DataabstractExtensive labeled data for image annotation systems, which learn to assign class labels to image regions, is difficult to obtain. We explore a hybrid model framework for utilizing partially labeled data that integrates a generative topic model for image appearance with discriminative label prediction. We propose three alternative formulations for imposing a spatial smoothness prior on the image labels. Tests of the new models and some baseline approaches on two real image datasets demonstrate the effectiveness of incorporating the latent structure. Xuming He 0001, Richard S. Zemel |
NIPS | 1 |
| 2008 | Learning Flexible Features for Conditional Random FieldsabstractExtending traditional models for discriminative labeling of structured data to include higher-order structure in the labels results in an undesirable exponential increase in model complexity. In this paper, we present a model that is capable of learning such structures using a random field of parameterized features. These features can be functions of arbitrary combinations of observations, labels and auxiliary hidden variables. We also present a simple induction scheme to learn these features, which can automatically determine the complexity needed for a given data set. We apply the model to two real-world tasks, information extraction and image labeling, and compare our results to several other methods for discriminative labeling. Liam Stewart, Xuming He 0001, Richard S. Zemel |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Learning and Incorporating Top-Down Cues in Image Segmentation
Xuming He 0001, Richard S. Zemel, Debajyoti Ray |
ECCV (1) | 1 |
| 2004 | Multiscale Conditional Random Fields for Image Labeling
Xuming He 0001, Richard S. Zemel, Miguel Á. Carreira-Perpiñán |
CVPR (2) | 1 |