VLDB 2026 Research / reviewers in the wild / expert
Ruiping Wang 0001
dblp:60/1529-1
· DBLP profile ↗
105ranked-venue papers
5as first author
44since 2021 · last 2026
0000-0003-1830-2595ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 77 · 4 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 66 · 4 first-author · 21 since 2021Human-computer interaction and ubiquitous computing · 3Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tell as You Want: Customizing Image Narrative with Knowledge and ThoughtsabstractWith the advancement of vision-language models, image captioning has made significant progress, leading to the generation of more accurate and detailed descriptions. Current image captioning primarily focuses on describing the apparent visual characteristics, which are easily observed by most humans, but less helpful in real-world scenarios. When users seek a deeper understanding of visual content, they may be concerned with fine-grained categories, function properties, and other background knowledge, rather than merely appearances. Additionally, as users' interests vary, there is a growing demand for customizable content generation. To address these challenges, we propose the task of image narrative generation, which aims to produce knowledge-rich natural language responses for input images, customized to the user preference. Furthermore, we propose T^4, an image narrative generation model progressing through cascade steps: Tailor, reTrieve, Think, and Tell. Specifically, it takes the image and various types of prompts as input, and first refines or predicts potentially interesting queries that are tailored to the user expertise level. Subsequently, the model enriches contextual knowledge through retrieval-augmentation and employs chain-of-thoughts to decompose the generation process step by step, thereby telling an accurate and logically coherent image narrative. In addition, we construct the ImgNarr-23K dataset to support task training and evaluation. Experimental results demonstrate that the proposed approach generates image narratives that better satisfy user requirements, and achieves state-of-the-art performance in knowledge-based VQA tasks without additional finetuning. T^4 presents a promising solution for customized content generation in specialized domains. Ziwei Yao, Ruiping Wang 0001, Xilin Chen 0001 |
AAAI | 3 |
| 2026 | M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal ModelsabstractMultilingual capability is a crucial requirement for large multimodal models, which are increasingly deployed across diverse countries and languages. However, most existing benchmarks for multilingual multimodal reasoning fail to effectively distinguish models of different strengths; in fact, even text-only language models without visual capabilities can often achieve high scores. As a result, the comprehensive evaluation of state-of-the-art multilingual multimodal models remains underexplored. In this work, we present M4U, a novel and challenging benchmark designed to evaluate multilingual, multi-discipline multimodal understanding and reasoning. M4U comprises 10k samples spanning 64 disciplines across 16 subfields in Science, Engineering, and Healthcare, covering six languages. Using this benchmark, we conduct extensive evaluations of leading Large Multimodal Models (LMMs) and Large Language Models (LLMs) augmented with external tools. Our results reveal that even the strongest LMMs exhibit pronounced language preferences and struggle with reasoning tasks that require integrating multilingual information across visual and textual modalities. In particular, performance drops markedly when models are prompted with cross-lingual multimodal questions, highlighting significant gaps in current multilingual multimodal reasoning capabilities.1 Senwei Xie, Ruiping Wang 0001, Zhaojie Xie, Chuyan Xiong, Xilin Chen 0001 |
WACV | 4 |
| 2026 | DyToS: Budget-aware dynamic token scheduling for efficient multi-modal large language models
Yifei Xing 0001, Ruiping Wang 0001, Dongmei Jiang, Xiangyuan Lan |
Neurocomputing | 3 |
| 2026 | A Survey on Interpretability in Visual RecognitionabstractVisual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand for deployment in safety-critical areas like autonomous driving and medical diagnostics has accelerated the development of eXplainable AI (XAI). Distinct from generic XAI, visual recognition XAI is positioned at the intersection of vision and language, which represent the two most fundamental human modalities and form the cornerstones of multimodal intelligence. This paper provides a systematic survey of XAI in visual recognition by establishing a multi-dimensional taxonomy from a human-centered perspective based on intent, object, presentation, and methodology. Beyond categorization, we summarize critical evaluation desiderata and metrics, conducting an extensive qualitative assessment across different categories and demonstrating quantitative benchmarks within specific dimensions. Furthermore, we explore the interpretability of Multimodal Large Language Models and practical applications, identifying emerging trends and opportunities. By synthesizing these diverse perspectives, this survey provides an insightful roadmap to inspire future research on the interpretability of visual recognition models. Qiyang Wan, Chengzhi Gao, Ruiping Wang 0001, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Learning interpretable binary codes via semantic alignment for customized image retrieval
Shishi Qiao, Ruiping Wang 0001, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2026 | Decoupled gradient-guided stratification for resource-efficient multi-modal data pruning
Yifei Xing 0001, Ruiping Wang 0001, Xiangyuan Lan, Yaowei Wang 0001 |
Pattern Recognit. Lett. | 3 |
| 2025 | R2C: Mapping Room to Chessboard to Unlock LLM As Low-Level Action PlannerabstractThis paper explores using large language models (LLMs) as low-level action planners for embodied tasks. While LLMs excel as the robot’s “brain” for high-level planning, they face challenges in directly controlling the “body” by generating precise low-level actions. This limitation arises from LLMs’ strength in high-level conceptual understanding but their inability to handle spatial perception effectively, restricting their potential in embodied tasks. To address this, we bridge the gap by enabling LLMs to not only comprehend complex instructions but also produce actionable, low-level plans. We introduce Room to Chessboard (R2C), a novel semantic representation that maps environmental states onto a grid-based chessboard, empowering LLMs to generate specific low-level coordinates and guide the robot in a manner akin to playing a game of chess. To further enhance decision-making, we propose the Chain-of-Thought Decision (CoT-D) paradigm, which improves LLMs’ interpretability and context-awareness in spatial reasoning. By jointly training LLMs for high-level task decomposition and low-level action generation, we create a unified "brain-body" system capable of handling complex, free-form instructions while producing precise low-level actions, allowing the robot to flexibly control its movements and adapt to varying tasks. We validate R2C using both fine-tuned open-source LLMs and GPT-4, demonstrating effectiveness on the challenging ALFRED benchmark. Results show that with our R2C framework, LLMs can effectively act as low-level planners, generalizing across diverse settings and open-vocabulary robotic tasks. The code and demonstrations are available at: https://vipl-vsu.github.io/Room2Chessboard. Ziyi Bai, Hanxuan Li, Chuyan Xiong, Ruiping Wang 0001, Xilin Chen 0001 |
CVPR | 5 |
| 2025 | OV3D-CG: Open-Vocabulary 3D Instance Segmentation with Contextual Guidance
Ruiping Wang 0001, Xilin Chen 0001 |
ICCV | 3 |
| 2025 | EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical AlignmentabstractMamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanced cross-modal alignment between visual and textural latents, negatively impacting performance on multi-modal tasks. In this work, we propose Empowering Multi-modal Mamba with Structural and Hierarchical Alignment (EMMA), which enables the MLLM to extract fine-grained visual information. Specifically, we propose a pixel-wise alignment module to autoregressively optimize the learning and processing of spatial image-level features along with textual tokens, enabling structural alignment at the image level. In addition, to prevent the degradation of visual information during the cross-model alignment process, we propose a multi-scale feature fusion (MFF) module to combine multi-scale visual features from intermediate layers, enabling hierarchical alignment at the feature level. Extensive experiments are conducted across a variety of multi-modal benchmarks. Our model shows lower latency than other Mamba-based MLLMs and is nearly four times faster than transformer-based MLLMs of similar scale during inference. Due to better cross-modal alignment, our model exhibits lower degrees of hallucination and enhanced sensitivity to visual details, which manifests in superior performance across diverse multi-modal benchmarks. Code provided at https://github.com/xingyifei2016/EMMA. Yifei Xing 0001, Xiangyuan Lan, Ruiping Wang 0001, Dongmei Jiang, Yaowei Wang 0001 |
ICLR | 3 |
| 2025 | Generic Scene Graph Generation Model with Hierarchical Prompt Learning
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan |
Int. J. Comput. Vis. | 3 |
| 2025 | UniFa: A unified feature hallucination framework for any-shot object detection
Hui Nie 0001, Ruiping Wang 0001, Xilin Chen 0001 |
Pattern Recognit. Lett. | 2 |
| 2025 | Cross-Domain Few-Shot 3D Point Cloud Semantic Segmentation
Jiwei Xiao, Ruiping Wang 0001, Xilin Chen 0001 |
Pattern Recognit. Lett. | 2 |
| 2024 | Point2Real: Bridging the Gap between Point Cloud and Realistic Image for Open-World 3D RecognitionabstractRecognition in open-world scenarios is an important and challenging field, where Vision-Language Pre-training paradigms have greatly impacted the 2D domain. This inspires a growing interest in introducing 2D pre-trained models, such as CLIP, into the 3D domain to enhance the ability of point cloud understanding. Considering the difference between discrete 3D point clouds and real-world 2D images, reducing the domain gap is crucial. Some recent works project point clouds onto a 2D plane to enable 3D zero-shot capabilities without training. However, this simplistic approach leads to an unclear or even distorted geometric structure, limiting the potential of 2D pre-trained models in 3D. To address the domain gap, we propose Point2Real, a training-free framework based on the realistic rendering technique to automate the transformation of the 3D point cloud domain into the Vision-Language domain. Specifically, Point2Real leverages a shape recovery module that devises an iterative ball-pivoting algorithm to convert point clouds into meshes, narrowing the gap in shape at first. To simulate photo-realistic images, a set of refined textures as candidates is applied for rendering, where the CLIP confidence is utilized to select the suitable one. Moreover, to tackle the viewpoint challenge, a heuristic multi-view adapter is implemented for feature aggregation, which exploits the depth surface as an effective indicator of view-specific discriminability for recognition. We conduct experiments on ModelNet10, ModelNet40, and ScanObjectNN datasets, and the results demonstrate that Point2Real outperforms other approaches in zero-shot and few-shot tasks by a large margin. Hanxuan Li, Ruiping Wang 0001, Xilin Chen 0001 |
AAAI | 3 |
| 2024 | Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models
Qiyang Wan, Ruiping Wang 0001, Xilin Chen 0001 |
BMVC | 4 |
| 2024 | Hierarchical Prompt Learning for Scene Graph Generation
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan |
BMVC | 3 |
| 2024 | Think Before Placement: Common Sense Enhanced Transformer for Object Placement
Yaxuan Qin, Ruiping Wang 0001, Xilin Chen 0001 |
ECCV (73) | 3 |
| 2024 | HiFi-Score: Fine-Grained Image Description Evaluation with Hierarchical Parsing Graphs
Ziwei Yao, Ruiping Wang 0001, Xilin Chen 0001 |
ECCV (62) | 2 |
| 2024 | Calibration for Long-tailed Scene Graph GenerationabstractMiscalibrated models tend to be unreliable and insecure for downstream applications. In this work, we attempt to highlight and remedy miscalibration in current scene graph generation (SGG) models, which has been overlooked by previous works. We discover that obtaining well-calibrated models for SGG is more challenging than conventional calibration settings, as long-tailed SGG training data exacerbates miscalibration with overconfidence in head classes and underconfidence in tail classes. We further analyze which components are explicitly impacted by the long-tailed data during optimization, thereby exacerbating miscalibration and unbalanced learning, including biased parameters, deviated boundaries, and distorted target distribution. To address the above issues, we propose the Compositional Optimization Calibration (COC) method, comprising three modules: i. A parameter calibration module that utilizes a hyperspherical classifier to eliminate the bias introduced by biased parameters. ii. A boundary calibration module that disperses features of majority classes to consolidate the decision boundaries of minority classes and mitigate deviated boundaries. iii. A target distribution calibration module that addresses distorted target distribution, leverages within-triplet prior to guide confidence-aware and label-aware target calibration, and applies curriculum regulation to constrain learning focus from easy to hard classes. Extensive evaluation on popular benchmarks demonstrates the effectiveness of our proposed method in improving model calibration and resolving unbalanced learning for long-tailed SGG. Finally, our proposed method performs best on model calibration compared to different types of calibration methods and achieves state-of-the-art trade-off performance on balanced SGG learning. Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan |
ACM Multimedia | 3 |
| 2024 | Interpretable Object Recognition by Semantic Prototype AnalysisabstractPeople can usually give reasons for recognizing a particular object as a specific category, using various means such as body language (by pointing out) and natural language (by telling). This inspires us to develop a recognition model with such principles to explain the recognition process to enhance human trust. We propose Semantic Prototype Analysis Network (SPANet), an interpretable object recognition approach that enables models to explicate the decision process more lucidly and comprehensibly to humans by "pointing out where to focus" and "telling about why it is" simultaneously. With the proposed method, some part prototypes with semantic concepts will be provided to elaborate on the classification together with a group of visualized samples to achieve both part-wise and semantic interpretability. The results of extensive experiments demonstrate that SPANet is able to recognize objects almost as well as the non-interpretable models, at the same time generating intelligible explanations for its decision process. Qiyang Wan, Ruiping Wang 0001, Xilin Chen 0001 |
WACV | 2 |
| 2024 | Introspective GAN: Learning to grow a GAN for incremental generation and classification
Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2024 | Hierarchical image-to-image translation with nested distributions modeling
Shishi Qiao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2024 | Local context attention learning for fine-grained scene graph generationabstractFine-grained scene graph generation aims to parse the objects and their fine-grained relationships within scenes. Despite the significant progress in recent years, their performance is still limited by two major issues: (1) ambiguous perception under a global view; (2) the lack of reliable, fine-grained annotations. We argue that understanding the local context is important in addressing the two issues. However, previous works often overlook it, which limits their effectiveness in fine-grained scene graph generation. To tackle this challenge, we introduce a Local-context Attention Learning method that concentrates on local context and can generate high-reliability, fine-grained annotations. It comprises two components: (1) The Fine-grained Location Attention Network (FLAN), a multi-branch network that encompasses global and local branches, can attend to local informative context and perceive granularity levels in different regions, thereby adaptively enhancing the learning of fine-grained locations. (2) The Fine-grained Location Label Transfer (FLLT) method identifies coarse-grained labels inconsistent with the local context and determines which labels should be transferred through the global confidence thresholding strategy, finally transferring them to reliable local context-consistent fine-grained ones. Experiments conducted on the Visual Genome, OpenImage, and GQA-200 datasets show that the proposed methods achieve significant improvements on the fine-grained scene graph generation task. By addressing the challenge mentioned above, our method also achieves state-of-the-art performances on the three datasets. Xuhan Zhu, Ruiping Wang 0001, Xiangyuan Lan, Yaowei Wang 0001 |
Pattern Recognit. | 2 |
| 2024 | Data-efficient 3D instance segmentation by transferring knowledge from synthetic scans
Ruiping Wang 0001, Xilin Chen 0001 |
Pattern Recognit. Lett. | 2 |
| 2024 | Mind the Gap: Open Set Domain Adaptation via Mutual-to-Separate FrameworkabstractUnsupervised domain adaptation aims to leverage labeled data from a source domain to learn a classifier for an unlabeled target domain. Amongst its many variants, open set domain adaptation (OSDA) is perhaps the most challenging one, as it further assumes the presence of unknown classes in the target domain. In this paper, we study OSDA with a particular focus on enriching its ability to traverse across larger domain gaps, and we show that existing state-of-the-art methods suffer a considerable performance drop in the presence of larger domain gaps, especially on a new dataset (PACS) that we re-purposed for OSDA. Exploring this is pivotal for OSDA as with increasing domain shift, identifying unknown samples in the target domain becomes harder for the model, thus making negative transfer between source and target domains more challenging. Accordingly, we propose a Mutual-to-Separate (MTS) framework to address the larger domain gaps. Essentially we design two networks – (a) Sample Separation Network (SSN): which is trained to learn a hyperplane for separating unknown samples from known ones, and (b) Distribution Matching Network (DMN): which is trained to maximise domain confusion between source and target domains without unknown samples under the guidance of the SSN. The key insight lies in how we exploit the mutually beneficial information between these two networks. On closer observation, we see that SSN can reveal which samples in the target domain belong to the unknown class by instance weighting whereas, DMN pushes apart the samples that most likely belong to the unknown class in the target domain, which in turn reduces the difficulty of SSN in identifying unknown samples. It follows that (a) and (b) will mutually supervise each other and alternate until convergence, which can better align the source and target domains in the shared label space. Extensive experiments on five datasets (Office-31, Office-Home, PACS, VisDA, andmini_DomainNet) demonstrate the efficiency of the proposed method. Detailed ablation experiments also validate the effectiveness of each component and the generality of the proposed framework. Codes are available at: https://github.com/PRIS-CV/Mutual-to-Separate. Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Ruiping Wang 0001, Jun Guo 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Event Graph Guided Compositional Spatial-Temporal Reasoning for Video Question AnsweringabstractVideo question answering (VideoQA) is challenging since it requires the model to extract and combine multi-level visual concepts from local objects to global actions from complex events for compositional reasoning. Existing works represent the video with fixed-duration clip features that make the model struggle in capturing the crucial concepts in multiple granularities. To overcome this shortcoming, we propose to represent the video with an Event Graph in a hierarchical structure whose nodes correspond to visual concepts of different levels (object, relation, scene and action) and edges indicate their spatial-temporal relationships. We further propose a H ierarchical S patial- T emporal T ransformer (HSTT) which takes nodes from the graph as visual input to realize compositional reasoning guided by the event graph. To fully exploit the spatial-temporal context delivered from the graph structure, on the one hand, we encode the nodes in the order of their semantic hierarchy (depth) and occurrence time (breadth) with our improved graph search algorithm; On the other hand, we introduce edge-guided attention to combine the spatial-temporal context among nodes according to their edge connections. HSTT then performs QA by cross-modal interactions guaranteed by the hierarchical correspondence between the multi-level event graph and the cross-level question. Experiments on the recent challenging AGQA and STAR datasets show that the proposed method clearly outperforms the existing VideoQA models by a large margin, including those pre-trained with large-scale external data. Our code is available at https://github.com/ByZ0e/HSTT. Ziyi Bai, Ruiping Wang 0001, Difei Gao, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | Glance and Focus: Memory Prompting for Multi-Event Video Question AnsweringabstractVideo Question Answering (VideoQA) has emerged as a vital tool to evaluate agents’ ability to understand human daily behaviors. Despite the recent success of large vision language models in many multi-modal tasks, complex situation reasoning over videos involving multiple human-object interaction events still remains challenging. In contrast, humans can easily tackle it by using a series of episode memories as anchors to quickly locate question-related key moments for reasoning. To mimic this effective reasoning strategy, we propose the Glance- Focus model. One simple way is to apply an action detection model to predict a set of actions as key memories. However, these actions within a closed set vocabulary are hard to generalize to various video domains. Instead of that, we train an Encoder-Decoder to generate a set of dynamic event memories at the glancing stage. Apart from using supervised bipartite matching to obtain the event memories, we further design an unsupervised memory generation method to get rid of dependence on event annotations. Next, at the focusing stage, these event memories act as a bridge to establish the correlation between the questions with high-level event concepts and low-level lengthy video content. Given the question, the model first focuses on the generated key event memory, then focuses on the most relevant moment for reasoning through our designed multi-level cross- attention mechanism. We conduct extensive experiments on four Multi-Event VideoQA benchmarks including STAR, EgoTaskQA, AGQA, and NExT-QA. Our proposed model achieves state-of-the-art results, surpassing current large models in various challenging reasoning tasks. The code and models are available at https://github.com/ByZ0e/Glance-Focus. Ziyi Bai, Ruiping Wang 0001, Xilin Chen 0001 |
NeurIPS | 2 |
| 2023 | Semantic Guided Latent Parts Embedding for Few-Shot LearningabstractThe ability of few-shot learning (FSL) is a basic requirement of intelligent agent learning in the open visual world. However, existing deep learning systems rely too heavily on large numbers of training samples, making it hard to learn new categories efficiently from limited size of training data. Two key challenges of FSL are insufficient comprehension and imperfect modeling of the few-shot novel class. For insufficient visual comprehension, semantic knowledge which is information from other modalities can help replenish the understanding of novel classes. But even so, most works still suffer from the second challenge because the single global class prototype they adopted is extremely unstable and imperfect given the larger intra-class variation and harder inter-class discrimination in FSL scenario. Thus, we propose to represent each class by its several different parts with the help of class semantic knowledge. Since we can never pre-define parts for unknown novel classes, we embed them in a latent manner. Concretely, we train a generator that takes the class semantic knowledge as input and outputs several filters of class-specific semantic latent parts. By applying each part filter, our model can pay attention to corresponding local regions containing each part. At the inference stage, the classification is conducted by comparing the similarities between those parts. Experiments on several FSL benchmarks demonstrate the effectiveness of our proposed method and show its potential to go beyond class recognition to class understanding. Furthermore, we also find when semantic knowledge is more visualized and customized, it will be more helpful in the FSL task. Fengyuan Yang 0002, Ruiping Wang 0001, Xilin Chen 0001 |
WACV | 2 |
| 2023 | Importance First: Generating Scene Graph of Human Interest
Wenbin Wang 0001, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | CRIC: A VQA Dataset for Compositional Reasoning on Vision and CommonsenseabstractAlternatively inferring on the visual facts and commonsense is fundamental for an advanced visual question answering (VQA) system. This ability requires models to go beyond the literal understanding of commonsense. The system should not just treat objects as the entrance to query background knowledge, but fully ground commonsense to the visual world and imagine the possible relationships between objects, e.g., "fork, can lift, food". To comprehensively evaluate such abilities, we propose a VQA benchmark, Compositional Reasoning on vIsion and Commonsense(CRIC), which introduces new types of questions about CRIC, and an evaluation metric integrating the correctness of answering and commonsense grounding. To collect such questions and rich additional annotations to support the metric, we also propose an automatic algorithm to generate question samples from the scene graph associated with the images and the relevant knowledge graph. We further analyze several representative types of VQA models on the CRIC dataset. Experimental results show that grounding the commonsense to the image region and joint reasoning on vision and commonsense are still challenging for current approaches. The dataset is available at https://cricvqa.github.io. Difei Gao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Hierarchical disentangling network for object representation learningabstractAn object can be described as the combination of primary visual attributes. Disentangling such underlying primitives is the long-term objective of representation learning . It is observed that categories have natural hierarchical characteristics, i.e., any two objects can share some common primitives at a particular category level while possess unique traits at another. However, previous works usually operate in a flat manner (i.e., at a particular level) to disentangle the representations of objects. Even though they may obtain the primitives to constitute objects as the categories at that level, their results are obviously not efficient and complete. In this paper, we propose a Hierarchical Disentangling Network (HDN) to exploit the rich hierarchical characteristics among categories to divide the disentangling process in a coarse-to-fine manner (i.e., level-wise), such that each level only focuses on learning the specific representations and finally the common and unique representations at all levels jointly constitute the raw object. Specifically, HDN is designed based on an encoder-decoder architecture. To simultaneously ensure the level-wise disentanglement and interpretability of the encoded representations, a novel hierarchical Generative Adversarial Network (GAN) is introduced. Quantitative and qualitative evaluations on popular object datasets validate the effectiveness of our method. Shishi Qiao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2023 | RingMo: A Remote Sensing Foundation Model With Masked Image ModelingabstractDeep learning approaches have contributed to the rapid development of remote sensing (RS) image interpretation. The most widely used training paradigm is to use ImageNet pretrained models to process RS data for specified tasks. However, there are issues such as domain gap between natural and RS scenes and the poor generalization capacity of RS models. It makes sense to develop a foundation model with general RS feature representation. Since a large amount of unlabeled data is available, the self-supervised method has more development significance than the fully supervised method in RS. However, most of the current self-supervised methods use contrastive learning, whose performance is sensitive to data augmentation, additional information, and selection of positive and negative pairs. In this article, we leverage the benefits of generative self-supervised learning (SSL) for RS images and propose an RS foundationmodel framework called RingMo, which consists of two parts. First, a large-scale dataset is constructed by collecting two million RS images from satellite and aerial platforms, covering multiple scenes and objects around the world. Second, we propose an RS foundation model training method designed for dense and small objects in complicated RS scenes. We show that the foundation model trained on our dataset with RingMo method achieves state-of-the-art (SOTA) on eight datasets across four downstream tasks, demonstrating the effectiveness of the proposed framework. Through in-depth exploration, we believe it is time for RS researchers to embrace generative SSL and leverage its general representation capabilities to speed up the development of RS applications. Xian Sun 0001, Peijin Wang, Wanxuan Lu, Zicong Zhu, Qibin He 0001, Junxi Li, Xuee Rong, Zhujun Yang, Qinglin He, Ruiping Wang 0001, Jiwen Lu, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 13 |
| 2022 | Implicit-Part Based Context Aggregation for Point Cloud Instance SegmentationabstractContext information is important for instance segmentation on point clouds. Existing methods either only use local surroundings by stacking multiple convolution layers or use non-local methods to model long-range interactions. However, they usually directly operate on points which is an unstructured and low-level representation and is highly dependent on context. To address this issue, we propose an effective framework named Implicit-Part Context Aggregation (IPCA), which adopts implicit parts as an intermediate representation and achieves context aggregation through message passing along the implicit part graph. Specifically, we first organize unstructured points into geometrically consistent implicit parts and construct the implicit part graph according to the geometric adjacency. Then, an initial part embedding is extracted using the proposed Implicit Part Network (IPN) which can aggregate point features and capture the intrinsic geometric shape of the part. We further refine the part embedding by a graph reasoning module named Context Aggregation Network (CAN), which helps to make a more precise prediction by well exploiting the context information. Instance proposals are then generated by grouping implicit parts. Finally, we propose an additional step to attribute the entire instance proposal to a Semantic Criterion Net (SCN) to infer the semantics of the instance. The purpose is to correct the semantic prediction errors caused by not knowing the boundary and overall shape of the object in the previous steps. Extensive experiments on two large datasets, ScanNet and 3RScan, demonstrate the effectiveness of our method. To our knowledge, it yields the highest performance on the ScanNet test benchmark and its AP@50 is 9.5 points higher than the baseline. The code is available at https://github.com/xiaodongww/IPCA Ruiping Wang 0001, Xilin Chen 0001 |
IROS | 2 |
| 2022 | From Node to Graph: Joint Reasoning on Visual-Semantic Relational Graph for Zero-Shot DetectionabstractZero-Shot Detection (ZSD), which aims at localizing and recognizing unseen objects in a complicated scene, usually leverages the visual and semantic information of individual objects alone. However, scene understanding of human exceeds recognizing individual objects separately: the contextual information among multiple objects such as visual relational information (e.g. visually similar objects) and semantic relational information (e.g. co-occurrences) is helpful for understanding of visual scene. In this paper, we verify that contextual information plays a more important role in ZSD than in traditional object detection. To make full use of such information, we propose a new end-to-end ZSD method GRaph Aligning Network (GRAN) based on graph modeling and reasoning which simultaneously considers visual and semantic information of multiple objects instead of individual objects. Specifically, we formulate a Visual Relational Graph (VRG) and a Semantic Relational Graph (SRG), where the nodes are the objects in the image and the semantic representations of classes respectively and the edges are the relevance between nodes in each graph. To characterize mutual effect between two modalities, the two graphs are further merged into a heterogeneous Visual-Semantic Relational Graph (VSRG), where modal translators are designed for the two subgraphs to enable modal information to transform into a common space for communication, and message passing among nodes is enforced to refine their representations. Comprehensive experiments on MSCOCO dataset demonstrate the advantage of our method over state-of-the-arts, and qualitative analysis suggests the validity of using contextual information. Hui Nie 0001, Ruiping Wang 0001, Xilin Chen 0001 |
WACV | 2 |
| 2022 | SEGA: Semantic Guided Attention on Visual Prototype for Few-Shot LearningabstractTeaching machines to recognize a new category based on few training samples especially only one remains challenging owing to the incomprehensive understanding of the novel category caused by the lack of data. However, human can learn new classes quickly even given few samples since human can tell what discriminative features should be focused on about each category based on both the visual and semantic prior knowledge. To better utilize those prior knowledge, we propose the SEmantic Guided Attention (SEGA) mechanism where the semantic knowledge is used to guide the visual perception in a top-down manner about what visual features should be paid attention to when distinguishing a category from the others. As a result, the embedding of the novel class even with few samples can be more discriminative. Concretely, a feature extractor is trained to embed few images of each novel class into a visual prototype with the help of transferring visual prior knowledge from base classes. Then we learn a network that maps semantic knowledge to category-specific attention vectors which will be used to perform feature selection to enhance the visual prototypes. Extensive experiments on miniImageNet, tieredImageNet, CIFAR-FS, and CUB indicate that our semantic guided attention realizes anticipated function and outperforms state-of-the-art results. Fengyuan Yang 0002, Ruiping Wang 0001, Xilin Chen 0001 |
WACV | 2 |
| 2022 | CVPR 2020 continual learning in computer vision competition: Approaches, results, current challenges and future directions
Vincenzo Lomonaco, Lorenzo Pellegrini, Pau Rodríguez, Massimo Caccia, Qi She, Quentin Jodelet, Ruiping Wang 0001, Zheda Mai, David Vázquez 0001, German Ignacio Parisi, Nikhil Churamani, Marc Pickett, Issam H. Laradji, Davide Maltoni |
Artif. Intell. | 8 |
| 2022 | Rethinking class orders and transferability in class incremental learning
Ruiping Wang 0001, Xilin Chen 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | FAIEr: Fidelity and Adequacy Ensured Image Caption EvaluationabstractImage caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. When humans evaluate a caption, they usually consider multiple aspects, such as whether it is related to the target image without distortion, how much image gist it conveys, as well as how fluent and beautiful the language and wording is. The above three different evaluation orientations can be summarized as fidelity, adequacy, and fluency. The former two rely on the image content, while fluency is purely related to linguistics and more subjective. Inspired by human judges, we propose a learning-based metric named FAIEr to ensure evaluating the fidelity and adequacy of the captions. Since image captioning involves two different modalities, we employ the scene graph as a bridge between them to represent both images and captions. FAIEr mainly regards the visual scene graph as the criterion to measure the fidelity. Then for evaluating the adequacy of the candidate caption, it high-lights the image gist on the visual scene graph under the guidance of the reference captions. Comprehensive experimental results show that FAIEr has high consistency with human judgment as well as high stability, low reference dependency, and the capability of reference-free evaluation. Sijin Wang, Ziwei Yao, Ruiping Wang 0001, Zhongqin Wu, Xilin Chen 0001 |
CVPR | 3 |
| 2021 | Local Feature Enhancement Network for Set-based Face RecognitionabstractSet-based Face Recognition is widely applied in scenarios like law enforcement and online media data management. Compared with face recognition using a single image, the faces in the set often contain abundant appearance changes. Therefore, how to make full use of the rich information from the set and integrate them into a unified set representation become the key to set-based face recognition. Inspired by the fact that humans usually complete this fine-grained task through integrating the information from the congruent local regions (e.g. an eye to an eye) of multiple faces in a set, we propose a novel method called Local Feature Enhancement Network (LFENet), which can automatically enhance the local feature through transferring the local information across the images. Specifically, we retain the spatial semantic information of the feature maps and apply different relational functions to establish the correlation among the local features. The contained local information will be transferred to the relevant local features to enhance their discriminability. By doing so, the valuable local information carried in some local features can complement those with incomplete information. Besides, the various local information is aligned across faces under different conditions to help the model learn intra-set-compact face representations. Our method achieves state-of-the-art performances on two mainstream set-based face recognition benchmarks: IJB-A and IJB-C, which fully reflects the rationality and effectiveness of our local feature enhancement mechanism. Ziyi Bai, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
FG | 2 |
| 2021 | Env-QA: A Video Question Answering Benchmark for Comprehensive Understanding of Dynamic EnvironmentsabstractVisual understanding goes well beyond the study of images or videos on the web. To achieve complex tasks in volatile situations, the human can deeply understand the environment, quickly perceive events happening around, and continuously track objects’ state changes, which are still challenging for current AI systems. To equip AI system with the ability to understand dynamic ENVironments, we build a video Question Answering dataset named Env-QA. Env-QA contains 23K egocentric videos, where each video is composed of a series of events about exploring and interacting in the environment. It also provides 85K questions to evaluate the ability of understanding the composition, layout, and state changes of the environment presented by the events in videos. Moreover, we propose a video QA model, Temporal Segmentation and Event Attention network (TSEA), which introduces event-level video representation and corresponding attention mechanisms to better extract environment information and answer questions. Comprehensive experiments demonstrate the effectiveness of our framework and show the formidable challenges of Env-QA in terms of long-term state tracking, multi-event temporal reasoning and event counting, etc. Difei Gao, Ruiping Wang 0001, Ziyi Bai, Xilin Chen 0001 |
ICCV | 2 |
| 2021 | Topic Scene Graph Generation by Attention Distillation from CaptionabstractIf an image tells a story, the image caption is the briefest narrator. Generally, a scene graph prefers to be an omniscient "generalist", while the image caption is more willing to be a "specialist", which outlines the gist. Lots of previous studies have found that a scene graph is not as practical as expected unless it can reduce the trivial contents and noises. In this respect, the image caption is a good tutor. To this end, we let the scene graph borrow the ability from the image caption so that it can be a specialist on the basis of remaining all-around, resulting in the socalled Topic Scene Graph. What an image caption pays attention to is distilled and passed to the scene graph for estimating the importance of partial objects, relationships, and events. Specifically, during the caption generation, the attention about individual objects in each time step is collected, pooled, and assembled to obtain the attention about relationships, which serves as weak supervision for regularizing the estimated importance scores of relationships. In addition, as this attention distillation process provides an opportunity for combining the generation of image caption and scene graph together, we further transform the scene graph into linguistic form with rich and free-form expressions by sharing a single generation model with image caption. Experiments show that attention distillation brings significant improvements in mining important relationships without strong supervision, and the topic scene graph shows great potential in subsequent applications. Wenbin Wang 0001, Ruiping Wang 0001, Xilin Chen 0001 |
ICCV | 2 |
| 2021 | Holistic Pose Graph: Modeling Geometric Structure among Objects in a Scene using Graph Inference for 3D Object PredictionabstractDue to the missing depth cues, it is essentially ambiguous to detect 3D objects from a single RGB image. Existing methods predict the 3D pose for each object independently or merely by combining local relationships within limited surroundings, but rarely explore the inherent geometric relationships from a global perspective. To address this issue, we argue that modeling geometric structure among objects in a scene is very crucial, and thus elaborately devise the Holistic Pose Graph (HPG) that explicitly integrates all geometric poses including the object pose treated as nodes and the relative pose treated as edges. The inference of the HPG uses GRU to encode the pose features from their corresponding regions in a single RGB image, and passes messages along the graph structure iteratively to improve the predicted poses. To further enhance the correspondence between the object pose and the relative pose, we propose a novel consistency loss to explicitly measure the deviations between them. Finally, we apply Holistic Pose Estimation (HPE) to jointly evaluate both the independent object pose and the relative pose. Our experiments on the SUN RGB-D dataset demonstrate that the proposed method provides a significant improvement on 3D object prediction. Jiwei Xiao, Ruiping Wang 0001, Xilin Chen 0001 |
ICCV | 2 |
| 2021 | Neural computing and applications (NCAA) special issue on best of DICTA 2019 papers
Ajmal Mian, Lei Wang 0108, Ruiping Wang 0001, Hamid Laga, Naveed Akhtar |
Neural Comput. Appl. | 3 |
| 2021 | What is a Tabby? Interpretable Model Decisions by Learning Attribute-Based Classification CriteriaabstractState-of-the-art classification models are usually considered as black boxes since their decision processes are implicit to humans. On the contrary, human experts classify objects according to a set of explicit hierarchical criteria. For example, "tabby is a domestic cat with stripes, dots, or lines", where tabby is defined by combining its superordinate category (domestic cat) and some certain attributes (e.g., has stripes). Inspired by this mechanism, we propose an interpretable Hierarchical Criteria Network (HCN) by additionally learning such criteria. To achieve this goal, images and semantic entities (e.g., taxonomies and attributes) are embedded into a common space, where each category can be represented by the linear combination of its superordinate category and a set of learned discriminative attributes. Specifically, a two-stream convolutional neural network (CNN) is elaborately devised, which embeds images and taxonomies with the two streams respectively. The model is trained by minimizing the prediction error of hierarchy labels on both streams. Extensive experiments on two widely studied datasets (CIFAR-100 and ILSVRC) demonstrate that HCN can learn meaningful attributes as well as reasonable and interpretable classification criteria. Therefore, the proposed method enables further human feedback for model correction as an additional benefit. Haomiao Liu, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Deep video code for efficient face video retrieval
Shishi Qiao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2020 | Multi-Modal Graph Neural Network for Joint Reasoning on Vision and Scene TextabstractAnswering questions that require reading texts in an image is challenging for current models. One key difficulty of this task is that rare, polysemous, and ambiguous words frequently appear in images, e.g., names of places, products, and sports teams. To overcome this difficulty, only resorting to pre-trained word embedding models is far from enough. A desired model should utilize the rich information in multiple modalities of the image to help understand the meaning of scene texts, e.g., the prominent text on a bottle is most likely to be the brand. Following this idea, we propose a novel VQA approach, Multi-Modal Graph Neural Network (MM-GNN). It first represents an image as a graph consisting of three sub-graphs, depicting visual, semantic, and numeric modalities respectively. Then, we introduce three aggregators which guide the message passing from one graph to another to utilize the contexts in various modalities, so as to refine the features of nodes. The updated nodes have better features for the downstream question answering module. Experimental evaluations show that our MM-GNN represents the scene texts better and obviously facilitates the performances on two VQA tasks that require reading scene texts. Difei Gao, Kenneth Li 0002, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 3 |
| 2020 | Sketching Image Gist: Human-Mimetic Hierarchical Scene Graph Generation
Wenbin Wang 0001, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ECCV (13) | 2 |
| 2020 | Hybrid Video and Image Hashing for Robust Face RetrievalabstractVideo face retrieval (VFR) is an appealing and practical computer vision task, which aims to search particular character from masses of videos like in TV-Series. The challenges of this task mainly lie in two aspects, i.e. faces in such videos contain complex appearance variations with uncontrollable shooting environment and searching from big data usually requires high efficiency in both space and time. To fulfill this task, current works typically proceed in a learning to hash manner by fusing single-frame features within a video to obtain the video representation and further embedding it into Hamming space to yield video binary codes. The feature fusion stage has inevitably discarded too much frame information and leads to less discriminative video codes. In this paper, we propose Hybrid Video and Image Hashing (HVIH) to learn more effective binary codes for face videos. Specifically, we fully exploit the dense frame features rather than simply discarding them after the video level fusion and jointly optimize binary codes for the video and its composed frames in adapted supervised manners. To achieve more robust video representation, we introduce a module of video center alignment to ensure the binary codes location of the video and its frames to be as compact and consistent as possible in the Hamming space, which naturally facilitates both tasks of video-to-video retrieval and image-to-video retrieval. Extensive experiments on two challenging video face databases demonstrate the superiority of our approach over the state-of-the-art. Ruikui Wang, Shishi Qiao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
FG | 3 |
| 2020 | Deep Position-Aware Hashing for Semantic Continuous Image RetrievalabstractPreserving the semantic similarity is one of the most important goals of hashing. Most existing deep hashing methods employ pairs or triplets of samples in training stage, which only consider the semantic similarity within a minibatch and depict the local positional relationship in Hamming space, leading to intermittent semantic similarity preservation. In this paper, we propose Deep Position-Aware Hashing (DPAH) to ensure continuous semantic similarity in Hamming space by modeling global positional relationship. Specifically, we introduce a set of learnable class centers as the global proxies to represent the global information and generate discriminative binary codes by constraining the distance between data points and class centers. In addition, in order to reduce the information loss caused by relaxing the binary codes to real-values in optimization, we propose kurtosis loss (KT loss) to handle the distribution of real-valued features before thresholding to be double-peak, and then enable the real-valued features to be more binarylike. Comprehensive experiments on three datasets show that our DPAH outperforms state-of-the-art methods. Ruikui Wang, Ruiping Wang 0001, Shishi Qiao, Shiguang Shan, Xilin Chen 0001 |
WACV | 2 |
| 2020 | Cross-modal Scene Graph Matching for Relationship-aware Image-Text RetrievalabstractImage-text retrieval of natural scenes has been a popular research topic. Since image and text are heterogeneous cross-modal data, one of the key challenges is how to learn comprehensive yet unified representations to express the multi-modal data. A natural scene image mainly involves two kinds of visual concepts, objects and their relationships, which are equally essential to image-text retrieval. Therefore, a good representation should account for both of them. In the light of recent success of scene graph in many CV and NLP tasks for describing complex natural scenes, we propose to represent image and text with two kinds of scene graphs: visual scene graph (VSG) and textual scene graph (TSG), each of which is exploited to jointly characterize objects and relationships in the corresponding modality. The image-text retrieval task is then naturally formulated as cross-modal scene graph matching. Specifically, we design two particular scene graph encoders in our model for VSG and TSG, which can refine the representation of each node on the graph by aggregating neighborhood information. As a result, both object-level and relationship-level cross-modal features can be obtained, which favorably enables us to evaluate the similarity of image and text in the two levels in a more plausible way. We achieve state-of-the-art results on Flickr30k and MS COCO, which verifies the advantages of our graph matching based approach for image-text retrieval. Sijin Wang, Ruiping Wang 0001, Ziwei Yao, Shiguang Shan, Xilin Chen 0001 |
WACV | 2 |
| 2020 | Learning Multifunctional Binary Codes for Personalized Image Retrieval
Haomiao Liu, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Deep Heterogeneous Hashing for Face Video RetrievalabstractRetrieving videos of a particular person with face image as query via hashing technique has many important applications. While face images are typically represented as vectors in Euclidean space, characterizing face videos with some robust set modeling techniques (e.g. covariance matrices as exploited in this study, which reside on Riemannian manifold), has recently shown appealing advantages. This hence results in a thorny heterogeneous spaces matching problem. Moreover, hashing with handcrafted features as done in many existing works is clearly inadequate to achieve desirable performance for this task. To address such problems, we present an end-toend Deep Heterogeneous Hashing (DHH) method that integrates three stages including image feature learning, video modeling, and heterogeneous hashing in a single framework, to learn unified binary codes for both face images and videos. To tackle the key challenge of hashing on manifold, a well-studied Riemannian kernel mapping is employed to project data (i.e. covariance matrices) into Euclidean space and thus enables to embed the two heterogeneous representations into a common Hamming space, where both intra-space discriminability and inter-space compatibility are considered. To perform network optimization, the gradient of the kernel mapping is innovatively derived via structured matrix backpropagation in a theoretically principled way. Experiments on three challenging datasets show that our method achieves quite competitive performance compared with existing hashing methods. Shishi Qiao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Exploring Context and Visual Pattern of Relationship for Scene Graph GenerationabstractRelationship is the core of scene graph, but its prediction is far from satisfying because of its complex visual diversity. To alleviate this problem, we treat relationship as an abstract object, exploring not only significative visual pattern but contextual information for it, which are two key aspects when considering object recognition. Our observation on current datasets reveals that there exists intimate association among relationships. Therefore, inspired by the successful application of context to object-oriented tasks, we especially construct context for relationships where all of them are gathered so that the recognition could benefit from their association. Moreover, accurate recognition needs discriminative visual pattern for object, and so does relationship. In order to discover effective pattern for relationship, traditional relationship feature extraction methods such as using union region or combination of subject-object feature pairs are replaced with our proposed intersection region which focuses on more essential parts. Therefore, we present our so-called Relationship Context - InterSeCtion Region (CISC) method. Experiments for scene graph generation on Visual Genome dataset and visual relationship prediction on VRD dataset indicate that both the relationship context and intersection region improve performances and realize anticipated functions. Wenbin Wang 0001, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2019 | Transferable Contrastive Network for Generalized Zero-Shot LearningabstractZero-shot learning (ZSL) is a challenging problem that aims to recognize the target categories without seen data, where semantic information is leveraged to transfer knowledge from some source classes. Although ZSL has made great progress in recent years, most existing approaches are easy to overfit the sources classes in generalized zero-shot learning (GZSL) task, which indicates that they learn little knowledge about target classes. To tackle such problem, we propose a novel Transferable Contrastive Network (TCN) that explicitly transfers knowledge from the source classes to the target classes. It automatically contrasts one image with different classes to judge whether they are consistent or not. By exploiting the class similarities to make knowledge transfer from source images to similar target classes, our approach is more robust to recognize the target images. Experiments on five benchmark datasets show the superiority of our approach for GZSL. Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2019 | Prior Knowledge Guided Small Object Detection on High-Resolution ImagesabstractWhen applying common object detection algorithms to detect small objects on high-resolution images, the down-sampling operation of the input images is inevitable due to the limitation of GPU memory. Accordingly, the details for characterizing small objects are lost. To resolve this contradiction, a small object detection method in a coarse-to-fine manner is presented. Specifically, some rough regions of interest (ROI) are firstly computed from low-resolution images. The prior knowledge of the positions of objects is used to guide the generation of ROIs. Then the features of small ROIs are recomputed from high-resolution images, and the features of large ROIs are obtained from the feature maps used to generate ROIs. The proposed method is validated on two datasets. One is a plant phenotyping dataset and the other is a public traffic sign dataset. Experimental results convincingly show the effectiveness of the proposed method. Xiujuan Chai, Ruiping Wang 0001, Weijun Guo, Li Pu, Xilin Chen 0001 |
ICIP | 3 |
| 2019 | Deep Supervised Hashing for Fast Image Retrieval
Haomiao Liu, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Int. J. Comput. Vis. | 2 |
| 2019 | Adaptive Metric Learning For Zero-Shot RecognitionabstractZero-shot learning (ZSL) has enjoyed great popularity in recent years due to its ability to recognize novel objects, where semantic information is exploited to build up relations among different categories. Traditional ZSL approaches usually focus on learning more robust visual-semantic embeddings among seen classes and directly apply them to the unseen classes without considering whether they are suitable. It is well known that domain gap exists between seen and unseen classes. In order to tackle such problem, we propose a novel adaptive metric learning approach to measure the compatibility between visual samples and class semantics, where class similarities are utilized to adapt the visual-semantic embedding to the unseen classes. Extensive experiments on four benchmark ZSL datasets show the effectiveness of the proposed approach. Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Signal Process. Lett. | 2 |
| 2018 | COSONet: Compact Second-Order Network for Video Face Recognition
Yirong Mao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 2 |
| 2018 | Exemplar-Supported Generative Reproduction for Class Incremental Learning
Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 2 |
| 2018 | Structure Inference Net: Object Detection Using Scene-Level Context and Instance-Level RelationshipsabstractContext is important for accurate visual recognition. In this work we propose an object detection algorithm that not only considers object visual appearance, but also makes use of two kinds of context including scene contextual information and object relationships within a single image. Therefore, object detection is regarded as both a cognition problem and a reasoning problem when leveraging these structured information. Specifically, this paper formulates object detection as a problem of graph structure inference, where given an image the objects are treated as nodes in a graph and relationships between the objects are modeled as edges in such graph. To this end, we present a so-called Structure Inference Network (SIN), a detector that incorporates into a typical detection framework (e.g. Faster R-CNN) with a graphical model which aims to infer object state. Comprehensive experiments on PASCAL VOC and MS COCO datasets indicate that scene context and object relationships truly improve the performance of object detection with more desirable and reasonable outputs. Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2018 | Learning Class Prototypes via Structure Alignment for Zero-Shot Recognition
Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ECCV (10) | 2 |
| 2018 | Fusing magnitude and phase features with multiple face models for robust face recognition
Yan Li 0014, Shiguang Shan, Ruiping Wang 0001, Zhen Cui 0001, Xilin Chen 0001 |
Frontiers Comput. Sci. | 3 |
| 2018 | Attribute annotation on large-scale image database by active knowledge transfer
Huajie Jiang, Ruiping Wang 0001, Yan Li 0014, Haomiao Liu, Shiguang Shan, Xilin Chen 0001 |
Image Vis. Comput. | 2 |
| 2018 | Cross Euclidean-to-Riemannian Metric Learning with Application to Face Recognition from VideoabstractRiemannian manifolds have been widely employed for video representations in visual classification tasks including video-based face recognition. The success mainly derives from learning a discriminant Riemannian metric which encodes the non-linear geometry of the underlying Riemannian manifolds. In this paper, we propose a novel metric learning framework to learn a distance metric across a Euclidean space and a Riemannian manifold to fuse average appearance and pattern variation of faces within one video. The proposed metric learning framework can handle three typical tasks of video-based face recognition: Video-to-Still, Still-to-Video and Video-to-Video settings. To accomplish this new framework, by exploiting typical Riemannian geometries for kernel embedding, we map the source Euclidean space and Riemannian manifold into a common Euclidean subspace, each through a corresponding high-dimensional Reproducing Kernel Hilbert Space (RKHS). With this mapping, the problem of learning a cross-view metric between the two source heterogeneous spaces can be converted to learning a single-view Euclidean distance metric in the target common Euclidean space. By learning information on heterogeneous data with the shared label, the discriminant metric in the common space improves face recognition from videos. Extensive experiments on four challenging video face databases demonstrate that the proposed framework has a clear advantage over the state-of-the-art methods in the three classical video-based face recognition scenarios. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Luc Van Gool, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2018 | Distance metric learning for pattern recognition
Jiwen Lu, Ruiping Wang 0001, Ajmal Mian, Sudeep Sarkar |
Pattern Recognit. | 2 |
| 2018 | Geometry-Aware Similarity Learning on SPD Manifolds for Visual RecognitionabstractSymmetric positive definite (SPD) matrices have been employed for data representation in many visual recognition tasks. The success is mainly attributed to learning discriminative SPD matrices encoding the Riemannian geometry of the underlying SPD manifolds. In this paper, we propose a geometry-aware SPD similarity learning (SPDSL) framework to learn discriminative SPD features by directly pursuing a manifold-manifold transformation matrix of full column rank. Specifically, by exploiting the Riemannian geometry of the manifolds of fixed-rank positive semidefinite (PSD) matrices, we present a new solution to reduce optimization over the space of column full-rank transformation matrices to optimization on the PSD manifold, which has a well-established Riemannian structure. Under this solution, we exploit a new supervised SPDSL technique to learn the manifold-manifold transformation by regressing the similarities of selected SPD data pairs to their ground-truth similarities on the target SPD manifold. To optimize the proposed objective function, we further derive an optimization algorithm on the PSD manifold. Evaluations on three visual classification tasks show the advantages of the proposed approach over the existing SPD-based discriminant learning methods. Zhiwu Huang, Ruiping Wang 0001, Xianqiu Li, Wenxian Liu, Shiguang Shan, Luc Van Gool, Xilin Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Discriminant Analysis on Riemannian Manifold of Gaussian Distributions for Face Recognition With Image SetsabstractTo address the problem of face recognition with image sets, we aim to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as the Gaussian mixture model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. Since in the light of information geometry, the Gaussians lie on a specific Riemannian manifold, this paper presents a method named discriminant analysis on Riemannian manifold of Gaussian distributions (DARG). We investigate several distance metrics between Gaussians and accordingly two discriminative learning frameworks are presented to meet the geometric and statistical characteristics of the specific manifold. The first framework derives a series of provably positive definite probabilistic kernels to embed the manifold to a high-dimensional Hilbert space, where conventional discriminant analysis methods developed in Euclidean space can be applied, and a weighted Kernel discriminant analysis is devised which learns discriminative representation of the Gaussian components in GMMs with their prior probabilities as sample weights. Alternatively, the other framework extends the classical graph embedding method to the manifold by utilizing the distance metrics between Gaussians to construct the adjacency graph, and hence the original manifold is embedded to a lower-dimensional and discriminative target manifold with the geometric structure preserved and the interclass separability maximized. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB, and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art.To address the problem of face recognition with image sets, we aim to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as the Gaussian mixture model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. Since in the light of information geometry, the Gaussians lie on a specific Riemannian manifold, this paper presents a method named discriminant analysis on Riemannian manifold of Gaussian distributions (DARG). We investigate several distance metrics between Gaussians and accordingly two discriminative learning frameworks are presented to meet the geometric and statistical characteristics of the specific manifold. The first framework derives a series of provably positive definite probabilistic kernels to embed the manifold to a high-dimensional Hilbert space, where conventional discriminant analysis methods developed in Euclidean space can be applied, and a weighted Kernel discriminant analysis is devised which learns discriminative representation of the Gaussian components in GMMs with their prior probabilities as sample weights. Alternatively, the other framework extends the classical graph embedding method to the manifold by utilizing the distance metrics between Gaussians to construct the adjacency graph, and hence the original manifold is embedded to a lower-dimensional and discriminative target manifold with the geometric structure preserved and the interclass separability maximized. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB, and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art. Wen Wang 0019, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Visual Textbook Network: Watch Carefully before Answering Visual Questions
Difei Gao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 2 |
| 2017 | Learning Multifunctional Binary Codes for Both Category and Attribute Oriented Retrieval TasksabstractIn this paper we propose a unified framework to address multiple realistic image retrieval tasks concerning both category and attributes. Considering the scale of modern datasets, hashing is favorable for its low complexity. However, most existing hashing methods are designed to preserve one single kind of similarity, thus incapable of dealing with the different tasks simultaneously. To overcome this limitation, we propose a new hashing method, named Dual Purpose Hashing (DPH), which jointly preserves the category and attribute similarities by exploiting the convolutional networks (CNN) to hierarchically capture the correlations between category and attributes. Since images with both category and attribute labels are scarce, our method is designed to take the abundant partially labelled images on the Internet as training inputs. With such a framework, the binary codes of new-coming images can be readily obtained by quantizing the network outputs of a binary-like layer, and the attributes can be recovered from the codes easily. Experiments on two large-scale datasets show that our dual purpose hash codes can achieve comparable or even better performance than those state-of-the-art methods specifically designed for each individual retrieval task, while being more compact than the compared methods. Haomiao Liu, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2017 | Discriminative Covariance Oriented Representation Learning for Face Recognition with Image SetsabstractFor face recognition with image sets, while most existing works mainly focus on building robust set models with hand-crafted feature, it remains a research gap to learn better image representations which can closely match the subsequent image set modeling and classification. Taking sample covariance matrix as set model in the light of its recent promising success, we present a Discriminative Covariance oriented Representation Learning (DCRL) framework to bridge the above gap. The framework constructs a feature learning network (e.g. a CNN) to project the face images into a target representation space, and the network is trained towards the goal that the set covariance matrix calculated in the target space has maximum discriminative ability. To encode the discriminative ability of set covariance matrices, we elaborately design two different loss functions, which respectively lead to two different representation learning schemes, i.e., the Graph Embedding scheme and the Softmax Regression scheme. Both schemes optimize the whole network containing both image representation mapping and set model classification in a joint learning manner. The proposed method is extensively validated on three challenging and large scale databases for the task of face recognition with image sets, i.e., YouTube Celebrities, YouTube Face DB and Point-and-Shoot Challenge. Wen Wang 0019, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2017 | Learning Discriminative Latent Attributes for Zero-Shot Classification
Huajie Jiang, Ruiping Wang 0001, Shiguang Shan, Yi Yang 0001, Xilin Chen 0001 |
ICCV | 2 |
| 2017 | Prototype Discriminative Learning for Image Set ClassificationabstractThis letter presents a prototype discriminative learning (PDL) method for image set classification. We aim to simultaneously learn prototypes and a linear discriminative projection to drive that in the target subspace each image set can be discriminated with its nearest neighbor prototype. To reveal the unseen appearance variations implicitly in an image set, the prototypes are actually “virtual,” which do not certainly appear in the set but are searched in the corresponding affine hull. Moreover, to enhance the stability and robustness of the learned target subspace, an orthogonality constraint is imposed on the projection. Thus, to optimize the prototypes and the projection jointly, we design a specific gradient descent mechanism by updating the projection on Stiefel manifold and the prototypes in Euclidean space in an alternative optimization manner. Experimental results on four challenging databases demonstrate the superiority of the proposed PDL method. Wen Wang 0019, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Signal Process. Lett. | 2 |
| 2016 | Deep Video Code for Efficient Face Video Retrieval
Shishi Qiao, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 2 |
| 2016 | Prototype Discriminative Learning for Face Image Set Classification
Wen Wang 0019, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 2 |
| 2016 | Deep Supervised Hashing for Fast Image RetrievalabstractIn this paper, we present a new hashing method to learn compact binary codes for highly efficient image retrieval on large-scale datasets. While the complex image appearance variations still pose a great challenge to reliable retrieval, in light of the recent progress of Convolutional Neural Networks (CNNs) in learning robust image representation on various vision tasks, this paper proposes a novel Deep Supervised Hashing (DSH) method to learn compact similarity-preserving binary code for the huge body of image data. Specifically, we devise a CNN architecture that takes pairs of images (similar/dissimilar) as training inputs and encourages the output of each image to approximate discrete values (e.g. +1/-1). To this end, a loss function is elaborately designed to maximize the discriminability of the output space by encoding the supervised information from the input image pairs, and simultaneously imposing regularization on the real-valued outputs to approximate the desired discrete values. For image retrieval, new-coming query images can be easily encoded by propagating through the network and then quantizing the network outputs to binary codes representation. Extensive experiments on two large scale datasets CIFAR-10 and NUS-WIDE show the promising performance of our method compared with the state-of-the-arts. Haomiao Liu, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2016 | Learning prototypes and similes on Grassmann manifold for spontaneous expression recognition
Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Comput. Vis. Image Underst. | 2 |
| 2016 | Deep and Structured Robust Information Theoretic Learning for Image AnalysisabstractThis paper presents a robust information theoretic (RIT) model to reduce the uncertainties, i.e., missing and noisy labels, in general discriminative data representation tasks. The fundamental pursuit of our model is to simultaneously learn a transformation function and a discriminative classifier that maximize the mutual information of data and their labels in the latent space. In this general paradigm, we, respectively, discuss three types of the RIT implementations with linear subspace embedding, deep transformation, and structured sparse learning. In practice, the RIT and deep RIT are exploited to solve the image categorization task whose performances will be verified on various benchmark data sets. The structured sparse RIT is further applied to a medical image analysis task for brain magnetic resonance image segmentation that allows group-level feature selections on the brain tissues. Yue Deng 0001, Feng Bao 0002, XueSong Deng, Ruiping Wang 0001, Youyong Kong, Qionghai Dai |
IEEE Trans. Image Process. | 4 |
| 2016 | Spatial Pyramid Covariance-Based Compact Video Code for Robust Face Retrieval in TV-SeriesabstractWe address the problem of face video retrieval in TV-series, which searches video clips based on the presence of specific character, given one face track of his/her. This is tremendously challenging because on one hand, faces in TV-series are captured in largely uncontrolled conditions with complex appearance variations, and on the other hand, retrieval task typically needs efficient representation with low time and space complexity. To handle this problem, we propose a compact and discriminative representation for the huge body of video data, named compact video code (CVC). Our method first models the face track by its sample (i.e., frame) covariance matrix to capture the video data variations in a statistical manner. To incorporate discriminative information and obtain more compact video signature suitable for retrieval, the high-dimensional covariance representation is further encoded as a much lower dimensional binary vector, which finally yields the proposed CVC. Specifically, each bit of the code, i.e., each dimension of the binary vector, is produced via supervised learning in a max margin framework, which aims to make a balance between the discriminability and stability of the code. Besides, we further extend the descriptive granularity of covariance matrix from traditional pixel-level to more general patch-level, and proceed to propose a novel hierarchical video representation named spatial pyramid covariance along with a fast calculation method. Face retrieval experiments on two challenging TV-series video databases, i.e., the Big Bang Theory and Prison Break, demonstrate the competitiveness of the proposed CVC over the state-of-the-art retrieval methods. In addition, as a general video matching algorithm, CVC is also evaluated in traditional video face recognition task on a standard Internet database, i.e., YouTube Celebrities, showing its quite promising performance by using an extremely compact code with only 128 bits. Yan Li 0014, Ruiping Wang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 2016 | Learning Expressionlets via Universal Manifold Model for Dynamic Facial Expression RecognitionabstractFacial expression is a temporally dynamic event which can be decomposed into a set of muscle motions occurring in different facial regions over various time intervals. For dynamic expression recognition, two key issues, temporal alignment and semantics-aware dynamic representation, must be taken into account. In this paper, we attempt to solve both problems via manifold modeling of videos based on a novel mid-level representation, i.e., expressionlet. Specifically, our method contains three key stages: 1) each expression video clip is characterized as a spatial-temporal manifold (STM) formed by dense low-level features; 2) a universal manifold model (UMM) is learned over all low-level features and represented as a set of local modes to statistically unify all the STMs; and 3) the local modes on each STM can be instantiated by fitting to the UMM, and the corresponding expressionlet is constructed by modeling the variations in each local mode. With the above strategy, expression videos are naturally aligned both spatially and temporally. To enhance the discriminative power, the expressionlet-based STM representation is further processed with discriminant embedding. Our method is evaluated on four public expression databases, CK+, MMI, Oulu-CASIA, and FERA. In all cases, our method outperforms the known state of the art by a large margin. Shiguang Shan, Ruiping Wang 0001, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2015 | Projection Metric Learning on Grassmann Manifold with Application to Video based Face RecognitionabstractIn video based face recognition, great success has been made by representing videos as linear subspaces, which typically lie in a special type of non-Euclidean space known as Grassmann manifold. To leverage the kernel-based methods developed for Euclidean space, several recent methods have been proposed to embed the Grassmann manifold into a high dimensional Hilbert space by exploiting the well established Project Metric, which can approximate the Riemannian geometry of Grassmann manifold. Nevertheless, they inevitably introduce the drawbacks from traditional kernel-based methods such as implicit map and high computational cost to the Grassmann manifold. To overcome such limitations, we propose a novel method to learn the Projection Metric directly on Grassmann manifold rather than in Hilbert space. From the perspective of manifold learning, our method can be regarded as performing a geometry-aware dimensionality reduction from the original Grassmann manifold to a lower-dimensional, more discriminative Grassmann manifold where more favorable classification can be achieved. Experiments on several real-world video face datasets demonstrate that the proposed method yields competitive performance compared with the state-of-the-art algorithms. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2015 | Face video retrieval with image query via hashing across Euclidean space and Riemannian manifoldabstractRetrieving videos of a specific person given his/her face image as query becomes more and more appealing for applications like smart movie fast-forwards and suspect searching. It also forms an interesting but challenging computer vision task, as the visual data to match, i.e., still image and video clip are usually represented quite differently. Typically, face image is represented as point (i.e., vector) in Euclidean space, while video clip is seemingly modeled as a point (e.g., covariance matrix) on some particular Riemannian manifold in the light of its recent promising success. It thus incurs a new hashing-based retrieval problem of matching two heterogeneous representations, respectively in Euclidean space and Riemannian manifold. This work makes the first attempt to embed the two heterogeneous spaces into a common discriminant Hamming space. Specifically, we propose Hashing across Euclidean space and Riemannian manifold (HER) by deriving a unified framework to firstly embed the two spaces into corresponding reproducing kernel Hilbert spaces, and then iteratively optimize the intra- and inter-space Hamming distances in a max-margin framework to learn the hash functions for the two spaces. Extensive experiments demonstrate the impressive superiority of our method over the state-of-the-art competitive hash learning methods. Yan Li 0014, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2015 | Discriminant analysis on Riemannian manifold of Gaussian distributions for face recognition with image setsabstractThis paper presents a method named Discriminant Analysis on Riemannian manifold of Gaussian distributions (DARG) to solve the problem of face recognition with image sets. Our goal is to capture the underlying data distribution in each set and thus facilitate more robust classification. To this end, we represent image set as Gaussian Mixture Model (GMM) comprising a number of Gaussian components with prior probabilities and seek to discriminate Gaussian components from different classes. In the light of information geometry, the Gaussians lie on a specific Riemannian manifold. To encode such Riemannian geometry properly, we investigate several distances between Gaussians and further derive a series of provably positive definite probabilistic kernels. Through these kernels, a weighted Kernel Discriminant Analysis is finally devised which treats the Gaussians in GMMs as samples and their prior probabilities as sample weights. The proposed method is evaluated by face identification and verification tasks on four most challenging and largest databases, YouTube Celebrities, COX, YouTube Face DB and Point-and-Shoot Challenge, to demonstrate its superiority over the state-of-the-art. Wen Wang 0019, Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2015 | Two Birds, One Stone: Jointly Learning Binary Code for Large-Scale Face Image Retrieval and Attributes PredictionabstractWe address the challenging large-scale content-based face image retrieval problem, intended as searching images based on the presence of specific subject, given one face image of him/her. To this end, one natural demand is a supervised binary code learning method. While the learned codes might be discriminating, people often have a further expectation that whether some semantic message (e.g., visual attributes) can be read from the human-incomprehensible codes. For this purpose, we propose a novel binary code learning framework by jointly encoding identity discriminability and a number of facial attributes into unified binary code. In this way, the learned binary codes can be applied to not only fine-grained face image retrieval, but also facial attributes prediction, which is the very innovation of this work, just like killing two birds with one stone. To evaluate the effectiveness of the proposed method, extensive experiments are conducted on a new purified large-scale web celebrity database, named CFW 60K, with abundant manual identity and attributes annotation, and experimental results exhibit the superiority of our method over state-of-the-art. Yan Li 0014, Ruiping Wang 0001, Haomiao Liu, Huajie Jiang, Shiguang Shan, Xilin Chen 0001 |
ICCV | 2 |
| 2015 | Log-Euclidean Metric Learning on Symmetric Positive Definite Manifold with Application to Image Set ClassificationabstractThe manifold of Symmetric Positive Definite (SPD) matrices has been successfully used for data representation in image set classification. By endowing the SPD manifold with Log-Euclidean Metric, existing methods typically work on vector-forms of SPD matrix logarithms. This however not only inevitably distorts the geometrical structure of the space of SPD matrix logarithms but also brings low efficiency especially when the dimensionality of SPD matrix is high. To overcome this limitation, we propose a novel metric learning approach to work directly on logarithms of SPD matrices. Specifically, our method aims to learn a tangent map that can directly transform the matrix logarithms from the original tangent space to a new tangent space of more discriminability. Under the tangent map framework, the novel metric learning can then be formulated as an optimization problem of seeking a Mahalanobis-like matrix, which can take the advantage of traditional metric learning techniques. Extensive evaluations on several image set classification tasks demonstrate the effectiveness of our proposed metric learning method. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xianqiu Li, Xilin Chen 0001 |
ICML | 2 |
| 2015 | Sparsely encoded local descriptor for face verification
Zhen Cui 0001, Shiguang Shan, Ruiping Wang 0001, Lei Zhang 0006, Xilin Chen 0001 |
Neurocomputing | 3 |
| 2015 | Face recognition on large-scale video in the wild with hybrid Euclidean-and-Riemannian metric learning
Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
Pattern Recognit. | 2 |
| 2015 | A Benchmark and Comparative Study of Video-Based Face Recognition on COX Face DatabaseabstractFace recognition with still face images has been widely studied, while the research on video-based face recognition is inadequate relatively, especially in terms of benchmark datasets and comparisons. Real-world video-based face recognition applications require techniques for three distinct scenarios: 1) Videoto-Still (V2S); 2) Still-to-Video (S2V); and 3) Video-to-Video (V2V), respectively, taking video or still image as query or target. To the best of our knowledge, few datasets and evaluation protocols have benchmarked for all the three scenarios. In order to facilitate the study of this specific topic, this paper contributes a benchmarking and comparative study based on a newly collected still/video face database, named COX(1) Face DB. Specifically, we make three contributions. First, we collect and release a largescale still/video face database to simulate video surveillance with three different video-based face recognition scenarios (i.e., V2S, S2V, and V2V). Second, for benchmarking the three scenarios designed on our database, we review and experimentally compare a number of existing set-based methods. Third, we further propose a novel Point-to-Set Correlation Learning (PSCL) method, and experimentally show that it can be used as a promising baseline method for V2S/S2V face recognition on COX Face DB. Extensive experimental results clearly demonstrate that video-based face recognition needs more efforts, and our COX Face DB is a good benchmark database for evaluation. Zhiwu Huang, Shiguang Shan, Ruiping Wang 0001, Haihong Zhang, Shihong Lao, Alifu Kuerban, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2014 | Hybrid Euclidean-and-Riemannian Metric Learning for Image Set Classification
Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 2 |
| 2014 | Deeply Learning Deformable Facial Action Parts Model for Dynamic Expression Analysis
Shaoxin Li 0001, Shiguang Shan, Ruiping Wang 0001, Xilin Chen 0001 |
ACCV (4) | 4 |
| 2014 | Compact Video Code and Its Application to Robust Face Retrieval in TV-Series
Yan Li 0014, Ruiping Wang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 2 |
| 2014 | Learning Euclidean-to-Riemannian Metric for Point-to-Set ClassificationabstractIn this paper, we focus on the problem of point-to-set classification, where single points are matched against sets of correlated points. Since the points commonly lie in Euclidean space while the sets are typically modeled as elements on Riemannian manifold, they can be treated as Euclidean points and Riemannian points respectively. To learn a metric between the heterogeneous points, we propose a novel Euclidean-to-Riemannian metric learning framework. Specifically, by exploiting typical Riemannian metrics, the Riemannian manifold is first embedded into a high dimensional Hilbert space to reduce the gaps between the heterogeneous spaces and meanwhile respect the Riemannian geometry of the manifold. The final distance metric is then learned by pursuing multiple transformations from the Hilbert space and the original Euclidean space (or its corresponding Hilbert space) to a common Euclidean subspace, where classical Euclidean distances of transformed heterogeneous points can be measured. Extensive experiments clearly demonstrate the superiority of our proposed approach over the state-of-the-art methods. Zhiwu Huang, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 2 |
| 2014 | Learning Expressionlets on Spatio-temporal Manifold for Dynamic Facial Expression RecognitionabstractFacial expression is temporally dynamic event which can be decomposed into a set of muscle motions occurring in different facial regions over various time intervals. For dynamic expression recognition, two key issues, temporal alignment and semantics-aware dynamic representation, must be taken into account. In this paper, we attempt to solve both problems via manifold modeling of videos based on a novel mid-level representation, i.e. expressionlet. Specifically, our method contains three key components: 1) each expression video clip is modeled as a spatio-temporal manifold (STM) formed by dense low-level features, 2) a Universal Manifold Model (UMM) is learned over all low-level features and represented as a set of local ST modes to statistically unify all the STMs. 3) the local modes on each STM can be instantiated by fitting to UMM, and the corresponding expressionlet is constructed by modeling the variations in each local ST mode. With above strategy, expression videos are naturally aligned both spatially and temporally. To enhance the discriminative power, the expressionlet-based STM representation is further processed with discriminant embedding. Our method is evaluated on four public expression databases, CK+, MMI, Oulu-CASIA, and AFEW. In all cases, our method reports results better than the known state-of-the-art. Shiguang Shan, Ruiping Wang 0001, Xilin Chen 0001 |
CVPR | 3 |
| 2014 | Combining Multiple Kernel Methods on Riemannian Manifold for Emotion Recognition in the WildabstractIn this paper, we present the method for our submission to the Emotion Recognition in the Wild Challenge (EmotiW 2014). The challenge is to automatically classify the emotions acted by human subjects in video clips under real-world environment. In our method, each video clip can be represented by three types of image set models (i.e. linear subspace, covariance matrix, and Gaussian distribution) respectively, which can all be viewed as points residing on some Riemannian manifolds. Then different Riemannian kernels are employed on these set models correspondingly for similarity/distance measurement. For classification, three types of classifiers, i.e. kernel SVM, logistic regression, and partial least squares, are investigated for comparisons. Finally, an optimal fusion of classifiers learned from different kernels and different modalities (video and audio) is conducted at the decision level for further boosting the performance. We perform an extensive evaluation on the challenge data (including validation set and blind test set), and evaluate the effects of different strategies in our pipeline. The final recognition accuracy achieved 50.4% on test set, with a significant gain of 16.7% above the challenge baseline 33.7%. Ruiping Wang 0001, Shaoxin Li 0001, Shiguang Shan, Zhiwu Huang, Xilin Chen 0001 |
ICMI | 2 |
| 2014 | Robust Head-Shoulder Detection Using a Two-Stage Cascade FrameworkabstractHead-shoulder detection is widely used in many applications, and robust image descriptors are crucial to the detection performance. In this paper, by exploiting the second-order region covariance descriptor as a complement to widely-used histogram-based descriptors, we propose a new two-stage coarse-to-fine cascade framework to make full use of both types of descriptors for robust head-shoulder detection. Specifically, in the first stage, two histogram-based descriptors, i.e., local Histogram of Oriented Gradients (HOG) and histogram of Local Binary Pattern (LBP), are utilized by a Viola-Jones classifier to rapidly reject most non-head-shoulder candidate windows. In contrast, the second stage further boost the performance via multiple kernel learning on Riemannian manifold formed by Region Covariance Matrix (RCM), a second-order statistic descriptor with stronger discriminative power. Experimental results on a public dataset demonstrate that our method improves detection rate significantly with satisfactory detection speed. Ronghang Hu, Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001 |
ICPR | 2 |
| 2014 | A Parametric Model for Describing the Correlation Between Single Color Images and Depth MapsabstractThis letter introduces a new approach for modeling the correlation between a single color image and its depth map with a set of parameters. The proposed model treats the color image as a set of patches and describes the correlation with a kernel function in a non-linear mapping space. We also present how to estimate the model parameters from sampled color image patches as well as the corresponding depth values. The proposed approach is tested on different color images and experimental results are comparable to the state-of-the-art, which demonstrates the power of the proposed method. Furthermore, we validate the efficiency of the proposed parametric model by evaluating each of its component, including the filters optimization, the choice of the patches and the kernel function. Yangang Wang 0001, Ruiping Wang 0001, Qionghai Dai |
IEEE Signal Process. Lett. | 2 |
| 2014 | A Data-Driven Approach for Facial Expression Retargeting in VideoabstractThis paper presents a data-driven approach for facial expression retargeting in video, i.e., synthesizing a face video of a target subject that mimics the expressions of a source subject in the input video. Our approach takes advantage of a pre-existing facial expression database of the target subject to achieve realistic synthesis. First, for each frame of the input video, a new facial expression similarity metric is proposed for querying the expression database of the target person to select multiple candidate images that are most similar to the input. The similarity metric is developed using a metric learning approach to reliably handle appearance difference between different subjects. Secondly, we employ an optimization approach to choose the best candidate image for each frame, resulting in a retrieved sequence that is temporally coherent. Finally, a spatio-temporal expression mapping method is employed to further improve the synthesized sequence. Experimental results show that our system is capable of generating high quality facial expression videos that match well with the input sequences, even when the source and target subjects have big identity difference. In addition, extensive evaluations demonstrate the high accuracy of the learned expression similarity metric and the effectiveness of our retrieval strategy. Kai Li 0016, Qionghai Dai, Ruiping Wang 0001, Yebin Liu, Feng Xu 0005, Jue Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | Coupling Alignments with Recognition for Still-to-Video Face RecognitionabstractThe Still-to-Video (S2V) face recognition systems typically need to match faces in low-quality videos captured under unconstrained conditions against high quality still face images, which is very challenging because of noise, image blur, low face resolutions, varying head pose, complex lighting, and alignment difficulty. To address the problem, one solution is to select the frames of `best quality' from videos (hereinafter called quality alignment in this paper). Meanwhile, the faces in the selected frames should also be geometrically aligned to the still faces offline well-aligned in the gallery. In this paper, we discover that the interactions among the three tasks-quality alignment, geometric alignment and face recognition-can benefit from each other, thus should be performed jointly. With this in mind, we propose a Coupling Alignments with Recognition (CAR) method to tightly couple these tasks via low-rank regularized sparse representation in a unified framework. Our method makes the three tasks promote mutually by a joint optimization in an Augmented Lagrange Multiplier routine. Extensive experiments on two challenging S2V datasets demonstrate that our method outperforms the state-of-the-art methods impressively. Zhiwu Huang, Shiguang Shan, Ruiping Wang 0001, Xilin Chen 0001 |
ICCV | 4 |
| 2013 | Partial least squares regression on grassmannian manifold for emotion recognitionabstractIn this paper, we propose a method for video-based human emotion recognition. For each video clip, all frames are represented as an image set, which can be modeled as a linear subspace to be embedded in Grassmannian manifold. After feature extraction, Class-specific One-to-Rest Partial Least Squares (PLS) is learned on video and audio data respectively to distinguish each class from the other confusing ones. Finally, an optimal fusion of classifiers learned from both modalities (video and audio) is conducted at decision level. Our method is evaluated on the Emotion Recognition In The Wild Challenge (EmotiW 2013). The experimental results on both validation set and blind test set are presented for comparison. The final accuracy achieved on test set outperforms the baseline by 26%. Ruiping Wang 0001, Zhiwu Huang, Shiguang Shan, Xilin Chen 0001 |
ICMI | 2 |
| 2012 | Covariance discriminative learning: A natural and efficient approach to image set classificationabstractWe propose a novel discriminative learning approach to image set classification by modeling the image set with its natural second-order statistic, i.e. covariance matrix. Since nonsingular covariance matrices, a.k.a. symmetric positive definite (SPD) matrices, lie on a Riemannian manifold, classical learning algorithms cannot be directly utilized to classify points on the manifold. By exploring an efficient metric for the SPD matrices, i.e., Log-Euclidean Distance (LED), we derive a kernel function that explicitly maps the covariance matrix from the Riemannian manifold to a Euclidean space. With this explicit mapping, any learning method devoted to vector space can be exploited in either its linear or kernel formulation. Linear Discriminant Analysis (LDA) and Partial Least Squares (PLS) are considered in this paper for their feasibility for our specific problem. We further investigate the conventional linear subspace based set modeling technique and cast it in a unified framework with our covariance matrix based modeling. The proposed method is evaluated on two tasks: face recognition and object categorization. Extensive experimental results show not only the superiority of our method over state-of-the-art ones in both accuracy and efficiency, but also its stability to two real challenges: noisy set data and varying set size. Ruiping Wang 0001, Huimin Guo, Larry Davis 0001, Qionghai Dai |
CVPR | 1 |
| 2012 | Commute time guided transformation for feature extraction
Yue Deng 0001, Qionghai Dai, Ruiping Wang 0001, Zengke Zhang |
Comput. Vis. Image Underst. | 3 |
| 2012 | Manifold-Manifold Distance and its Application to Face Recognition With Image SetsabstractIn this paper, we address the problem of classifying image sets for face recognition, where each set contains images belonging to the same subject and typically covering large variations. By modeling each image set as a manifold, we formulate the problem as the computation of the distance between two manifolds, called manifold-manifold distance (MMD). Since an image set can come in three pattern levels, point, subspace, and manifold, we systematically study the distance among the three levels and formulate them in a general multilevel MMD framework. Specifically, we express a manifold by a collection of local linear models, each depicted by a subspace. MMD is then converted to integrate the distances between pairs of subspaces from one of the involved manifolds. We theoretically and experimentally study several configurations of the ingredients of MMD. The proposed method is applied to the task of face recognition with image sets, where identification is achieved by seeking the minimum MMD from the probe to the gallery of image sets. Our experiments demonstrate that, as a general set similarity measure, MMD consistently outperforms other competing nondiscriminative methods and is also promisingly comparable to the state-of-the-art discriminative methods. Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001, Qionghai Dai, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Maximal Linear Embedding for Dimensionality ReductionabstractOver the past few decades, dimensionality reduction has been widely exploited in computer vision and pattern analysis. This paper proposes a simple but effective nonlinear dimensionality reduction algorithm, named Maximal Linear Embedding (MLE). MLE learns a parametric mapping to recover a single global low-dimensional coordinate space and yields an isometric embedding for the manifold. Inspired by geometric intuition, we introduce a reasonable definition of locally linear patch, Maximal Linear Patch (MLP), which seeks to maximize the local neighborhood in which linearity holds. The input data are first decomposed into a collection of local linear models, each depicting an MLP. These local models are then aligned into a global coordinate space, which is achieved by applying MDS to some randomly selected landmarks. The proposed alignment method, called Landmarks-based Global Alignment (LGA), can efficiently produce a closed-form solution with no risk of local optima. It just involves some small-scale eigenvalue problems, while most previous aligning techniques employ time-consuming iterative optimization. Compared with traditional methods such as ISOMAP and LLE, our MLE yields an explicit modeling of the intrinsic variation modes of the observation data. Extensive experiments on both synthetic and real data indicate the effectivity and efficiency of the proposed algorithm. Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001, Jie Chen 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Manifold Discriminant AnalysisabstractThis paper presents a novel discriminative learning method, called manifold discriminant analysis (MDA), to solve the problem of image set classification. By modeling each image set as a manifold, we formulate the problem as classification-oriented multi-manifolds learning. Aiming at maximizing “manifold margin”, MDA seeks to learn an embedding space, where manifolds with different class labels are better separated, and local data compactness within each manifold is enhanced. As a result, new testing manifold can be more reliably classified in the learned embedding space. The proposed method is evaluated on the tasks of object recognition with image sets, including face recognition and object categorization. Comprehensive comparisons and extensive experiments demonstrate the effectiveness of our method. Ruiping Wang 0001, Xilin Chen 0001 |
CVPR | 1 |
| 2009 | Optimization of a training set for more robust face detection
Jie Chen 0001, Xilin Chen 0001, Jie Yang 0001, Shiguang Shan, Ruiping Wang 0001, Wen Gao 0001 |
Pattern Recognit. | 5 |
| 2008 | Manifold-Manifold Distance with application to face recognition based on image setabstractIn this paper, we address the problem of classifying image sets, each of which contains images belonging to the same class but covering large variations in, for instance, viewpoint and illumination. We innovatively formulate the problem as the computation of Manifold-Manifold Distance (MMD), i.e., calculating the distance between nonlinear manifolds each representing one image set. To compute MMD, we also propose a novel manifold learning approach, which expresses a manifold by a collection of local linear models, each depicted by a subspace. MMD is then converted to integrating the distances between pair of subspaces respectively from one of the involved manifolds. The proposed MMD method is evaluated on the task of Face Recognition based on Image Set (FRIS). In FRIS, each known subject is enrolled with a set of facial images and modeled as a gallery manifold, while a testing subject is modeled as a probe manifold, which is then matched against all the gallery manifolds by MMD. Identification is achieved by seeking the minimum MMD. Experimental results on two public face databases, Honda/UCSD and CMU MoBo, demonstrate that the proposed MMD method outperforms the competing methods. Ruiping Wang 0001, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
CVPR | 1 |
| 2007 | Enhancing Human Face Detection by Resampling Examples Through ManifoldsabstractAs a large-scale database of hundreds of thousands of face images collected from the Internet and digital cameras becomes available, how to utilize it to train a well-performed face detector is a quite challenging problem. In this paper, we propose a method to resample a representative training set from a collected large-scale database to train a robust human face detector. First, in a high-dimensional space, we estimate geodesic distances between pairs of face samples/examples inside the collected face set by isometric feature mapping (Isomap) and then subsample the face set. After that, we embed the face set to a low-dimensional manifold space and obtain the low-dimensional embedding. Subsequently, in the embedding, we interweave the face set based on the weights computed by locally linear embedding (LLE). Furthermore, we resample nonfaces by Isomap and LLE likewise. Using the resulting face and nonface samples, we train an AdaBoost-based face detector and run it on a large database to collect false alarms. We then use the false detections to train a one-class support vector machine (SVM). Combining the AdaBoost and one-class SVM-based face detector, we obtain a stronger detector. The experimental results on the MIT + CMU frontal face test set demonstrated that the proposed method significantly outperforms the other state-of-the-art methods. Jie Chen 0001, Ruiping Wang 0001, Shengye Yan, Shiguang Shan, Xilin Chen 0001, Wen Gao 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 2 |