EDBT 2026 Demo / reviewers in the wild / expert
Yang Wu 0001
dblp:56/1428-1
· DBLP profile ↗
93ranked-venue papers
16as first author
49since 2021 · last 2026
0000-0001-8010-6857ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 10 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 58 · 11 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical frequency adaptation for all-in-one image restoration
Yang Wu 0001, Ye Deng 0005, Siqi Hui, Yuhan Liu 0006, Kangyi Wu, Wenli Huang 0004, Jinjun Wang |
Knowl. Based Syst. | 1 |
| 2026 | DictCR-former: Content-aware dictionary transformer for cloud removal
Wenli Huang 0004, Yang Wu 0001, Sanping Zhou, Xiaomeng Xin, Xiaobo Jia, Ye Deng 0005 |
Pattern Recognit. | 2 |
| 2026 | Frequency-guided generalizable representation learning for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Wenli Huang 0004, Yang Wu 0001, Jinjun Wang |
Pattern Recognit. | 5 |
| 2025 | Event-Equalized Dense Video CaptioningabstractDense video captioning aims to localize and caption all events in arbitrary untrimmed videos. Although previous methods have achieved appealing results, they still face the issue of temporal bias, i.e, models tend to focus more on events with certain temporal characteristics. Specifically, 1) the temporal distribution of events in training datasets is uneven. Models trained on these datasets will pay less attention to out-of-distribution events. 2) long-duration events have more frame features than short ones and will attract more attention. To address this, we argue that events, with varying temporal characteristics, should be treated equally when it comes to dense video captioning. Intuitively, different events tend to have distinct visual differences due to varied camera views, backgrounds, or subjects. Inspired by that, we intend to utilize visual features to have an approximate perception of possible events and pay equal attention to them. In this paper, we introduce a simple but effective framework, called Event-Equalized Dense Video Captioning (E2DVC) to overcome the temporal bias and treat all possible events equally. Experimental results on ActivityNet Captions and YouCook2 dataset validate the effectiveness of the proposed methods and show State-of-the-art (SOTA) performance on dense video captioning. Kangyi Wu, Pengna Li, Jingwen Fu, Yang Wu 0001, Yuhan Liu 0006, Jinjun Wang, Sanping Zhou |
CVPR | 5 |
| 2025 | Correlative3D: Inter-Object Correlation-Aware 3D Scene UnderstandingabstractHolistic 3D scene understanding from a single image is challenging due to the information loss in 2D-to-3D reconstruction. Existing approaches either explore object properties independently or overlook their levels of correlation, leading to inaccurate estimations in complex scenes. To this end, we propose Correlative3D, a novel inter-object correlation-aware method for 3D scene understanding. Our method integrates a Scene Graph Attention Network to implicitly enhance features through a correlation-driven weighting strategy, selectively prioritizing relationships among objects. In addition, an auxiliary task pertaining to the relative arrangement of objects is formulated to impose explicit constraints. Furthermore, we introduce a novel 3DCIoU loss that sensitively responds to geometric variations in 3D bounding boxes. Extensive experiments demonstrate that our method produces more coherent scene layouts compared to existing methods. Tingxuan Gao, Wenming Yang, Yang Wu 0001, Yehu Shen |
ICASSP | 3 |
| 2025 | Mind the Gap: Aligning Vision Foundation Models to Image Feature MatchingabstractLeveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy between the foundation models focusing on single-image understanding and the cross-image understanding requirement of feature matching. Specifically, 1) the embeddings derived from commonly used foundation models exhibit discrepancies with the optimal embeddings required for feature matching; 2) lacking an effective mechanism to leverage the single-image understanding ability into cross-image understanding. A significant consequence of the misalignment is they struggle when addressing multi-instance feature matching problems. To address this, we introduce a simple but effective framework, called IMD (Image feature Matching with a pre-trained Diffusion model) with two parts: 1) Unlike the dominant solutions employing contrastive-learning based foundation models that emphasize global semantics, we integrate the generative-based diffusion models to effectively capture instance-level details. 2) We leverage the prompt mechanism in generative model as a natural tunnel, propose a novel cross-image interaction prompting module to facilitate bidirectional information interaction between image pairs. To more accurately measure the misalignment, we propose a new benchmark called IMIM, which focuses on multi-instance scenarios. Our proposed IMD establishes a new state-of-the-art in commonly evaluated benchmarks, and the superior improvement 12% in IMIM indicates our method efficiently mitigates the misalignment. Yuhan Liu 0006, Jingwen Fu, Yang Wu 0001, Kangyi Wu, Pengna Li, Jiayi Wu 0002, Sanping Zhou, Jingmin Xin |
ICCV | 3 |
| 2025 | Auxiliary Loss Reweighting for Image InpaintingabstractImage inpainting aims to reconstruct missing regions in corrupted images with semantically consistent content. While modern methods employ perceptual and style losses to enhance inpainting quality by supervising deep feature representations, two key challenges persist: (i) existing approaches necessitate time-consuming grid searches to determine optimal loss weights, and (ii) heterogeneous auxiliary loss terms are assigned fixed weights, limiting their adaptive contributions. To address these limitations, we propose a framework featuring dynamically weighted auxiliary losses and an automated weight adaptation mechanism. Specifically, we introduce Tunable Perceptual Loss (TPL) and Tunable Style Loss (TSL), which generalize traditional perceptual and style losses by incorporating tunable weights that independently scale distinct loss components according to their auxiliary potential. These are optimized via our Adaptive Weight Adjustment (AWA) algorithm, which dynamically reweights TPL and TSL during training by prioritizing loss terms that maximally improve inpainting performance. Empirical evaluations on public datasets demonstrate that our framework enhances state-of-the-art inpainting performance while eliminating manual weight tuning. Wenli Huang 0004, Siqi Hui, Ye Deng 0005, Xiaomeng Xin, Yang Wu 0001, Jinjun Wang |
IECON | 5 |
| 2025 | UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual ReasoningabstractRecent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention has been given to scaling fine-grained pixel-level understanding capabilities, where the models are expected to realize pixel-level alignment between visual signals and language semantics. Some previous studies have applied LMMs to related tasks such as region-level captioning and referring expression segmentation. However, these models are limited to performing either referring or segmentation tasks independently and fail to integrate these fine-grained perception capabilities into visual reasoning. To bridge this gap, we propose UniPixel, a large multi-modal model capable of flexibly comprehending visual prompt inputs and generating mask-grounded responses. Our model distinguishes itself by seamlessly integrating pixel-level perception with general visual understanding capabilities. Specifically, UniPixel processes visual prompts and generates relevant masks on demand, and performs subsequent reasoning conditioning on these intermediate pointers during inference, thereby enabling fine-grained pixel-level reasoning. The effectiveness of our approach has been verified on 10 benchmarks across a diverse set of tasks, including pixel-level referring/segmentation and object-centric understanding in images/videos. A novel PixelQA task that jointly requires referring, segmentation, and question answering is also designed to verify the flexibility of our method. Ye Liu 0002, Zongyang Ma, Junfu Pu, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
NeurIPS | 5 |
| 2024 | Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion ModelabstractText-guided motion synthesis aims to generate 3D human motion that not only precisely reflects the textual description but reveals the motion details as much as possible. Pioneering methods explore the diffusion model for text-to-motion synthesis and obtain significant superiority. However, these methods conduct diffusion processes either on the raw data distribution or the low-dimensional latent space, which typically suffer from the problem of modality inconsistency or detail-scarce. To tackle this problem, we propose a novel Basic-to-Advanced Hierarchical Diffusion Model, named B2A-HDM, to collaboratively exploit low-dimensional and high-dimensional diffusion models for high quality detailed motion synthesis. Specifically, the basic diffusion model in low-dimensional latent space provides the intermediate denoising result that to be consistent with the textual description, while the advanced diffusion model in high-dimensional latent space focuses on the following detail-enhancing denoising process. Besides, we introduce a multi-denoiser framework for the advanced diffusion model to ease the learning of high-dimensional model and fully explore the generative potential of the diffusion model. Quantitative and qualitative experiment results on two text-to-motion benchmarks (HumanML3D and KIT-ML) demonstrate that B2A-HDM can outperform existing state-of-the-art methods in terms of fidelity, modality consistency, and diversity. Zhenyu Xie, Yang Wu 0001, Xuehao Gao, Zhongqian Sun, Wei Yang 0019, Xiaodan Liang |
AAAI | 2 |
| 2024 | Learning Pseudo 3D Guidance for View-Consistent Texturing with 2D Diffusion
Kehan Li 0002, Yanbo Fan, Yang Wu 0001, Zhongqian Sun, Wei Yang 0019, Xiangyang Ji, Li Yuan 0007, Jie Chen 0001 |
ECCV (86) | 3 |
| 2024 | DEGAN: Discrimination Enhanced GAN for Perceptual-Oriented Super-ResolutionabstractRecent years, generative adversarial networks (GANs) have gained significant prominence in single image super-resolution (SISR) tasks. This can mainly be attributed to their exceptional ability to generate intricate details. However, the instability and lack of realism in the details generated by GANs have been challenges. Existing methods mainly concentrate on improving the generator and designing complex loss functions, often overlooking the important role of discrimination. To this end, we propose our discrimination enhanced GAN (DEGAN) by improving the discriminator and simplify the discrimination task. We introduce an efficient wide activation UNet to enhance the discriminator, enabling a more comprehensive and nuanced analysis of the input image. Additionally, we introduce a texture aware mask that provides more precise guidance and alleviates the difficulty of discrimination. Our DEGAN is simple yet effective. Quantitative and visual comparisons with state-of-the-art methods on benchmark datasets demonstrate the superiority of our method. Xiaoyu Jin, Wenqi Huang 0002, Lingyu Liang, Yang Wu 0001, Qunsheng Zeng, Ruiye Zhou, Zhuojun Cai, Jianing Shang, Wenming Yang |
ICASSP | 4 |
| 2024 | PortraitNeRF: A Single Neural Radiance Field for Complete and Coordinated Talking Portrait GenerationabstractWe present a novel framework named PortraitNeRF to generate high-fidelity talking portrait videos for performing faithful identity-preserving reenactment of source videos. This is a challenging task because the generated results should be natural and match the speaker’s head movement, expression, eye blinks and speech audio. To acquire sufficient guidance from source video, the proposed PortraitNeRF exploits not only speech audio but also detailed motion information derived from visual data, including facial expressions, head pose and head position information. By adopting only a single neural radiance field, PortraitNeRF is able to generate complete and coordinated portrait video without bells and whistles. The completeness is ensured by the single-NeRF structure, and the superior head-torso coordination ability comes from using head pose and position information as its conditional input. Moreover, a simple yet effective mouth region emphasis strategy that fits well with the NeRF mechanism helps improving the accuracy of mouth shape. Experimental results and ablation studies demonstrate the superiority and effectiveness of PortraitNeRF. Xiuzhe Wu, Yang Wu 0001, Wenming Yang |
ICME | 3 |
| 2024 | E.T. Bench: Towards Open-Ended Event-Level Video-Language UnderstandingabstractRecent advances in Video Large Language Models (Video-LLMs) have demonstrated their great potential in general-purpose video understanding. To verify the significance of these models, a number of benchmarks have been proposed to diagnose their capabilities in different scenarios. However, existing benchmarks merely evaluate models through video-level question-answering, lacking fine-grained event-level assessment and task diversity. To fill this gap, we introduce E.T. Bench (Event-Level & Time-Sensitive Video Understanding Benchmark), a large-scale and high-quality benchmark for open-ended event-level video understanding. Categorized within a 3-level task taxonomy, E.T. Bench encompasses 7.3K samples under 12 tasks with 7K videos (251.4h total length) under 8 domains, providing comprehensive evaluations. We extensively evaluated 8 Image-LLMs and 12 Video-LLMs on our benchmark, and the results reveal that state-of-the-art models for coarse-level (video-level) understanding struggle to solve our fine-grained tasks, e.g., grounding event-of-interests within videos, largely due to the short video context length, improper time representations, and lack of multi-event training data. Focusing on these issues, we further propose a strong baseline model, E.T. Chat, together with an instruction-tuning dataset E.T. Instruct 164K tailored for fine-grained event-level understanding. Our simple but effective solution demonstrates superior performance in multiple scenarios. Ye Liu 0002, Zongyang Ma, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
NeurIPS | 4 |
| 2024 | Gradient-guided channel masking for cross-domain few-shot learning
Siqi Hui, Sanping Zhou, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
Knowl. Based Syst. | 4 |
| 2024 | Sparse self-attention transformer for image inpainting
Wenli Huang 0004, Ye Deng 0005, Siqi Hui, Yang Wu 0001, Sanping Zhou, Jinjun Wang |
Pattern Recognit. | 4 |
| 2024 | Fully Unsupervised Domain-Agnostic Image RetrievalabstractRecent research in cross-domain image retrieval has focused on addressing two challenging issues: handling domain variations in the data and dealing with the lack of sufficient training labels. However, these problems have often been studied separately, limiting the practicality and significance of the research outcomes. The existing cross-domain setting is also restricted to cases where domain labels are known during training, and all samples have semantic category information or instance correspondences. In this paper, we propose a novel approach to address a more general and practical problem:fully unsupervised domain-agnostic image retrievalunder the domain-unknown setting, where no annotations are provided. Our approach tackles both thedomain variationandmissing labelschallenges simultaneously. We introduce a new fully unsupervised One-Shot Synthesis-based Contrastive learning method (termed OSSCo) to project images from different data distributions into a shared feature space for similarity measurement. To handle the domain-unknown setting, we propose One-Shot unpaired image-to-image Translation (OST) between a randomly selected one-shot image and the rest of the training images. By minimizing the global distance between the original images and the generated images from OST, the model learns domain-agnostic representations. To address the label-unknown setting, we employ contrastive learning with a synthesis-based transform module from the OST training. This allows for effective representation learning without any annotations or external constraints. We evaluate our proposed method on diverse datasets, and the results demonstrate its effectiveness. Notably, our approach achieves comparable performance to current state-of-the-art supervised methods. Ziqiang Zheng, Hao Ren 0002, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Attentive Contextual Attention for Cloud RemovalabstractCloud cover can significantly hinder the use of remote sensing images for Earth observation, prompting urgent advancements in cloud removal technology. Recently, deep learning strategies, especially convolutional neural networks (CNNs) with attention mechanisms, have shown strong potential in restoring cloud-obscured areas. These methods utilize convolution to extract intricate local features and attention mechanisms to gather long-range information, improving the overall comprehension of the scene. However, a common drawback of these approaches is that the resulting images often suffer from blurriness, artifacts, and inconsistencies. This is partly because attention mechanisms apply weights to all features based on generalized similarity scores, which can inadvertently introduce noise and irrelevant details from cloud-covered areas. To overcome this limitation and better capture relevant distant context, we introduce a novel approach named attentive contextual attention (AC-Attention). This method enhances conventional attention mechanisms by dynamically learning data-driven attentive selection scores, enabling it to filter out noise and irrelevant features effectively. By integrating the AC-Attention module into the DSen2-CR cloud removal framework, we significantly improve the model’s ability to capture essential distant information, leading to more effective cloud removal. Our extensive evaluation of various datasets shows that our method outperforms existing ones regarding image reconstruction quality. In addition, we conducted ablation studies by integrating AC-Attention into multiple existing methods and widely used network architectures. These studies demonstrate the effectiveness and adaptability of AC-Attention and reveal its ability to focus on relevant features, thereby improving the overall performance of the networks. The code is available athttps://github.com/huangwenwenlili/ACA-CRNet. Wenli Huang 0004, Ye Deng 0005, Yang Wu 0001, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | CR-former: Single-Image Cloud Removal With Focused Taylor AttentionabstractCloud removal aims to restore high-quality images from cloud-contaminated captures, which is essential in remote sensing applications. Effectively modeling the long-range relationships between image features is key to achieving high-quality cloud-free images. While self-attention mechanisms excel at modeling long-distance relationships, their computational complexity scales quadratically with image resolution, limiting their applicability to high-resolution remote sensing images. Current cloud removal methods have mitigated this issue by restricting the global receptive field to smaller regions or adopting channel attention to model long-range relationships. However, these methods either compromise pixel-level long-range dependencies or lose spatial information, potentially leading to structural inconsistencies in restored images. In this work, we propose the focused Taylor attention (FT-Attention), which captures pixel-level long-range relationships without limiting the spatial extent of attention and achieves the$\mathcal {O}(N)$computational complexity, where N represents the image resolution. Specifically, we utilize Taylor series expansions to reduce the computational complexity of the attention mechanism from$\mathcal {O}(N^{2})$to$\mathcal {O}(N)$, enabling efficient capture of pixel relationships directly in high-resolution images. Additionally, to fully leverage the informative pixel, we develop a new normalization function for the query and key, which produces more distinguishable attention weights, enhancing focus on important features. Building on FT-Attention, we design a U-net style network, termed the CR-former, specifically for cloud removal. Extensive experimental results on representative cloud removal datasets demonstrate the superior performance of our CR-former. The code is available athttps://github.com/wuyang2691/CR-former. Yang Wu 0001, Ye Deng 0005, Sanping Zhou, Yuhan Liu 0006, Wenli Huang 0004, Jinjun Wang |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion PredictionabstractInferring 3D human motion is fundamental in many applications, including understanding human activity and analyzing one's intention. While many fruitful efforts have been made to human motion prediction, most approaches focus on pose-driven prediction and inferring human motion in isolation from the contextual environment, thus leaving the body location movement in the scene behind. However, real-world human movements are goal-directed and highly influenced by the spatial layout of their surrounding scenes. In this paper, instead of planning future human motion in a "dark" room, we propose a Multi-Condition Latent Diffusion network (MCLD) that reformulates the human motion prediction task as a multi-condition joint inference problem based on the given historical 3D body motion and the current 3D scene contexts. Specifically, instead of directly modeling joint distribution over the raw motion sequences, MCLD performs a conditional diffusion process within the latent embedding space, characterizing the cross-modal mapping from the past body movement and current scene context condition embeddings to the future human motion embedding. Extensive experiments on large-scale human motion prediction datasets demonstrate that our MCLD achieves significant improvements over the state-of-the-art methods on both realistic and diverse predictions. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du, Guo-Jun Qi |
IEEE Trans. Image Process. | 3 |
| 2024 | Learning Heterogeneous Spatial-Temporal Context for Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have been widely used and achieved fruitful progress in the skeleton-based action recognition task. In GCNs, node interaction modeling dominates the context aggregation and, therefore, is crucial for a graph-based convolution kernel to extract representative features. In this article, we introduce a closer look at a powerful graph convolution formulation to capture rich movement patterns from these skeleton-based graphs. Specifically, we propose a novel heterogeneous graph convolution (HetGCN) that can be considered as the middle ground between the extremes of (2 + 1)-D and 3-D graph convolution. The core observation of HetGCN is that multiple information flows are jointly intertwined in a 3-D convolution kernel, including spatial, temporal, and spatial-temporal cues. Since spatial and temporal information flows characterize different cues for action recognition, HetGCN first dynamically analyzes pairwise interactions between each node and its cross-space-time neighbors and then encourages heterogeneous context aggregation among them. Considering the HetGCN as a generic convolution formulation, we further develop it into two specific instantiations (i.e., intra-scale and inter-scale HetGCN) that significantly facilitate cross-space-time and cross-scale learning on skeleton graphs. By integrating these modules, we propose a strong human action recognition system that outperforms state-of-the-art methods with the accuracy of 93.1% on NTU-60 cross-subject (X-Sub) benchmark, 88.9% on NTU-120 X-Sub benchmark, and 38.4% on kinetics skeleton. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | DAGCN: Dynamic and Adaptive Graph Convolutional Network for Salient Object DetectionabstractDeep-learning-based salient object detection (SOD) has achieved significant success in recent years. The SOD focuses on the context modeling of the scene information, and how to effectively model the context relationship in the scene is the key. However, it is difficult to build an effective context structure and model it. In this article, we propose a novel SOD method called dynamic and adaptive graph convolutional network (DAGCN) that is composed of two parts, adaptive neighborhood-wise graph convolutional network (AnwGCN) and spatially restricted K-nearest neighbors (SRKNN). The AnwGCN is novel adaptive neighborhood-wise graph convolution, which is used to model and analyze the saliency context. The SRKNN constructs the topological relationship of the saliency context by measuring the non-Euclidean spatial distance within a limited range. The proposed method constructs the context relationship as a topological graph by measuring the distance of the features in the non-Euclidean space, and conducts comparative modeling of context information through AnwGCN. The model has the ability to learn the metrics from features and can adapt to the hidden space distribution of the data. The description of the feature relationship is more accurate. Through the convolutional kernel adapted to the neighborhood, the model obtains the structure learning ability. Therefore, the graph convolution process can adapt to different graph data. Experimental results demonstrate that our solution achieves satisfactory performance on six widely used datasets and can also effectively detect camouflaged objects. Our code will be available at: https://github.com/CSIM-LUT/DAGCN.git. Ce Li 0001, Fenghua Liu, Shaoyi Du, Yang Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion GenerationabstractIn this article, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up generation objectives by grouping body joints of detailed skeletons in close semantic proximity together and then replacing each of such joint group with a single body-part node. Such an operation recursively abstracts a human pose to coarser and coarser skeletons at multiple granularity levels. With gradually increasing the abstraction level, human motion becomes more and more concise and stable, significantly benefiting the cross-modal motion synthesis task. The whole text-driven human motion synthesis problem is then divided into multiple abstraction levels and solved with a multi-stage generation framework with a cascaded latent diffusion model: an initial generator first generates the coarsest human motion guess from a given text description; then, a series of successive generators gradually enrich the motion details based on the textual description and the previous synthesized results. Notably, we further integrate GUESS with the proposed dynamic multi-condition fusion mechanism to dynamically balance the cooperative effects of the given textual condition and synthesized coarse motion prompt in different generation stages. Extensive experiments on large-scale datasets verify that GUESS outperforms existing state-of-the-art methods by large margins in terms of accuracy, realisticness, and diversity. Xuehao Gao, Yang Yang 0066, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, Yang Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Decompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion PredictionabstractEncouraged by the effectiveness of encoding temporal dynamics within the frequency domain, recent human motion prediction systems prefer to first convert the motion representation from the original pose space into the frequency space. In this paper, we introduce two closer looks at effective frequency representation learning for robust motion prediction and summarize them as: decompose more and aggregate better. Motivated by these two insights, we develop two powerful units that factorize the frequency representation learning task with a novel decomposition-aggregation two-stage strategy: (1) frequency decomposition unit unweaves multi-view frequency representations from an input body motion by embedding its frequency features into multiple spaces; (2) feature aggregation unit deploys a series of intra-space and inter-space feature aggregation layers to collect comprehensive frequency representations from these spaces for robust human motion prediction. As evaluated on large-scale datasets, we develop a strong baseline model for the human motion prediction task that outperforms state-of-the-art methods by large margins: 8%∼12% on Human3.6M, 3%∼7% on CMU MoCap, and 7%∼10% on 3DPW. Xuehao Gao, Shaoyi Du, Yang Wu 0001, Yang Yang 0066 |
CVPR | 3 |
| 2023 | Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoabstractSynthesizing realistic videos according to a given speech is still an open challenge. Previous works have been plagued by issues such as inaccurate lip shape generation and poor image quality. The key reason is that only motions and appearances on limited facial areas (e.g., lip area) are mainly driven by the input speech. Therefore, directly learning a mapping function from speech to the entire head image is prone to ambiguity, particularly when using a short video for training. We thus propose a decomposition-synthesis-composition framework named Speech to Lip (Speech2Lip) that disentangles speech-sensitive and speech-insensitive motion/appearance to facilitate effective learning from limited training data, resulting in the generation of natural-looking videos. First, given a fixed head pose (i.e., canonical space), we present a speech-driven implicit model for lip image generation which concentrates on learning speech-sensitive motion and appearance. Next, to model the major speech-insensitive motion (i.e., head movement), we introduce a geometry-aware mutual explicit mapping (GAMEM) module that establishes geometric mappings between different head poses. This allows us to paste generated lip images at the canonical space onto head images with arbitrary poses and synthesize talking videos with natural head movements. In addition, a Blend-Net and a contrastive sync loss are introduced to enhance the overall synthesis performance. Quantitative and qualitative results on three benchmarks demonstrate that our model can be trained by a video of just a few minutes in length and achieve state-of-the-art performance in both visual quality and speechvisual synchronization. Code: https://github.com/CVMILab/Speech2Lip. Xiuzhe Wu, Yang Wu 0001, Xiaoyang Lyu, Yan-Pei Cao 0001, Ying Shan, Wenming Yang, Zhongqian Sun, Xiaojuan Qi 0001 |
ICCV | 3 |
| 2023 | Cross-Domain Autonomous Driving Perception Using Contrastive Appearance AdaptationabstractAddressing domain shifts for complex perception tasks in autonomous driving has long been a challenging problem. In this paper, we show that existing domain adaptation methods pay little attention to the content mismatch issue between source and target domains, thus weakening the domain adaptation per-formance and the decoupling of domain-invariant and domain-specific representations. To solve the aforementioned problems, we propose an image-level domain adaptation framework that aims at adapting source-domain images to the target domain with content-aligned source-target image pairs. Our framework consists of three mutually beneficial modules in a cycle: a cross-domain content alignment module to generate source-target pairs with consistent content representations in a self-supervised manner, a reference-guided image synthesis based on the generated content-aligned source-target image pairs, and a contrastive learning module to self-supervise domain-invariant feature extractor. Our contrastive appearance adaptation is task-agnostic and robust to complex perception tasks in autonomous driving. Our proposed method demonstrates state-of-the-art results in cross-domain object detection, semantic segmentation, and depth estimation as well as better image synthesis ability qualitatively and quantitatively. Ziqiang Zheng, Yingshu Chen, Binh-Son Hua, Yang Wu 0001, Sai-Kit Yeung |
IROS | 4 |
| 2023 | Toward Human Perception-Centric Video Thumbnail GenerationabstractVideo thumbnail plays an essential role in summarizing video content into a compact and concise image for users to browse efficiently. However, automatically generating attractive and informative video thumbnails remains an open problem due to the difficulty of formulating human aesthetic perception and the scarcity of paired training data. This work proposes a novel Human Perception-Centric Video Thumbnail Generation (HPCVTG) to address these challenges. Specifically, our framework first generates a set of thumbnails using a principle-based system, which conforms to established aesthetic and human perception principles, such as visual balance in the layout and avoiding overlapping elements. Then rather than designing from scratch, we ask human annotators to evaluate some of these thumbnails and select their preferred ones. A Transformer-based Variational Auto-Encoder (VAE) model is firstly pre-trained with Model-Agnostic Meta-Learning (MAML) and then fine-tuned on these human-selected thumbnails. The exploration of combining the MAML pre-training paradigm with human feedback in training can reduce human involvement and make the training process more efficient. Extensive experimental results show that our HPCVTG framework outperforms existing methods in objective and subjective evaluations, highlighting its potential to improve the user experience when browsing videos and inspire future research in human perception-centric content generation tasks. The code and dataset will be released via https://github.com/yangtao2019yt/HPCVTG. Junfan Lin, Zhongang Qi, Yang Wu 0001, Ying Shan, Chang Wen Chen |
ACM Multimedia | 5 |
| 2023 | Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsabstractMost text-driven human motion generation methods employ sequential modeling approaches, e.g., transformer, to extract sentence-level text representations automatically and implicitly for human motion synthesis. However, these compact text representations may overemphasize the action names at the expense of other important properties and lack fine-grained details to guide the synthesis of subtly distinct motion. In this paper, we propose hierarchical semantic graphs for fine-grained control over motion generation. Specifically, we disentangle motion descriptions into hierarchical semantic graphs including three levels of motions, actions, and specifics. Such global-to-local structures facilitate a comprehensive understanding of motion description and fine-grained control of motion generation. Correspondingly, to leverage the coarse-to-fine topology of hierarchical semantic graphs, we decompose the text-to-motion diffusion process into three semantic levels, which correspond to capturing the overall motion, local actions, and action specifics. Extensive experiments on two benchmark human motion datasets, including HumanML3D and KIT, with superior performances, justify the efficacy of our method. More encouragingly, by modifying the edge weights of hierarchical semantic graphs, our method can continuously refine the generated motion, which may have a far-reaching impact on the community. Code and pre-trained weights are available at https://github.com/jpthu17/GraphMotion. Peng Jin 0001, Yang Wu 0001, Yanbo Fan, Zhongqian Sun, Wei Yang 0019, Li Yuan 0007 |
NeurIPS | 2 |
| 2023 | CL-NeRF: Continual Learning of Neural Radiance Fields for Evolving Scene RepresentationabstractExisting methods for adapting Neural Radiance Fields (NeRFs) to scene changes require extensive data capture and model retraining, which is both time-consuming and labor-intensive. In this paper, we tackle the challenge of efficiently adapting NeRFs to real-world scene changes over time using a few new images while retaining the memory of unaltered areas, focusing on the continual learning aspect of NeRFs. To this end, we propose CL-NeRF, which consists of two key components: a lightweight expert adaptor for adapting to new changes and evolving scene representations and a conflict-aware knowledge distillation learning objective for memorizing unchanged parts. We also present a new benchmark for evaluating Continual Learning of NeRFs with comprehensive metrics. Our extensive experiments demonstrate that CL-NeRF can synthesize high-quality novel views of both changed and unchanged regions with high training efficiency, surpassing existing methods in terms of reducing forgetting and adapting to changes. Code and benchmark will be made available. Xiuzhe Wu, Peng Dai 0003, Weipeng Deng, Handi Chen, Yang Wu 0001, Yan-Pei Cao 0001, Ying Shan, Xiaojuan Qi 0001 |
NeurIPS | 5 |
| 2023 | DaCo: domain-agnostic contrastive learning for visual place recognition
Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001 |
Appl. Intell. | 3 |
| 2023 | Video Region Annotation with Sparse Bounding Boxes
Yuzheng Xu, Yang Wu 0001, Nur Sabrina binti Zuraimi, Shohei Nobuhara, Ko Nishino |
Int. J. Comput. Vis. | 2 |
| 2023 | Joint regularization and low-rank fusion for atmospheric turbulence removal
Yanyun Qu, Yuan Xie 0006, Yang Wu 0001, Hanzi Wang |
Neural Comput. Appl. | 5 |
| 2023 | Specialized re-ranking: A novel retrieval-verification framework for cloth changing person re-identification
Huaxin Song, Fangbin Wan, Yanwei Fu 0001, Hirokazu Kato 0001, Yang Wu 0001 |
Pattern Recognit. | 7 |
| 2023 | ACNet: Approaching-and-Centralizing Network for Zero-Shot Sketch-Based Image RetrievalabstractThe huge domain gap between sketches and photos poses huge challenges for Sketch-Based Image Retrieval (SBIR). The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) is more generic and practical but brings an even greater challenge: the additional knowledge gap between the seen and unseen categories. In order to simultaneously mitigate both gaps, we propose an Approaching-and-Centralizing Network (termed “ACNet”) to jointly optimize sketch-to-photo synthesis and image retrieval. The retrieval module guides the synthesis module to generate large amounts of diverse photo-like images that help the sketch domain gradually approach the photo domain to eliminate the domain gap, and thus better serves retrieval. Meanwhile, the retrieval module itself centralizes the embeddings of training samples for learning a similarity measurement to eliminate the knowledge gap. Our approach is simple yet effective, which achieves state-of-the-art performance on two widely used ZS-SBIR datasets and surpasses previous methods by a large margin (eg, 8.2% improvement in terms of mAP@all on TU-Berlin Extended dataset). Hao Ren 0002, Ziqiang Zheng, Yang Wu 0001, Hong Lu 0001, Yang Yang 0002, Ying Shan, Sai-Kit Yeung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Asynchronous Generative Adversarial Network for Asymmetric Unpaired Image-to-Image TranslationabstractThe unpaired image-to-image translation aims to translate input images from one source domain to some desired outputs in a target domain by learning from unpaired training data. Cycle-consistency constraint provides a general principle to estimate and measure forward and backward mapping functions between two domains. In many cases, the information entropy of images from the two domains is not equal, resulting in an information-rich domain and an information-poor domain. However, existing solutions based on cycle-consistency either completely discard the information asymmetry between the two domains (a common choice), which leads to inferior translation performance for the asymmetric unpaired image-to-image translation, or have to rely on special task-specific designs and introduce extra loss components. These elaborative designs especially for the relatively harder translation direction from the information-poor domain to the information-rich domain (poor-to-rich translation) require extra labor and are limited to some specific tasks. In this paper, we propose a novel asynchronous generative adversarial network named Async-GAN, which provides a model-agnostic framework for easily turning symmetrical models into powerful asymmetric counterparts that can handle asymmetric unpaired image-to-image translation much better. The key innovation is to iteratively build gradually improving intermediate domains for generating pseudo paired training samples, which provide stronger full supervision for assisting the poor-to-rich translation. Extensive experiments on various asymmetric unpaired translation tasks demonstrate the superiority of the proposal. Furthermore, the proposed training framework could be extended to various Cycle-GAN solutions and achieve a performance gain. Ziqiang Zheng, Yi Bin, Xiaoou Lv, Yang Wu 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Multim. | 4 |
| 2022 | UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight DetectionabstractFinding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment retrieval and highlight detection is an emerging research topic, even though its component problems and some related tasks have already been studied for a while. In this paper, we present the first unified framework, named Unified Multi-modal Transformers (UMT), capable of realizing such joint optimization while can also be easily degenerated for solving individual problems. As far as we are aware, this is the first scheme to integrate multi-modal (visual-audio) learning for either joint optimization or the individual moment retrieval task, and tackles moment retrieval as a keypoint detection problem using a novel query generator and query decoder. Extensive comparisons with existing methods and ablation studies on QVHighlights, Charades-STA, YouTube Highlights, and TVSum datasets demonstrate the effectiveness, superiority, and flexibility of the proposed method under various settings. Source code and pre-trained models are available at https://github.com/TencentARC/UMT. Ye Liu 0002, Siyuan Li 0026, Yang Wu 0001, Chang Wen Chen, Ying Shan, Xiaohu Qie |
CVPR | 3 |
| 2022 | Learning Knowledge Graph Embedding with Batch Circle LossabstractKnowledge Graph Embedding (KGE) is the process to learn low-dimension representations for entities and relations in knowledge graphs. It is a critical component in Knowledge Graph (KG) for link prediction and knowledge discovery. Many works focus on designing proper score function for KGE, while the study of loss function has attracted relatively less attention. In this paper, we focus on improving the loss function when learning KGE. Specifically, we find that the frequently used margin-based loss in KGE models seeks to maximize the gap between the true facts score fpand the false facts score fnand only cares about the relative order of scores. Since its optimization objective is fp- fn= m, increasing fpis equivalent to decreasing fn. Its optimization objective creates an ambiguous convergence status which impairs the separability of positive and negative facts in embedding space. Inspired by the circle loss that offers a more flexible optimization manner with definite convergence targets and is widely used in computer vision tasks, we further extend it into the KGE model with the presented Batch Circle Loss (BCL). BCL allows multiple positives to be considered per anchor (h, r) (or (r, t)) in addition to multiple negatives (as opposed to a single positive sample as used before in KGE models). By comparing with other approaches, the obtained KGE models using our proposed loss function and training method shows superior performance. Yang Wu 0001, Wenli Huang 0004, Siqi Hui, Jinjun Wang |
IJCNN | 1 |
| 2022 | UoLMM'22: 2nd International Workshop on Robust Understanding of Low-quality Multimedia Data: Unitive Enhancement, Analysis and EvaluationabstractLow-quality multimedia data (including low resolution, low illumination, defects, blurriness, etc.) often pose a challenge for content understanding, as algorithms are typically developed under ideal conditions (high resolution and good visibility). To alleviate this problem, data enhancement techniques (e.g., super-resolution, low-light enhancement, derain, and inpainting) have been proposed to restore low-quality multimedia data. Efforts are also being made to develop robust content understanding algorithms in adverse weather and lighting conditions. Some quality assessment techniques aiming at evaluating the analytical quality of data have also emerged. Even though these topics are mostly studied independently, they are tightly related in terms of ensuring a robust understanding of multimedia content. For example, enhancement should maintain the semantic consistency of the analysis, while quality assessment should consider the comprehensibility of the multimedia data. The purpose of this workshop is to bring together individuals in three areas: enhancement, analysis, and evaluation, for sharing ideas and discussion on current developments and future directions. Yang Wu 0001, Xiao Wang 0029, Jing Xiao 0004 |
ACM Multimedia | 3 |
| 2022 | Identification of Bird's Nest Hazard Level of Transmission Line Based on Improved Yolov5 and Location Constraints
Yang Wu 0001, Qunsheng Zeng, Wenqi Huang 0002, Lingyu Liang |
PRCV (4) | 1 |
| 2022 | Depth-Aware Shadow RemovalabstractAbstract Shadow removal from a single image is an ill‐posed problem because shadow generation is affected by the complex interactions of geometry, albedo, and illumination. Most recent deep learning‐based methods try to directly estimate the mapping between the non‐shadow and shadow image pairs to predict the shadow‐free image. However, they are not very effective for shadow images with complex shadows or messy backgrounds. In this paper, we propose a novel end‐to‐end depth‐aware shadow removal method without using depth images, which estimates depth information from RGB images and leverages the depth feature as guidance to enhance shadow removal and refinement. The proposed framework consists of three components, including depth prediction, shadow removal, and boundary refinement. First, the depth prediction module is used to predict the corresponding depth map of the input shadow image. Then, we propose a new generative adversarial network (GAN) method integrated with depth information to remove shadows in the RGB image. Finally, we propose an effective boundary refinement framework to alleviate the artifact around boundaries after shadow removal by depth cues. We conduct experiments on several public datasets and real‐world shadow images. The experimental results demonstrate the efficiency of the proposed method and superior performance against state‐of‐the‐art methods. Yanping Fu, Zhenyu Gai, Haifeng Zhao 0001, Shaojie Zhang 0002, Ying Shan, Yang Wu 0001, Jin Tang 0001 |
Comput. Graph. Forum | 6 |
| 2022 | ResLNet: deep residual LSTM network with longer input for action recognition
Tian Wang 0002, Huai-Ning Wu, Ce Li 0001, Hichem Snoussi, Yang Wu 0001 |
Frontiers Comput. Sci. | 6 |
| 2022 | Tackling multiple object tracking with complicated motions - Re-designing the integration of motion and appearance
Fan Yang 0032, Zheng Wang 0007, Yang Wu 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
Image Vis. Comput. | 3 |
| 2021 | Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkabstractLocalizing individuals in crowds is more in accordance with the practical demands of subsequent high-level crowd analysis tasks than simply counting. However, existing localization based methods relying on intermediate representations (i.e., density maps or pseudo boxes) serving as learning targets are counter-intuitive and error-prone. In this paper, we propose a purely point-based framework for joint crowd counting and individual localization. For this framework, instead of merely reporting the absolute counting error at image level, we propose a new metric, called density Normalized Average Precision (nAP), to provide more comprehensive and more precise performance evaluation. Moreover, we design an intuitive solution under this framework, which is called Point to Point Network (P2PNet). P2PNet discards superfluous steps and directly predicts a set of point proposals to represent heads in an image, being consistent with the human annotation results. By thorough analysis, we reveal the key step towards implementing such a novel idea is to assign optimal learning targets for these proposals. Therefore, we propose to conduct this crucial association in an one-to-one matching manner using the Hungarian algorithm. The P2PNet not only significantly surpasses state-of-the-art methods on popular counting benchmarks, but also achieves promising localization accuracy. The codes will be available at: TencentYoutuResearch/CrowdCounting-P2PNet. Qingyu Song 0001, Changan Wang, Zhengkai Jiang 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yang Wu 0001 |
ICCV | 9 |
| 2021 | Uniformity in Heterogeneity: Diving Deep into Count Interval Partition for Crowd CountingabstractRecently, the problem of inaccurate learning targets in crowd counting draws increasing attention. Inspired by a few pioneering work, we solve this problem by trying to predict the indices of pre-defined interval bins of counts instead of the count values themselves. However, an inappropriate interval setting might make the count error contributions from different intervals extremely imbalanced, leading to inferior counting performance. Therefore, we propose a novel count interval partition criterion called Uniform Error Partition (UEP), which always keeps the expected counting error contributions equal for all intervals to minimize the prediction risk. Then to mitigate the inevitably introduced discretization errors in the count quantization process, we propose another criterion called Mean Count Proxies (MCP). The MCP criterion selects the best count proxy for each interval to represent its count value during inference, making the overall expected discretization error of an image nearly negligible. As far as we are aware, this work is the first to delve into such a classification task and ends up with a promising solution for count interval partition. Following the above two theoretically demonstrated criterions, we propose a simple yet effective model termed Uniform Error Partition Network (UEPNet), which achieves state-of-the-art performance on several challenging datasets. The codes will be available at: TencentYoutuResearch/CrowdCounting-UEPNet. Changan Wang, Qingyu Song 0001, Boshen Zhang, Yabiao Wang, Ying Tai, Xuyi Hu, Chengjie Wang 0001, Jiayi Ma 0001, Yang Wu 0001 |
ICCV | 10 |
| 2021 | SiamRCR: Reciprocal Classification and Regression for Visual Object TrackingabstractRecently, most siamese network based trackers locate targets via object classification and bounding-box regression. Generally, they select the bounding-box with maximum classification confidence as the final prediction. This strategy may miss the right result due to the accuracy misalignment between classification and regression. In this paper, we propose a novel siamese tracking algorithm called SiamRCR, addressing this problem with a simple, light and effective solution. It builds reciprocal links between classification and regression branches, which can dynamically re-weight their losses for each positive sample. In addition, we add a localization branch to predict the localization accuracy, so that it can work as the replacement of the regression assistance link during inference. This branch makes the training and inference more consistent. Extensive experimental results demonstrate the effectiveness of SiamRCR and its superiority over the state-of-the-art competitors on GOT-10k, LaSOT, TrackingNet, OTB-2015, VOT-2018 and VOT-2019. Moreover, our SiamRCR runs at 65 FPS, far above the real-time requirement. Jinlong Peng, Zhengkai Jiang 0001, Yueyang Gu, Yang Wu 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Weiyao Lin |
IJCAI | 4 |
| 2021 | Human-object interaction detection with missing objects
Kaen Kogashi, Yang Wu 0001, Shohei Nobuhara, Ko Nishino |
Image Vis. Comput. | 2 |
| 2021 | ReMOT: A model-agnostic refinement for multiple object tracking
Fan Yang 0032, Sakriani Sakti, Yang Wu 0001, Satoshi Nakamura 0001 |
Image Vis. Comput. | 4 |
| 2021 | Generative Adversarial Network with Multi-branch Discriminator for imbalanced cross-species image-to-image translation
Ziqiang Zheng, Zhibin Yu 0002, Yang Wu 0001, Haiyong Zheng, Minho Lee 0001 |
Neural Networks | 3 |
| 2021 | Local minima found in the subparameter space can be effective for ensembles of deep convolutional neural networks
Yongquan Yang, Haijun Lv, Yang Wu 0001, Zhongxi Zheng |
Pattern Recognit. | 4 |
| 2021 | Instance-Level Heterogeneous Domain Adaptation for Limited-Labeled Sketch-to-Photo RetrievalabstractAlthough sketch-to-photo retrieval has a wide range of applications, it is costly to obtain paired and rich-labeled ground truth. Differently, photo retrieval data is easier to acquire. Therefore, previous works pre-train their models on rich-labeled photo retrieval data (i.e., source domain) and then fine-tune them on the limited-labeled sketch-to-photo retrieval data (i.e., target domain). However, without co-training source and target data, source domain knowledge might be forgotten during the fine-tuning process, while simply co-training them may cause negative transfer due to domain gaps. Moreover, identity label spaces of source data and target data are generally disjoint and therefore conventional category-level Domain Adaptation (DA) is not directly applicable. To address these issues, we propose an Instance-level Heterogeneous Domain Adaptation (IHDA) framework. We apply the fine-tuning strategy for identity label learning, aiming to transfer the instance-level knowledge in an inductive transfer manner. Meanwhile, labeled attributes from the source data are selected to form a shared label space for source and target domains. Guided by shared attributes, DA is utilized to bridge cross-dataset domain gaps and heterogeneous domain gaps, which transfers instance-level knowledge in a transductive transfer manner. Experiments show that our method has set a new state of the art on three sketch-to-photo image retrieval benchmarks without extra annotations, which opens the door to train more effective models on limited-labeled heterogeneous image retrieval tasks. Fan Yang 0032, Yang Wu 0001, Zheng Wang 0007, Sakriani Sakti, Satoshi Nakamura 0001 |
IEEE Trans. Multim. | 2 |
| 2020 | Video Region Annotation with Sparse Bounding Boxes
Yuzheng Xu, Yang Wu 0001, Nur Sabrina binti Zuraimi, Shohei Nobuhara, Ko Nishino |
BMVC | 2 |
| 2020 | Dynamic Face Video Segmentation via Reinforcement LearningabstractFor real-time semantic video segmentation, most recent works utilised a dynamic framework with a key scheduler to make online key/non-key decisions. Some works used a fixed key scheduling policy, while others proposed adaptive key scheduling methods based on heuristic strategies, both of which may lead to suboptimal global performance. To overcome this limitation, we model the online key decision process in dynamic video segmentation as a deep reinforcement learning problem and learn an efficient and effective scheduling policy from expert information about decision history and from the process of maximising global return. Moreover, we study the application of dynamic video segmentation on face videos, a field that has not been investigated before. By evaluating on the 300VW dataset, we show that the performance of our reinforcement key scheduler outperforms that of various baselines in terms of both effective key selections and running speed. Further results on the Cityscapes dataset demonstrate that our proposed method can also generalise to other scenarios. To the best of our knowledge, this is the first work to use reinforcement learning for online key-frame decision in dynamic video segmentation, and also the first work on its application on face videos. Yujiang Wang 0001, Mingzhi Dong, Jie Shen 0008, Yang Wu 0001, Shiyang Cheng 0001, Maja Pantic |
CVPR | 4 |
| 2020 | Chained-Tracker: Chaining Paired Attentive Regression Results for End-to-End Joint Multiple-Object Detection and Tracking
Jinlong Peng, Changan Wang, Fangbin Wan, Yang Wu 0001, Yabiao Wang, Ying Tai, Chengjie Wang 0001, Feiyue Huang, Yanwei Fu 0001 |
ECCV (4) | 4 |
| 2020 | ForkGAN: Seeing into the Rainy Night
Ziqiang Zheng, Yang Wu 0001, Xinran Han, Jianbo Shi |
ECCV (3) | 2 |
| 2020 | Using Panoramic Videos for Multi-Person Localization and Tracking In A 3D Panoramic Coordinateabstract3D panoramic multi-person localization and tracking are prominent in many applications, however, conventional methods using LiDAR equipment could be economically expensive and also computationally inefficient due to the processing of point cloud data. In this work, we propose an effective and efficient approach at a low cost. First, we obtain panoramic videos with four normal cameras. Then, we transform human locations from a 2D panoramic image coordinate to a 3D panoramic camera coordinate using camera geometry and human bio-metric property (i.e., height). Finally, we generate 3D tracklets by associating human appearance and 3D trajectory. We verify the effectiveness of our method on three datasets including a new one built by us, in terms of 3D single-view multi-person localization, 3D single-view multi-person tracking, and 3D panoramic multi-person localization and tracking. Our code and dataset are available at https://github.com/fandulu/MPLT. Fan Yang 0032, Feiran Li, Yang Wu 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
ICASSP | 3 |
| 2020 | Beyond Intra-modality: A Survey of Heterogeneous Person Re-identificationabstractAn efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous person re-identification (Hetero-ReID). In this paper, we provide a comprehensive review of state-of-the-art Hetero-ReID methods that address the challenge of inter-modality discrepancies. According to the application scenario, we classify the methods into four categories --- low-resolution, infrared, sketch, and text. We begin with an introduction of ReID, and make a comparison between Homogeneous ReID (Homo-ReID) and Hetero-ReID tasks. Then, we describe and compare existing datasets for performing evaluations, and survey the models that have been widely employed in Hetero-ReID. We also summarize and compare the representative approaches from two perspectives, i.e., the application scenario and the learning pipeline. We conclude by a discussion of some future research directions. Follow-up updates are available at https://github.com/lightChaserX/Awesome-Hetero-reID Zheng Wang 0007, Zhixiang Wang 0001, Yinqiang Zheng, Yang Wu 0001, Wenjun Zeng 0001, Shin'ichi Satoh 0001 |
IJCAI | 4 |
| 2020 | FTBME: feature transferring based multi-model ensemble
Yongquan Yang, Haijun Lv, Yang Wu 0001, Zhongxi Zheng |
Multim. Tools Appl. | 4 |
| 2020 | Compressing 3DCNNs based on tensor train decomposition
Dingheng Wang, Guang-She Zhao, Guoqi Li 0002, Lei Deng 0003, Yang Wu 0001 |
Neural Networks | 5 |
| 2020 | Learning Sparse and Identity-Preserved Hidden Attributes for Person Re-IdentificationabstractPerson re-identification (Re-ID) aims at matching person images captured in non-overlapping camera views. To represent person appearance, low-level visual features are sensitive to environmental changes, while high-level semantic attributes, such as "short-hair" or "long-hair", are relatively stable. Hence, researches have started to design semantic attributes to reduce the visual ambiguity. However, to train a prediction model for semantic attributes, it requires plenty of annotations, which are hard to obtain in practical large-scale applications. To alleviate the reliance on annotation efforts, we propose to incrementally generate Deep Hidden Attribute (DHA) based on baseline deep network for newly uncovered annotations. In particular, we propose an auto-encoder model that can be plugged into any deep network to mine latent information in an unsupervised manner. To optimize the effectiveness of DHA, we reform the auto-encoder model with additional orthogonal generation module, along with identity-preserving and sparsity constraints. 1) Orthogonally generating: In order to make DHAs different from each other, Singular Vector Decomposition (SVD) is introduced to generate DHAs orthogonally. 2) Identity-preserving constraint: The generated DHAs should be distinct for telling different persons, so we associate DHAs with person identities. 3) Sparsity constraint: To enhance the discriminability of DHAs, we also introduce the sparsity constraint to restrict the number of effective DHAs for each person. Experiments conducted on public datasets have validated the effectiveness of the proposed network. On two large-scale datasets, i.e., Market-1501 and DukeMTMC-reID, the proposed method outperforms the state-of-the-art methods. Zheng Wang 0007, Junjun Jiang, Yang Wu 0001, Mang Ye, Xiang Bai, Shin'ichi Satoh 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Make Skeleton-based Action Recognition Model Smaller, Faster and BetterabstractAlthough skeleton-based action recognition has achieved great success in recent years, most of the existing methods may suffer from a large model size and slow execution speed. To alleviate this issue, we analyze skeleton sequence properties to propose a Double-feature Double-motion Network (DD-Net) for skeleton-based action recognition. By using a lightweight network structure (i.e., 0.15 million parameters), DD-Net can reach a super fast speed, as 3,500 FPS on an ordinary GPU (e.g., GTX 1080Ti), or, 2,000 FPS on an ordinary CPU (e.g., Intel E5-2620). By employing robust features, DD-Net achieves state-of-the-art performance on our experiment datasets: SHREC (i.e., hand actions) and JHMDB (i.e., body actions). Our code is on https://github.com/fandulu/DD-Net. Fan Yang 0032, Yang Wu 0001, Sakriani Sakti, Satoshi Nakamura 0001 |
MMAsia | 2 |
| 2019 | Explorations on visual localization from active to passive
Yongquan Yang, Yang Wu 0001 |
Multim. Tools Appl. | 2 |
| 2018 | Temporal-Enhanced Convolutional Network for Person Re-IdentificationabstractWe propose a new neural network called Temporal-enhanced Convolutional Network (T-CN) for video-based person reidentification. For each video sequence of a person, a spatial convolutional subnet is first applied to each frame for representing appearance information, and then a temporal convolutional subnet links small ranges of continuous frames to extract local motion information. Such spatial and temporal convolutions together construct our T-CN based representation. Finally, a recurrent network is utilized to further explore global dynamics, followed by temporal pooling to generate an overall feature vector for the whole sequence. In the training stage, a Siamese network architecture is adopted to jointly optimize all the components with losses covering both identification and verification. In the testing stage, our network generates an overall discriminative feature representation for each input video sequence (whose length may vary a lot) in a feed-forward way, and even a simple Euclidean distance based matching can generate good re-identification results. Experiments on the most widely used benchmark datasets demonstrate the superiority of our proposal, in comparison with the state-of-the-art. Yang Wu 0001, Jun Takamatsu, Tsukasa Ogasawara |
AAAI | 1 |
| 2018 | Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future GoalsabstractIn this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints. Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001 |
CVPR | 16 |
| 2018 | Pose-Normalized Image Generation for Person Re-identification
Xuelin Qian, Yanwei Fu 0001, Tao Xiang 0002, Wenxuan Wang 0003, Yang Wu 0001, Yu-Gang Jiang 0001, Xiangyang Xue 0001 |
ECCV (9) | 6 |
| 2018 | Dynamic Ensemble Active Learning: A Non-Stationary Bandit with Expert AdviceabstractActive learning aims to reduce annotation cost by predicting which samples are useful for a human teacher to label. However it has become clear there is no best active learning algorithm. Inspired by various philosophies about what constitutes a good criteria, different algorithms perform well on different datasets. This has motivated research into ensembles of active learners that learn what constitutes a good criteria in a given scenario, typically via multi-armed bandit algorithms. Though algorithm ensembles can lead to better results, they overlook the fact that not only does algorithm efficacy vary across datasets, but also during a single active learning session. That is, the best criteria is non-stationary. This breaks existing algorithms' guarantees and hampers their performance in practice. In this paper, we propose dynamic ensemble active learning as a more general and promising research direction. We develop a dynamic ensemble active learner based on a non-stationary multi-armed bandit with expert advice algorithm. Our dynamic ensemble selects the right criteria at each step of active learning. It has theoretical guarantees, and shows encouraging results on 13 popular datasets. Kunkun Pang, Mingzhi Dong, Yang Wu 0001, Timothy M. Hospedales |
ICPR | 3 |
| 2017 | Transferring CNNS to multi-instance multi-label classification on small datasetsabstractImage tagging is a well known challenge in image processing. It is typically addressed through multi-instance multi-label (MIML) classification methodologies. Convolutional Neural Networks (CNNs) possess great potential to perform well on MIML tasks, since multi-level convolution and max pooling coincide with the multi-instance setting and the sharing of hidden representation may benefit multi-label modeling. However, CNNs usually require a large amount of carefully labeled data for training, which is hard to obtain in many real applications. In this paper, we propose a new approach for transferring pre-trained deep networks such as VGG16 on Imagenet to small MIML tasks. We extract features from each group of the network layers and apply multiple binary classifiers to them for multi-label prediction. Moreover, we adopt an L1-norm regularized Logistic Regression (L1LR) to find the most effective features for learning the multi-label classifiers. The experiment results on two most-widely used and relatively small benchmark MIML image datasets demonstrate that the proposed approach can substantially outperform the state-of-the-art algorithms, in terms of all popular performance metrics. Mingzhi Dong, Kunkun Pang, Yang Wu 0001, Jing-Hao Xue, Timothy M. Hospedales, Tsukasa Ogasawara |
ICIP | 3 |
| 2017 | Re-identification by neighborhood structure metric learning
Wei Li 0049, Yang Wu 0001, Jianqing Li 0006 |
Pattern Recognit. | 2 |
| 2017 | Joint Hierarchical Category Structure Learning and Large-Scale Image ClassificationabstractWe investigate the scalable image classification problem with a large number of categories. Hierarchical visual data structures are helpful for improving the efficiency and performance of large-scale multi-class classification. We propose a novel image classification method based on learning hierarchical inter-class structures. Specifically, we first design a fast algorithm to compute the similarity metric between categories, based on which a visual tree is constructed by hierarchical spectral clustering. Using the learned visual tree, a test sample label is efficiently predicted by searching for the best path over the entire tree. The proposed method is extensively evaluated on the ILSVRC2010 and Caltech 256 benchmark datasets. The experimental results show that our method obtains significantly better category hierarchies than other state-of-the-art visual tree-based methods and, therefore, much more accurate classification. Yanyun Qu, Li Lin 0005, Fumin Shen, Yang Wu 0001, Yuan Xie 0006, Dacheng Tao |
IEEE Trans. Image Process. | 5 |
| 2015 | Saturation-preserving specular reflection separationabstractSpecular reflection generally decreases the saturation of surface colors, which will be possibly confused with other colors that have the same hue but lower saturation. Traditional methods for specular reflection separation suffer this problem of hue-saturation ambiguity, producing over-saturated specular-free images quite often. We proposed a two-step approach to solve this problem. In the first step, we produce an over-saturated specular-free image by global chromaticity propagation from specular-free pixels to highlighted ones. Then we recover the saturation based on priors of the piecewise constancy of diffuse chromaticity as well as the spatial sparsity and smoothness of specular reflection. We achieve this through increasing the achromatic component of diffuse chromaticity, while the magnitudes of increments are determined by linear programming under the constraints derived from the priors. Experiments on both laboratory and natural images show that our method can separate the specular reflection while preserving the saturation of the underlying surface colors. Yuanliu Liu, Zejian Yuan, Nanning Zheng 0001, Yang Wu 0001 |
CVPR | 4 |
| 2015 | Learning local Gaussian process regression for image super-resolution
Yanyun Qu, Cuihua Li, Yuan Xie 0006, Yang Wu 0001, Jianping Fan 0001 |
Neurocomputing | 5 |
| 2015 | Locality based discriminative measure for multiple-shot human re-identification
Wei Li 0049, Yang Wu 0001, Masayuki Mukunoki, Yinghui Kuang, Michihiko Minoh |
Neurocomputing | 2 |
| 2014 | Discriminative Collaborative Representation for Classification
Yang Wu 0001, Wei Li 0049, Masayuki Mukunoki, Michihiko Minoh, Shihong Lao |
ACCV (4) | 1 |
| 2014 | Description-Discrimination Collaborative Tracking
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Yang Wu 0001, Nanning Zheng 0001 |
ECCV (1) | 4 |
| 2013 | Locality based discriminative measure for multiple-shot person re-identificationabstractMultiple-shot person re-identification tackles the problem to build the correspondences between sets of human images obtained from distributed cameras. It is challenging due to large within-class variations and small between-class differences, caused by the changing of human appearance and environment. Existing methods for addressing this issue include designing the representation to capture the within-set correlation, or crafting the measure to explore the between-set separation. This paper proposes a novel set based matching model called “Locality Based Discriminative Measure (LBDM)”, in which the discriminative potentiality of a new set-to-set distance is exploited by using the learned local metric field. As experimentally demonstrated, the proposal remarkably outperforms state-of-the-art schemes on public benchmark datasets. Wei Li 0049, Yang Wu 0001, Masayuki Mukunoki, Michihiko Minoh |
AVSS | 2 |
| 2013 | Collaboratively Regularized Nearest Points for Set Based RecognitionabstractSet based recognition has been attracting more and more attention in recent years, benefitting from two facts: the difficulty of collecting sets of images for recognition fades quickly, and set based recognition models generally outperform the ones for single instance based recognition. In this paper, we propose a novel model called collaboratively regularized nearest points (CRNP) for solving this problem. The proposal inherits the merits of simplicity, robustness, and high-efficiency from the very recently introduced regularized nearest points (RNP) method on finding the set-to-set distance using the l2-norm regularized affine hulls. Meanwhile, CRNP makes use of the powerful discriminative ability induced by collaborative representation, following the same idea as that in sparse recognition for classification (SRC) for image-based recognition and collaborative sparse approximation (CSA) for set-based recognition. However, CRNP uses l2-norm instead of the expensive l1-norm for coefficients regularization, which makes it much more efficient. Extensive experiments on five benchmark datasets for face recognition and person re-identification demonstrate that CRNP is not only more effective but also significantly faster than other state-of-the-art methods, including RNP and CSA. Yang Wu 0001, Michihiko Minoh, Masayuki Mukunoki |
BMVC | 1 |
| 2013 | Salient Object Detection: A Discriminative Regional Feature Integration ApproachabstractSalient object detection has been attracting a lot of interest, and recently various heuristic computational models have been designed. In this paper, we regard saliency map computation as a regression problem. Our method, which is based on multi-level image segmentation, uses the supervised learning approach to map the regional feature vector to a saliency score, and finally fuses the saliency scores across multiple levels, yielding the saliency map. The contributions lie in two-fold. One is that we show our approach, which integrates the regional contrast, regional property and regional background ness descriptors together to form the master saliency map, is able to produce superior saliency maps to existing algorithms most of which combine saliency maps heuristically computed from different types of features. The other is that we introduce a new regional feature vector, background ness, to characterize the background, which can be regarded as a counterpart of the objectness descriptor [2]. The performance evaluation on several popular benchmark data sets validates that our approach outperforms existing state-of-the-arts. Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001, Shipeng Li 0001 |
CVPR | 4 |
| 2013 | Constructing Adaptive Complex Cells for Robust Visual TrackingabstractRepresentation is a fundamental problem in object tracking. Conventional methods track the target by describing its local or global appearance. In this paper we present that, besides the two paradigms, the composition of local region histograms can also provide diverse and important object cues. We use cells to extract local appearance, and construct complex cells to integrate the information from cells. With different spatial arrangements of cells, complex cells can explore various contextual information at multiple scales, which is important to improve the tracking performance. We also develop a novel template-matching algorithm for object tracking, where the template is composed of temporal varying cells and has two layers to capture the target and background appearance respectively. An adaptive weight is associated with each complex cell to cope with occlusion as well as appearance variation. A fusion weight is associated with each complex cell type to preserve the global distinctiveness. Our algorithm is evaluated on 25 challenging sequences, and the results not only confirm the contribution of each component in our tracking system, but also outperform other competing trackers. Dapeng Chen, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001 |
ICCV | 3 |
| 2013 | Probabilistic salient object contour detection based on superpixelsabstractIn this paper, we propose a data-driven approach to detect the probabilistic salient object contour, which is formulated as predicting the probability of superpixel boundaries being on the object contour based on the learned regressor. Each superpixel boundary is jointly described by the superpixel saliency, superpixel contrast, and boundary geometry features. Experimental results on the benchmark data set validate the effectiveness of our approach. Furthermore, we demonstrate that the predicted probabilistic salient object contour is useful for improving the multiple segmentations for salient object detection. Huaizu Jiang, Yang Wu 0001, Zejian Yuan |
ICIP | 2 |
| 2013 | Can feature-based inductive transfer learning help person re-identification?abstractPerson re-identification concerns about the problem of recognizing people across space (captured by different cameras) and/or over time gaps. Though recently the literature on it grows rapidly, all the proposed solutions have treated it as a normal classification or ranking problem. In this paper, however, we argue that it is in fact a natural transfer learning problem, thus it's valuable and also necessary to investigate how the progress on transfer learning could benefit the research on it. We present so far the first study on justifying the effectiveness of a representative transfer learning methodology: feature-based inductive transfer learning, for person re-identification. Extensive experiments on standard datasets with typical methods result in several important findings. Yang Wu 0001, Wei Li 0049, Michihiko Minoh, Masayuki Mukunoki |
ICIP | 1 |
| 2012 | Collaborative Sparse Approximation for Multiple-Shot Across-Camera Person Re-identificationabstractIn this paper we propose a simple and effective solution to the important and challenging problem of across-camera person re-identification. We focus on the common case in video surveillance where multiple images or video frames are available for each person. Instead of exploring new features, the proposed approach aims at making a better use of such images/frames. It builds a collaborative representation over all the gallery images (of known person individuals) to best approximate the query images (containing an unknown person) via affine combinations. The approximation is measured by the nearest point distance between the two affine hulls constructed by the query images and gallery images, respectively. By enforcing the sparsity of the samples used for approximating the two nearest points, the relative importance of the gallery images belonging to different persons has the ability to reveal the identity of the querying person. Extensive experiments on public benchmark datasets demonstrate that the proposed approach greatly outperforms the state-of-the-art methods. Yang Wu 0001, Michihiko Minoh, Masayuki Mukunoki, Wei Li 0049, Shihong Lao |
AVSS | 1 |
| 2012 | Set Based Discriminative Ranking for Recognition
Yang Wu 0001, Michihiko Minoh, Masayuki Mukunoki, Shihong Lao |
ECCV (3) | 1 |
| 2012 | Common-near-neighbor analysis for person re-identificationabstractPerson re-identification tackles the problem whether an observed person of interest reappears in a network of cameras. The difficulty primarily originates from few samples per class but large amounts of intra-class variations in real scenarios: illumination, pose and viewpoint changes across cameras. So far, proposals in the literature have treated this either as a matching problem focusing on feature representation or as a classification/ranking problem relying on metric optimization. This paper presents a new way called Common-Near-Neighbor Analysis, which to some extent combines the strengths of these two methodologies. It analyzes the commonness of the near neighbors of each pair of samples in a learned metric space, measured by a novel rank-order based dissimilarity. Our method, using only color cue, has been tested on widely-used benchmark datasets, showing significant performance improvement over the state-of-the-art. Wei Li 0049, Yang Wu 0001, Masayuki Mukunoki, Michihiko Minoh |
ICIP | 2 |
| 2012 | Robust object recognition via third-party collaborative representation
Yang Wu 0001, Michihiko Minoh, Masayuki Mukunoki, Shihong Lao |
ICPR | 1 |
| 2012 | Fine-Grained and Layered Object RecognitionabstractThis paper presents a novel research on promoting the performance and enriching the functionalities of object recognition. Instead of simply fitting various data to a few predefined semantic object categories, we propose to generate proper results for different object instances based on their actual visual appearances. The results can be fine-grained and layered categorization along with absolute or relative localization. We present a generic model based on structured prediction and an efficient online learning algorithm to solve it. Experiments on a new benchmark dataset demonstrate the effectiveness of our model and its superiority against traditional recognition methods. Yang Wu 0001, Nanning Zheng 0001, Yuanliu Liu, Zejian Yuan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2012 | IAIR-CarPed: A psychophysically annotated dataset with fine-grained and layered semantic labels for object recognition
Yang Wu 0001, Yuanliu Liu, Zejian Yuan, Nanning Zheng 0001 |
Pattern Recognit. Lett. | 1 |
| 2011 | Optimizing Mean Reciprocal Rank for person re-identificationabstractPerson re-identification is one of the most challenging issues in network-based surveillance. The difficulties mainly come from the great appearance variations induced by illumination, camera view and body pose changes. Maybe influenced by the research on face recognition and general object recognition, this problem is habitually treated as a verification or classification problem, and much effort has been put on optimizing standard recognition criteria. However, we found that in practical applications the users usually have different expectations. For example, in a real surveillance system, we may expect that a visual user interface can show us the relevant images in the first few (e.g. 20) candidates, but not necessarily before all the irrelevant ones. In other words, there is no problem to leave the final judgement to the users. Based on such an observation, this paper treats the re-identification problem as a ranking problem and directly optimizes a listwise ranking function named Mean Reciprocal Rank (MRR), which is considered by us to be able to generate results closest to human expectations. Using a maximum-margin based structured learning model, we are able to show improved re-identification results on widely-used benchmark datasets. Yang Wu 0001, Masayuki Mukunoki, Takuya Funatomi, Michihiko Minoh, Shihong Lao |
AVSS | 1 |
| 2011 | Object detection using discriminative photogrammetric contextabstractPhotogrammetric context captures the relationship between object heights and camera viewpoint, and can be used to reject false detections that appear in wrong locations or scales. In this work, we address the problem of using photogrammetric constraints in object detection when camera poses are unknown. We propose a model to capture both local appearance features and global photogrammetric context, in which the camera pose is treated as a latent variable. We use latent Structural SVM to learn the model parameters. To solve the NP-hard problem in structured prediction, we propose a branch-bound-and-cut algorithm, where cuts of the latent variable are embedded into a branch-and-bound process. The model is experimentally evaluated on INRIA pedestrian dataset. The results show that our model can get significantly better detection performance than models using only appearance features or using photogrammetric context in a graphical model. Yuanliu Liu, Yang Wu 0001, Zejian Yuan |
ICIP | 2 |
| 2008 | Saliency Based Opportunistic Search for Object Part Extraction and Labeling
Yang Wu 0001, Qihui Zhu, Jianbo Shi, Nanning Zheng 0001 |
ECCV (4) | 1 |
| 2008 | Contour Context Selection for Object Detection: A Set-to-Set Contour Matching Approach
Qihui Zhu, Yang Wu 0001, Jianbo Shi |
ECCV (2) | 3 |
| 2008 | Analysis of Solution for Supervised Graph EmbeddingabstractRecently, Graph Embedding Framework has been proposed for feature extraction. However, it is still an open issue on how to compute robust discriminant transformation for this purpose. In this paper, we show that supervised graph embedding algorithms share a general criterion. Based on the analysis of this criterion, we propose a general solution, called General Solution for Supervised Graph Embedding (GSSGE), for extracting the robust discriminant transformation of Supervised Graph Embedding. Then, we analyze the superiority of our algorithm over traditional algorithms. Extensive experiments on both artificial and real-world data are performed to demonstrate the effectiveness and robustness of our proposed GSSGE. Qubo You, Nanning Zheng 0001, Shaoyi Du, Yang Wu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2007 | General Solution for Supervised Graph Embedding
Qubo You, Nanning Zheng 0001, Shaoyi Du, Yang Wu 0001 |
ECML | 4 |
| 2007 | AN Extension of the ICP Algorithm Considering Scale FactorabstractThe ICP algorithm is accurate and fast for registration between two point sets in a same scale, but it doesn't handle the case with different scales. This paper instead introduces a novel approach named the scaling iterative closest point (SICP) algorithm which integrates a scale matrix with boundaries into the original ICP algorithm for scaling registration. This method uses a simple iterative algorithm with the SVD algorithm and the properties of parabola incorporated to compute the translation, rotation and scale transformations at each iterative step, and its convergence is rapid with only a few iterations. The SICP algorithm is independent of shape representation and feature extraction; thereby it is general for scaling registration. Experimental results demonstrate its robustness and fast speed compared with the standard ICP algorithm. Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Qubo You, Yang Wu 0001 |
ICIP (5) | 5 |
| 2007 | Object Recognition by Learning Informative, Biologically Inspired Visual FeaturesabstractThis paper presents a novel, effective way to improve the object recognition performance of a biologically-motivated model by learning informative visual features. The original model has an obvious bottleneck when learning features. Therefore, we propose a circumspect algorithm to solve this problem. First, a novel information factor was designed to find the most informative feature for each image, and then complementary features were selected based on additional information. Finally, an intra-class clustering strategy was used to select the most typical features for each category. By integrating two other improvements, our algorithm performs better than any other system so far based on the same model. Yang Wu 0001, Nanning Zheng 0001, Qubo You, Shaoyi Du |
ICIP (1) | 1 |
| 2007 | Neighborhood discriminant projection for face recognition
Qubo You, Nanning Zheng 0001, Shaoyi Du, Yang Wu 0001 |
Pattern Recognit. Lett. | 4 |