Jingdong Wang 0001

dblp:49/3441 · DBLP profile ↗
← Back
342ranked-venue papers
25as first author
161since 2021 · last 2026
0000-0002-4888-4445ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 241 · 13 first-author · 142 since 2021Graphics, computer vision, multimedia, augmented reality and games · 225 · 8 first-author · 97 since 2021Databases, data management, data science and information retrieval · 13 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2
YearPublicationVenuePosition
2026 EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
abstract
Efficient Multimodal Large Language Models (MLLMs) compress vision tokens to reduce resource consumption, but the loss of visual information can degrade comprehension capabilities. Although some priors introduce Knowledge Distillation to enhance student models, they overlook the fundamental differences in fine-grained vision comprehension caused by unbalanced vision tokens between the efficient student and vanilla teacher. In this paper, we propose EM-KD, a novel paradigm that enhances the Efficient MLLMs with Knowledge Distillation. To overcome the challenge of unbalanced vision tokens, we first calculate the Manhattan distance between the vision logits of teacher and student, and then align them in the spatial dimension with the Hungarian matching algorithm. After alignment, EM-KD introduces two distillation strategies: 1) Vision-Language Affinity Distillation (VLAD) and 2) Vision Semantic Distillation (VSD). Specifically, VLAD calculates the affinity matrix between text tokens and aligned vision tokens, and minimizes the smooth L1 distance of the student and the teacher affinity matrices. Considering the semantic richness of vision logits in the final layer, VSD employs the reverse KL divergence to measure the discrete probability distributions of the aligned vision logits over the vocabulary space. Comprehensive evaluation on diverse benchmarks demonstrates that EM-KD trained model outperforms prior Efficient MLLMs on both accuracy and efficiency with a large margin, validating its effectiveness. Compared with previous distillation methods, which are equipped with our proposed vision token matching strategy for fair comparison, EM-KD also achieves better performance.
Ze Feng, Boqiang Duan, Wankou Yang, Jingdong Wang 0001
AAAI5
2026 MagicGeo: Training-free text-guided geometric diagram generation
abstract
While text-to-image generation has made strides in photorealistic imagery, creating accurate geometric diagrams remains a challenge due to the need for precise spatial relationships and the scarcity of geometry-specific datasets. This paper presents MagicGeo, a training-free framework for generating geometric diagrams from textual descriptions. MagicGeo formulates the diagram generation process as a coordinate optimization problem, ensuring geometric correctness through a formal language solver, and then employs coordinate-aware generation. The framework leverages the strong language translation capability of large language models, while formal mathematical solving ensures geometric correctness. We further introduce MagicGeoBench, a benchmark dataset of 220 geometric diagram descriptions, and demonstrate that MagicGeo outperforms current methods in both qualitative and quantitative evaluations. This work provides a scalable, accurate solution for automated diagram generation, with significant implications for educational and academic applications.
Ting Zhang 0002, Qunyi Xie, Jingdong Wang 0001, Hua Huang 0001
Graph. Model.5
2026 Cosh-DiT: Co-Speech Gesture Video Synthesis via Hybrid Audio-Visual Diffusion Transformers
Yasheng Sun, Hang Zhou 0009, Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Borong Liang, Haocheng Feng, Jingdong Wang 0001, Ziwei Liu 0002, Hideki Koike
Int. J. Comput. Vis.10
2026 Improving Generalized Visual Grounding With Instance-Aware Joint Learning
abstract
Generalized visual grounding tasks, including Generalized Referring Expression Comprehension (GREC) and Segmentation (GRES), extend the classical visual grounding paradigm by accommodating multi-target and non-target scenarios. Specifically, GREC focuses on accurately identifying all referential objects at the coarse bounding box level, while GRES aims for achieve fine-grained pixel-level perception. However, existing approaches typically treat these tasks independently, overlooking the benefits of jointly training GREC and GRES to ensure consistent multi-granularity predictions and streamline the overall process. Moreover, current methods often treat GRES as a semantic segmentation task, neglecting the crucial role of instance-aware capabilities and the necessity of ensuring consistent predictions between instance-level boxes and masks. To address these limitations, we propose InstanceVG, a multi-task generalized visual grounding framework equipped with instance-aware capabilities, which leverages instance queries to unify the joint and consistency predictions of instance-level boxes and masks. To the best of our knowledge, InstanceVG is the first framework to simultaneously tackle both GREC and GRES while incorporating instance-aware capabilities into generalized visual grounding. To instantiate the framework, we assign each instance query a prior reference point, which also serves as an additional basis for target matching. This design facilitates consistent predictions of points, boxes, and masks for the same instance. Extensive experiments obtained on ten datasets across four tasks demonstrate that InstanceVG achieves state-of-the-art performance, significantly surpassing the existing methods in various evaluation metrics. The code and model will be publicly available at https://github.com/Dmmm1997/InstanceVG.
Wenxuan Cheng, Jiang-Jiang Liu 0001, Lingfeng Yang, Zhenhua Feng 0001, Wankou Yang, Jingdong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Toward Efficient Semi-Supervised Object Detection With Detection Transformer
abstract
Semi-supervised object detection (SSOD) mitigates the annotation burden in object detection by leveraging unlabeled data, providing a scalable solution for modern perception systems. Concurrently, detection transformers (DETRs) have emerged as a popular end-to-end framework, offering advantages such as non-maximum suppression (NMS)-free inference. However, existing SSOD methods are predominantly designed for conventional detectors, leaving the exploration of DETR-based SSOD largely uncharted. This paper presents a systematic study to bridge this gap. We begin by identifying two principal obstacles in semi-supervised DETR training: (1) the inherent one-to-one assignment mechanism of DETRs is highly sensitive to noisy pseudo-labels, which impedes training efficiency; and (2) the query-based decoder architecture complicates the design of an effective consistency regularization scheme, limiting further performance gains. To address these challenges, we propose Semi-DETR++, a novel framework for efficient SSOD with DETRs. Our approach introduces a stage-wise hybrid matching strategy that enhances robustness to noisy pseudo-labels by synergistically combining one-to-many and one-to-one assignments while preserving NMS-free inference. Furthermore, based on our observation of the unique layer-wise decoding behavior in DETRs, we develop a simple yet effective re-decode query consistency training method to regularize the decoder. Extensive experiments demonstrate that Semi-DETR++ enables more efficient semi-supervised learning across various DETR architectures, outperforming existing methods by significant margins. The proposed components are also flexible and versatile, showing superior generalization by readily extending to semi-supervised segmentation tasks.
Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Hongbo Gao 0001, Jingdong Wang 0001, Guanbin Li
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 I2V-Adapter: Fast adapting image pre-trained models for video correspondence
Hannan Lu, Xinyu Zhang 0015, Zhi Tian, Xiaohe Wu, Wangmeng Zuo, Jingdong Wang 0001
Pattern Recognit.6
2026 MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001
Pattern Recognit.7
2026 Corrigendum to "MA-FSAR: Multimodal Adaptation of CLIP for few-shot action recognition" [Pattern Recognition 169 (2026) 111902]
Jiazheng Xing, Jian Zhao 0006, Chao Xu 0023, Mengmeng Wang 0005, Guang Dai, Yong Liu 0007, Jingdong Wang 0001, Xuelong Li 0001
Pattern Recognit.7
2026 BUGS: Universal 3D Gaussian Splatting With a Bi-Directional Gaussian Growing Mechanism
abstract
3D Gaussian Splatting has the capability for realtime and high-quality scene reconstruction, bringing the efficiency and accuracy of this task to a new level. Most previous methods improve the vanilla 3DGS by incorporating external priors such as depth priors and shape priors. However, these methods often suffer from high memory cost, distorted geometric structures, and limitations in scene density. To solve these issues, this paper proposes a new 3DGS-based method for scene reconstruction, showing stable reconstruction and better geometry in both dense and sparse scenes. We re-model the growing of the Gaussians as a bi-directional process without any external priors, adaptively increasing and decreasing the density of the Gaussians across various scenes. For this purpose, we introduce a bi-directional Gaussian growing mechanism that determines the cloning, splitting and merging operations based on the region consistency within a local receptive field. We utilize these three operations to enable adaptive bi-directional control of the density of the Gaussian points across different regions, effectively improving the geometric structure and simplifying the Gaussian points. Extensive experiments on several open benchmarks have demonstrated that our method could universally achieve comparable results with much fewer Gaussian points and better geometry structures in both dense and sparse scenes with a concise framework. We will release our codes after this paper is accepted.
Fan Duan, Xiao Tan 0001, Jingdong Wang 0001, Li Chen 0031
IEEE Trans. Multim.5
2025 XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
abstract
Comprehensive testing of autonomous systems through simulation is essential to ensure the safety of autonomous driving vehicles. This requires the generation of safety-critical scenarios that extend beyond the limitations of real-world data collection, as many of these scenarios are rare or rarely encountered on public roads. However, evaluating most existing novel view synthesis (NVS) methods relies on sporadic sampling of image frames from the training data, comparing the rendered images with ground-truth images. Unfortunately, this evaluation protocol falls short of meeting the actual requirements in closed-loop simulations. Specifically, the true application demands the capability to render novel views that extend beyond the original trajectory (such as cross-lane views), which are challenging to capture in the real world. To address this, this paper presents a synthetic dataset for novel driving view synthesis evaluation, which is specifically designed for autonomous driving simulations. This unique dataset includes testing images captured by deviating from the training trajectory by 1–4 meters. It comprises six sequences that cover various times and weather conditions. Each sequence contains 450 training images, 120 testing images, and their corresponding camera poses and intrinsic parameters. Leveraging this novel dataset, we establish the first realistic benchmark for evaluating existing NVS approaches under frontonly and multicamera settings. The experimental findings underscore the significant gap in current approaches, revealing their inadequate ability to fulfill the demanding prerequisites of cross-lane or closed-loop simulation. Our dataset and code are released publicly on the project page: https://3d-aigc.github.io/XLD.
Hao Li 0075, Chenming Wu, Chen Zhao 0011, Chunyu Song, Haocheng Feng, Errui Ding, Dingwen Zhang, Jingdong Wang 0001
3DV10
2025 SpotActor: Training-Free Layout-Controlled Consistent Image Generation
abstract
Text-to-image diffusion models significantly enhance the efficiency of artistic creation with high-fidelity image generation. However, in typical application scenarios like comic book production, they can neither place each subject into its expected spot nor maintain the consistent appearance of each subject across images. For these issues, we pioneer a novel task, Layout-to-Consistent-Image (L2CI) generation, which produces consistent and compositional images in accordance with the given layout conditions and text prompts. To accomplish this challenging task, we present a new formalization of dual energy guidance with optimization in a dual semantic-latent space and thus propose a training-free pipeline, SpotActor, which features a layout-conditioned optimizing stage and a consistent sampling stage. In the optimizing stage, we innovate a nuanced layout energy function to mimic the attention activations with a sigmoid-like objective. While in the sampling stage, we design Regional Interconnection Self-Attention (RISA) and Semantic Fusion Cross-Attention (SFCA) mechanisms that allow mutual interactions across images. To evaluate the performance, we present ActorBench, a specified benchmark with hundreds of reasonable prompt-box pairs stemming from object detection datasets. Comprehensive experiments are conducted to demonstrate the effectiveness of our method. The results prove that SpotActor fulfills the expectations of this task and showcases the potential for practical applications with superior layout alignment, subject consistency, prompt conformity and background diversity.
Jiahao Wang 0004, Caixia Yan, Weizhan Zhang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Tieliang Gong, Hao Sun 0015, Jingdong Wang 0001
AAAI9
2025 Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models
abstract
Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They exhibit limited generalization in out-of-domain scenarios, such as new environments or unseen spoofing types. In this work, we introduce a multimodal large language model (MLLM) framework for FAS, termed Interpretable Face Anti-Spoofing (I-FAS), which transforms the FAS task into an interpretable visual question answering (VQA) paradigm. Specifically, we propose a Spoof-aware Captioning and Filtering (SCF) strategy to generate high-quality captions for FAS images, enriching the model's supervision with natural language interpretations. To mitigate the impact of noisy captions during training, we develop a Lopsided Language Model (L-LM) loss function that separates loss calculations for judgment and interpretation, prioritizing the optimization of the former. Furthermore, to enhance the model's perception of global visual features, we design a Globally Aware Connector (GAC) to align multi-level visual representations with the language model. Extensive experiments on standard and newly devised One to Eleven cross-domain benchmarks, comprising 12 public datasets, demonstrate that our method significantly outperforms state-of-the-art methods.
Keyao Wang, Haixiao Yue, Ajian Liu 0001, Errui Ding, Jingdong Wang 0001
AAAI8
2025 Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
abstract
Existing methodologies for animating portrait images face significant challenges, particularly in handling non-frontal perspectives, rendering dynamic objects around the portrait, and generating immersive, realistic backgrounds. In this paper, we introduce the first application of a pretrained transformer-based video generative model that demonstrates strong generalization capabilities and generates highly dynamic, realistic videos for portrait animation, effectively addressing these challenges. The adoption of a new video backbone model makes previous U-Net-Based methods for identity maintenance, audio conditioning, and video extrapolation inapplicable. To address this limitation, we design an identity reference network consisting of a causal 3D VAE combined with a stacked series of transformer layers, ensuring consistent facial identity across video sequences. Additionally, we investigate various speech audio conditioning and motion frame mechanisms to enable the generation of continuous video driven by speech audio. Our method is validated through experiments on benchmark and newly proposed wild datasets, demonstrating substantial improvements over prior methods in generating realistic portraits characterized by diverse orientations within dynamic and immersive scenes.
Jiahao Cui 0003, Yun Zhan, Hanlin Shang, Kaihui Cheng, Shan Mu, Hang Zhou 0009, Jingdong Wang 0001, Siyu Zhu 0001
CVPR9
2025 Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
abstract
Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g., objects) have not been well investigated. Despite human hand synthesis already being an intricate problem, generating objects in contact with hands and their interactions presents an even more challenging task, especially when the objects exhibit obvious variations in size and shape. To tackle these issues, we present a novel video reenactment framework focusing on Human-Object Interaction (HOI) via an adaptive Layout-instructed Diffusion model (Re-HOLD). Our key insight is to employ specialized layout representation for hands and objects, respectively. Such representations enable effective disentanglement of hand modeling and object adaptation to diverse motion sequences. To further improve the quality of the HOI generation, we design an interactive textural enhancement module for both hands and objects by introducing two independent memory banks. We also propose a layout adjustment strategy for the cross-object reenactment scenario to adaptively adjust unreasonable layouts caused by diverse object sizes during inference. Comprehensive qualitative and quantitative evaluations demonstrate that our proposed framework significantly outperforms existing methods. Project page: https://fyycs.github.io/Re-HOLD.
Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Haocheng Feng, Errui Ding, Yu Wu 0011, Jingdong Wang 0001
CVPR9
2025 AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
abstract
Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic human videos with both accurate lip-sync and delicate co-speech gestures w.r.t. given audio. In this work, we propose AudCast, a generalized audio-driven human video generation framework adopting a cascade Diffusion-Transformers (DiTs) paradigm, which synthesizes holistic human videos based on a reference image and a given audio. 1) Firstly, an audio-conditioned Holistic Human DiT architecture is proposed to directly drive the movements of any human body with vivid gesture dynamics. 2) Then to enhance hand and face details that are well-knownly difficult to handle, a Regional Refinement DiT leverages regional 3D fitting as the bridge to reform the signals, producing the final results. Extensive experiments demonstrate that our framework generates high-fidelity audio-driven holistic human videos with temporal coherence and fine facial and hand details. Resources can be found at https://guanjz20.github.io/projects/AudCast.
Jiazhi Guan, Kaisiyuan Wang, Quanwei Yang, Yasheng Sun, Shengyi He, Borong Liang, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Youjian Zhao, Hang Zhou 0009, Ziwei Liu 0002
CVPR12
2025 OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
abstract
Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality, human-centric video datasets presents a challenge to progress in this field. To bridge this gap, we introduce OpenHumanVid, a large-scale and high-quality humancentric video dataset characterized by precise and detailed captions that encompass both human appearance and motion states, along with supplementary human motion conditions, including skeleton sequences and speech audio. To validate the efficacy of this dataset and the associated training strategies, we propose an extension of existing classical diffusion transformer architectures and conduct further pretraining of our models on the proposed dataset. Our findings yield two critical insights: First, the incorporation of a large-scale, high-quality dataset substantially enhances evaluation metrics for generated human videos while preserving performance in general video generation tasks. Second, the effective alignment of text with human appearance, human motion, and facial motion is essential for producing high-quality video outputs. Based on these insights and corresponding methodologies, the straightforward extended network trained on the proposed dataset demonstrates an obvious improvement in the generation of human-centric videos. The source code and the dataset are available at: https://fudan-generative-vision.github.io/OpenHumanVid.
Mingwang Xu, Yun Zhan, Shan Mu, Kaihui Cheng, Jingdong Wang 0001, Siyu Zhu 0001
CVPR10
2025 TexGarment: Consistent Garment UV Texture Generation via Efficient 3D Structure-Guided Diffusion Transformer
abstract
This paper introduces TexGarment, an efficient method for synthesizing high-quality, 3D-consistent garment textures in UV space. Traditional approaches based on 2D-to-3D mapping often suffer from 3D inconsistency, while methods learning from limited 3D data lack sufficient texture diversity. These limitations are particularly problematic in garment texture generation, where high demands exist for both detail and variety. To address these challenges, TexGarment leverages a pre-trained text-to-image diffusion Transformer model with robust generalization capabilities, introducing structural information to guide the model in generating 3D-consistent garment textures in a single inference step. Specifically, We utilize the 2D UV position map to guide the layout during the UV texture generation process, ensuring a coherent texture arrangement and enhancing it by integrating global 3D structural information from the mesh surface point cloud. This combined guidance effectively aligns 3D structural integrity with 2D layout. Our method efficiently generates high-quality, diverse UV textures in a single inference step while maintaining 3D consistency. Experimental results validate the effectiveness of TexGarment, achieving state-of-the-art performance in 3D garment texture generation.
Jialun Liu, Xiaobo Gao, Bojun Xiong, Chen Zhao 0011, Hongbin Pei, Haocheng Feng, Errui Ding, Jingdong Wang 0001
CVPR12
2025 Action Detail Matters: Refining Video Recognition with Local Action Queries
abstract
Video action recognition involves interpreting both global context and specific details to accurately identify actions. While previous models are effective at capturing spatiotemporal features, they often lack a focused representation of key action details. To address this, we introduce FocusVideo, a framework designed for refining video action recognition through integrated global and local feature learning. Inspired by human visual cognition theory, our approach balances the focus on both broad contextual changes and action-specific details, minimizing the influence of irrelevant background noise. We first employ learnable action queries to selectively emphasize action-relevant regions without requiring region-specific labels. Next, these queries are learned by a local action streaming branch that enables progressive query propagation. Moreover, we introduce a parameter-free feature interaction mechanism for effective multi-scale interaction between global and local features with minimal additional overhead. Extensive experiments demonstrate that FocusVideo achieves state-of-the-art performance across multiple action recognition datasets, validating its effectiveness and robustness in handling action-relevant details.
Mengmeng Wang 0005, Zeyi Huang, Xiangjie Kong 0001, Guojiang Shen, Guang Dai, Jingdong Wang 0001, Yong Liu 0007
CVPR6
2025 Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?
abstract
The ultimate goal of generative models is to perfectly capture the data distribution. For image generation, common metrics of visual quality (e.g., FID) and the perceived truthfulness of generated images seem to suggest that we are nearing this goal. However, through distribution classification tasks, we reveal that, from the perspective of neural network-based classifiers, even advanced diffusion models are still far from this goal. Specifically, classifiers are able to consistently and effortlessly distinguish real images from generated ones across various settings. Moreover, we uncover an intriguing discrepancy: classifiers can easily differentiate between diffusion models with comparable performance (e.g., U-ViTH vs. DiT-XL), but struggle to distinguish between models within the same family but of different scales (e.g., EDM2-XS vs. EDM2-XXL). Our methodology carries several important implications. First, it naturally serves as a diagnostic tool for diffusion models by analyzing specific features of generated data. Second, it sheds light on the model autophagy disorder and offers insights into the use of generated data: augmenting real data with generated data is more effective than replacing it. Third, classifier guidance can significantly enhance the realism of generated images.
Zebin You, Xinyu Zhang 0017, Hanzhong Guo, Jingdong Wang 0001, Chongxuan Li
CVPR4
2025 Continual SFT Matches Multimodal RLHF with Negative Supervision
abstract
Multimodal RLHF usually happens after supervised fine-tuning (SFT) stage to continually improve vision-language models’ (VLMs) comprehension. Conventional wisdom holds its superiority over continual SFT during this preference alignment stage. In this paper, we observe that the inherent value of multimodal RLHF lies in its negative supervision, the logit of the rejected responses. We thus propose a novel negative supervised fine-tuning (nSFT) approach that fully excavates these information resided. Our nSFT disentangles this negative supervision in RLHF paradigm, and continually aligns VLMs with a simple SFT loss. This is more memory efficient than multimodal RLHF where 2 (e.g., DPO) or 4 (e.g., PPO) large VLMs are strictly required. The effectiveness of nSFT is rigorously proved by comparing it with various multimodal RLHF approaches, across different dataset sources, base VLMs and evaluation metrics. Besides, fruitful of ablations are provided to support our hypothesis. Code will be found in https://github.com/Kevinzcode/nSFT/.
Yanpeng Sun, Qiang Chen 0007, Jiangjiang Liu 0006, Jingdong Wang 0001
CVPR7
2025 VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
abstract
Recent advancements in camera-based occupancy prediction have focused on the simultaneous prediction of 3D semantics and scene flow, a task that presents significant challenges due to specific difficulties, e.g., occlusions and unbalanced dynamic environments. In this paper, we analyze these challenges and their underlying causes. To address them, we propose a novel regularization framework called VoxelSplat. This framework leverages recent developments in 3D Gaussian Splatting to enhance model performance in two key ways: (i) Enhanced Semantics Supervision through 2D Projection: During training, our method decodes sparse semantic 3D Gaussians from 3D representations and projects them onto the 2D camera view. This provides additional supervision signals in the camera-visible space, allowing 2D labels to improve the learning of 3D semantics. (ii) Scene Flow Learning: Our framework uses the predicted scene flow to model the motion of Gaussians, and is thus able to learn the scene flow of moving objects in a self-supervised manner using the labels of adjacent frames. Our method can be seamlessly integrated into various existing occupancy models, enhancing performance without increasing inference time. Extensive experiments on benchmark datasets demonstrate the effectiveness of Voxel-Splat in improving the accuracy of both semantic occupancy and scene flow estimation. The project page and codes are available at https://zzy816.github.io/VoxelSplat-Demo/.
Ziyue Zhu, Shenlong Wang, Jin Xie 0001, Jiang-jiang Liu, Jingdong Wang 0001, Jian Yang 0003
CVPR5
2025 Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
abstract
Recent advances in latent diffusion-based generative models for portrait image animation, such as Hallo, have achieved impressive results in short-duration video synthesis. In this paper, we present updates to Hallo, introducing several design enhancements to extend its capabilities.First, we extend the method to produce long-duration videos. To address substantial challenges such as appearance drift and temporal artifacts, we investigate augmentation strategies within the image space of conditional motion frames. Specifically, we introduce a patch-drop technique augmented with Gaussian noise to enhance visual consistency and temporal coherence over long duration.Second, we achieve 4K resolution portrait video generation. To accomplish this, we implement vector quantization of latent codes and apply temporal alignment techniques to maintain coherence across the temporal dimension. By integrating a high-quality decoder, we realize visual synthesis at 4K resolution.Third, we incorporate adjustable semantic textual labels for portrait expressions as conditional inputs. This extends beyond traditional audio cues to improve controllability and increase the diversity of the generated content. To the best of our knowledge, Hallo2, proposed in this paper, is the first method to achieve 4K resolution and generate hour-long, audio-driven portrait image animations enhanced with textual prompts. We have conducted extensive experiments to evaluate our method on publicly available datasets, including HDTF, CelebV, and our introduced ''Wild'' dataset. The experimental results demonstrate that our approach achieves state-of-the-art performance in long-duration portrait video animation, successfully generating rich and controllable content at 4K resolution for duration extending up to tens of minutes.
Jiahao Cui 0003, Yao Yao 0008, Hao Zhu 0004, Hanlin Shang, Kaihui Cheng, Hang Zhou 0009, Siyu Zhu 0001, Jingdong Wang 0001
ICLR9
2025 MGMapNet: Multi-Granularity Representation Learning for End-to-End Vectorized HD Map Construction
abstract
The construction of vectorized high-definition map typically requires capturing both category and geometry information of map elements. Current state-of-the-art methods often adopt solely either point-level or instance-level representation, overlooking the strong intrinsic relationship between points and instances. In this work, we propose a simple yet efficient framework named MGMapNet (multi-granularity map network) to model map elements with multi-granularity representation, integrating both coarse-grained instance-level and fine-grained point-level queries. Specifically, these two granularities of queries are generated from the multi-scale bird's eye view features using a proposed multi-granularity aggregator. In this module, instance-level query aggregates features over the entire scope covered by an instance, and the point-level query aggregates features locally. Furthermore, a point-instance interaction module is designed to encourage information exchange between instance-level and point-level queries. Experimental results demonstrate that the proposed MGMapNet achieves state-of-the-art performances, surpassing MapTRv2 by 5.3 mAP on the nuScenes dataset and 4.4 mAP on the Argoverse2 dataset, respectively.
Minyue Jiang, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Hanli Wang
ICLR7
2025 Manifold Constraint Reduces Exposure Bias in Accelerated Diffusion Sampling
abstract
Diffusion models have demonstrated significant potential for generating high-quality images, audio, and videos. However, their iterative inference process entails substantial computational costs, limiting practical applications. Recently, researchers have introduced accelerated sampling methods that enable diffusion models to generate samples with far fewer timesteps than those used during training. Nonetheless, as the number of sampling steps decreases, the prediction errors significantly degrade the quality of generated outputs. Additionally, the exposure bias in diffusion models further amplifies these errors. To address these challenges, we leverage a manifold hypothesis to explore the exposure bias problem in depth. Based on this geometric perspective, we propose a manifold constraint that effectively reduces exposure bias during accelerated sampling of diffusion models. Notably, our method involves no additional training and requires only minimal hyperparameter tuning. Extensive experiments demonstrate the effectiveness of our approach, achieving a FID score of 15.60 with 10-step SDXL on MS-COCO, surpassing the baseline by a reduction of 2.57 in FID.
Yuzhe Yao, Jun Chen 0023, Zeyi Huang, Haonan Lin, Mengmeng Wang 0005, Guang Dai, Jingdong Wang 0001
ICLR7
2025 DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-Making
abstract
Integrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extracting key spatiotemporal patterns from video that capture visual saliency and temporal evolution, referred to as dynamic representation. Building on this, we introduce DynaMind, a framework that enhances decision-making through dynamic reasoning. Specifically, we design an adaptive FrameScorer to evaluate video frames based on semantic consistency and visual saliency, assigning each frame an importance score. These scores are used to filter redundant video content and synthesize compact dynamic representations. Leveraging these representations, we predict critical future dynamics and apply a dynamic-guided policy to generate coherent and context-aware actions. Extensive results demonstrate that DynaMind significantly outperforms the baselines across several simulation benchmarks and real-world scenarios.
Ziru Wang, Mengmeng Wang 0005, Jade Dai, Teli Ma, Guo-Jun Qi, Yong Liu 0007, Guang Dai, Jingdong Wang 0001
ICML8
2025 DGTR: Distributed Gaussian Turbo-Reconstruction for Sparse-View Vast Scenes
abstract
Novel-view synthesis approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are limited. Additionally, few-shot methods often struggle with poor reconstruction quality in vast environments. This paper presents DGTR, a novel distributed framework for efficient Gaussian reconstruction for sparse-view vast scenes. Our approach divides the scene into regions, processed independently by drones with sparse image inputs. Using a feed-forward Gaussian model, we predict high-quality Gaussian primitives, followed by a global alignment algorithm to ensure geometric consistency. Depth priors is incorporated to further enhance training, while a distillation-based model aggregation mechanism enables efficient reconstruction. Our method achieves high-quality large-scale scene reconstruction and novel-view synthesis in significantly reduced training times, outperforming existing approaches in both speed and scalability. We demonstrate the effectiveness of our framework on vast aerial scenes, achieving high-quality results within minutes. Code will released on our project page https://3d-aigc.github.io/DGTR.
Hao Li 0075, Haosong Peng, Chenming Wu, Weicai Ye, Yufeng Zhan, Chen Zhao 0011, Dingwen Zhang, Jingdong Wang 0001, Junwei Han 0001
ICRA9
2025 Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving
abstract
The autoregressive world model exhibits robust generalization capabilities in vectorized scene understanding but encounters difficulties in deriving actions due to insufficient uncertainty modeling and self-delusion. In this paper, we explore the feasibility of deriving decisions from an autoregres-sive world model by addressing these challenges through the formulation of multiple probabilistic hypotheses. We propose LatentDriver, a framework models the environment's next states and the ego vehicle's possible actions as a mixture distribution, from which a deterministic control signal is then derived. By incorporating mixture modeling, the stochastic nature of decision-making is captured. Additionally, the self-delusion problem is mitigated by providing intermediate actions sampled from a distribution to the world model. Experimen-tal results on the recently released closed-loop benchmark Waymax demonstrate that LatentDriver surpasses state-of-the-art reinforcement learning and imitation learning methods, achieving expert-level performance. The code and models will be made available at https://github.com/Sephirex-X/LatentDriver.
Lingyu Xiao, Jiang-Jiang Liu 0001, Xiaoqing Ye, Wankou Yang, Jingdong Wang 0001
ICRA7
2025 VidEvo: Evolving Video Editing through Exhaustive Temporal Modeling
abstract
Text-guided video editing (TGVE) has become a recent hotspot due to its entertainment value and practical applications. To reduce overhead, existing methods primarily extend from text-to-image diffusion models and typically involve reconstruction and editing phases. However, challenges persist, particularly in enhancing temporal consistency of a video while adhering to textual alignment requirements. A crucial factor leading to the aforementioned issue is the inadequate and implicit tuning of the attention module within existing methods, which is specifically designed to capture temporal information. In light of this, we introduce VidEvo, a novel one-shot video editing method that leverages explicit cues derived from the original video to enhance temporal modeling. By integrating null-video embedding (NVE) and window-frame attention (WFA) components, VidEvo facilitates the smooth and coherent generation of videos from global and local perspectives simultaneously. To be specific, NVE learns a set of multi-scale temporal embeddings within the visual space during the reconstruction phase. These embeddings are subsequently directly injected into the attention module of the editing phase, explicitly augmenting the temporal consistency of the entire video. On the other hand, WFA enhances local temporal modeling by dynamically optimizing attention mechanisms between adjacent frames, which improves temporal coherence with reduced computational costs. Experimental evaluations show that VidEvo enhances frame-to-frame temporal consistency. Ablation studies confirm NVE and WFA’s effectiveness and their plug-and-play capability with other methods.
Sizhe Dang, Huan Liu 0001, Mengmeng Wang 0005, Guang Dai, Jingdong Wang 0001
IJCAI6
2025 TriCLIP-3D: A Unified Parameter-Efficient Framework for Tri-Modal 3D Visual Grounding based on CLIP
abstract
3D visual grounding allows an embodied agent to understand visual information in real-world 3D environments based on human instructions, which is crucial for embodied intelligence. Existing 3D visual grounding methods typically rely on separate encoders for different modalities (e.g., RGB images, text, and 3D point clouds), resulting in large and complex models that are inefficient to train. While some approaches use pre-trained 2D multi-modal models like CLIP for 3D tasks, they still struggle with aligning point cloud data to 2D encoders. As a result, these methods continue to depend on 3D encoders for feature extraction, further increasing model complexity and training inefficiency. In this paper, we propose a unified 2D pre-trained multi-modal network to process all three modalities (RGB images, text, and point clouds), significantly simplifying the architecture. By leveraging a 2D CLIP bi-modal model with adapter-based fine-tuning, this framework effectively adapts to the tri-modal setting, improving both adaptability and performance across modalities. Our Geometric-Aware 2D-3D Feature Recovery and Fusion (GARF) module is designed to fuse geometric multi-scale features from point clouds and images. We then integrate textual features for final modality fusion and introduce a multi-modal decoder to facilitate deep cross-modal understanding. Together, our method achieves unified feature extraction and fusion across the three modalities, enabling an end-to-end 3D visual grounding model. Compared to the baseline, our method reduces the number of trainable parameters by approximately 58%, while achieving a 6.52% improvement in the 3D detection task and a 6.25% improvement in the 3D visual grounding task.
Zanyi Wang, Zeyi Huang, Guang Dai, Jingdong Wang 0001, Mengmeng Wang 0005
ACM Multimedia5
2025 MonoLift: Learning 3D Manipulation Policies from Monocular RGB via Distillation
abstract
Although learning 3D manipulation policies from monocular RGB images is lightweight and deployment-friendly, the lack of structural information often leads to inaccurate action estimation. While explicit 3D inputs can mitigate this issue, they typically require additional sensors and introduce data acquisition overhead. An intuitive alternative is to incorporate a pre-trained depth estimator; however, this often incurs substantial inference-time cost. To address this, we propose MonoLift, a tri-level knowledge distillation framework that transfers spatial, temporal, and action-level knowledge from a depth-guided teacher to a monocular RGB student. By jointly distilling geometry-aware features, temporal dynamics, and policy behaviors during training, MonoLift enables the student model to perform 3D-aware reasoning and precise control at deployment using only monocular RGB input. Extensive experiments on both simulated and real-world manipulation tasks show that MonoLift not only outperforms existing monocular approaches but even surpasses several methods that rely on explicit 3D input, offering a resource-efficient and effective solution for vision-based robotic control. The video demonstration is available on our project page: https://robotasy.github.io/MonoLift/.
Ziru Wang, Mengmeng Wang 0005, Guang Dai, Yongliu Long, Jingdong Wang 0001
NeurIPS5
2025 High-Fidelity Dynamic Portrait Animation via Direct Preference Optimization and Temporal Motion Modulation
abstract
Generating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity body motion dynamics. We propose a human-preference-aligned diffusion framework that addresses these challenges through two key innovations. First, we introduce direct preference optimization tailored for human-centric animation, leveraging a curated dataset of human preferences to align generated outputs with perceptual metrics for portrait motion-video alignment and naturalness of expression. Second, the proposed temporal motion modulation resolves spatiotemporal resolution mismatches by reshaping motion conditions into dimensionally aligned latent features through temporal channel redistribution and proportional feature expansion, preserving the fidelity of high-frequency motion details in diffusion-based synthesis. The proposed mechanism is complementary to existing UNet and DiT-based portrait diffusion approaches, and experiments demonstrate obvious improvements in lip-audio synchronization, expression vividness, body motion coherence over baseline methods, alongside notable gains in human preference metrics. Code and data for this paper are at https://github.com/fudan-generative-vision/hallo4.
Jiahao Cui 0003, Baoyou Chen, Mingwang Xu, Hanlin Shang, Qinkun Su, Zilong Dong, Yao Yao 0008, Jingdong Wang 0001, Siyu Zhu 0001
SIGGRAPH Asia9
2025 Unified Frequency-Assisted Transformer Framework for Detecting and Grounding Multi-modal Manipulation
Huan Liu 0030, Zichang Tan, Qiang Chen 0007, Yunchao Wei, Yao Zhao 0001, Jingdong Wang 0001
Int. J. Comput. Vis.6
2025 Mutual Prompt Leaning for Vision Language Models
Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Jingyuan Feng, Sheng-Sheng Wang 0001, Jingdong Wang 0001
Int. J. Comput. Vis.8
2025 Real-Time Neural Radiance Talking Portrait Synthesis via Audio-Spatial Decomposition
Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou 0009, Xiaokang Chen, Dongliang He, Tianshu Hu, Jingtuo Liu, Ziwei Liu 0002, Jingdong Wang 0001
Int. J. Comput. Vis.10
2025 TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
Jiazheng Xing, Chao Xu 0023, Yijie Qian, Yang Liu 0356, Guang Dai, Baigui Sun, Yong Liu 0007, Jingdong Wang 0001
Int. J. Comput. Vis.8
2025 Fusion4DAL: Offline Multi-modal 3D Object Detection for 4D Auto-labeling
Xuekuan Wang, Wei Zhang 0197, Xiao Tan 0001, Jincheng Lu, Jingdong Wang 0001, Errui Ding, Cairong Zhao
Int. J. Comput. Vis.6
2025 Skim then Focus: Integrating Contextual and Fine-grained Views for Repetitive Action Counting
Zhengqi Zhao, Xiaohu Huang, Hao Zhou 0039, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
Int. J. Comput. Vis.6
2025 Cap4Video++: Enhancing Video Understanding With Auxiliary Captions
abstract
Understanding videos, especially aligning them with textual data, presents a significant challenge in computer vision. The advent of vision-language models (VLMs) like CLIP has sparked interest in leveraging their capabilities for enhanced video understanding, showing marked advancements in both performance and efficiency. However, current methods often neglect vital user-generated metadata such as video titles. In this paper, we present Cap4Video++, a universal framework that leverages auxiliary captions to enrich video understanding. More recently, we witness the flourishing of large language models (LLMs) like ChatGPT. Cap4Video++ harnesses the synergy of vision-language models (VLMs) and large language models (LLMs) to generate video captions, utilized in three key phases: (i) Input stage employs Semantic Pair Sampling to extract beneficial samples from captions, aiding contrastive learning. (ii) Intermediate stage sees Video-Caption Cross-modal Interaction and Adaptive Caption Selection work together to bolster video and caption representations. (iii) Output stage introduces a Complementary Caption-Text Matching branch, enhancing the primary video branch by improving similarity calculations. Our comprehensive experiments on text-video retrieval and video action recognition across nine benchmarks clearly demonstrate Cap4Video++'s superiority over existing models, highlighting its effectiveness in utilizing automatically generated captions to advance video understanding.
Jingdong Wang 0001, Yi Yang 0001, Wanli Ouyang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Unsupervised Pre-Training With Language-Vision Prompts for Low-Data Instance Segmentation
abstract
In recent times, following the paradigm of DETR (DEtection TRansformer), query-based end-to-end instance segmentation (QEIS) methods have exhibited superior performance compared to CNN-based models, particularly when trained on large-scale datasets. Nevertheless, the effectiveness of these QEIS methods diminishes significantly when confronted with limited training data. This limitation arises from their reliance on substantial data volumes to effectively train the pivotal queries/kernels that are essential for acquiring localization and shape priors. To address this problem, we propose a novel method for unsupervised pre-training in low-data regimes. Inspired by the recently successful prompting technique, we introduce a new method, Unsupervised Pre-training with Language-Vision Prompts (UPLVP), which improves QEIS models' instance segmentation by bringing language-vision prompts to queries/kernels. Our method consists of three parts: (1) Masks Proposal: Utilizes language-vision models to generate pseudo masks based on unlabeled images. (2) Prompt-Kernel Matching: Converts pseudo masks into prompts and injects the best-matched localization and shape features to their corresponding kernels. (3) Kernel Supervision: Formulates supervision for pre-training at the kernel level to ensure robust learning. With the help of our pre-training method, QEIS models can converge faster and perform better than CNN-based models in low-data regimes. Experimental evaluations conducted on MS COCO, Cityscapes, and CTW1500 datasets indicate that the QEIS models' performance can be significantly improved when pre-trained with our method.
Dingwen Zhang, Hao Li 0075, Diqi He, Nian Liu 0002, Lechao Cheng, Jingdong Wang 0001, Junwei Han 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Surfel-Based Gaussian Inverse Rendering for Fast and Relightable Dynamic Human Reconstruction From Monocular Videos
abstract
Efficient and accurate reconstruction of a relightable, dynamic clothed human avatar from a monocular video is crucial for the entertainment industry. This article presents SGIA (Surfel-based Gaussian Inverse Avatar), which introduces efficient training and rendering for relightable dynamic human reconstruction. SGIA advances previous Gaussian Avatar methods by comprehensively modeling Physically-Based Rendering (PBR) properties for clothed human avatars, allowing for the manipulation of avatars into novel poses under diverse lighting conditions. Specifically, our approach integrates pre-integration and image-based lighting for fast light calculations that surpass the performance of existing implicit-based techniques. To address challenges related to material lighting disentanglement and accurate geometry reconstruction, we propose an innovative occlusion approximation strategy and a progressive training approach. Extensive experiments demonstrate that SGIA not only achieves highly accurate physical properties but also significantly enhances the realistic relighting of dynamic human avatars, providing a substantial speed advantage.
Yiqun Zhao, Chenming Wu, Binbin Huang 0004, Yihao Zhi, Chen Zhao 0011, Jingdong Wang 0001, Shenghua Gao
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 PSDiff: Diffusion Model for Person Search With Iterative and Collaborative Refinement
abstract
Dominant Person Search methods aim to localize and recognize query persons in a unified network, which jointly optimizes the two sub-tasks of pedestrian detection and Re-Identification (ReID). Despite significant progress, current methods face two primary challenges: 1) the pedestrian candidates learned within detectors are suboptimal for the ReID task. 2) the potential for collaboration between two sub-tasks is overlooked. To address these issues, we present a novel Person Search framework based on the Diffusion model, PSDiff. PSDiff formulates the person search as a dual denoising process from noisy boxes and ReID embeddings to ground truths. Distinct from the conventional Detection-to-ReID approach, our denoising paradigm discards prior pedestrian candidates generated by detectors, thereby avoiding the local optimum problem of the ReID task. Following the new paradigm, we further design a new Collaborative Denoising Layer (CDL) to optimize detection and ReID sub-tasks in an iterative and collaborative way, which makes two sub-tasks mutually beneficial. Extensive experiments on the standard benchmarks show that PSDiff achieves state-of-the-art performance with fewer parameters and elastic computing overhead.
Chengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai, Xiaojun Chang, Jingdong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Disentangled Noisy Correspondence Learning
abstract
Cross-modal retrieval is crucial in understanding latent correspondences across modalities. However, existing methods implicitly assume well-matched training data, which is impractical as real-world data inevitably involves imperfect alignments, i.e., noisy correspondences. Although some works explore similarity-based strategies to address such noise, they suffer from sub-optimal similarity predictions influenced by modality-exclusive information (MEI), e.g., background noise in images and abstract definitions in texts. This issue arises as MEI is not shared across modalities, thus aligning it in training can markedly mislead similarity predictions. Moreover, although intuitive, directly applying previous cross-modal disentanglement methods suffers from limited noise tolerance and disentanglement efficacy. Inspired by the robustness of information bottlenecks against noise, we introduce DisNCL, a novel information-theoretic framework for feature Disentanglement in Noisy Correspondence Learning, to adaptively balance the extraction of modality-invariant information (MII) and MEI with certifiable optimal cross-modal disentanglement efficacy. DisNCL then enhances similarity predictions in modality-invariant subspace, thereby greatly boosting similarity-based alleviation strategy for noisy correspondences. Furthermore, DisNCL introduces soft matching targets to model noisy many-to-many relationships inherent in multi-modal inputs for noise-robust and accurate cross-modal alignment. Extensive experiments confirm DisNCL's efficacy by 2% average recall improvement. Mutual information estimation and visualization results show that DisNCL learns meaningful MII/MEI subspaces, validating our theoretical analyses.
Zhuohang Dang, Minnan Luo, Jihong Wang 0003, Chengyou Jia, Haochen Han, Herun Wan, Guang Dai, Xiaojun Chang, Jingdong Wang 0001
IEEE Trans. Image Process.9
2025 Exploring Effective Factors for Improving Visual In-Context Learning
abstract
The In-Context Learning (ICL) is to understand a new task via a few demonstrations (aka. prompt) and predict new inputs without tuning the models. While it has been widely studied in NLP, it is still a relatively new area of research in computer vision. To reveal the factors influencing the performance of visual in-context learning, this paper shows that Prompt Selection and Prompt Fusion are two major factors that have a direct impact on the inference performance of visual in-context learning. Prompt selection is the process of selecting the most suitable prompt for query image. This is crucial because high-quality prompts assist large-scale visual models in rapidly and accurately comprehending new tasks. Prompt fusion involves combining prompts and query images to activate knowledge within large-scale visual models. However, altering the prompt fusion method significantly impacts its performance on new tasks. Based on these findings, we propose a simple framework prompt-SelF to improve visual in-context learning. Specifically, we first use the pixel-level retrieval method to select a suitable prompt, and then use different prompt fusion methods to activate diverse knowledge stored in the large-scale vision model, and finally, ensemble the prediction results obtained from different prompt fusion methods to obtain the final prediction results. We conducted extensive experiments on single-object segmentation and detection tasks to demonstrate the effectiveness of prompt-SelF. Remarkably, prompt-SelF has outperformed OSLSM method-based meta-learning in 1-shot segmentation for the first time. This indicated the great potential of visual in-context learning. The source code and models will be available at https://github.com/syp2ysy/prompt-SelF.
Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Jingdong Wang 0001, Zechao Li
IEEE Trans. Image Process.4
2024 Noisy Correspondence Learning with Self-Reinforcing Errors Mitigation
abstract
Cross-modal retrieval relies on well-matched large-scale datasets that are laborious in practice. Recently, to alleviate expensive data collection, co-occurring pairs from the Internet are automatically harvested for training. However, it inevitably includes mismatched pairs, i.e., noisy correspondences, undermining supervision reliability and degrading performance. Current methods leverage deep neural networks' memorization effect to address noisy correspondences, which overconfidently focus on similarity-guided training with hard negatives and suffer from self-reinforcing errors. In light of above, we introduce a novel noisy correspondence learning framework, namely Self-Reinforcing Errors Mitigation (SREM). Specifically, by viewing sample matching as classification tasks within the batch, we generate classification logits for the given sample. Instead of a single similarity score, we refine sample filtration through energy uncertainty and estimate model's sensitivity of selected clean samples using swapped classification entropy, in view of the overall prediction distribution. Additionally, we propose cross-modal biased complementary learning to leverage negative matches overlooked in hard-negative training, further improving model optimization stability and curbing self-reinforcing errors. Extensive experiments on challenging benchmarks affirm the efficacy and efficiency of SREM.
Zhuohang Dang, Minnan Luo, Chengyou Jia, Guang Dai, Xiaojun Chang, Jingdong Wang 0001
AAAI6
2024 SSMG: Spatial-Semantic Map Guided Diffusion Model for Free-Form Layout-to-Image Generation
abstract
Despite significant progress in Text-to-Image (T2I) generative models, even lengthy and complex text descriptions still struggle to convey detailed controls. In contrast, Layout-to-Image (L2I) generation, aiming to generate realistic and complex scene images from user-specified layouts, has risen to prominence. However, existing methods transform layout information into tokens or RGB images for conditional control in the generative process, leading to insufficient spatial and semantic controllability of individual instances. To address these limitations, we propose a novel Spatial-Semantic Map Guided (SSMG) diffusion model that adopts the feature map, derived from the layout, as guidance. Owing to rich spatial and semantic information encapsulated in well-designed feature maps, SSMG achieves superior generation quality with sufficient spatial and semantic controllability compared to previous works. Additionally, we propose the Relation-Sensitive Attention (RSA) and Location-Sensitive Attention (LSA) mechanisms. The former aims to model the relationships among multiple objects within scenes while the latter is designed to heighten the model's sensitivity to the spatial information embedded in the guidance. Extensive experiments demonstrate that SSMG achieves highly promising results, setting a new state-of-the-art across a range of metrics encompassing fidelity, diversity, and controllability.
Chengyou Jia, Minnan Luo, Zhuohang Dang, Guang Dai, Xiaojun Chang, Mengmeng Wang 0005, Jingdong Wang 0001
AAAI7
2024 A Multimodal, Multi-Task Adapting Framework for Video Action Recognition
abstract
Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named M2-CLIP to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals, including the original contrastive learning head, a cross-modal classification head, a cross-modal masked language modeling head, and a visual classification head. This multi-task decoder adeptly satisfies the need for strong supervised performance within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios.
Mengmeng Wang 0005, Jiazheng Xing, Boyuan Jiang, Jun Chen 0023, Jianbiao Mei, Xingxing Zuo 0001, Guang Dai, Jingdong Wang 0001, Yong Liu 0007
AAAI8
2024 Multi-Domain Incremental Learning for Face Presentation Attack Detection
abstract
Previous face Presentation Attack Detection (PAD) methods aim to improve the effectiveness of cross-domain tasks. However, in real-world scenarios, the original training data of the pre-trained model is not available due to data privacy or other reasons. Under these constraints, general methods for fine-tuning single-target domain data may lose previously learned knowledge, leading to a catastrophic forgetting problem. To address these issues, we propose a multi-domain incremental learning (MDIL) method for PAD, which not only learns knowledge well from the new domain but also maintains the performance of previous domains stably. Specifically, we propose an adaptive domain-specific experts (ADE) framework based on the vision transformer to preserve the discriminability of previous domains. Furthermore, an asymmetric classifier is designed to keep the output distribution of different classifiers consistent, thereby improving the generalization ability. Extensive experiments show that our proposed method achieves state-of-the-art performance compared to prior methods of incremental learning. Excitingly, under more stringent setting conditions, our method approximates or even outperforms the DA/DG-based methods.
Keyao Wang, Haixiao Yue, Ajian Liu 0001, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001
AAAI9
2024 Learning to Rematch Mismatched Pairs for Robust Cross-Modal Retrieval
abstract
Collecting well-matched multimedia datasets is crucial for training cross-modal retrieval models. However, in real-world scenarios, massive multimodal data are harvested from the Internet, which inevitably contains Partially Mis-matched Pairs (PMPs). Undoubtedly, such semantical irrelevant data will remarkably harm the cross-modal retrieval performance. Previous efforts tend to mitigate this problem by estimating a soft correspondence to down-weight the contribution of PMPs. In this paper, we aim to address this challenge from a new perspective: the potential semantic similarity among unpaired samples makes it possible to excavate useful knowledge from mismatched pairs. To achieve this, we propose L2RM, a general framework based on Optimal Transport (OT) that learns to rematch mismatched pairs. In detail, L2RM aims to generate refined alignments by seeking a minimal-cost transport plan across different modalities. To formalize the rematching idea in OT, first, we propose a self-supervised cost function that automatically learns from explicit similarity-cost mapping relation. Second, we present to model a partial OT problem while restricting the transport among false positives to further boost refined alignments. Extensive experiments on three benchmarks demonstrate our L2RM significantly improves the robustness against PMPs for existing models. The code is available at https://github.com/hhc1997/L2RM.
Haochen Han, Guang Dai, Minnan Luo, Jingdong Wang 0001
CVPR5
2024 GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
abstract
Applying Neural Radiance Fields (NeRF) to downstream perception tasks for scene understanding and representation is becoming increasingly popular. Most existing methods treat semantic prediction as an additional rendering task, i.e., the “label rendering” task, to build semantic NeRFs. However, by rendering semantic/instance labels per pixel without considering the contextual information of the rendered image, these methods usually suffer from unclear boundary segmentation and abnormal segmentation of pixels within an object. To solve this problem, we propose Generalized Perception NeRF (GP-NeRF), a novel pipeline that makes the widely used segmentation model and NeRF work compatibly under a unified framework, for facilitating context-aware 3D scene perception. To accomplish this goal, we introduce transformers to aggregate radiance as well as semantic embedding fields jointly for novel views and facilitate the joint volumetric rendering of both fields. In addition, we propose two self-distillation mechanisms, i.e., the Semantic Distill Loss and the Depth-Guided Semantic Distill Loss, to enhance the discrimination and quality of the semantic field and the maintenance of geometric consistency. In evaluation, as shown in Fig. 1 we conduct experimental comparisons under two perception tasks (i.e. semantic and instance segmentation) using both synthetic and real-world datasets. Notably, our method outperforms SOTA approaches by 6.94%,11.76%, and 8.47% on generalized semantic segmentation, finetuning semantic segmentation, and instance segmentation, respectively. Project.
Hao Li 0075, Dingwen Zhang, Yalun Dai, Nian Liu 0002, Lechao Cheng, Jingfeng Li, Jingdong Wang 0001, Junwei Han 0001
CVPR7
2024 Forgery-aware Adaptive Transformer for Generalizable Synthetic Image Detection
abstract
In this paper, we study the problem of generalizable syn-thetic image detection, aiming to detect forgery images from diverse generative methods, e.g., GANs and diffusion mod-els. Cutting-edge solutions start to explore the benefits of pre-trained models, and mainly follow the fixed paradigm of solely training an attached classifier, e.g., combining frozen CLIP-ViT with a learnable linear layer in UniFD [43]. However, our analysis shows that such a fixed paradigm is prone to yield detectors with insufficient learning regarding forgery representations. We attribute the key challenge to the lack of forgery adaptation, and present a novel forgery-aware adaptive transformer approach, namely FatFormer. Based on the pre-trained vision-language spaces of CLIP, FatFormer introduces two core designs for the adaption to build generalized forgery representations. First, motivated by the fact that both image and frequency analysis are es-sential for synthetic image detection, we develop a forgery-aware adapter to adapt image features to discern and inte-grate local forgery traces within image and frequency do-mains. Second, we find that considering the contrastive ob-jectives between adapted image features and text prompt embeddings, a previously overlooked aspect, results in a nontrivial generalization improvement. Accordingly, we in-troduce language-guided alignment to supervise the forgery adaptation with image and text prompts in FatFormer. Ex-periments show that, by coupling these two designs, our approach tuned on 4-class ProGAN data attains a remarkable detection performance, achieving an average of 98% accu-racy to unseen GANs, and surprisingly generalizes to un-seen diffusion models with 95% accuracy.
Huan Liu 0030, Zichang Tan, Chuangchuang Tan, Yunchao Wei, Jingdong Wang 0001, Yao Zhao 0001
CVPR5
2024 VRP-SAM: SAM with Visual Reference Prompt
abstract
In this paper, we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Any-thing Model (SAM) to utilize annotated reference images as prompts for segmentation, creating the VRP-SAM model. In essence, VRP-SAM can utilize annotated reference images to comprehend specific objects and perform segmen-tation of specific objects in target image. It is note that the VRP encoder can support a variety of annotation for-mats for reference images, including point, box, scribble, and mask. VRP-SAM achieves a breakthrough within the SAM framework by extending its versatility and applicabil-ity while preserving SAM's inherent strengths, thus enhancing user-friendliness. To enhance the generalization abil-ity of VRP-SAM, the VRP encoder adopts a meta-learning strategy. To validate the effectiveness of VRP-SAM, we con-ducted extensive empirical studies on the Pascal and COCO datasets. Remarkably, VRP-SAM achieved state-of-the-art performance in visual reference segmentation with mini-mal learnable parameters. Furthermore, VRP-SAM demon-strates strong generalization capabilities, allowing it to per-form segmentation of unseen objects and enabling cross-domain segmentation. The source code and models will be available at https://github.com/syp2ysy/VRP-SAM
Yanpeng Sun, Shan Zhang 0002, Xinyu Zhang 0017, Qiang Chen 0007, Errui Ding, Jingdong Wang 0001, Zechao Li
CVPR8
2024 BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object Detection
abstract
Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it en-compasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping, ignoring the position approximation error in the voxel pooling process. Inspired by this insight, we propose a novel voxel pooling strategy to reduce such error, dubbed BEVSpread. Specifically, instead of bringing the image features contained in a frustum point to a single BEV grid, BEVSpread considers each frustum point as a source and spreads the image features to the surrounding BEV grids with adaptive weights. To achieve superior prop- agation performance, a specific weight function is designed to dynamically control the decay speed of the weights according to distance and depth. Aided by customized CUDA parallel acceleration, BEVSpread achieves comparable inference time as the original voxel pooling. Extensive experiments on two large-scale roadside benchmarks demonstrate that, as a plug-in, BEVSpread can significantly improve the performance of existing frustum-based BEV methods by a large margin of (1.12, 5.26, 3.01) AP in vehicle, pedestrian and cyclist. The source code will be made publicly available at BEVSpread.
Yehao Lu, Guangcong Zheng, Shuigen Zhan, Xiaoqing Ye, Zichang Tan, Jingdong Wang 0001, Gaoang Wang, Xi Li 0001
CVPR7
2024 Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection
abstract
We delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM30D) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy, leading to significant optimization conflicts with other re-liable forms of supervision. To tackle these issues, we introduce a novel decoupled pseudo-labeling (DPL) approach for SSM30D. Our approach features a Decoupled Pseudo-label Generation (DPG) module, designed to efficiently generate pseudo-labels by separately processing 2D and 3D attributes. This module incorporates a unique homography-based method for identifying dependable pseudo-labels in Bird's Eye View (BEV) space, specifically for 3D attributes. Additionally, we present a Depth Gradient Projection (DGP) module to mitigate optimization conflicts caused by noisy depth supervision of pseudo-labels, effectively decoupling the depth gradient and re-moving conflicting gradients. This dual decoupling strat-egy-at both the pseudo-label generation and gradient lev-els-significantly improves the utilization of pseudo-labels in SSM30D. Our comprehensive experiments on the KITTI benchmark demonstrate the superiority of our method over existing approaches.
Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li
CVPR8
2024 MS-DETR: Efficient DETR Training with Mixed Supervision
abstract
DETR accomplishes end-to-end object detection through iteratively generating multiple object candidates based on image features and promoting one candidate for each ground-truth object. The traditional training procedure using one-to-one supervision in the original DETR lacks di-rect supervision for the object detection candidates. We aim at improving the DETR training efficiency by explicitly supervising the candidate generation procedure through mixing one-to-one supervision and one-to-many su-pervision. Our approach, namely MS-DETR, is simple, and places one-to-many supervision to the object queries of the primary decoder that is used for inference. In comparison to existing DETR variants with one-to-many supervision, such as Group DETR and Hybrid DETR, our approach does not need additional decoder branches or object queries; the object queries of the primary decoder in our approach di-rectly benefit from one-to-many supervision and thus are superior in object candidate prediction. Experimental results show that our approach outperforms related DETR variants, such as DN-DETR, Hybrid DETR, and Group DETR, and the combination with related DETR variants further improves the performance. Code is available at: https://github.com/Atten4Vis/MS-DETR.
Chuyang Zhao, Yifan Sun 0003, Qiang Chen 0007, Errui Ding, Yi Yang 0001, Jingdong Wang 0001
CVPR7
2024 LaMI-DETR: Open-Vocabulary Detection with Language Model Instruction
Penghui Du, Yifan Sun 0003, Luting Wang 0001, Yue Liao, Errui Ding, Yan Wang 0059, Jingdong Wang 0001, Si Liu 0001
ECCV (23)9
2024 ReSyncer: Rewiring Style-Based Generator for Unified Audio-Visually Synced Facial Performer
Jiazhi Guan, Hang Zhou 0009, Kaisiyuan Wang, Shengyi He, Zhanwang Zhang, Borong Liang, Haocheng Feng, Errui Ding, Jingtuo Liu, Jingdong Wang 0001, Youjian Zhao, Ziwei Liu 0002
ECCV (41)11
2024 OPEN: Object-Wise Position Embedding for Multi-view 3D Object Detection
Jinghua Hou, Xiaoqing Ye, Zhe Liu 0033, Shi Gong, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai
ECCV (26)8
2024 GGRt: Towards Pose-Free Generalizable 3D Gaussian Splatting in Real-Time
Hao Li 0075, Chenming Wu, Dingwen Zhang, Yalun Dai, Chen Zhao 0011, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Junwei Han 0001
ECCV (71)9
2024 SEED: A Simple and Effective 3D DETR in Point Clouds
Zhe Liu 0033, Jinghua Hou, Xiaoqing Ye, Jingdong Wang 0001, Xiang Bai
ECCV (11)5
2024 3D-Aware Text-Driven Talking Avatar Generation
Xiuzhe Wu, Yang-Tian Sun, Handi Chen, Hang Zhou 0009, Jingdong Wang 0001, Zhengzhe Liu, Xiaojuan Qi 0001
ECCV (88)5
2024 Timestep-Aware Correction for Quantized Diffusion Models
Yuzhe Yao, Feng Tian 0002, Jun Chen 0023, Haonan Lin, Guang Dai, Yong Liu 0007, Jingdong Wang 0001
ECCV (66)7
2024 Make Your ViT-Based Multi-view 3D Detectors Faster via Token Compression
Dingyuan Zhang, Dingkang Liang, Zichang Tan, Xiaoqing Ye, Cheng Zhang 0020, Jingdong Wang 0001, Xiang Bai
ECCV (47)6
2024 Interactive 3D Object Detection with Prompts
Rui Zhang 0003, Xiangru Lin, Wei Zhang 0197, Jincheng Lu, Xuekuan Wang, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Guanbin Li
ECCV (17)9
2024 IRGen: Generative Modeling for Image Retrieval
Ting Zhang 0002, Dong Chen 0003, Yujing Wang 0002, Qi Chen 0009, Xing Xie 0001, Hao Sun 0015, Qi Zhang 0066, Fan Yang 0024, Mao Yang 0004, Qingmin Liao, Jingdong Wang 0001, Baining Guo
ECCV (15)13
2024 Towards Unified Multi-granularity Text Detection with Interactive Attention
abstract
Existing OCR engines or document image analysis systems typically rely on training separate models for text detection in varying scenarios and granularities, leading to significant computational complexity and resource demands. In this paper, we introduce "Detect Any Text" (DAT), an advanced paradigm that seamlessly unifies scene text detection, layout analysis, and document page detection into a cohesive, end-to-end model. This design enables DAT to efficiently manage text instances at different granularities, including word, line, paragraph and page. A pivotal innovation in DAT is the across-granularity interactive attention module, which significantly enhances the representation learning of text instances at varying granularities by correlating structural information across different text queries. As a result, it enables the model to achieve mutually beneficial detection performances across multiple text granularities. Additionally, a prompt-based segmentation module refines detection outcomes for texts of arbitrary curvature and complex layouts, thereby improving DAT’s accuracy and expanding its real-world applicability. Experimental results demonstrate that DAT achieves state-of-the-art performances across a variety of text-related benchmarks, including multi-oriented/arbitrarily-shaped scene text detection, document layout analysis and page detection tasks.
Xingyu Wan, Chengquan Zhang, Pengyuan Lv, Sen Fan, Zihan Ni, Errui Ding, Jingdong Wang 0001
ICML8
2024 Mobile Attention: Mobile-Friendly Linear-Attention for Vision Transformers
abstract
Vision Transformers (ViTs) excel in computer vision tasks due to their ability to capture global context among tokens. However, their quadratic complexity $\mathcal{O}(N^2D)$ in terms of token number $N$ and feature dimension $D$ limits practical use on mobile devices, necessitating more mobile-friendly ViTs with reduced latency. Multi-head linear-attention is emerging as a promising alternative with linear complexity $\mathcal{O}(NDd)$, where $d$ is the per-head dimension. Still, more compute is needed as $d$ gets large for model accuracy. Reducing $d$ improves mobile friendliness at the expense of excessive small heads weak at learning valuable subspaces, ultimately impeding model capability. To overcome this efficiency-capability dilemma, we propose a novel Mobile-Attention design with a head-competition mechanism empowered by information flow, which prevents overemphasis on less important subspaces upon trivial heads while preserving essential subspaces to ensure Transformer's capability. It enables linear-time complexity on mobile devices by supporting a small per-head dimension $d$ for mobile efficiency. By replacing the standard attention of ViTs with Mobile-Attention, our optimized ViTs achieved enhanced model capacity and competitive performance in a range of computer vision tasks. Specifically, we have achieved remarkable reductions in latency on the iPhone 12. Code is available at https://github.com/thuml/MobileAttention.
Zhiyu Yao, Jian Wang 0066, Haixu Wu, Jingdong Wang 0001, Mingsheng Long
ICML4
2024 Uni4DAL: A Unified Baseline for Multi-dataset 4D Auto-Labeling
Xuekuan Wang, Wei Zhang 0197, Xiao Tan 0001, Jinchen Lu, Jingdong Wang 0001, Errui Ding, Cairong Zhao
ICPR (30)6
2024 Generating Action-conditioned Prompts for Open-vocabulary Video Action Recognition
Chengyou Jia, Minnan Luo, Xiaojun Chang, Zhuohang Dang, Mingfei Han 0002, Mengmeng Wang 0005, Guang Dai, Sizhe Dang, Jingdong Wang 0001
ACM Multimedia9
2024 LION: Linear Group RNN for 3D Object Detection in Point Clouds
abstract
The benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward this goal, we propose a simple and effective window-based framework built on Linear group RNN (i.e., perform linear RNN for grouped features) for accurate 3D object detection, called LION. The key property is to allow sufficient feature interaction in a much larger group than transformer-based methods. However, effectively applying linear group RNN to 3D object detection in highly sparse point clouds is not trivial due to its limitation in handling spatial modeling. To tackle this problem, we simply introduce a 3D spatial feature descriptor and integrate it into the linear group RNN operators to enhance their spatial features rather than blindly increasing the number of scanning orders for voxel features. To further address the challenge in highly sparse point clouds, we propose a 3D voxel generation strategy to densify foreground features thanks to linear group RNN as a natural property of auto-regressive models. Extensive experiments verify the effectiveness of the proposed components and the generalization of our LION on different linear group RNN operators including Mamba, RWKV, and RetNet. Furthermore, it is worth mentioning that our LION-Mamba achieves state-of-the-art on Waymo, nuScenes, Argoverse V2, and ONCE datasets. Last but not least, our method supports kinds of advanced linear RNN operators (e.g., RetNet, RWKV, Mamba, xLSTM and TTT) on small but popular KITTI dataset for a quick experience with our linear RNN-based framework.
Zhe Liu 0033, Jinghua Hou, Xinyu Wang 0024, Xiaoqing Ye, Jingdong Wang 0001, Hengshuang Zhao, Xiang Bai
NeurIPS5
2024 Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
abstract
Comprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is an essential dimension measuring the visual vividness and the honesty of video content to text prompts. In this study, we propose an effective evaluation protocol, termed DEVIL, which centers on the dynamics dimension to evaluate T2V generation models, as well as improving existing evaluation metrics. In practice, we define a set of dynamics scores corresponding to multiple temporal granularities, and a new benchmark of text prompts under multiple dynamics grades. Upon the text prompt benchmark, we assess the generation capacity of T2V models, characterized by metrics of dynamics ranges and T2V alignment. Moreover, we analyze the relevance of existing metrics to dynamics metrics, improving them from the perspective of dynamics. Experiments show that DEVIL evaluation metrics enjoy up to about 90\% consistency with human ratings, demonstrating the potential to advance T2V generation models.
Mingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo, Fang Wan 0001, Tianyu Wang 0028, Yuzhong Zhao, Jingdong Wang 0001, Xinyu Zhang 0017
NeurIPS8
2024 Schedule Your Edit: A Simple yet Effective Diffusion Noise Schedule for Image Editing
abstract
Text-guided diffusion models have significantly advanced image editing, enabling high-quality and diverse modifications driven by text prompts. However, effective editing requires inverting the source image into a latent space, a process often hindered by prediction errors inherent in DDIM inversion. These errors accumulate during the diffusion process, resulting in inferior content preservation and edit fidelity, especially with conditional inputs. We address these challenges by investigating the primary contributors to error accumulation in DDIM inversion and identify the singularity problem in traditional noise schedules as a key issue. To resolve this, we introduce the *Logistic Schedule*, a novel noise schedule designed to eliminate singularities, improve inversion stability, and provide a better noise space for image editing. This schedule reduces noise prediction errors, enabling more faithful editing that preserves the original content of the source image. Our approach requires no additional retraining and is compatible with various existing editing methods. Experiments across eight editing tasks demonstrate the Logistic Schedule's superior performance in content preservation and edit fidelity compared to traditional noise schedules, highlighting its adaptability and effectiveness. The project page is available at https://lonelvino.github.io/SYE/.
Haonan Lin, Yan Chen 0031, Jiahao Wang 0004, Wenbin An, Mengmeng Wang 0005, Feng Tian 0002, Yong Liu 0007, Guang Dai, Jingdong Wang 0001, Qianying Wang 0002
NeurIPS9
2024 Flipped Classroom: Aligning Teacher Attention with Student in Generalized Category Discovery
abstract
Recent advancements have shown promise in applying traditional Semi-Supervised Learning strategies to the task of Generalized Category Discovery (GCD). Typically, this involves a teacher-student framework in which the teacher imparts knowledge to the student to classify categories, even in the absence of explicit labels. Nevertheless, GCD presents unique challenges, particularly the absence of priors for new classes, which can lead to the teacher's misguidance and unsynchronized learning with the student, culminating in suboptimal outcomes. In our work, we delve into why traditional teacher-student designs falter in generalized category discovery as compared to their success in closed-world semi-supervised learning. We identify inconsistent pattern learning as the crux of this issue and introduce FlipClass—a method that dynamically updates the teacher to align with the student's attention, instead of maintaining a static teacher reference. Our teacher-attention-update strategy refines the teacher's focus based on student feedback, promoting consistent pattern recognition and synchronized learning across old and new classes. Extensive experiments on a spectrum of benchmarks affirm that FlipClass significantly surpasses contemporary GCD methods, establishing new standards for the field.
Haonan Lin, Wenbin An, Jiahao Wang 0004, Yan Chen 0031, Feng Tian 0002, Mengmeng Wang 0005, Qianying Wang 0002, Guang Dai, Jingdong Wang 0001
NeurIPS9
2024 OpenGaussian: Towards Point-Level 3D Gaussian-based Open Vocabulary Understanding
abstract
This paper introduces OpenGaussian, a method based on 3D Gaussian Splatting (3DGS) that possesses the capability for 3D point-level open vocabulary understanding. Our primary motivation stems from observing that existing 3DGS-based open vocabulary methods mainly focus on 2D pixel-level parsing. These methods struggle with 3D point-level tasks due to weak feature expressiveness and inaccurate 2D-3D feature associations. To ensure robust feature presentation and 3D point-level understanding, we first employ SAM masks without cross-frame associations to train instance features with 3D consistency. These features exhibit both intra-object consistency and inter-object distinction. Then, we propose a two-stage codebook to discretize these features from coarse to fine levels. At the coarse level, we consider the positional information of 3D points to achieve location-based clustering, which is then refined at the fine level. Finally, we introduce an instance-level 3D-2D feature association method that links 3D points to 2D masks, which are further associated with 2D CLIP features. Extensive experiments, including open vocabulary-based 3D object selection, 3D point cloud understanding, click-based 3D object selection, and ablation studies, demonstrate the effectiveness of our proposed method. The source code is available at our project page https://3d-aigc.github.io/OpenGaussian.
Yanmin Wu, Jiarui Meng, Haijie Li, Chenming Wu, Yahao Shi, Xinhua Cheng, Chen Zhao 0011, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Jian Zhang 0018
NeurIPS10
2024 ShowMaker: Creating High-Fidelity 2D Human Video via Fine-Grained Diffusion Modeling
abstract
Although significant progress has been made in human video generation, most previous studies focus on either human facial animation or full-body animation, which cannot be directly applied to produce realistic conversational human videos with frequent hand gestures and various facial movements simultaneously. To address these limitations, we propose a 2D human video generation framework, named ShowMaker, capable of generating high-fidelity half-body conversational videos via fine-grained diffusion modeling. We leverage dual-stream diffusion models as the backbone of our framework and carefully design two novel components for crucial local regions (i.e., hands and face) that can be easily integrated into our backbone. Specifically, to handle the challenging hand generation caused by sparse motion guidance, we propose a novel Key Point-based Fine-grained Hand Modeling module by amplifying positional information from raw hand key points and constructing a corresponding key point-based codebook. Moreover, to restore richer facial details in generated results, we introduce a Face Recapture module, which extracts facial texture features and global identity features from the aligned human face and integrates them into the diffusion process for face enhancement. Extensive quantitative and qualitative experiments demonstrate the superior visual quality and temporal consistency of our method.
Quanwei Yang, Jiazhi Guan, Kaisiyuan Wang, Lingyun Yu 0002, Wenqing Chu, Hang Zhou 0009, ZhiQiang Feng, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001
NeurIPS10
2024 Dense Connector for MLLMs
abstract
*Do we fully leverage the potential of visual encoder in Multimodal Large Language Models (MLLMs)?* The recent outstanding performance of MLLMs in multimodal understanding has garnered broad attention from both academia and industry. In the current MLLM rat race, the focus seems to be predominantly on the linguistic side. We witness the rise of larger and higher-quality instruction datasets, as well as the involvement of larger-sized LLMs. Yet, scant attention has been directed towards the visual signals utilized by MLLMs, often assumed to be the final high-level features extracted by a frozen visual encoder. In this paper, we introduce the **Dense Connector** - a simple, effective, and plug-and-play vision-language connector that significantly enhances existing MLLMs by leveraging multi-layer visual features, with minimal additional computational overhead. Building on this, we also propose the Efficient Dense Connector, which achieves performance comparable to LLaVA-v1.5 with only 25% of the visual tokens. Furthermore, our model, trained solely on images, showcases remarkable zero-shot capabilities in video understanding as well. Experimental results across various vision encoders, image resolutions, training dataset scales, varying sizes of LLMs (2.7B→70B), and diverse architectures of MLLMs (e.g., LLaVA-v1.5, LLaVA-NeXT and Mini-Gemini) validate the versatility and scalability of our approach, achieving state-of-the-art performance across 19 image and video benchmarks. We hope that this work will provide valuable experience and serve as a basic module for future MLLM development. Code is available at https://github.com/HJYao00/DenseConnector.
Huanjin Yao, Taojiannan Yang, YuXin Song 0001, Haocheng Feng, Yifan Sun 0003, Wanli Ouyang, Jingdong Wang 0001
NeurIPS10
2024 Automated Multi-level Preference for MLLMs
abstract
Current multimodal Large Language Models (MLLMs) suffer from ''hallucination'', occasionally generating responses that are not grounded in the input images. To tackle this challenge, one promising path is to utilize reinforcement learning from human feedback (RLHF), which steers MLLMs towards learning superior responses while avoiding inferior ones. We rethink the common practice of using binary preferences (*i.e.*, superior, inferior), and find that adopting multi-level preferences (*e.g.*, superior, medium, inferior) is better for two benefits: 1) It narrows the gap between adjacent levels, thereby encouraging MLLMs to discern subtle differences. 2) It further integrates cross-level comparisons (beyond adjacent-level comparisons), thus providing a broader range of comparisons with hallucination examples. To verify our viewpoint, we present the Automated Multi-level Preference (**AMP**) framework for MLLMs. To facilitate this framework, we first develop an automated dataset generation pipeline that provides high-quality multi-level preference datasets without any human annotators. Furthermore, we design the Multi-level Direct Preference Optimization (MDPO) algorithm to robustly conduct complex multi-level preference learning. Additionally, we propose a new hallucination benchmark, MRHal-Bench. Extensive experiments across public hallucination and general benchmarks, as well as our MRHal-Bench, demonstrate the effectiveness of our proposed method. Code is available at https://github.com/takomc/amp.
YuXin Song 0001, Kang Rong, Huanjin Yao, Fanglong Liu, Haocheng Feng, Jingdong Wang 0001, Yifan Sun 0003
NeurIPS10
2024 Octopus: A Multi-modal LLM with Parallel Recognition and Sequential Understanding
abstract
A mainstream of Multi-modal Large Language Models (MLLMs) have two essential functions, i.e., visual recognition (e.g., grounding) and understanding (e.g., visual question answering). Presently, all these MLLMs integrate visual recognition and understanding in a same sequential manner in the LLM head, i.e., generating the response token-by-token for both recognition and understanding. We think unifying them in the same sequential manner is not optimal for two reasons: 1) parallel recognition is more efficient than sequential recognition and is actually prevailing in deep visual recognition, and 2) the recognition results can be integrated to help high-level cognition (while the current manner does not). Such motivated, this paper proposes a novel “parallel recognition → sequential understanding” framework for MLLMs. The bottom LLM layers are utilized for parallel recognition and the recognition results are relayed into the top LLM layers for sequential understanding. Specifically, parallel recognition in the bottom LLM layers is implemented via object queries, a popular mechanism in DEtection TRansformer, which we find to harmonize well with the LLM layers. Empirical studies show our MLLM named Octopus improves accuracy on popular MLLM tasks and is up to 5× faster on visual grounding tasks.
Chuyang Zhao, YuXin Song 0001, Kang Rong, Haocheng Feng, Shufan Ji, Jingdong Wang 0001, Errui Ding, Yifan Sun 0003
NeurIPS8
2024 MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
abstract
Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, the results often fall short of naturalness due to insufficient training priors. We alleviate the issue in this work from two perspectives. 1) From the data aspect, we carefully collect a human-centric dataset comprising over one million high-quality human-in-the-scene images and two specific sets of close-up images of faces and hands. These datasets collectively provide a rich prior knowledge base to enhance the human-centric image generation capabilities of the diffusion model. 2) On the methodological front, we propose a simple yet effective method called Mixture of Low-rank Experts (MoLE) by considering low-rank modules trained on close-up hand and face images respectively as experts. This concept draws inspiration from our observation of low-rank refinement, where a low-rank module trained by a customized close-up dataset has the potential to enhance the corresponding image part when applied at an appropriate scale. To validate the superiority of MoLE in the context of human-centric image generation compared to state-of-the-art, we construct two benchmarks and perform evaluations with diverse metrics and human studies. Datasets, model, and code are released at https://sites.google.com/view/mole4diffuser/.
Yixiong Chen, Mingyu Ding, Ping Luo 0002, Leye Wang, Jingdong Wang 0001
NeurIPS6
2024 PLIP: Language-Image Pre-training for Person Representation Learning
abstract
Language-image pre-training is an effective technique for learning powerful representations in general domains. However, when directly turning to person representation learning, these general pre-training methods suffer from unsatisfactory performance. The reason is that they neglect critical person-related characteristics, i.e., fine-grained attributes and identities. To address this issue, we propose a novel language-image pre-training framework for person representation learning, termed PLIP. Specifically, we elaborately design three pretext tasks: 1) Text-guided Image Colorization, aims to establish the correspondence between the person-related image regions and the fine-grained color-part textual phrases. 2) Image-guided Attributes Prediction, aims to mine fine-grained attribute information of the person body in the image; and 3) Identity-based Vision-Language Contrast, aims to correlate the cross-modal representations at the identity level rather than the instance level. Moreover, to implement our pre-train framework, we construct a large-scale person dataset with image-text pairs named SYNTH-PEDES by automatically generating textual annotations. We pre-train PLIP on SYNTH-PEDES and evaluate our models by spanning downstream person-centric tasks. PLIP not only significantly improves existing methods on all these tasks, but also shows great ability in the zero-shot and domain generalization settings. The code, dataset and weight will be made publicly available.
Jialong Zuo, Jiahao Hong, Feng Zhang 0039, Changqian Yu, Hanyu Zhou, Changxin Gao, Nong Sang, Jingdong Wang 0001
NeurIPS8
2024 TALK-Act: Enhance Textural-Awareness for 2D Speaking Avatar Reenactment with Diffusion Model
Jiazhi Guan, Quanwei Yang, Kaisiyuan Wang, Hang Zhou 0009, Shengyi He, Haocheng Feng, Errui Ding, Jingdong Wang 0001, Hongtao Xie 0001, Youjian Zhao, Ziwei Liu 0002
SIGGRAPH Asia9
2024 Context Autoencoder for Self-supervised Representation Learning
Xiaokang Chen, Mingyu Ding, Shentong Mo, Shumin Han, Ping Luo 0002, Jingdong Wang 0001
Int. J. Comput. Vis.10
2024 CSDG-FAS: Closed-Space Domain Generalization for Face Anti-spoofing
Keyao Wang, Haixiao Yue, Yanyan Liang 0001, Mouxiao Huang, Junyu Han, Errui Ding, Jingdong Wang 0001
Int. J. Comput. Vis.9
2024 Transferring Vision-Language Models for Visual Recognition: A Classifier Perspective
abstract
Abstract Transferring knowledge from pre-trained deep models for downstream tasks, particularly with limited labeled samples, is a fundamental problem in computer vision research. Recent advances in large-scale, task-agnostic vision-language pre-trained models, which are learned with billions of samples, have shed new light on this problem. In this study, we investigate how to efficiently transfer aligned visual and textual knowledge for downstream visual recognition tasks. We first revisit the role of the linear classifier in the vanilla transfer learning framework, and then propose a new paradigm where the parameters of the classifier are initialized with semantic targets from the textual encoder and remain fixed during optimization. To provide a comparison, we also initialize the classifier with knowledge from various resources. In the empirical study, we demonstrate that our paradigm improves the performance and training speed of transfer learning tasks. With only minor modifications, our approach proves effective across 17 visual datasets that span three different data domains: image, video, and 3D point cloud.
Zhun Sun, YuXin Song 0001, Jingdong Wang 0001, Wanli Ouyang
Int. J. Comput. Vis.4
2024 A Survey of Multimodal Controllable Diffusion Models
Guangcong Zheng, Tian-Rui Yang, Jingdong Wang 0001, Xi Li 0001
J. Comput. Sci. Technol.5
2024 Recent advances in artificial intelligence generated content
abstract
人工智能生成内容(AIGC)是近年来人工智能(AI)领域一个研究热点,它有望取代人类以较低成本高效率执行内容生成工作,如音乐、绘画、多模态内容生成、新闻文章、总结报告、股评摘要,以至元宇宙中的内容生成和数字人。AIGC为未来AI发展和实现提供了一条新的技术路径。 在此背景下,《信息与电子工程前沿(英文)》期刊组织了一期关于AIGC最新进展的特刊。本期特刊关注AIGC理论、算法、应用及相关领域。通过吸引高质量论文,我们希望帮助学术界和工业界研究人员更深入了解AIGC背后的基本理论及其潜在应用,激励更多研究人员加入并推进AIGC领域的研究。因此,我们就以下主题(但不限于)征集论文:(1)AI生成音乐;(2)AI生成绘画;(3)AI对话模型;(4)AI新闻摘要;(5)AI与元宇宙;(6)AI与数字人;(7)AI图像编辑;(8)AI生成短视频;(9)AI生成多媒体内容;(10)ChatGPT相关工作。经严格评审,选出12篇论文,包括1篇评论、1篇观点、3篇综述、6篇研究和1篇通讯。我们将其划分为3个主要部分:ChatGPT、扩散模型、提示学习和多模态。 总体而言,本期特刊涵盖了与AIGC开发和应用相关的广泛研究主题,包括人工智能图像/文本生成、三维内容创建、以用户为中心的图形设计、特定风格的音乐生成,以及与因果表征学习、高阶扩散模型相关的工作。此外,还详细调研了概率扩散模型、提示学习和ChatGPT。 最后,感谢所有作者对本期特刊的支持,特别感谢所有评审人对专刊投稿富有见地的意见和有益建议。
Junping Zhang, Lingyun Sun, Junbin Gao, Jiebo Luo 0001, Jingdong Wang 0001
Frontiers Inf. Technol. Electron. Eng.9
2024 Towards Lightweight Super-Resolution With Dual Regression Learning
abstract
Deep neural networks have exhibited remarkable performance in image super-resolution (SR) tasks by learning a mapping from low-resolution (LR) images to high-resolution (HR) images. However, the SR problem is typically an ill-posed problem and existing methods would come with several limitations. First, the possible mapping space of SR can be extremely large since there may exist many different HR images that can be super-resolved from the same LR image. As a result, it is hard to directly learn a promising SR mapping from such a large space. Second, it is often inevitable to develop very large models with extremely high computational cost to yield promising SR performance. In practice, one can use model compression techniques to obtain compact models by reducing model redundancy. Nevertheless, it is hard for existing model compression methods to accurately identify the redundant components due to the extremely large SR mapping space. To alleviate the first challenge, we propose a dual regression learning scheme to reduce the space of possible SR mappings. Specifically, in addition to the mapping from LR to HR images, we learn an additional dual regression mapping to estimate the downsampling kernel and reconstruct LR images. In this way, the dual mapping acts as a constraint to reduce the space of possible mappings. To address the second challenge, we propose a dual regression compression (DRC) method to reduce model redundancy in both layer-level and channel-level based on channel pruning. Specifically, we first develop a channel number search method that minimizes the dual regression loss to determine the redundancy of each layer. Given the searched channel numbers, we further exploit the dual regression manner to evaluate the importance of channels and prune the redundant ones. Extensive experiments show the effectiveness of our method in obtaining accurate and efficient SR models.
Mingkui Tan, Zeshuai Deng, Jingdong Wang 0001, Qi Chen 0014, Jiezhang Cao, Yanwu Xu 0001, Jian Chen 0011
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Disentangled Representation Learning With Transmitted Information Bottleneck
abstract
Encoding only the task-related information from the raw data, i.e., disentangled representation learning, can greatly contribute to the robustness and generalizability of models. Although significant advances have been made by regularizing the information in representations with information theory, two major challenges remain: 1) the representation compression inevitably leads to performance drop; 2) the disentanglement constraints on representations are in complicated optimization. To these issues, we introduce Bayesian networks with transmitted information to formulate the interaction among input and representations during disentanglement. Building upon this framework, we propose DisTIB (Transmitted Information Bottleneck for Disentangled representation learning), a novel objective that navigates the balance between information compression and preservation. We employ variational inference to derive a tractable estimation for DisTIB. This estimation can be simply optimized via standard gradient descent with a reparameterization trick. Moreover, we theoretically prove that DisTIB can achieve optimal disentanglement, underscoring its superior efficacy. To solidify our claims, we conduct extensive experiments on various downstream tasks to demonstrate the appealing efficacy of DisTIB and validate our theoretical analyses.
Zhuohang Dang, Minnan Luo, Chengyou Jia, Guang Dai, Jihong Wang 0003, Xiaojun Chang, Jingdong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Multi-Modal 3D Object Detection by Box Matching
abstract
Multi-modal 3D object detection has received growing attention as the information from different sensors like LiDAR and cameras are complementary. Most fusion methods for 3D detection rely on an accurate alignment and calibration between 3D point clouds and RGB images. However, such an assumption is not reliable in a real-world self-driving system, as the alignment between different modalities is easily affected by asynchronous sensors and disturbed sensor placement. We propose a novel Fusion network by Box Matching (FBMNet) for multi-modal 3D detection, which provides an alternative way for cross-modal feature alignment by learning the correspondence at the bounding box level to free up the dependency of calibration during inference. With the learned assignments between 3D and 2D object proposals, the fusion for detection can be effectively performed by combining their ROI features. Extensive experiments on the nuScenes dataset demonstrate that our method is much more robust in dealing with challenging cases such as asynchronous sensors, misaligned sensor placement, and degenerated camera images than existing fusion methods. We hope that our FBMNet could provide an available solution to dealing with these challenging cases for safety in real autonomous driving scenarios.
Zhe Liu 0033, Xiaoqing Ye, Zhikang Zou, Xinwei He 0001, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001, Xiang Bai
IEEE Trans. Intell. Transp. Syst.7
2023 Robust Video Portrait Reenactment via Personalized Representation Quantization
abstract
While progress has been made in the field of portrait reenactment, the problem of how to produce high-fidelity and robust videos remains. Recent studies normally find it challenging to handle rarely seen target poses due to the limitation of source data. This paper proposes the Video Portrait via Non-local Quantization Modeling (VPNQ) framework, which produces pose- and disturbance-robust reenactable video portraits. Our key insight is to learn position-invariant quantized local patch representations and build a mapping between simple driving signals and local textures with non-local spatial-temporal modeling. Specifically, instead of learning a universal quantized codebook, we identify that a personalized one can be trained to preserve desired position-invariant local details better. Then, a simple representation of projected landmarks can be used as sufficient driving signals to avoid 3D rendering. Following, we employ a carefully designed Spatio-Temporal Transformer to predict reasonable and temporally consistent quantized tokens from the driving signal. The predicted codes can be decoded back to robust and high-quality videos. Comprehensive experiments have been conducted to validate the effectiveness of our approach.
Kaisiyuan Wang, Changcheng Liang, Hang Zhou 0009, Jiaxiang Tang, Qianyi Wu, Dongliang He, Zhibin Hong, Jingtuo Liu, Errui Ding, Ziwei Liu 0002, Jingdong Wang 0001
AAAI11
2023 Cyclically Disentangled Feature Translation for Face Anti-spoofing
abstract
Current domain adaptation methods for face anti-spoofing leverage labeled source domain data and unlabeled target domain data to obtain a promising generalizable decision boundary. However, it is usually difficult for these methods to achieve a perfect domain-invariant liveness feature disentanglement, which may degrade the final classification performance by domain differences in illumination, face category, spoof type, etc. In this work, we tackle cross-scenario face anti-spoofing by proposing a novel domain adaptation method called cyclically disentangled feature translation network (CDFTN). Specifically, CDFTN generates pseudo-labeled samples that possess: 1) source domain-invariant liveness features and 2) target domain-specific content features, which are disentangled through domain adversarial training. A robust classifier is trained based on the synthetic pseudo-labeled images under the supervision of source domain labels. We further extend CDFTN for multi-target domain adaptation by leveraging data from more unlabeled target domains. Extensive experiments on several public datasets demonstrate that our proposed approach significantly outperforms the state of the art. Code and models are available at https://github.com/vis-face/CDFTN.
Haixiao Yue, Keyao Wang, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001
AAAI7
2023 StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-Based Generator
abstract
Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous studies either require long-term data for training or produce a similar movement pattern on all subjects with low quality. In this paper, we propose StyleSync, an effective framework that enables high-fidelity lip synchronization. We identify that a style-based generator would sufficiently enable such a charming property on both one-shot and few-shot scenarios. Specifically, we design a mask-guided spatial information encoding module that preserves the details of the given face. The mouth shapes are accurately modified by audio through modulated convolutions. Moreover, our design also enables personalized lip-sync by introducing style space and generator refinement on only limited frames. Thus the identity and talking style of a target person could be accurately preserved. Extensive experiments demonstrate the effectiveness of our method in producing high-fidelity results on a variety of scenes. Resources can be found at https:/hangz-nju-cuhk.github.io/projects/StyleSync.
Jiazhi Guan, Zhanwang Zhang, Hang Zhou 0009, Tianshu Hu, Kaisiyuan Wang, Dongliang He, Haocheng Feng, Jingtuo Liu, Errui Ding, Ziwei Liu 0002, Jingdong Wang 0001
CVPR11
2023 Ambiguity-Resistant Semi-Supervised Learning for Dense Object Detection
abstract
With basic Semi-Supervised Object Detection (SSOD) techniques, one-stage detectors generally obtain limited promotions compared with two-stage clusters. We experimentally find that the root lies in two kinds of ambiguities: (1) Selection ambiguity that selected pseudo labels are less accurate, since classification scores cannot properly represent the localization quality. (2) Assignment ambiguity that samples are matched with improper labels in pseudo-label assignment, as the strategy is misguided by missed objects and inaccurate pseudo boxes. To tackle these problems, we propose a Ambiguity-Resistant Semi-supervised Learning (ARSL) for one-stage detectors. Specifically, to alleviate the selection ambiguity, Joint-Confidence Estimation (JCE) is proposed to jointly quantifies the classification and localization quality of pseudo labels. As for the assignment ambiguity, Task-Separation Assignment (TSA) is introduced to assign labels based on pixel-level predictions rather than unreliable pseudo boxes. It employs a ‘divide-and-conquer’ strategy and separately exploits positives for the classification and localization task, which is more robust to the assignment ambiguity. Comprehensive experiments demonstrate that ARSL effectively mitigates the ambiguities and achieves state-of-the-art SSOD performance on MS COCO and PASCAL VOC. Codes can be found at https://github.com/PaddlePaddle/PaddleDetection.
Chang Liu 0082, Weiming Zhang 0006, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Xiaomao Li, Errui Ding, Jingdong Wang 0001
CVPR9
2023 Beyond Attentive Tokens: Incorporating Token Importance and Diversity for Efficient Vision Transformers
abstract
Vision transformers have achieved significant improvements on various vision tasks but their quadratic interactions between tokens significantly reduce computational efficiency. Many pruning methods have been proposed to remove redundant tokens for efficient vision transformers recently. However, existing studies mainly focus on the token importance to preserve local attentive tokens but completely ignore the global token diversity. In this paper, we emphasize the cruciality of diverse global semantics and propose an efficient token decoupling and merging method that can jointly consider the token importance and diversity for token pruning. According to the class token attention, we decouple the attentive and inattentive tokens. In addition to preserving the most discriminative local tokens, we merge similar inattentive tokens and match homogeneous attentive tokens to maximize the token diversity. Despite its simplicity, our method obtains a promising trade-off between model complexity and classification accuracy. On DeiT-S, our method reduces the FLOPs by 35% with only a 0.2% accuracy drop. Notably, benefiting from maintaining the token diversity, our method can even improve the accuracy of DeiT-T by 0.1% after reducing its FLOPs by 40%.
Sifan Long 0001, Zhen Zhao 0001, Jimin Pi, Sheng-Sheng Wang 0001, Jingdong Wang 0001
CVPR5
2023 PSVT: End-to-End Multi-Person 3D Pose and Shape Estimation with Progressive Video Transformers
abstract
Existing methods of multi-person video 3D human Pose and Shape Estimation (PSE) typically adopt a two-stage strategy, which first detects human instances in each frame and then performs single-person PSE with temporal model. However, the global spatio-temporal context among spatial instances can not be captured. In this paper, we propose a new end-to-end multi-person 3D Pose and Shape estimation framework with progressive Video Transformer, termed PSVT. In PSVT, a spatio-temporal encoder (STE) captures the global feature dependencies among spatial objects. Then, spatio-temporal pose decoder (STPD) and shape decoder (STSD) capture the global dependencies between pose queries and feature tokens, shape queries and feature tokens, respectively. To handle the variances of objects as time proceeds, a novel scheme of progressive decoding is used to update pose and shape queries at each frame. Besides, we propose a novel pose-guided attention (PGA) for shape decoder to better predict shape parameters. The two components strengthen the decoder of PSVT to improve performance. Extensive experiments on the four datasets show that PSVT achieves stage-of-the-art results.
Zhongwei Qiu, Qiansheng Yang, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Chang Xu 0002, Dongmei Fu, Jingdong Wang 0001
CVPR9
2023 Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval?
abstract
Most existing text-video retrieval methods focus on cross-modal matching between the visual content of videos and textual query sentences. However, in real-world scenarios, online videos are often accompanied by relevant text information such as titles, tags, and even subtitles, which can be utilized to match textual queries. This insight has motivated us to propose a novel approach to text-video retrieval, where we directly generate associated captions from videos using zero-shot video captioning with knowledge from web-scale pre-trained models (e.g., CLIP and GPT-2). Given the generated captions, a natural question arises: what benefits do they bring to text-video retrieval? To answer this, we introduce Cap4Video, a new framework that leverages captions in three ways: i) Input data: video-caption pairs can augment the training data. ii) Intermediate feature interaction: we perform cross-modal feature interaction between the video and caption to produce enhanced video representations. iii) Output score: the Query-Caption matching branch can complement the original Query-Video matching branch for text-video retrieval. We conduct comprehensive ablation studies to demonstrate the effectiveness of our approach. Without any post-processing, Cap4Video achieves state-of-the-art performance on four standard text-video retrieval benchmarks: MSR-VTT (51.4%), VATEX (66.6%), MSVD (51.8%), and DiDeMo (52.0%). The code is available at https://github.com/whwu95/Cap4Video.
Bo Fang 0003, Jingdong Wang 0001, Wanli Ouyang
CVPR4
2023 Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models
abstract
Vision-language models (VLMs) pre-trained on large- scale image-text pairs have demonstrated impressive transferability on various visual tasks. Transferring knowledge from such powerful VLMs is a promising direction for building effective video recognition models. However, current exploration in this field is still limited. We believe that the greatest value of pre-trained VLMs lies in building a bridge between visual and textual domains. In this paper, we propose a novel framework called BIKE, which utilizes the cross-modal bridge to explore bidirectional knowledge: i) We introduce the Video Attribute Association mechanism, which leverages the Video-to-Text knowledge to generate textual auxiliary attributes for complementing video recognition. ii) We also present a Temporal Concept Spotting mechanism that uses the Text-to-Video expertise to capture temporal saliency in a parameter-free manner, leading to enhanced video representation. Extensive studies on six popular video datasets, including Kinetics-400 & 600, UCF-101, HMDB-51, ActivityNet and Charades, show that our method achieves state-of-the-art performance in various recognition scenarios, such as general, zero-shot, and few-shot video recognition. Our best model achieves a state-of-the-art accuracy of 88.6% on the challenging Kinetics-400 using the released CLIP model. The code is available at https://github.com/whwu95/BIKE.
Jingdong Wang 0001, Yi Yang 0001, Wanli Ouyang
CVPR4
2023 CAPE: Camera View Position Embedding for Multi-View 3D Object Detection
abstract
In this paper, we address the problem of detecting 3D ob-jects from multi-view images. Current query-based methods rely on global 3D position embeddings (PE) to learn the ge-ometric correspondence between images and 3D space. We claim that directly interacting 2D image features with global 3D PE could increase the difficulty of learning view trans-formation due to the variation of camera extrinsics. Thus we propose a novel method based on CAmera view Position Embedding, called CAPE. We form the 3D position embed-dings under the local camera-view coordinate system instead of the global coordinate system, such that 3D position em-bedding is free of encoding camera extrinsic parameters. Furthermore, we extend our CAPE to temporal modeling by exploiting the object queries of previous frames and encoding the ego motion for boosting 3D object detection. CAPE achieves the state-of-the-art performance (61.0% NDS and 52.5% mAP) among all LiDAR-free methods on nuScenes dataset. Codes and models are available.11Codes of Paddle3D and PyTorch Implementation.
Kaixin Xiong, Shi Gong, Xiaoqing Ye, Xiao Tan 0001, Ji Wan, Errui Ding, Jingdong Wang 0001, Xiang Bai
CVPR7
2023 Semi-DETR: Semi-Supervised Object Detection with Detection Transformers
abstract
We analyze the DETR-based framework on semi-supervised object detection (SSOD) and observe that (1) the one-to-one assignment strategy generates incorrect matching when the pseudo ground-truth bounding box is inaccurate, leading to training inefficiency; (2) DETR-based detectors lack deterministic correspondence between the input query and its prediction output, which hinders the applicability of the consistency-based regularization widely used in current SSOD methods. We present Semi-DETR, the first transformer-based end-to-end semi-supervised object detector, to tackle these problems. Specifically, we propose a Stage-wise Hybrid Matching strategy that combines the one-to-many assignment and one-to-one assignment strategies to improve the training efficiency of the first stage and thus provide high-quality pseudo labels for the training of the second stage. Besides, we introduce a Cross-view Query Consistency method to learn the semantic feature invariance of object queries from different views while avoiding the need to find deterministic query correspondence. Furthermore, we propose a Cost-based Pseudo Label Mining module to dynamically mine more pseudo boxes based on the matching cost of pseudo ground truth bounding boxes for consistency training. Extensive experiments on all SSOD settings of both COCO and Pascal VOC benchmark datasets show that our Semi-DETR method outperforms all state-of-the-art methods by clear margins.
Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li
CVPR8
2023 Instance-Specific and Model-Adaptive Supervision for Semi-Supervised Semantic Segmentation
abstract
Recently, semi-supervised semantic segmentation has achieved promising performance with a small fraction of labeled data. However, most existing studies treat all unlabeled data equally and barely consider the differences and training difficulties among unlabeled instances. Differentiating unlabeled instances can promote instance-specific supervision to adapt to the model's evolution dynamically. In this paper, we emphasize the cruciality of instance differences and propose an instance-specific and model-adaptive supervision for semi-supervised semantic segmentation, named iMAS. Relying on the model's performance, iMAS employs a class-weighted symmetric intersection-over-union to evaluate quantitative hardness of each unlabeled instance and supervises the training on unlabeled data in a model-adaptive manner. Specifically, iMAS learns from unlabeled instances progressively by weighing their corresponding consistency losses based on the evaluated hardness. Besides, iMAS dynamically adjusts the augmentation for each instance such that the distortion degree of augmented instances is adapted to the model's generalization capability across the training course. Not integrating additional losses and training procedures, iMAS can obtain remarkable performance gains against current state-of-the-art approaches on segmentation benchmarks under different semi-supervised partition protocols11Code and logs: https://github.com/zhenzhao/iMAS.
Zhen Zhao 0001, Sifan Long 0001, Jimin Pi, Jingdong Wang 0001, Luping Zhou
CVPR4
2023 Augmentation Matters: A Simple-Yet-Effective Approach to Semi-Supervised Semantic Segmentation
abstract
Recent studies on semi-supervised semantic segmentation (SSS) have seen fast progress. Despite their promising performance, current state-of-the-art methods tend to increasingly complex designs at the cost of introducing more network components and additional training procedures. Differently, in this work, we follow a standard teacher-student framework and propose AugSeg, a simple and clean approach that focuses mainly on data perturbations to boost the SSS performance. We argue that various data augmentations should be adjusted to better adapt to the semi-supervised scenarios instead of directly applying these techniques from supervised learning. Specifically, we adopt a simplified intensity-based augmentation that selects a random number of data transformations with uniformly sampling distortion strengths from a continuous space. Based on the estimated confidence of the model on different unlabeled samples, we also randomly inject labelled information to augment the unlabeled samples in an adaptive manner. Without bells and whistles, our simple AugSeg can readily achieve new state-of-the-art performance on SSS benchmarks under different partition protocols11Code and logs: https://github.com/zhenzhao/AugSeg..
Zhen Zhao 0001, Lihe Yang, Sifan Long 0001, Jimin Pi, Luping Zhou, Jingdong Wang 0001
CVPR6
2023 Group DETR: Fast DETR Training with Group-Wise One-to-Many Assignment
abstract
Detection transformer (DETR) relies on one-to-one assignment, assigning one ground-truth object to one prediction, for end-to-end detection without NMS post-processing. It is known that one-to-many assignment, assigning one ground-truth object to multiple predictions, succeeds in detection methods such as Faster R-CNN and FCOS. While the naive one-to-many assignment does not work for DETR, and it remains challenging to apply one-to-many assignment for DETR training. In this paper, we introduce Group DETR, a simple yet efficient DETR training approach that introduces a group-wise way for one-to-many assignment. This approach involves using multiple groups of object queries, conducting one-to-one assignment within each group, and performing decoder self-attention separately. It resembles data augmentation with automatically-learned object query augmentation. It is also equivalent to simultaneously training parameter-sharing networks of the same architecture, introducing more supervision and thus improving DETR training. The inference process is the same as DETR trained normally and only needs one group of queries without any architecture modification. Group DETR is versatile and is applicable to various DETR variants. The experiments show that Group DETR signifi-cantly speeds up the training convergence and improves the performance of various DETR-based models. Code will be available at https://github.com/Atten4Vis/GroupDETR.
Qiang Chen 0007, Xiaokang Chen, Jian Wang 0066, Shan Zhang 0002, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001
ICCV10
2023 σ-Adaptive Decoupled Prototype for Few-Shot Object Detection
abstract
Meta-learning-based few-shot detectors use one K-average-pooled prototype (averaging along K-shot dimension) in both Region Proposal Network (RPN) and Detection head (DH) for query detection. Such plain operation would harm the FSOD performance in two aspects: 1) the poor quality of the prototype, and 2) the equivocal guidance due to the contradictions between RPN and DH. In this paper, we look closely into those critical issues and propose the σ-Adaptive Decoupled Prototype (σ-ADP) as a solution. To generate the high-quality prototype, we prioritize salient representations and deemphasize trivial variations by accessing both angle distance and magnitude dispersion (σ) across K-support samples. To provide precise information for the query image, the prototype is decoupled into task-specific ones, which provide tailored guidance for ‘where to look’ and ‘what to look for’, respectively.Beyond that, we find our σ-ADP can gradually strengthen the generalization power of encoding network during meta-training. So it can robustly deal with intra-class variations and a simple K- average pooling is enough to generate a high-quality prototype at meta-testing. We provide theoretical analysis to support its rationality. Extensive experiments on Pascal VOC, MS-COCO and FSOD datasets demonstrate that the proposed method achieves new state-of-the-art performance. Notably, our method surpasses the baseline model by a large margin – up to around 5.0% AP50and 8.0% AP75on novel classes.
Jinhao Du, Shan Zhang 0002, Qiang Chen 0007, Haifeng Le, Yanpeng Sun, Yao Ni, Jian Wang 0066, Jingdong Wang 0001
ICCV9
2023 UATVR: Uncertainty-Adaptive Text-Video Retrieval
abstract
With the explosive growth of web videos and emerging large-scale vision-language pre-training models, e.g., CLIP, retrieving videos of interest with text instructions has attracted increasing attention. A common practice is to transfer text-video pairs to the same embedding space and craft cross-modal interactions with certain entities in specific granularities for semantic correspondence. Unfortunately, the intrinsic uncertainties of optimal entity combinations in appropriate granularities for cross-modal queries are understudied, which is especially critical for modalities with hierarchical semantics, e.g., video, text, etc. In this paper, we propose an Uncertainty-Adaptive Text-Video Retrieval approach, termed UATVR, which models each lookup as a distribution matching procedure. Concretely, we add additional learnable tokens in the encoders to adaptively aggregate multi-grained semantics for flexible high-level reasoning. In the refined embedding space, we represent text-video pairs as probabilistic distributions where prototypes are sampled for matching evaluation. Comprehensive experiments on four benchmarks justify the superiority of our UATVR, which achieves new state-of-the-art results on MSR-VTT (50.8%), VATEX (64.5%), MSVD (49.7%), and DiDeMo (45.8%). The code is available at https://github.com/bofang98/UATVR.
Bo Fang 0003, Yu Zhou 0015, YuXin Song 0001, Weiping Wang 0005, Xiangbo Shu, Xiangyang Ji, Jingdong Wang 0001
ICCV9
2023 Forward Flow for Novel View Synthesis of Dynamic Scenes
abstract
This paper proposes a neural radiance field (NeRF) approach for novel view synthesis of dynamic scenes using forward warping. Existing methods often adopt a static NeRF to represent the canonical space, and render dynamic images at other time steps by mapping the sampled 3D points back to the canonical space with the learned backward flow field. However, this backward flow field is non-smooth and discontinuous, which is difficult to be fitted by commonly used smooth motion models. To address this problem, we propose to estimate the forward flow field and directly warp the canonical radiance field to other time steps. Such forward flow field is smooth and continuous within the object region, which benefits the motion model learning. To achieve this goal, we represent the canonical radiance field with voxel grids to enable efficient forward warping, and propose a differentiable warping process, including an average splatting operation and an inpaint network, to resolve the many-to-one and one-to-many mapping issues. Thorough experiments show that our method outperforms existing methods in both novel view rendering and motion modeling, demonstrating the effectiveness of our forward flow motion modeling. Project page: https://npucvr.github.io/ForwardFlowDNeRF.
Jiadai Sun, Yuchao Dai, Guanying Chen, Xiaoqing Ye, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001
ICCV9
2023 CFCG: Semi-Supervised Semantic Segmentation via Cross-Fusion and Contour Guidance Supervision
abstract
Current state-of-the-art semi-supervised semantic segmentation (SSSS) methods typically adopt pseudo labeling and consistency regularization between multiple learners with different perturbations. Although the performance is desirable, many issues remain: (1) supervisions from a single learner tend to be noisy which causes unreliable consistency regularization (2) existing pixel-wise confidence-score-based reliability measurement causes potential error accumulation as the training proceeds. In this paper, we propose a novel SSSS framework, called CFCG, which combines cross-fusion and contour guidance supervision to tackle these issues. Concretely, we adopt both image-level and feature-level perturbations to expand feature distribution thus pushing the potential limits of consistency regularization. Then, two particular modules are proposed to enable effective semi-supervised learning under heavy coherent perturbations. Firstly, Cross-Fusion Supervision (CFS) mechanism leverages multiple learners to enhance the quality of pseudo labels. Secondly, we introduce an adaptive contour guidance module (ACGM) to effectively identify unreliable spatial regions in pseudo labels. Finally, our proposed CFCG achieves gains of mIoU +1.40%, +0.89% with a single learner and +1.85%, +1.33% by fusion inference on PASCAL VOC 2012 and on Cityscapes respectively under 1/8 protocols, clearly surpassing previous methods and reaching the state-of-the-art.
Shuo Li 0012, Weiming Zhang 0006, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001
ICCV8
2023 Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection
abstract
Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MS-COCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the performance of semi-supervised object detectors far from satisfactory. Besides, the research for this problem in SSOD is severely under-explored. To bridge this research gap, we comprehensively study the class imbalance problem for SSOD under more challenging scenarios, thus forming the first experimental setting for class imbalanced SSOD (CI-SSOD). Moreover, we propose a simple yet effective gradient-based sampling framework that tackles the class imbalance problem from the perspective of two types of confirmation biases. To tackle confirmation bias towards majority classes, the gradient-based reweighting and gradient-based thresholding modules leverage the gradients from each class to fully balance the influence of the majority and minority classes. To tackle the confirmation bias from incorrect pseudo labels of minority classes, the class-rebalancing sampling module resamples unlabeled data following the guidance of the gradient-based reweighting module. Experiments on three proposed sub-tasks, namely MS-COCO, MS-COCO → Object365 and LVIS, suggest that our method outperforms current class imbalanced object detectors by clear margins, serving as a baseline for future research in CI-SSOD. Code will be available at https://github.com/nightkeepers/CI-SSOD.
Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li
ICCV8
2023 Group Pose: A Simple Baseline for End-to-End Multi-person Pose Estimation
abstract
In this paper, we study the problem of end-to-end multi-person pose estimation. State-of-the-art solutions adopt the DETR-like framework, and mainly develop the complex decoder, e.g., regarding pose estimation as keypoint box detection and combining with human detection in ED-Pose [38], hierarchically predicting with pose decoder and joint (keypoint) decoder in PETR [27].We present a simple yet effective transformer approach, named Group Pose. We simply regard K-keypoint pose estimation as predicting a set of N × K keypoint positions, each from a keypoint query, as well as representing each pose with an instance query for scoring N pose predictions.Motivated by the intuition that the interaction, among across-instance queries of different types, is not directly helpful, we make a simple modification to decoder self-attention. We replace single self-attention over all the N × (K + 1) queries with two subsequent group self-attentions: (i) N within-instance self-attention, with each over K keypoint queries and one instance query, and (ii) (K +1) same-type across-instance self-attention, each over N queries of the same type. The resulting decoder removes the interaction among across-instance type-different queries, easing the optimization and thus improving the performance. Experimental results on MS COCO and Crowd-Pose show that our approach without human box supervision is superior to previous methods with complex decoders, and even is slightly better than ED-Pose that uses human box supervision. Paddle1and PyTorch2codes are available.
Huan Liu 0030, Qiang Chen 0007, Zichang Tan, Jiang-Jiang Liu 0001, Jian Wang 0066, Xiangbo Su, Xiaolong Li 0001, Junyu Han, Errui Ding, Yao Zhao 0001, Jingdong Wang 0001
ICCV12
2023 CPCM: Contextual Point Cloud Modeling for Weakly-supervised Point Cloud Semantic Segmentation
abstract
We study the task of weakly-supervised point cloud semantic segmentation with sparse annotations (e.g., less than 0.1% points are labeled), aiming to reduce the expensive cost of dense annotations. Unfortunately, with extremely sparse annotated points, it is very difficult to extract both contextual and object information for scene understanding such as semantic segmentation. Motivated by masked modeling (e.g., MAE) in image and video representation learning, we seek to endow the power of masked modeling to learn contextual information from sparsely-annotated points. However, directly applying MAE to 3D point clouds with sparse annotations may fail to work. First, it is nontrivial to effectively mask out the informative visual context from 3D point clouds. Second, how to fully exploit the sparse annotations for context modeling remains an open question. In this paper, we propose a simple yet effective Contextual Point Cloud Modeling (CPCM) method that consists of two parts: a region-wise masking (Region-Mask) strategy and a contextual masked training (CMT) method. Specifically, RegionMask masks the point cloud continuously in geometric space to construct a meaningful masked prediction task for subsequent context learning. CMT disentangles the learning of supervised segmentation and unsupervised masked context prediction for effectively learning the very limited labeled points and mass unlabeled points, respectively. Extensive experiments on the widely-tested ScanNet V2 and S3DIS benchmarks demonstrate the superiority of CPCM over the state-of-the-art.
Lizhao Liu, Zhuangwei Zhuang, Shangxin Huang, Xunlong Xiao, Tianhang Xiang, Cen Chen 0002, Jingdong Wang 0001, Mingkui Tan
ICCV7
2023 Task-Oriented Multi-Modal Mutual Learning for Vision-Language Models
abstract
Prompt learning has become one of the most efficient paradigms for adapting large pre-trained vision-language models to downstream tasks. Current state-of-the-art methods, like CoOp and ProDA, tend to adopt soft prompts to learn an appropriate prompt for each specific task. Recent CoCoOp further boosts the base-to-new generalization performance via an image-conditional prompt. However, it directly fuses identical image semantics to prompts of different labels and significantly weakens the discrimination among different classes as shown in our experiments. Motivated by this observation, we first propose a class-aware text prompt (CTP) to enrich generated prompts with label-related image information. Unlike CoCoOp, CTP can effectively involve image semantics and avoid introducing extra ambiguities into different prompts. On the other hand, instead of reserving the complete image representations, we propose text-guided feature tuning (TFT) to make the image branch attend to class-related representation. A contrastive loss is employed to align such augmented text and image representations on downstream tasks. In this way, the image-to-text CTP and text-to-image TFT can be mutually promoted to enhance the adaptation of VLMs for downstream tasks. Extensive experiments demonstrate that our method outperforms the existing methods by a significant margin. Especially, compared to CoCoOp, we achieve an average improvement of 4.03% on new classes and 3.19% on harmonic-mean over eleven classification benchmarks.
Sifan Long 0001, Zhen Zhao 0001, Junkun Yuan, Zichang Tan, Jiangjiang Liu 0006, Luping Zhou, Sheng-Sheng Wang 0001, Jingdong Wang 0001
ICCV8
2023 Unified Pre-training with Pseudo Texts for Text-To-Image Person Re-identification
abstract
The pre-training task is indispensable for the text-to-image person re-identification (T2I-ReID) task. However, there are two underlying inconsistencies between these two tasks that may impact the performance: i) Data inconsistency. A large domain gap exists between the generic images/texts used in public pre-trained models and the specific person data in the T2I-ReID task. This gap is especially severe for texts, as general textual data are usually unable to describe specific people in fine-grained detail. ii) Training inconsistency. The processes of pre-training of images and texts are independent, despite cross-modality learning being critical to T2I-ReID. To address the above issues, we present a new unified pre-training pipeline (UniPT) designed specifically for the T2I-ReID task. We first build a large-scale text-labeled person dataset "LUPerson-T", in which pseudo-textual descriptions of images are automatically generated by the CLIP paradigm using a divide-conquer-combine strategy. Benefiting from this dataset, we then utilize a simple vision-and-language pre-training framework to explicitly align the feature space of the image and text modalities during pre-training. In this way, the pre-training task and the T2I-ReID task are made consistent with each other on both data and training levels. Without the need for any bells and whistles, our UniPT achieves competitive Rank-1 accuracy of, i.e., 68.50%, 60.09%, and 51.85% on CUHK-PEDES, ICFG-PEDES and RSTPReid, respectively. Both the LUPerson-T dataset and code are available at https://github.com/ZhiyinShao-H/UniPT.
Zhiyin Shao, Xinyu Zhang 0015, Changxing Ding, Jian Wang 0066, Jingdong Wang 0001
ICCV5
2023 Delicate Textured Mesh Recovery from NeRF via Adaptive Surface Refinement
abstract
Neural Radiance Fields (NeRF) have constituted a remarkable breakthrough in image-based 3D reconstruction. However, their implicit volumetric representations differ significantly from the widely-adopted polygonal meshes and lack support from common 3D software and hardware, making their rendering and manipulation inefficient. To overcome this limitation, we present a novel framework that generates textured surface meshes from images. Our approach begins by efficiently initializing the geometry and view-dependency decomposed appearance with a NeRF. Subsequently, a coarse mesh is extracted, and an iterative surface refinement algorithm is developed to adaptively adjust both vertex positions and face density based on reprojected rendering errors. We jointly refine the appearance with geometry and bake it into texture images for real-time rendering. Extensive experiments demonstrate that our method achieves superior mesh quality and competitive rendering quality.
Jiaxiang Tang, Hang Zhou 0009, Xiaokang Chen, Tianshu Hu, Errui Ding, Jingdong Wang 0001
ICCV6
2023 What Can Simple Arithmetic Operations Do for Temporal Modeling?
abstract
Temporal modeling plays a crucial role in understanding video content. To tackle this problem, previous studies built complicated temporal relations through time sequence thanks to the development of computationally powerful devices. In this work, we explore the potential of four simple arithmetic operations for temporal modeling. Specifically, we first capture auxiliary temporal cues by computing addition, subtraction, multiplication, and division between pairs of extracted frame features. Then, we extract corresponding features from these cues to benefit the original temporal-irrespective domain. We term such a simple pipeline as an Arithmetic Temporal Module (ATM), which operates on the stem of a visual backbone with a plug-and-play style. We conduct comprehensive ablation studies on the instantiation of ATMs and demonstrate that this module provides powerful temporal modeling capability at a low computational cost. Moreover, the ATM is compatible with both CNNs- and ViTs-based architectures. Our results show that ATM achieves superior performance over several popular video benchmarks. Specifically, on Something-Something V1, V2 and Kinetics-400, we reach top-1 accuracy of 65.6%, 74.6%, and 89.4% respectively. The code is available at https://github.com/whwu95/ATM.
YuXin Song 0001, Zhun Sun, Jingdong Wang 0001, Chang Xu 0002, Wanli Ouyang
ICCV4
2023 Boosting Few-shot Action Recognition with Graph-guided Hybrid Matching
abstract
Class prototype construction and matching are core aspects of few-shot action recognition. Previous methods mainly focus on designing spatiotemporal relation modeling modules or complex temporal alignment algorithms. Despite the promising results, they ignored the value of class prototype construction and matching, leading to unsatisfactory performance in recognizing similar categories in every task. In this paper, we propose GgHM, a new framework with Graph-guided Hybrid Matching. Concretely, we learn task-oriented features by the guidance of a graph neural network during class prototype construction, optimizing the intra- and inter-class feature correlation explicitly. Next, we design a hybrid matching strategy, combining frame-level and tuple-level matching to classify videos with multivariate styles. We additionally propose a learnable dense temporal modeling module to enhance the video feature temporal representation to build a more solid foundation for the matching process. GgHM shows consistent improvements over other challenging baselines on several few-shot datasets, demonstrating the effectiveness of our method. The code will be publicly available at https://github.com/jiazheng-xing/GgHM.
Jiazheng Xing, Mengmeng Wang 0005, Yudi Ruan, Bofan Chen, Yaowei Guo, Boyu Mu, Guang Dai, Jingdong Wang 0001, Yong Liu 0007
ICCV8
2023 ICDAR 2023 Competition on Structured Text Extraction from Visually-Rich Document Images
Wenwen Yu, Chengquan Zhang, Haoyu Cao 0001, Wei Hua 0005, Bohan Li 0010, Mingrui Chen 0001, Jianfeng Kuang, Mengjun Cheng, Yuning Du, Shikun Feng, Xiaoguang Hu, Pengyuan Lv, Yuechen Yu, Wanxiang Che, Errui Ding, Cheng-Lin Liu 0001, Jiebo Luo 0001, Shuicheng Yan, Min Zhang 0005, Dimosthenis Karatzas, Xing Sun 0001, Jingdong Wang 0001, Xiang Bai
ICDAR (2)26
2023 Graph Contrastive Learning for Skeleton-based Action Recognition
Xiaohu Huang, Hao Zhou 0039, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001, Xinggang Wang, Wenyu Liu 0001, Bin Feng 0001
ICLR7
2023 StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
Yuechen Yu, Yulin Li 0004, Chengquan Zhang, Xiaoqiang Zhang 0006, Zengyuan Guo, Xiameng Qin, Junyu Han, Errui Ding, Jingdong Wang 0001
ICLR10
2023 GridFormer: Towards Accurate Table Structure Recognition via Grid Prediction
abstract
All tables can be represented as grids. Based on this observation, we propose GridFormer, a novel approach for interpreting unconstrained table structures by predicting the vertex and edge of a grid. First, we propose a flexible table representation in the form of an M X N grid. In this representation, the vertexes and edges of the grid store the localization and adjacency information of the table. Then, we introduce a DETR-style table structure recognizer to efficiently predict this multi-objective information of the grid in a single shot. Specifically, given a set of learned row and column queries, the recognizer directly outputs the vertexes and edges information of the corresponding rows and columns. Extensive experiments on five challenging benchmarks which include wired, wireless, multi-merge-cell, oriented, and distorted tables demonstrate the competitive performance of our model over other methods.
Pengyuan Lv, Weihong Ma, Hongyi Wang 0008, Yuechen Yu, Chengquan Zhang, Yang Xue 0001, Jingdong Wang 0001
ACM Multimedia8
2023 DAC-DETR: Divide the Attention Layers and Conquer
abstract
This paper reveals a characteristic of DEtection Transformer (DETR) that negatively impacts its training efficacy, i.e., the cross-attention and self-attention layers in DETR decoder have contrary impacts on the object queries (though both impacts are important). Specifically, we observe the cross-attention tends to gather multiple queries around the same object, while the self-attention disperses these queries far away. To improve the training efficacy, we propose a Divide-And-Conquer DETR (DAC-DETR) that divides the cross-attention out from this contrary for better conquering. During training, DAC-DETR employs an auxiliary decoder that focuses on learning the cross-attention layers. The auxiliary decoder, while sharing all the other parameters, has NO self-attention layers and employs one-to-many label assignment to improve the gathering effect. Experiments show that DAC-DETR brings remarkable improvement over popular DETRs. For example, under the 12 epochs training scheme on MS-COCO, DAC-DETR improves Deformable DETR (ResNet-50) by +3.4 AP and achieves 50.9 (ResNet-50) / 58.1 AP (Swin-Large) based on some popular methods (i.e., DINO and an IoU-related loss). Our code will be made available at https://github.com/huzhengdongcs/DAC-DETR.
Zhengdong Hu, Yifan Sun 0003, Jingdong Wang 0001, Yi Yang 0001
NeurIPS3
2023 Leveraging Vision-Centric Multi-Modal Expertise for 3D Object Detection
abstract
Current research is primarily dedicated to advancing the accuracy of camera-only 3D object detectors (apprentice) through the knowledge transferred from LiDAR- or multi-modal-based counterparts (expert). However, the presence of the domain gap between LiDAR and camera features, coupled with the inherent incompatibility in temporal fusion, significantly hinders the effectiveness of distillation-based enhancements for apprentices. Motivated by the success of uni-modal distillation, an apprentice-friendly expert model would predominantly rely on camera features, while still achieving comparable performance to multi-modal models. To this end, we introduce VCD, a framework to improve the camera-only apprentice model, including an apprentice-friendly multi-modal expert and temporal-fusion-friendly distillation supervision. The multi-modal expert VCD-E adopts an identical structure as that of the camera-only apprentice in order to alleviate the feature disparity, and leverages LiDAR input as a depth prior to reconstruct the 3D scene, achieving the performance on par with other heterogeneous multi-modal experts. Additionally, a fine-grained trajectory-based distillation module is introduced with the purpose of individually rectifying the motion misalignment for each object in the scene. With those improvements, our camera-only apprentice VCD-A sets new state-of-the-art on nuScenes with a score of 63.1% NDS. The code will be released at https://github.com/OpenDriveLab/Birds-eye-view-Perception.
Linyan Huang, Chonghao Sima, Wenhai Wang, Jingdong Wang 0001, Yu Qiao 0001, Hongyang Li 0001
NeurIPS5
2023 HAP: Structure-Aware Masked Image Modeling for Human-Centric Perception
abstract
Model pre-training is essential in human-centric perception. In this paper, we first introduce masked image modeling (MIM) as a pre-training approach for this task. Upon revisiting the MIM training strategy, we reveal that human structure priors offer significant potential. Motivated by this insight, we further incorporate an intuitive human structure prior - human parts - into pre-training. Specifically, we employ this prior to guide the mask sampling process. Image patches, corresponding to human part regions, have high priority to be masked out. This encourages the model to concentrate more on body structure information during pre-training, yielding substantial benefits across a range of human-centric perception tasks. To further capture human characteristics, we propose a structure-invariant alignment loss that enforces different masked views, guided by the human part prior, to be closely aligned for the same image. We term the entire method as HAP. HAP simply uses a plain ViT as the encoder yet establishes new state-of-the-art performance on 11 human-centric benchmarks, and on-par result on one dataset. For example, HAP achieves 78.1% mAP on MSMT17 for person re-identification, 86.54% mA on PA-100K for pedestrian attribute recognition, 78.2% AP on MS COCO for 2D pose estimation, and 56.0 PA-MPJPE on 3DPW for 3D pose and shape estimation.
Junkun Yuan, Xinyu Zhang 0015, Hao Zhou 0039, Jian Wang 0066, Zhongwei Qiu, Zhiyin Shao, Shaofeng Zhang, Sifan Long 0001, Kun Kuang 0001, Junyu Han, Errui Ding, Lanfen Lin, Fei Wu 0001, Jingdong Wang 0001
NeurIPS15
2023 Guest editorial: special issue on human pose estimation and its applications
Zhou Ren, Jingdong Wang 0001
Mach. Vis. Appl.3
2023 Structured Knowledge Distillation for Dense Prediction
abstract
In this work, we consider transferring the structure information from large networks to compact ones for dense prediction tasks in computer vision. Previous knowledge distillation strategies used for dense prediction tasks often directly borrow the distillation scheme for image classification and perform knowledge distillation for each pixel separately, leading to sub-optimal performance. Here we propose to distill structured knowledge from large networks to compact networks, taking into account the fact that dense prediction is a structured prediction problem. Specifically, we study two structured distillation schemes: i) pair-wise distillation that distills the pair-wise similarities by building a static graph; and ii) holistic distillation that uses adversarial training to distill holistic knowledge. The effectiveness of our knowledge distillation approaches is demonstrated by experiments on three dense prediction tasks: semantic segmentation, depth estimation and object detection. Code is available at https://git.io/StructKD.
Yifan Liu 0001, Changyong Shu, Jingdong Wang 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 MixFormer: Mixing Features across Windows and Dimensions
abstract
While local-window self-attention performs notably in vision tasks, it suffers from limited receptive field and weak modeling capability issues. This is mainly because it performs self-attention within non-overlapped windows and shares weights on the channel dimension. We propose Mix-Former to find a solution. First, we combine local-window self-attention with depth-wise convolution in a parallel design, modeling cross-window connections to enlarge the receptive fields. Second, we propose bi-directional interactions across branches to provide complementary clues in the channel and spatial dimensions. These two designs are integrated to achieve efficient feature mixing among windows and dimensions. Our MixFormer provides competitive results on image classification with EfficientNet and shows better results than RegNet and Swin Transformer. Performance in downstream tasks outperforms its alternatives by significant margins with less computational costs in 5 dense prediction tasks on MS COCO, ADE20k, and LVIS. Code is available at https://github.com/PaddlePaddle/PaddleClas.
Qiang Chen 0007, Qiman Wu, Jian Wang 0066, Qinghao Hu 0001, Errui Ding, Jian Cheng 0001, Jingdong Wang 0001
CVPR8
2022 ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
abstract
Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of scene text information and directly adding this information may lead to performance degradation in scene text free scenarios. To address this issue, we propose a full transformer architecture to unify these cross-modal retrieval scenarios in a single Vision and Scene Text Aggregation framework (ViSTA). Specifically, ViSTA utilizes transformer blocks to directly encode image patches and fuse scene text embedding to learn an aggregated visual representation for cross-modal retrieval. To tackle the modality missing problem of scene text, we propose a novel fusion token based transformer aggregation approach to exchange the necessary scene text information only through the fusion token and concentrate on the most important features in each modality. To further strengthen the visual modality, we develop dual contrastive learning losses to embed both image-text pairs and fusion-text pairs into a common cross-modal space. Compared to existing methods, ViSTA enables to aggregate relevant scene text semantics with visual appearance, and hence improve results under both scene text free and scene text aware scenarios. Experimental results show that ViSTA outperforms other methods by at least 8.4% at Recall@ 1 for scene text aware retrieval task. Compared with state-of-the-art scene text free retrieval methods, ViSTA can achieve better accuracy on Flicker30K and MSCOCO while running at least three times faster during the inference stage, which validates the effectiveness of the proposed framework.
Mengjun Cheng, Yipeng Sun, Longchao Wang, Xiongwei Zhu, Jie Chen 0001, Guoli Song, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
CVPR11
2022 Expressive Talking Head Generation with Granular Audio-Visual Control
abstract
Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking Heads (GC-AVT), which controls lip movements, head poses, and facial expressions of a talking head in a granular manner. Our insight is to decouple the audio-visual driving sources through prior-based pre-processing designs. Detailedly, we disassemble the driving image into three complementary parts including: 1) a cropped mouth that facilitates lip-sync; 2) a masked head that implicitly learns pose; and 3) the upper face which works corporately and complementarily with a time-shifted mouth to contribute the expression. Interestingly, the encoded features from the three sources are integrally balanced through reconstruction training. Extensive experiments show that our method generates expressive faces with not only synced mouth shapes, controllable poses, but precisely animated emotional expressions as well.
Borong Liang, Yan Pan 0019, Zhizhi Guo, Hang Zhou 0009, Zhibin Hong, Xiaoguang Han 0001, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
CVPR10
2022 Few-Shot Head Swapping in the Wild
abstract
The head swapping task aims at flawlessly placing a source head onto a target body, which is of great importance to various entertainment scenarios. While face swapping has drawn much attention, the task of head swapping has rarely been explored, particularly under the few-shot setting. It is inherently challenging due to its unique needs in head modeling and background blending. In this paper, we present the Head Swapper (HeSer), which achieves few-shot head swapping in the wild through two delicately de-signed modules. Firstly, a Head2Head Aligner is devised to holistically migrate pose and expression information from the target to the source head by examining multi-scale in-formation. Secondly, to tackle the challenges of skin color variations and head-background mismatches in the swapping procedure, a Head2Scene Blender is introduced to si-multaneously modify facial skin color and fill mismatched gaps on the background around the head. Particularly, seamless blending is achieved with the help of a Semantic-Guided Color Reference Creation procedure and a Blending UNet. Extensive experiments demonstrate that the proposed method produces superior head swapping results on a variety of scenes.
Changyong Shu, Hemao Wu, Hang Zhou 0009, Jiaming Liu 0003, Zhibin Hong, Changxing Ding, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
CVPR10
2022 Few-Shot Font Generation by Learning Fine-Grained Local Styles
abstract
Few-shot font generation (FFG), which aims to generate a new font with a few examples, is gaining increasing attention due to the significant reduction in labor cost. A typical FFG pipeline considers characters in a standard font library as content glyphs and transfers them to a new target font by extracting style information from the reference glyphs. Most existing solutions explicitly disentangle content and style of reference glyphs globally or component-wisely. However, the style of glyphs mainly lies in the local details, i.e. the styles of radicals, components, and strokes together depict the style of a glyph. Therefore, even a single character can contain different styles distributed over spatial locations. In this paper, we propose a new font generation approach by learning 1) the fine-grained local styles from references, and 2) the spatial correspondence between the content and reference glyphs. Therefore, each spatial location in the content glyph can be assigned with the right fine-grained style. To this end, we adopt cross-attention over the representation of the content glyphs as the queries and the representations of the reference glyphs as the keys and values. Instead of explicitly disentangling global or component-wise modeling, the cross-attention mechanism can attend to the right local styles in the reference glyphs and aggregate the reference styles into a fine-grained style representation for the given content glyphs. The experiments show that the proposed method outperforms the state-of-the-art methods in FFG. In particular, the user studies also demonstrate the style consistency of our approach significantly outperforms previous methods.
Licheng Tang, Yiyang Cai, Jiaming Liu 0003, Zhibin Hong, Mingming Gong, Minhu Fan, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
CVPR10
2022 Implicit Sample Extension for Unsupervised Person Re-Identification
abstract
Most existing unsupervised person re-identification (Re-ID) methods use clustering to generate pseudo labels for model training. Unfortunately, clustering sometimes mixes different true identities together or splits the same identity into two or more sub clusters. Training on these noisy clusters substantially hampers the Re-ID accuracy. Due to the limited samples in each identity, we suppose there may lack some underlying information to well reveal the accurate clusters. To discover these information, we propose an Implicit Sample Extension (ISE) method to generate what we call support samples around the cluster boundaries. Specifically, we generate support samples from actual samples and their neighbouring clusters in the embedding space through a progressive linear interpolation (PLI) strategy. PLI controls the generation with two critical factors, i.e., 1) the direction from the actual sample towards its K-nearest clusters and 2) the degree for mixing up the context information from the K-nearest clusters. Meanwhile, given the support samples, ISE further uses a label-preserving loss to pull them towards their corresponding actual samples, so as to compact each cluster. Consequently, ISE reduces the “sub and mixed” clustering errors, thus improving the Re-ID performance. Extensive experiments demonstrate that the proposed method is effective and achieves state-of-the-art performance for unsupervised person Re-ID. Code is available at: https://github.com/PaddlePaddle/PaddleClas.
Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Errui Ding, Qinfeng Shi, Zhaoxiang Zhang 0001, Jingdong Wang 0001
CVPR8
2022 Human-Object Interaction Detection via Disentangled Transformer
abstract
Human-Object Interaction Detection tackles the problem of joint localization and classification of human object interactions. Existing HOI transformers either adopt a single decoder for triplet prediction, or utilize two parallel decoders to detect individual objects and interactions separately, and compose triplets by a matching process. In contrast, we decouple the triplet prediction into human-object pair detection and interaction classification. Our main motivation is that detecting the human-object instances and classifying interactions accurately needs to learn representations that focus on different regions. To this end, we present Disentangled Transformer, where both encoder and decoder are disentangled to facilitate learning of two sub-tasks. To associate the predictions of disentangled decoders, we first generate a unified representation for HOI triplets with a base decoder, and then utilize it as input feature of each disentangled decoder. Extensive experiments show that our method outperforms prior work on two public HOI benchmarks by a sizeable margin. Code will be available.
Desen Zhou, Jian Wang 0066, Leshan Wang, Errui Ding, Jingdong Wang 0001
CVPR7
2022 Action Quality Assessment with Temporal Parsing Transformer
Yang Bai 0011, Desen Zhou, Songyang Zhang 0001, Jian Wang 0066, Errui Ding, Yu Guan 0001, Yang Long 0001, Jingdong Wang 0001
ECCV (4)8
2022 DaViT: Dual Attention Vision Transformers
Mingyu Ding, Bin Xiao 0004, Noel Codella, Ping Luo 0002, Jingdong Wang 0001, Lu Yuan 0001
ECCV (24)5
2022 GitNet: Geometric Prior-Based Transformation for Birds-Eye-View Segmentation
Shi Gong, Xiaoqing Ye, Xiao Tan 0001, Jingdong Wang 0001, Errui Ding, Yu Zhou 0025, Xiang Bai
ECCV (1)4
2022 Diverse Learner: Exploring Diverse Supervision for Semi-supervised Object Detection
Minyue Jiang, Wei Zhang 0197, Xiangru Lin, Xiao Tan 0001, Jingdong Wang 0001, Errui Ding
ECCV (30)8
2022 CODER: Coupled Diversity-Sensitive Momentum Contrastive Learning for Image-Text Retrieval
Haoran Wang 0004, Dongliang He, Boyang Xia, Fu Li 0003, Zhong Ji, Errui Ding, Jingdong Wang 0001
ECCV (36)10
2022 UFO: Unified Feature Optimization
Teng Xi, Yifan Sun 0003, Deli Yu, Bi Li 0005, Nan Peng, Xinyu Zhang 0015, Zhigang Wang 0002, Jian Wang 0066, Haocheng Feng, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
ECCV (26)16
2022 StyleSwap: Style-Based Generator Empowers Robust Face Swapping
Hang Zhou 0009, Zhibin Hong, Ziwei Liu 0002, Jiaming Liu 0003, Zhizhi Guo, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang 0001
ECCV (14)10
2022 Learning Versatile Neural Architectures by Propagating Network Codes
Mingyu Ding, Yuqi Huo, Haoyu Lu, Zhe Wang 0006, Zhiwu Lu 0001, Jingdong Wang 0001, Ping Luo 0002
ICLR7
2022 On the Connection between Local Attention and Dynamic Depth-wise Convolution
Qi Han 0007, Zejia Fan, Qi Dai 0001, Ming-Ming Cheng, Jiaying Liu 0001, Jingdong Wang 0001
ICLR7
2022 Self-Guided Hard Negative Generation for Unsupervised Person Re-Identification
abstract
Recent unsupervised person re-identification (reID) methods mostly apply pseudo labels from clustering algorithms as supervision signals. Despite great success, this fashion is very likely to aggregate different identities with similar appearances into the same cluster. In result, the hard negative samples, playing important role in training reID models, are significantly reduced. To alleviate this problem, we propose a self-guided hard negative generation method for unsupervised person re-ID. Specifically, a joint framework is developed which incorporates a hard negative generation network (HNGN) and a re-ID network. To continuously generate harder negative samples to provide effective supervisions in the contrastive learning, the two networks are alternately trained in an adversarial manner to improve each other, where the reID network guides HNGN to generate challenging data and HNGN enforces the re-ID network to enhance discrimination ability. During inference, the performance of re-ID network is improved without introducing any extra parameters. Extensive experiments demonstrate that the proposed method significantly outperforms a strong baseline and also achieves better results than state-of-the-art methods.
Zhigang Wang 0002, Jian Wang 0066, Xinyu Zhang 0015, Errui Ding, Jingdong Wang 0001, Zhaoxiang Zhang 0001
IJCAI6
2022 Paint and Distill: Boosting 3D Object Detection with Semantic Passing Network
abstract
3D object detection task from lidar or camera sensors is essential for autonomous driving. Pioneer attempts at multi-modality fusion complement the sparse lidar point clouds with rich semantic texture information from images at the cost of extra network designs and overhead. In this work, we propose a novel semantic passing framework, named SPNet, to boost the performance of existing lidar-based 3D detection models with the guidance of rich context painting, with no extra computation cost during inference. Our key design is to first exploit the potential instructive semantic knowledge within the ground-truth labels by training a semantic-painted teacher model and then guide the pure-lidar network to learn the semantic-painted representation via knowledge passing modules at different granularities: class-wise passing, pixel-wise passing and instance-wise passing. Experimental results show that the proposed SPNet can seamlessly cooperate with most existing 3D detection frameworks with 1$\sim$5% AP gain and even achieve new state-of-the-art 3D detection performance on the KITTI test benchmark. Code is available at: https://github.com/jb892/SPNet.
Bo Ju, Zhikang Zou, Xiaoqing Ye, Minyue Jiang, Xiao Tan 0001, Errui Ding, Jingdong Wang 0001
ACM Multimedia7
2022 RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer
abstract
Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in this field, due to the time-consuming computation mechanism of transformer. We propose RTFormer, an efficient dual-resolution transformer for real-time semantic segmenation, which achieves better trade-off between performance and efficiency than CNN-based models. To achieve high inference efficiency on GPU-like devices, our RTFormer leverages GPU-Friendly Attention with linear complexity and discards the multi-head mechanism. Besides, we find that cross-resolution attention is more efficient to gather global context information for high-resolution branch by spreading the high level knowledge learned from low-resolution branch. Extensive experiments on mainstream benchmarks demonstrate the effectiveness of our proposed RTFormer, it achieves state-of-the-art on Cityscapes, CamVid and COCOStuff, and shows promising results on ADE20K.
Jian Wang 0066, Chenhui Gou, Qiman Wu, Haocheng Feng, Junyu Han, Errui Ding, Jingdong Wang 0001
NeurIPS7
2022 Delving into Sequential Patches for Deepfake Detection
abstract
Recent advances in face forgery techniques produce nearly visually untraceable deepfake videos, which could be leveraged with malicious intentions. As a result, researchers have been devoted to deepfake detection. Previous studies have identified the importance of local low-level cues and temporal information in pursuit to generalize well across deepfake methods, however, they still suffer from robustness problem against post-processings. In this work, we propose the Local- & Temporal-aware Transformer-based Deepfake Detection (LTTD) framework, which adopts a local-to-global learning protocol with a particular focus on the valuable temporal information within local sequences. Specifically, we propose a Local Sequence Transformer (LST), which models the temporal consistency on sequences of restricted spatial regions, where low-level information is hierarchically enhanced with shallow layers of learned 3D filters. Based on the local temporal embeddings, we then achieve the final classification in a global contrastive way. Extensive experiments on popular datasets validate that our approach effectively spots local forgery cues and achieves state-of-the-art performance.
Jiazhi Guan, Hang Zhou 0009, Zhibin Hong, Errui Ding, Jingdong Wang 0001, Chengbin Quan, Youjian Zhao
NeurIPS5
2022 Singular Value Fine-tuning: Few-shot Segmentation requires Few-parameters Fine-tuning
abstract
Freezing the pre-trained backbone has become a standard paradigm to avoid overfitting in few-shot segmentation. In this paper, we rethink the paradigm and explore a new regime: {\em fine-tuning a small part of parameters in the backbone}. We present a solution to overcome the overfitting problem, leading to better model generalization on learning novel classes. Our method decomposes backbone parameters into three successive matrices via the Singular Value Decomposition (SVD), then {\em only fine-tunes the singular values} and keeps others frozen. The above design allows the model to adjust feature representations on novel classes while maintaining semantic clues within the pre-trained backbone. We evaluate our {\em Singular Value Fine-tuning (SVF)} approach on various few-shot segmentation methods with different backbones. We achieve state-of-the-art results on both Pascal-5$^i$ and COCO-20$^i$ across 1-shot and 5-shot settings. Hopefully, this simple baseline will encourage researchers to rethink the role of backbone fine-tuning in few-shot settings.
Yanpeng Sun, Qiang Chen 0007, Jian Wang 0066, Haocheng Feng, Junyu Han, Errui Ding, Jian Cheng 0001, Zechao Li, Jingdong Wang 0001
NeurIPS10
2022 Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers
abstract
Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facial area, leading to non-realistic results. In this work, we delve into the formulation of altering only the mouth shapes of the target person. This requires masking a large percentage of the original image and seamlessly inpainting it with the aid of audio and reference frames. To this end, we propose the Audio-Visual Context-Aware Transformer (AV-CAT) framework, which produces accurate lip-sync with photo-realistic quality by predicting the masked mouth shapes. Our key insight is to exploit desired contextual information provided in audio and visual modalities thoroughly with delicately designed Transformers. Specifically, we propose a convolution-Transformer hybrid backbone and design an attention-based fusion strategy for filling the masked parts. It uniformly attends to the textural information on the unmasked regions and the reference frame. Then the semantic audio information is involved in enhancing the self-attention computation. Additionally, a refinement network with audio injection improves both image and lip-sync quality. Extensive experiments validate that our model can generate high-fidelity lip-synced results for arbitrary subjects.
Yasheng Sun, Hang Zhou 0009, Kaisiyuan Wang, Qianyi Wu, Zhibin Hong, Jingtuo Liu, Errui Ding, Jingdong Wang 0001, Ziwei Liu 0002, Hideki Koike
SIGGRAPH Asia8
2022 Few-Shot Image and Sentence Matching via Aligned Cross-Modal Memory
abstract
Image and sentence matching has attracted much attention recently, and many effective methods have been proposed to deal with it. But even the current state-of-the-arts still cannot well associate those challenging pairs of images and sentences containing few-shot content in their regions and words. In fact, such a few-shot matching problem is seldom studied and has become a bottleneck for further performance improvement in real-world applications. In this work, we formulate this challenging problem as few-shot image and sentence matching, and accordingly propose an Aligned Cross-Modal Memory (ACMM) model to deal with it. The model can not only softly align few-shot regions and words in a weakly-supervised manner, but also persistently store and update cross-modal prototypical representations of few-shot classes as references, without using any groundtruth region-word correspondence. The model can also adaptively balance the relative importance between few-shot and common content in the image and sentence, which leads to better measurement of overall similarity. We perform extensive experiments in terms of both few-shot and conventional image and sentence matching, and demonstrate the effectiveness of the proposed model by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Jingdong Wang 0001, Liang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Guest Editorial: Introduction to the Special Section on Fine-Grained Visual Categorization
abstract
This special section on fine-grained visual categorization has attracted many research works on fine-grain related topics. We thank all authors for submitting their papers to the special section and all reviewers who have provided professional, insightful, and timely reviews, leading to the high quality of accepted papers. We also thank TPAMI EIC Sven Dickinson and the Associate EICs for recognizing the widespread interest in this field, which warrants this special section. The accepted papers are divided into four groups based on their different focuses: Fine-grained image recognition, Fine-grained human analysis, Fine-grained video action recognition, and Fine-grained vision-language reasoning. We briefly review the accepted papers in each of the groups.
Jingdong Wang 0001, Zhuowen Tu, Jianlong Fu, Nicu Sebe, Serge J. Belongie
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Distillation-Guided Residual Learning for Binary Convolutional Neural Networks
abstract
It is challenging to bridge the performance gap between binary convolutional neural network (BCNN) and floating-point CNN (FCNN). This performance gap is mainly caused by the inferior modeling capability and training strategy of BCNN, which leads to substantial residuals in intermediate feature maps between BCNN and FCNN. To minimize the performance gap, we enforce BCNN to produce similar intermediate feature maps with the ones of FCNN. This intuition leads to a more effective training strategy for BCNN, i.e., optimizing each binary convolutional block with blockwise distillation loss derived from FCNN. The goal of minimizing the residuals in intermediate feature maps also motivates us to update the binary convolutional block architecture to facilitate the optimization of blockwise distillation loss. Specifically, a lightweight shortcut branch is inserted into each binary convolutional block to complement residuals at each block. Benefited from its squeeze-and-interaction (SI) structure, this shortcut branch introduces a fraction of parameters, e.g., less than 10% overheads, but effectively boosts the modeling capability of binary convolution blocks in BCNN. Extensive experiments on ImageNet demonstrate the superior performance of our method in both classification efficiency and accuracy, e.g., BCNN trained with our methods achieves the accuracy of 60.45% on ImageNet, better than many state-of-the-art ones.
Jianming Ye, Jingdong Wang 0001, Shiliang Zhang
IEEE Trans. Neural Networks Learn. Syst.2
2021 Boosting Adversarial Transferability through Enhanced Momentum
Xiaosen Wang, Jiadong Lin, Han Hu 0001, Jingdong Wang 0001, Kun He 0001
BMVC4
2021 Semi-Supervised Semantic Segmentation With Cross Pseudo Supervision
abstract
In this paper, we study the semi-supervised semantic segmentation problem via exploring both labeled data and extra unlabeled data. We propose a novel consistency regularization approach, called cross pseudo supervision (CPS). Our approach imposes the consistency on two segmentation networks perturbed with different initialization for the same input image. The pseudo one-hot label map, output from one perturbed segmentation network, is used to supervise the other segmentation network with the standard cross-entropy loss, and vice versa. The CPS consistency has two roles: encourage high similarity between the predictions of two perturbed networks for the same input image, and expand training data by using the unlabeled data with pseudo labels. Experiment results show that our approach achieves the state-of-the-art semi-supervised segmentation performance on Cityscapes and PASCAL VOC 2012.
Xiaokang Chen, Yuhui Yuan, Jingdong Wang 0001
CVPR4
2021 Bottom-Up Human Pose Estimation via Disentangled Keypoint Regression
abstract
In this paper, we are interested in the bottom-up paradigm of estimating human poses from an image. We study the dense keypoint regression framework that is previously inferior to the keypoint detection and grouping framework. Our motivation is that regressing keypoint positions accurately needs to learn representations that focus on the keypoint regions.We present a simple yet effective approach, named disentangled keypoint regression (DEKR). We adopt adaptive convolutions through pixel-wise spatial transformer to activate the pixels in the keypoint regions and accordingly learn representations from them. We use a multi-branch structure for separate regression: each branch learns a representation with dedicated adaptive convolutions and regresses one keypoint. The resulting disentangled representations are able to attend to the keypoint regions, respectively, and thus the keypoint regression is spatially more accurate. We empirically show that the proposed direct regression method outperforms keypoint detection and grouping methods and achieves superior bottom-up pose estimation results on two benchmark datasets, COCO and CrowdPose. The code and models are available at https://github.com/HRNet/DEKR.
Zigang Geng, Ke Sun 0009, Bin Xiao 0004, Zhaoxiang Zhang 0001, Jingdong Wang 0001
CVPR5
2021 Lite-HRNet: A Lightweight High-Resolution Network
abstract
We present an efficient high-resolution network, Lite-HRNet, for human pose estimation. We start by simply applying the efficient shuffle block in ShuffleNet to HRNet (high-resolution network), yielding stronger performance over popular lightweight networks, such as MobileNet, ShuffleNet, and Small HRNet. We find that the heavily-used pointwise (1 × 1) convolutions in shuffle blocks become the computational bottleneck. We introduce a lightweight unit, conditional channel weighting, to replace costly pointwise (1 × 1) convolutions in shuffle blocks. The complexity of channel weighting is linear w.r.t the number of channels and lower than the quadratic time complexity for pointwise convolutions. Our solution learns the weights from all the channels and over multiple resolutions that are readily available in the parallel branches in HRNet. It uses the weights as the bridge to exchange information across channels and resolutions, compensating the role played by the pointwise (1 × 1) convolution. Lite-HRNet demonstrates superior results on human pose estimation over popular lightweight networks. Moreover, Lite-HRNet can be easily applied to semantic segmentation task in the same lightweight manner. The code and models have been publicly available at https://github.com/HRNet/Lite-HRNet.
Changqian Yu, Bin Xiao 0004, Changxin Gao, Lu Yuan 0001, Lei Zhang 0001, Nong Sang, Jingdong Wang 0001
CVPR7
2021 Conditional DETR for Fast Training Convergence
abstract
The recently-developed DETR approach applies the transformer encoder and decoder architecture to object detection and achieves promising performance. In this paper, we handle the critical issue, slow training convergence, and present a conditional cross-attention mechanism for fast DETR training. Our approach is motivated by that the cross-attention in DETR relies highly on the content embeddings for localizing the four extremities and predicting the box, which increases the need for high-quality content embeddings and thus the training difficulty.Our approach, named conditional DETR, learns a conditional spatial query from the decoder embedding for decoder multi-head cross-attention. The benefit is that through the conditional spatial query, each cross-attention head is able to attend to a band containing a distinct region, e.g., one object extremity or a region inside the object box. This narrows down the spatial range for localizing the distinct regions for object classification and box regression, thus relaxing the dependence on the content embeddings and easing the training. Empirical results show that conditional DETR converges 6.7× faster for the backbones R50 and R101 and 10× faster for stronger backbones DC5-R50 and DC5-R101. Code is available at https://github.com/Atten4Vis/ConditionalDETR.
Depu Meng, Xiaokang Chen, Zejia Fan, Houqiang Li, Yuhui Yuan, Jingdong Wang 0001
ICCV8
2021 Admix: Enhancing the Transferability of Adversarial Attacks
abstract
Deep neural networks are known to be extremely vulnerable to adversarial examples under white-box setting. Moreover, the malicious adversaries crafted on the surrogate (source) model often exhibit black-box transferability on other models with the same learning task but having different architectures. Recently, various methods are proposed to boost the adversarial transferability, among which the input transformation is one of the most effective approaches. We investigate in this direction and observe that existing transformations are all applied on a single image, which might limit the adversarial transferability. To this end, we propose a new input transformation based attack method called Admix that considers the input image and a set of images randomly sampled from other categories. Instead of directly calculating the gradient on the original input, Admix calculates the gradient on the input image admixed with a small portion of each add-in image while using the original label of the input to craft more transferable adversaries. Empirical evaluations on standard ImageNet dataset demonstrate that Admix could achieve significantly better transferability than existing input transformation methods under both single model setting and ensemble-model setting. By incorporating with existing input transformations, our method could further improve the transferability and outperforms the state-of-the-art combination of input transformations by a clear margin when attacking nine advanced defense models under ensemble-model setting. Code is available at https://github.com/JHL-HUST/Admix.
Xiaosen Wang, Xuanran He, Jingdong Wang 0001, Kun He 0001
ICCV3
2021 Hybrid Network Compression via Meta-Learning
abstract
Neural network pruning and quantization are two major lines of network compression. This raises a natural question that whether we can find the optimal compression by considering multiple network compression criteria in a unified framework. This paper incorporates two criteria and seeks layer-wise compression by leveraging the meta-learning framework. A regularization loss is applied to unify the constraint of input and output channel numbers, bit-width of network activations and weights, so that the compressed network can satisfy a given Bit-OPerations counts (BOPs) constraint. We further propose an iterative compression constraint for optimizing the compression procedure, which effectively achieves a high compression rate and maintains the original network performance. Extensive experiments on various networks and vision tasks show that the proposed method yields better performance and compression rates than recent methods. For instance, our method achieves better image classification accuracy and compactness than the recent DJPQ. It achieves similar performance with the recent DHP in image super-resolution, meanwhile saves about 50% computation.
Jianming Ye, Shiliang Zhang, Jingdong Wang 0001
ACM Multimedia3
2021 SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood Search
abstract
The in-memory algorithms for approximate nearest neighbor search (ANNS) have achieved great success for fast high-recall search, but are extremely expensive when handling very large scale database. Thus, there is an increasing request for the hybrid ANNS solutions with small memory and inexpensive solid-state drive (SSD). In this paper, we present a simple but efficient memory-disk hybrid indexing and search system, named SPANN, that follows the inverted index methodology. It stores the centroid points of the posting lists in the memory and the large posting lists in the disk. We guarantee both disk-access efficiency (low latency) and high recall by effectively reducing the disk-access number and retrieving high-quality posting lists. In the index-building stage, we adopt a hierarchical balanced clustering algorithm to balance the length of posting lists and augment the posting list by adding the points in the closure of the corresponding clusters. In the search stage, we use a query-aware scheme to dynamically prune the access of unnecessary posting lists. Experiment results demonstrate that SPANN is 2X faster than the state-of-the-art ANNS solution DiskANN to reach the same recall quality 90% with same memory cost in three billion-scale datasets. It can reach 90% recall@1 and recall@10 in just around one millisecond with only about 10% of original memory cost. Code is available at: https://github.com/microsoft/SPTAG.
Qi Chen 0009, Mingqin Li, Chuanjie Liu, Zengzhong Li, Mao Yang 0004, Jingdong Wang 0001
NeurIPS8
2021 HRFormer: High-Resolution Vision Transformer for Dense Predict
abstract
We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that produces low-resolution representations and has high memory and computational cost. We take advantage of the multi-resolution parallel design introduced in high-resolution convolutional networks (HRNet [45]), along with local-window self-attention that performs self-attention over small non-overlapping image windows [21], for improving the memory and computation efficiency. In addition, we introduce a convolution into the FFN to exchange information across the disconnected image windows. We demonstrate the effectiveness of the HighResolution Transformer on both human pose estimation and semantic segmentation tasks, e.g., HRFormer outperforms Swin transformer [27] by 1.3 AP on COCO pose estimation with 50% fewer parameters and 30% fewer FLOPs. Code is available at: https://github.com/HRNet/HRFormer
Yuhui Yuan, Lang Huang 0001, Weihong Lin, Chao Zhang 0001, Xilin Chen 0001, Jingdong Wang 0001
NeurIPS7
2021 OCNet: Object Context for Semantic Segmentation
Yuhui Yuan, Lang Huang 0001, Jianyuan Guo, Chao Zhang 0001, Xilin Chen 0001, Jingdong Wang 0001
Int. J. Comput. Vis.6
2021 Content-aware convolutional neural networks
Yaofo Chen, Mingkui Tan, Kui Jia, Jian Chen 0011, Jingdong Wang 0001
Neural Networks6
2021 Deep High-Resolution Representation Learning for Visual Recognition
abstract
High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection. Existing state-of-the-art frameworks first encode the input image as a low-resolution representation through a subnetwork that is formed by connecting high-to-low resolution convolutions in series (e.g., ResNet, VGGNet), and then recover the high-resolution representation from the encoded low-resolution representation. Instead, our proposed network, named as High-Resolution Network (HRNet), maintains high-resolution representations through the whole process. There are two key characteristics: (i) Connect the high-to-low resolution convolution streams in parallel and (ii) repeatedly exchange the information across resolutions. The benefit is that the resulting representation is semantically richer and spatially more precise. We show the superiority of the proposed HRNet in a wide range of applications, including human pose estimation, semantic segmentation, and object detection, suggesting that the HRNet is a stronger backbone for computer vision problems. All the codes are available at https://github.com/HRNet.
Jingdong Wang 0001, Ke Sun 0009, Tianheng Cheng, Borui Jiang, Chaorui Deng, Yang Zhao 0019, Dong Liu 0002, Yadong Mu, Mingkui Tan, Xinggang Wang, Wenyu Liu 0001, Bin Xiao 0004
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Group Reidentification with Multigrained Matching and Integration
abstract
The task of reidentifying groups of people under different camera views is an important yet less-studied problem. Group reidentification (Re-ID) is a very challenging task since it is not only adversely affected by common issues in traditional single-object Re-ID problems, such as viewpoint and human pose variations, but also suffers from changes in group layout and group membership. In this paper, we propose a novel concept of group granularity by characterizing a group image by multigrained objects: individual people and subgroups of two and three people within a group. To achieve robust group Re-ID, we first introduce multigrained representations which can be extracted via the development of two separate schemes, that is, one with handcrafted descriptors and another with deep neural networks. The proposed representation seeks to characterize both appearance and spatial relations of multigrained objects, and is further equipped with importance weights which capture variations in intragroup dynamics. Optimal group-wise matching is facilitated by a multiorder matching process which, in turn, dynamically updates the importance weights in iterative fashion. We evaluated three multicamera group datasets containing complex scenarios and large dynamics, with experimental results demonstrating the effectiveness of our approach.
Weiyao Lin, Yuxi Li 0009, John See, Junni Zou, Hongkai Xiong, Jingdong Wang 0001, Tao Mei 0001
IEEE Trans. Cybern.7
2021 Learning to Segment Video Object With Accurate Boundaries
abstract
Video object segmentation has attracted considerable research interest these years. Top-performing video object segmentation methods mainly rely on fully convolutional neural networks which are specifically trained for predicting high-performance masks, resulting in a lack of preciseness in boundary details. This paper tackles the problem of predicting both mask-accurate and boundary-precise segmentation masks in videos. To solve this problem, we propose a simple and efficient network structure: the Mask-boundAry-Consistent Network (MAC-Net). TheMAC-Netis an end-to-end fully convolutional network, where both mask and boundaries are jointly optimized during training, enabling it to predict masks along with accurate boundaries. An inner-net boundary-computing module is incorporated in theMAC-Netfor producing spontaneously mask-consistent boundaries. We analyze the influence of parameter settings, network constructions of theMAC-Net, and compare with state-of-the-art algorithms on three widely-adopted datasets. Experimental results show that theMAC-Netachieves state-of-the-art performance, demonstrating the effectiveness of its mask-boundary-consistent network structure. We also propose that the boundary module inMAC-Nethas high compatibility, and can be easily adapted to other segmentation-related techniques.
Jingchun Cheng, Yuhui Yuan, Yali Li 0001, Jingdong Wang 0001, Shengjin Wang
IEEE Trans. Multim.4
2020 HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation
abstract
Bottom-up human pose estimation methods have difficulties in predicting the correct pose for small persons due to challenges in scale variation. In this paper, we present HigherHRNet: a novel bottom-up human pose estimation method for learning scale-aware representations using high-resolution feature pyramids. Equipped with multi-resolution supervision for training and multi-resolution aggregation for inference, the proposed approach is able to solve the scale variation challenge in bottom-up multi-person pose estimation and localize keypoints more precisely, especially for small person. The feature pyramid in HigherHRNet consists of feature map outputs from HRNet and upsampled higher-resolution outputs through a transposed convolution. HigherHRNet outperforms the previous best bottom-up method by 2.5% AP for medium person on COCO test-dev, showing its effectiveness in handling scale variation. Furthermore, HigherHRNet achieves new state-of-the-art result on COCO test-dev (70.5% AP) without using refinement or other post-processing techniques, surpassing all existing bottom-up methods. HigherHRNet even surpasses all top-down methods on CrowdPose test (67.6% AP), suggesting its robustness in crowded scene.
Bowen Cheng, Bin Xiao 0004, Jingdong Wang 0001, Humphrey Shi, Thomas S. Huang, Lei Zhang 0001
CVPR3
2020 Closed-Loop Matters: Dual Regression Networks for Single Image Super-Resolution
abstract
Deep neural networks have exhibited promising performance in image super-resolution (SR) by learning a nonlinear mapping function from low-resolution (LR) images to high-resolution (HR) images. However, there are two underlying limitations to existing SR methods. First, learning the mapping function from LR to HR images is typically an ill-posed problem, because there exist infinite HR images that can be downsampled to the same LR image. As a result, the space of the possible functions can be extremely large, which makes it hard to find a good solution. Second, the paired LR-HR data may be unavailable in real-world applications and the underlying degradation method is often unknown. For such a more general case, existing SR models often incur the adaptation problem and yield poor performance. To address the above issues, we propose a dual regression scheme by introducing an additional constraint on LR data to reduce the space of the possible functions. Specifically, besides the mapping from LR to HR images, we learn an additional dual regression mapping estimates the down-sampling kernel and reconstruct LR images, which forms a closed-loop to provide additional supervision. More critically, since the dual regression process does not depend on HR images, we can directly learn from LR images. In this sense, we can easily adapt SR models to real-world data, e.g., raw video frames from YouTube. Extensive experiments with paired training data and unpaired real-world data demonstrate our superiority over existing methods.
Jian Chen 0011, Jingdong Wang 0001, Qi Chen 0014, Jiezhang Cao, Zeshuai Deng, Yanwu Xu 0001, Mingkui Tan
CVPR3
2020 Weakly-Supervised Action Localization by Generative Attention Modeling
abstract
Weakly-supervised temporal action localization is a problem of learning an action localization model with only video-level action labeling available. The general framework largely relies on the classification activation, which employs an attention model to identify the action-related frames and then categorizes them into different classes. Such method results in the action-context confusion issue: context frames near action clips tend to be recognized as action frames themselves, since they are closely related to the specific classes. To solve the problem, in this paper we propose to model the class-agnostic frame-wise probability conditioned on the frame attention using conditional Variational Auto-Encoder (VAE). With the observation that the context exhibits notable difference from the action at representation level, a probabilistic model, i.e., conditional VAE, is learned to model the likelihood of each frame given the attention. By maximizing the conditional probability with respect to the attention, the action and non-action frames are well separated. Experiments on THUMOS14 and ActivityNet1.2 demonstrate advantage of our method and effectiveness in handling action-context confusion problem. Code is now available on GitHub.
Baifeng Shi, Qi Dai 0001, Yadong Mu, Jingdong Wang 0001
CVPR4
2020 Efficient Semantic Video Segmentation with Per-Frame Inference
Yifan Liu 0001, Chunhua Shen, Changqian Yu, Jingdong Wang 0001
ECCV (10)4
2020 Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
Fangyun Wei, Xiao Sun 0001, Hongyang Li 0001, Jingdong Wang 0001, Stephen Lin 0001
ECCV (10)4
2020 Object-Contextual Representations for Semantic Segmentation
Yuhui Yuan, Xilin Chen 0001, Jingdong Wang 0001
ECCV (6)3
2020 SegFix: Model-Agnostic Boundary Refinement for Segmentation
Yuhui Yuan, Xilin Chen 0001, Jingdong Wang 0001
ECCV (12)4
2020 Informative Dropout for Robust Representation Learning: A Shape-bias Perspective
abstract
Convolutional Neural Networks (CNNs) are known to rely more on local texture rather than global shape when making decisions. Recent work also indicates a close relationship between CNN’s texture-bias and its robustness against distribution shift, adversarial perturbation, random corruption, etc. In this work, we attempt at improving various kinds of robustness universally by alleviating CNN’s texture bias. With inspiration from the human visual system, we propose a light-weight model-agnostic method, namely Informative Dropout (InfoDrop), to improve interpretability and reduce texture bias. Specifically, we discriminate texture from shape based on local self-information in an image, and adopt a Dropout-like algorithm to decorrelate the model output from the local texture. Through extensive experiments, we observe enhanced robustness under various scenarios (domain generalization, few-shot classification, image corruption, and adversarial perturbation). To the best of our knowledge, this work is one of the earliest attempts to improve different kinds of robustness in a unified model, shedding new light on the relationship between shape-bias and robustness, also on new approaches to trustworthy machine learning algorithms. Code is available at https://github.com/bfshi/InfoDrop.
Baifeng Shi, Dinghuai Zhang, Qi Dai 0001, Zhanxing Zhu, Yadong Mu, Jingdong Wang 0001
ICML6
2020 S4Net: Single stage salient-instance segmentation
abstract
In this paper, we consider salient instance segmentation. As well as producing bounding boxes, our network also outputs high-quality instance-level segments as initial selections to indicate the regions of interest. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also the surrounding context, enabling us to distinguish instances in the same scope even with partial occlusion. Our network is end-to-end trainable and is fast (running at 40 fps for images with resolution 320 × 320). We evaluate our approach on a publicly available benchmark and show that it outperforms alternative solutions. We also provide a thorough analysis of our design choices to help readers better understand the function of each part of our network. Source code can be found at https://github.com/RuochenFan/S4Net.
Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001
Comput. Vis. Media5
2020 Object Detection in Videos by High Quality Object Linking
abstract
Compared with object detection in static images, object detection in videos is more challenging due to degraded image qualities. An effective way to address this problem is to exploit temporal contexts by linking the same object across video to form tubelets and aggregating classification scores in the tubelets. In this paper, we focus on obtaining high quality object linking results for better classification. Unlike previous methods that link objects by checking boxes between neighboring frames, we propose to link in the same frame. To achieve this goal, we extend prior methods in following aspects: (1) a cuboid proposal network that extracts spatio-temporal candidate cuboids which bound the movement of objects; (2) a short tubelet detection network that detects short tubelets in short video segments; (3) a short tubelet linking algorithm that links temporally-overlapping short tubelets to form long tubelets. Experiments on the ImageNet VID dataset show that our method outperforms both the static image detector and the previous state of the art. In particular, our method improves results by 8.8 percent over the static image detector for fast moving objects.
Peng Tang 0005, Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Wenjun Zeng 0001, Jingdong Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2020 Improving Person Re-Identification With Iterative Impression Aggregation
abstract
Our impression about one person often updates after we see more aspects of him/her and this process keeps iterating given more meetings. We formulate such an intuition into the problem of person re-identification (re-ID), where the representation of a query (probe) image is iteratively updated with new information from the candidates in the gallery. Specifically, we propose a simple attentional aggregation formulation to instantiate this idea and showcase that such a pipeline achieves competitive performance on standard benchmarks including CUHK03, Market-1501 and DukeMTMC. Not only does such a simple method improve the performance of the baseline models, it also achieves comparable performance with latest advanced re-ranking methods. Another advantage of this proposal is its flexibility to incorporate different representations and similarity metrics. By utilizing stronger representations and metrics, we further demonstrate state-of-the-art person re-ID performance, which also validates the general applicability of the proposed method.
Dengpan Fu, Bo Xin, Jingdong Wang 0001, Dongdong Chen 0001, Jianmin Bao, Gang Hua 0001, Houqiang Li
IEEE Trans. Image Process.3
2020 Semantic Image Segmentation by Scale-Adaptive Networks
abstract
Semantic image segmentation is an important yet unsolved problem. One of the major challenges is the large variability of the object scales. To tackle this scale problem, we propose a Scale-Adaptive Network (SAN) which consists of multiple branches with each one taking charge of the segmentation of the objects of a certain range of scales. Given an image, SAN first computes a dense scale map indicating the scale of each pixel which is automatically determined by the size of the enclosing object. Then the features of different branches are fused according to the scale map to generate the final segmentation map. To ensure that each branch indeed learns the features for a certain scale, we propose a scale-induced ground-truth map and enforce a scale-aware segmentation loss for the corresponding branch in addition to the final loss. Extensive experiments over the PASCAL-Person-Part, the PASCAL VOC 2012, and the Look into Person datasets demonstrate that our SAN can handle the large variability of the object scales and outperforms the state-of-the-art semantic segmentation methods.
Chunyu Wang 0001, Xinggang Wang, Wenyu Liu 0001, Jingdong Wang 0001
IEEE Trans. Image Process.5
2020 Guest Editorial Multimedia Computing With Interpretable Machine Learning
abstract
The papers in this special section is to broadly engage the machine learning and multimedia communities on the emerging yet challenging interpretable machine learning. Multimedia is increasingly becoming the “biggest big data,” among the most important and valuable source for insight and information. Many powerful machine learning algorithms, especially deep learning models such as convolutional neural networks (CNNs), have recently achieved outstanding predictive performance in a wide range of multimedia applications, including visual object classification, scene understanding, speech recognition, and activity prediction. Nevertheless, most deep learning algorithms are generally conceived as blackbox methods, and it is difficult to intuitively and quantitatively understand the results of their prediction and inference. Since this lack of interpretability is a major bottleneck in designing more successful predictive models and exploring wider-range useful applications, there has been an explosion of interest in interpreting the representations learned by these models, with profound implications for research into interpretable machine learning in the multimedia community.
Yonghong Tian 0001, Cees Snoek, Jingdong Wang 0001, Zhu Liu 0001, Rainer Lienhart, Susanne Boll
IEEE Trans. Multim.3
2020 Parameter Distribution Balanced CNNs
abstract
Convolutional neural network (CNN) is the primary technique that has greatly promoted the development of computer vision technologies. However, there is little research on how to allocate parameters in different convolution layers when designing CNNs. We research mainly on revealing the relationship between CNN parameter distribution, i.e., the allocation of parameters in convolution layers, and the discriminative performance of CNN. Unlike previous works, we do not append more elements into the network, such as more convolution layers or denser short connections. We focus on enhancing the discriminative performance of CNN through varying its parameter distribution under strict size constraint. We propose an energy function to represent the CNN parameter distribution, which establishes the connection between the allocation of parameters and the discriminative performance of CNN. Extensive experiments with shallow CNNs on three public image classification data sets demonstrate that the CNN parameter distribution with a higher energy value will promote the model to obtain better performance. According to the motivated observation, the problem of finding the optimal parameter distribution can be transformed into an optimization problem of finding the biggest energy value. We present a simple yet effective guideline that uses balanced parameter distribution to design CNNs. Extensive experiments on ImageNet with three popular backbones, i.e., AlexNet, ResNet34, and ResNet101, demonstrate that the proposed guideline can make consistent improvements upon different baselines under strict size constraint.
Lixin Liao, Yao Zhao 0001, Shikui Wei, Yunchao Wei, Jingdong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.5
2019 Deep High-Resolution Representation Learning for Human Pose Estimation
abstract
In this paper, we are interested in the human pose estimation problem with a focus on learning reliable high-resolution representations. Most existing methods recover high-resolution representations from low-resolution representations produced by a high-to-low resolution network. Instead, our proposed network maintains high-resolution representations through the whole process. We start from a high-resolution subnetwork as the first stage, gradually add high-to-low resolution subnetworks one by one to form more stages, and connect the mutli-resolution subnetworks in parallel. We conduct repeated multi-scale fusions such that each of the high-to-low resolution representations receives information from other parallel representations over and over, leading to rich high-resolution representations. As a result, the predicted keypoint heatmap is potentially more accurate and spatially more precise. We empirically demonstrate the effectiveness of our network through the superior pose estimation results over two benchmark datasets: the COCO keypoint detection dataset and the MPII Human Pose dataset. In addition, we show the superiority of our network in pose tracking on the PoseTrack dataset. The code and models have been publicly available at https://github.com/leoxiaobin/deep-high-resolution-net.pytorch.
Ke Sun 0009, Bin Xiao 0004, Dong Liu 0002, Jingdong Wang 0001
CVPR4
2019 S4Net: Single Stage Salient-Instance Segmentation
abstract
We consider an interesting problem---salient instance segmentation. Other than producing approximate bounding boxes, our network also outputs high-quality instance-level segments. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also its surrounding context, enabling us to distinguish the instances in the same scope even with obstruction. Our network is end-to-end trainable and runs at a fast speed (40 fps when processing an image with resolution 320 x 320). We evaluate our approach on a public available benchmark and show that it outperforms other alternative solutions. We also provide a thorough analysis of the design choices to help readers better understand the functions of each part of our network. The source code can be found at https://github.com/RuochenFan/S4Net.
Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001
CVPR5
2019 Structured Knowledge Distillation for Semantic Segmentation
abstract
In this paper, we investigate the issue of knowledge distillation for training compact semantic segmentation networks by making use of cumbersome networks. We start from the straightforward scheme, pixel-wise distillation, which applies the distillation scheme originally introduced for image classification and performs knowledge distillation for each pixel separately. We further propose to distill the structured knowledge from cumbersome networks into compact networks, which is motivated by the fact that semantic segmentation is a structured prediction problem. We study two such structured distillation schemes: (i) pair-wise distillation that distills the pairwise similarities, and (ii) holistic distillation that uses adversarial training to distill holistic knowledge. The effectiveness of our knowledge distillation approaches is demonstrated by extensive experiments on three scene parsing datasets: Cityscapes, Camvid and ADE20K.
Yifan Liu 0001, Chris Liu, Zengchang Qin, Zhenbo Luo, Jingdong Wang 0001
CVPR6
2019 Global-Local Temporal Representations for Video Person Re-Identification
abstract
This paper proposes the Global-Local Temporal Representation (GLTR) to exploit the multi-scale temporal cues in video sequences for video person Re-Identification (ReID). GLTR is constructed by first modeling the short-term temporal cues among adjacent frames, then capturing the long-term relations among inconsecutive frames. Specifically, the short-term temporal cues are modeled by parallel dilated convolutions with different temporal dilation rates to represent the motion and appearance of pedestrian. The long-term relations are captured by a temporal self-attention model to alleviate the occlusions and noises in video sequences. The short and long-term temporal cues are aggregated as the final GLTR by a simple single-stream CNN. GLTR shows substantial superiority to existing features learned with body part cues or metric learning on four widely-used video ReID datasets. For instance, it achieves Rank-1 Accuracy of 87.02% on MARS dataset without re-ranking, better than current state-of-the art.
Jianing Li 0001, Shiliang Zhang, Jingdong Wang 0001, Wen Gao 0001, Qi Tian 0001
ICCV3
2019 Cross View Fusion for 3D Human Pose Estimation
abstract
We present an approach to recover absolute 3D human poses from multi-view images by incorporating multi-view geometric priors in our model. It consists of two separate steps: (1) estimating the 2D poses in multi-view images and (2) recovering the 3D poses from the multi-view 2D poses. First, we introduce a cross-view fusion scheme into CNN to jointly estimate 2D poses for multiple views. Consequently, the 2D pose estimation for each view already benefits from other views. Second, we present a recursive Pictorial Structure Model to recover the 3D pose from the multi-view 2D poses. It gradually improves the accuracy of 3D pose with affordable computational cost. We test our method on two public datasets H36M and Total Capture. The Mean Per Joint Position Errors on the two datasets are 26mm and 29mm, which outperforms the state-of-the-arts remarkably (26mm vs 52mm, 29mm vs 35mm).
Haibo Qiu, Chunyu Wang 0001, Jingdong Wang 0001, Naiyan Wang, Wenjun Zeng 0001
ICCV3
2019 Disparity-preserved Deep Cross-platform Association for Cross-platform Video Recommendation
abstract
Cross-platform recommendation aims to improve recommendation accuracy through associating information from different platforms. Existing cross-platform recommendation approaches assume all cross-platform information to be consistent with each other and can be aligned. However, there remain two unsolved challenges: i) there exist inconsistencies in cross-platform association due to platform-specific disparity, and ii) data from distinct platforms may have different semantic granularities. In this paper, we propose a cross-platform association model for cross-platform video recommendation, i.e., Disparity-preserved Deep Cross-platform Association (DCA), taking platform-specific disparity and granularity difference into consideration. The proposed DCA model employs a partially-connected multi-modal autoencoder, which is capable of explicitly capturing platform-specific information, as well as utilizing nonlinear mapping functions to handle granularity differences. We then present a cross-platform video recommendation approach based on the proposed DCA model. Extensive experiments for our cross-platform recommendation framework on real-world dataset demonstrate that the proposed DCA model significantly outperform existing cross-platform recommendation methods in terms of various evaluation metrics.
Shengze Yu, Xin Wang 0019, Wenwu Zhu 0001, Peng Cui 0001, Jingdong Wang 0001
IJCAI5
2019 Joint salient object detection and existence prediction
Huaizu Jiang, Ming-Ming Cheng, Shijie Li 0006, Ali Borji, Jingdong Wang 0001
Frontiers Comput. Sci.5
2019 Ordinal Constraint Binary Coding for Approximate Nearest Neighbor Search
abstract
Binary code learning,a.k.a.hashing, has been successfully applied to the approximate nearest neighbor search in large-scale image collections. The key challenge lies in reducing the quantization error from the original real-valued feature space to a discrete Hamming space. Recent advances in unsupervised hashing advocate the preservation of ranking information, which is achieved by constraining the binary code learning to be correlated with pairwise similarity. However, few unsupervised methods consider the preservation ofordinal relationsin the learning process, which serves as a more basic cue to learn optimal binary codes. In this paper, we propose a novel hashing scheme, termed Ordinal Constraint Hashing (OCH), which embeds the ordinal relation among data points to preserve ranking into binary codes. The core idea is to construct an ordinal graph via tensor product, and then train the hash function over this graph to preserve the permutation relations among data points in the Hamming space. Subsequently, an in-depth acceleration scheme, termed Ordinal Constraint Projection (OCP), is introduced, which approximates the$n$-pair ordinal graph by$L$-pair anchor-based ordinal graph, and reduce the corresponding complexity from$O(n^4)$to$O(L^3)$($L\ll n$). Finally, to make the optimization tractable, we further relax the discrete constrains and design a customized stochastic gradient decent algorithm on the Stiefel manifold. Experimental results on serval large-scale benchmarks demonstrate that the proposed OCH method can achieve superior performance over the state-of-the-art approaches.
Hong Liu 0009, Rongrong Ji, Jingdong Wang 0001, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2019 Composite Quantization
abstract
This paper studies the compact coding approach to approximate nearest neighbor search. We introduce a composite quantization framework. It uses the composition of several ($M$M) elements, each of which is selected from a different dictionary, to accurately approximate a $D$D-dimensional vector, thus yielding accurate search, and represents the data vector by a short code composed of the indices of the selected elements in the corresponding dictionaries. Our key contribution lies in introducing a near-orthogonality constraint, which makes the search efficiency is guaranteed as the cost of the distance computation is reduced to $O(M)$O(M) from $O(D)$O(D) through a distance table lookup scheme. The resulting approach is called near-orthogonal composite quantization. We theoretically justify the equivalence between near-orthogonal composite quantization and minimizing an upper bound of a function formed by jointly considering the quantization error and the search cost according to a generalized triangle inequality. We empirically show the efficacy of the proposed approach over several benchmark datasets. In addition, we demonstrate the superior performances in other three applications: combination with inverted multi-index, inner-product similarity search, and query compression for mobile search.
Jingdong Wang 0001, Ting Zhang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2019 A Bilinear Ranking SVM for Knowledge Based Relation Prediction and Classification
abstract
As an important and challenging problem, knowledge representation and inference are typically carried out in a knowledge embedding framework over a multi-relational knowledge graph, and thus have a wide range of applications such as semantic retrieval and question answering. In this paper, we propose a bilinear learning framework which performs cross-entity knowledge relation analysis in the continuous vector space (derived from knowledge embedding). In the framework, we effectively model the intrinsic correlations among different types of knowledge relations within a max-margin multi-relational ranking scheme, which jointly optimizes the tasks of entity embedding and cross-entity relation prediction in terms of multi-relational structures of the knowledge graph. Specifically, we devise a bilinear scoring function that aims to evaluate the confidence degree of semantic relation prediction for entity pairs through a multi-relational learning-to-rank pipeline. In essence, the pipeline formulates the problem of relation prediction for entity pairs as that of learning relation-specific ranking functions by max-margin optimization. Experimental results demonstrate the effectiveness of the proposed framework on two common benchmark datasets.
Shengkang Yu, Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Jingdong Wang 0001, Yueting Zhuang, Xuelong Li 0001
IEEE Trans. Big Data6
2019 Automatic Ensemble Diffusion for 3D Shape and Image Retrieval
abstract
As a post-processing procedure, the diffusion process has demonstrated its ability of substantially improving the performance of various visual retrieval systems. Whereas, great efforts are also devoted to similarity (or metric) fusion, seeing that only one individual type of similarity cannot fully reveal the intrinsic relationship between objects. This stimulates a great research interest of considering similarity fusion in the framework of the diffusion process (i.e., fusion with diffusion) for robust retrieval. In this paper, we first revisit representative methods about fusion with diffusion and provide new insights which are ignored by previous researchers. Then, observing that existing algorithms are susceptible to noisy similarities, the proposed regularized ensemble diffusion (RED) is bundled with an automatic weight learning paradigm, so that the negative impacts of noisy similarities are suppressed. Though formulated as a convex optimization problem, one advantage of RED is that it converts back into the iteration-based solver with the same computational complexity as the conventional diffusion process. At last, we integrate several recently-proposed similarities with the proposed framework. The experimental results suggest that we can achieve new state-of-the-art performances on various retrieval tasks, including 3D shape retrieval on the ModelNet data set, and image retrieval on the Holidays and Ukbench data sets.
Song Bai 0001, Jingdong Wang 0001, Xiang Bai, Longin Jan Latecki, Qi Tian 0001
IEEE Trans. Image Process.3
2019 Learning Attentional Recurrent Neural Network for Visual Tracking
abstract
Existing visual tracking methods face many challenges: 1) the changed size and number of targets over time, occlusion in discrete frames, and mis-identification for crossing targets. Long short-term memory (LSTM) has the advantage of modeling long-term tasks and is suitable for tracking. We propose a novel online attentional recurrent neural network (ARNN) model for visual tracking, whose core component is a two-layer bidirectional LSTM along the x-and y-axes. Several bidirectional LSTMs can be cascaded or parallelly connected together to exploit multiscale target features and can give more precise tracked object locations. Each bidirectional LSTM utilizes the convolutional features of a convolutional neural network inside two bounding boxes from two frames to check whether the target in the current frame is the one in previous frames. An attention mechanism is also adopted to enhance the proposed model to better express the patch-level features of the tracking targets. Interattention and intra-attention models are proposed to imitate the temporal and spatial tracking mechanism of primate visual cortex. Interattention learns to overcome the occlusion problem, and intra-attention is able to mark important regions to better trace the target. The bidirectional LSTM and the attention mechanism are jointly trained. The combination of them further improves the accuracy of target tracking in videos. The outstanding performances in the experiments demonstrate the effectiveness of our proposed online method ARNN and yield competitive results compared with the state-of-the-art tracking methods.
Qiurui Wang, Chun Yuan 0003, Jingdong Wang 0001, Wenjun Zeng 0001
IEEE Trans. Multim.3
2019 Balanced Decoupled Spatial Convolution for CNNs
abstract
In this paper, we are interested in designing lightweight CNNs by decoupling the convolution along the spatial and channel dimension. Most existing decoupling techniques focus on approximating the filter matrix through decomposition. In contrast, we provide a decoupled view of the standard convolution to separate the spatial information and the channel information. The resulting decoupled process is exactly equivalent to the standard convolution. Inspired from our decoupled view, we propose an effective structure, balanced decoupled spatial convolution (BDSC), to relax the sparsity of the filter in spatial aggregation by learning a spatial configuration and reduce the redundancy by reducing the number of intermediate channels. We also designed an adaptive spatial configuration, which is simply adding a nonlinear activation layer [rectified linear units (ReLU)] after the intermediate output. Our experiments verify that the adaptive spatial configuration can improve the classification performance without extra cost. In addition, our BDSC achieves comparable classification performance with the standard convolution but with a smaller model size on Canadian Institute for Advanced Research (CIFAR)-100, CIFAR-10, and ImageNet. To show the potential of further reducing the redundancy of across channel-domain convolution, we also show experiments of our models with a designed lightweight across channel-domain convolution. Finally, we show in our experiments that our models achieve superior performance than the state-of-the-art models.
Guotian Xie, Kuiyuan Yang, Ting Zhang 0002, Jingdong Wang 0001, Jian-Huang Lai
IEEE Trans. Neural Networks Learn. Syst.4
2018 Decoupled Convolutions for CNNs
abstract
In this paper, we are interested in designing small CNNs by decoupling the convolution along the spatial and channel domains. Most existing decoupling techniques focus on approximating the filter matrix through decomposition. In contrast, we provide a two-step interpretation of the standard convolution from the filter at a single location to all locations, which is exactly equivalent to the standard convolution. Motivated by the observations in our decoupling view, we propose an effective approach to relax the sparsity of the filter in spatial aggregation by learning a spatial configuration, and reduce the redundancy by reducing the number of intermediate channels. Our approach achieves comparable classification performance with the standard uncoupled convolution, but with a smaller model size over CIFAR-100, CIFAR-10 and ImageNet.
Guotian Xie, Ting Zhang 0002, Kuiyuan Yang, Jian-Huang Lai, Jingdong Wang 0001
AAAI5
2018 IGCV3: Interleaved Low-Rank Group Convolutions for Efficient Deep Neural Networks
Ke Sun 0009, Mingjie Li 0007, Dong Liu 0002, Jingdong Wang 0001
BMVC4
2018 Weakly-Supervised Semantic Segmentation Network With Deep Seeded Region Growing
abstract
This paper studies the problem of learning image semantic segmentation networks only using image-level labels as supervision, which is important since it can significantly reduce human annotation efforts. Recent state-of-the-art methods on this problem first infer the sparse and discriminative regions for each object class using a deep classification network, then train semantic a segmentation network using the discriminative regions as supervision. Inspired by the traditional image segmentation methods of seeded region growing, we propose to train a semantic segmentation network starting from the discriminative regions and progressively increase the pixel-level supervision using by seeded region growing. The seeded region growing module is integrated in a deep segmentation network and can benefit from deep features. Different from conventional deep networks which have fixed/static labels, the proposed weakly-supervised network generates new labels using the contextual information within an image. The proposed method significantly outperforms the weakly-supervised semantic segmentation methods using static labels, and obtains the state-of-the-art performance, which are 63.2% mIoU score on the PASCAL VOC 2012 test set and 26.0% mIoU score on the COCO dataset.
Xinggang Wang, Jiasi Wang, Wenyu Liu 0001, Jingdong Wang 0001
CVPR5
2018 Global Versus Localized Generative Adversarial Nets
abstract
In this paper, we present a novel localized Generative Adversarial Net (GAN) to learn on the manifold of real data. Compared with the classic GAN that globally parameterizes a manifold, the Localized GAN (LGAN) uses local coordinate charts to parameterize distinct local geometry of how data points can transform at different locations on the manifold. Specifically, around each point there exists a local generator that can produce data following diverse patterns of transformations on the manifold. The locality nature of LGAN enables local generators to adapt to and directly access the local geometry without need to invert the generator in a global GAN. Furthermore, it can prevent the manifold from being locally collapsed to a dimensionally deficient tangent subspace by imposing an orthonormality prior between tangents. This provides a geometric approach to alleviating mode collapse at least locally on the manifold by imposing independence between data transformations in different tangent directions. We will also demonstrate the LGAN can be applied to train a robust classifier that prefers locally consistent classification decisions on the manifold, and the resultant regularizer is closely related with the Laplace-Beltrami operator. Our experiments show that the proposed LGANs can not only produce diverse image transformations, but also deliver superior classification performances.
Guo-Jun Qi, Liheng Zhang, Hao Hu 0010, Marzieh Edraki, Jingdong Wang 0001, Xian-Sheng Hua 0001
CVPR5
2018 Interleaved Structured Sparse Convolutional Neural Networks
abstract
In this paper, we study the problem of designing efficient convolutional neural network architectures with the interest in eliminating the redundancy in convolution kernels. In addition to structured sparse kernels, low-rank kernels and the product of low-rank kernels, the product of structured sparse kernels, which is a framework for interpreting the recently-developed interleaved group convolutions (IGC) and its variants (e.g., Xception), has been attracting increasing interests. Motivated by the observation that the convolutions contained in a group convolution in IGC can be further decomposed in the same manner, we present a modularized building block, IGC-V2: interleaved structured sparse convolutions. It generalizes interleaved group convolutions, which is composed of two structured sparse kernels, to the product of more structured sparse kernels, further eliminating the redundancy. We present the complementary condition and the balance condition to guide the design of structured sparse kernels, obtaining a balance among three aspects: model size, computation complexity and classification accuracy. Experimental results demonstrate the advantage on the balance among these three aspects compared to interleaved group convolutions and Xception, and competitive performance compared to other state-of-the-art architecture design methods.
Guotian Xie, Jingdong Wang 0001, Ting Zhang 0002, Jian-Huang Lai, Richang Hong, Guo-Jun Qi
CVPR2
2018 Part-Aligned Bilinear Representations for Person Re-identification
Yumin Suh, Jingdong Wang 0001, Siyu Tang 0001, Tao Mei 0001, Kyoung Mu Lee
ECCV (14)2
2018 Rethinking ReLU to Train Better CNNs
abstract
Most of convolutional neural networks share the same characteristic: each convolutional layer is followed by a nonlinear activation layer where Rectified Linear Unit (ReLU) is the most widely used. In this paper, we argue that the designed structure with the equal ratio between these two layers may not be the best choice since it could result in the poor generalization ability. Thus, we try to investigate a more suitable method on using ReL U to explore the better network architectures. Specifically, we propose a proportional module to keep the ratio between convolution and ReLU amount to be N:m (n>m). The proportional module can be applied in almost all networks with no extra computational cost to improve the performance. Comprehensive experimental results indicate that the proposed method achieves better performance on different benchmarks with different network architectures, thus verify the superiority of our work.
Gangming Zhao, Zhaoxiang Zhang 0001, He Guan, Peng Tang 0005, Jingdong Wang 0001
ICPR5
2018 Deep Convolutional Neural Networks with Merge-and-Run Mappings
abstract
A deep residual network, built by stacking a sequence of residual blocks, is easy to train, because identity mappings skip residual branches and thus improve information flow. To further reduce the training difficulty, we present a simple network architecture, deep merge-and-run neural networks. The novelty lies in a modularized building block, merge-and-run block, which assembles residual branches in parallel through a merge-and-run mapping: average the inputs of these residual branches (Merge), and add the average to the output of each residual branch as the input of the subsequent residual branch (Run), respectively. We show that the merge-and-run mapping is a linear idempotent function in which the transformation matrix is idempotent, and thus improves information flow, making training easy. In comparison with residual networks, our networks enjoy compelling advantages: they contain much shorter paths and the width, i.e., the number of channels, is increased, and the time complexity remains unchanged. We evaluate the performance on the standard recognition tasks. Our approach demonstrates consistent improvements over ResNets with the comparable setup, and achieves competitive results (e.g., 3.06% testing error on CIFAR-10, 17.55% on CIFAR-100, 1.51% on SVHN).
Mingjie Li 0007, Depu Meng, Xi Li 0001, Zhaoxiang Zhang 0001, Yueting Zhuang, Zhuowen Tu, Jingdong Wang 0001
IJCAI8
2018 Deep Triplet Quantization
abstract
Deep hashing establishes efficient and effective image retrieval by end-to-end learning of deep representations and hash codes from similarity data. We present a compact coding solution, focusing on deep learning to quantization approach that has shown superior performance over hashing solutions for similarity retrieval. We propose Deep Triplet Quantization (DTQ), a novel approach to learning deep quantization models from the similarity triplets. To enable more effective triplet training, we design a new triplet selection approach, Group Hard, that randomly selects hard triplets in each image group. To generate compact binary codes, we further apply a triplet quantization with weak orthogonality during triplet training. The quantization loss reduces the codebook redundancy and enhances the quantizability of deep representations through back-propagation. Extensive experiments demonstrate that DTQ can generate high-quality and compact binary codes, which yields state-of-the-art image retrieval performance on three benchmark datasets, NUS-WIDE, CIFAR-10, and MS-COCO.
Bin Liu 0035, Yue Cao 0001, Mingsheng Long, Jianmin Wang 0001, Jingdong Wang 0001
ACM Multimedia5
2018 Group Re-Identification: Leveraging and Integrating Multi-Grain Information
abstract
This paper addresses an important yet less-studied problem: re-identifying groups of people in different camera views. Group re-identification (Re-ID) is very challenging since it is not only interfered by view-point and human pose variations in the traditional single-object Re-ID tasks, but also suffers from group layout and group member variations. To handle these issues, we propose to leverage the information of multi-grain objects: individual person and subgroups of two and three people inside a group image. We compute multi-grain representations to characterize the appearance and spatial features of multi-grain objects and evaluate the importance weight of each object for group Re-ID, so as to handle the interferences from group dynamics. We compute the optimal group-wise matching by using a multi-order matching process based on the multi-grain representation and importance weights. Furthermore, we dynamically update the importance weights according to the current matching results and then compute a new optimal group-wise matching. The two steps are iteratively conducted, yielding the final matching results.Experimental results on various datasets demonstrate the effectiveness of our approach.
Weiyao Lin, Bin Sheng 0001, Ke Lu 0002, Junchi Yan, Jingdong Wang 0001, Errui Ding, Hongkai Xiong
ACM Multimedia6
2018 Weakly Supervised Dense Event Captioning in Videos
abstract
Dense event captioning aims to detect and describe all events of interest contained in a video. Despite the advanced development in this area, existing methods tackle this task by making use of dense temporal annotations, which is dramatically source-consuming. This paper formulates a new problem: weakly supervised dense event captioning, which does not require temporal segment annotations for model training. Our solution is based on the one-to-one correspondence assumption, each caption describes one temporal segment, and each temporal segment has one caption, which holds in current benchmark datasets and most real world cases. We decompose the problem into a pair of dual problems: event captioning and sentence localization and present a cycle system to train our model. Extensive experimental results are provided to demonstrate the ability of our model on both dense event captioning and sentence localization in videos.
Xuguang Duan, Wenbing Huang 0001, Chuang Gan 0001, Jingdong Wang 0001, Wenwu Zhu 0001, Junzhou Huang
NeurIPS4
2018 Multi-Dimensional Sparse Models
abstract
Traditional synthesis/analysis sparse representation models signals in a one dimensional (1D) way, in which a multidimensional (MD) signal is converted into a 1D vector. 1D modeling cannot sufficiently handle MD signals of high dimensionality in limited computational resources and memory usage, as breaking the data structure and inherently ignores the diversity of MD signals (tensors). We utilize the multilinearity of tensors to establish the redundant basis of the space of multi linear maps with the sparsity constraint, and further propose MD synthesis/analysis sparse models to effectively and efficiently represent MD signals in their original form. The dimensional features of MD signals are captured by a series of dictionaries simultaneously and collaboratively. The corresponding dictionary learning algorithms and unified MD signal restoration formulations are proposed. The effectiveness of the proposed models and dictionary learning algorithms is demonstrated through experiments on MD signals denoising, image super-resolution and texture classification. Experiments show that the proposed MD models outperform state-of-the-art 1D models in terms of signal representation quality, computational overhead, and memory storage. Moreover, our proposed MD sparse models generalize the 1D sparse models and are flexible and adaptive to both homogeneous and inhomogeneous properties of MD signals.
Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Jingdong Wang 0001, Junbin Gao
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 A Survey on Learning to Hash
abstract
Nearest neighbor search is a problem of finding the data points from the database such that the distances from them to the query point are the smallest. Learning to hash is one of the major solutions to this problem and has been widely studied recently. In this paper, we present a comprehensive survey of the learning to hash algorithms, categorize them according to the manners of preserving the similarities into: pairwise similarity preserving, multiwise similarity preserving, implicit similarity preserving, as well as quantization, and discuss their relations. We separate quantization from pairwise similarity preserving as the objective function is very different though quantization, as we show, can be derived from preserving the pairwise similarities. In addition, we present the evaluation protocols, and the general performance analysis, and point out that the quantization algorithms perform superiorly in terms of search accuracy, search time cost, and space cost. Finally, we introduce a few emerging topics.
Jingdong Wang 0001, Ting Zhang 0002, Jingkuan Song, Nicu Sebe, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Face Alignment With Deep Regression
abstract
In this paper, we present a deep regression approach for face alignment. The deep regressor is a neural network that consists of a global layer and multistage local layers. The global layer estimates the initial face shape from the whole image, while the following local layers iteratively update the shape with local image observations. Combining standard derivations and numerical approximations, we make all layers able to backpropagate error differentials, so that we can apply the standard backpropagation to jointly learn the parameters from all layers. We show that the resulting deep regressor gradually and evenly approaches the true facial landmarks stage by stage, avoiding the tendency that often occurs in the cascaded regression methods and deteriorates the overall performance: yielding early stage regressors with high alignment accuracy gains but later stage regressors with low alignment accuracy gains. Experimental results on standard benchmarks demonstrate that our approach brings significant improvements over previous cascaded regression algorithms.
Baoguang Shi, Xiang Bai, Wenyu Liu 0001, Jingdong Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2017 Ensemble Diffusion for Retrieval
abstract
As a postprocessing procedure, diffusion process has demonstrated its ability of substantially improving the performance of various visual retrieval systems. Whereas, great efforts are also devoted to similarity (or metric) fusion, seeing that only one individual type of similarity cannot fully reveal the intrinsic relationship between objects. This stimulates a great research interest of considering similarity fusion in the framework of diffusion process (i.e., fusion with diffusion) for robust retrieval. In this paper, we firstly revisit representative methods about fusion with diffusion, and provide new insights which are ignored by previous researchers. Then, observing that existing algorithms are susceptible to noisy similarities, the proposed Regularized Ensemble Diffusion (RED) is bundled with an automatic weight learning paradigm, so that the negative impacts of noisy similarities are suppressed. At last, we integrate several recently-proposed similarities with the proposed framework. The experimental results suggest that we can achieve new state-of-the-art performances on various retrieval tasks, including 3D shape retrieval on ModelNet dataset, and image retrieval on Holidays and Ukbench dataset.
Song Bai 0001, Jingdong Wang 0001, Xiang Bai, Longin Jan Latecki, Qi Tian 0001
ICCV3
2017 Human Pose Estimation Using Global and Local Normalization
abstract
In this paper, we address the problem of estimating the positions of human joints, i.e., articulated pose estimation. Recent state-of-the-art solutions model two key issues, joint detection and spatial configuration refinement, together using convolutional neural networks. Our work mainly focuses on spatial configuration refinement by reducing variations of human poses statistically, which is motivated by the observation that the scattered distribution of the relative locations of joints (e.g., the left wrist is distributed nearly uniformly in a circular area around the left shoulder) makes the learning of convolutional spatial models hard. We present a two-stage normalization scheme, human body normalization and limb normalization, to make the distribution of the relative joint locations compact, resulting in easier learning of convolutional spatial models and more accurate pose estimation. In addition, our empirical results show that incorporating multi-scale supervision and multi-scale fusion into the joint detection network is beneficial. Experiment results demonstrate that our method consistently outperforms state-of-the-art methods on the benchmarks.
Ke Sun 0009, Cuiling Lan, Junliang Xing, Wenjun Zeng 0001, Dong Liu 0002, Jingdong Wang 0001
ICCV6
2017 Interleaved Group Convolutions
abstract
In this paper, we present a simple and modularized neural network architecture, named interleaved group convolutional neural networks (IGCNets). The main point lies in a novel building block, a pair of two successive interleaved group convolutions: primary group convolution and secondary group convolution. The two group convolutions are complementary: (i) the convolution on each partition in primary group convolution is a spatial convolution, while on each partition in secondary group convolution, the convolution is a point-wise convolution; (ii) the channels in the same secondary partition come from different primary partitions. We discuss one representative advantage: Wider than a regular convolution with the number of parameters and the computation complexity preserved. We also show that regular convolutions, group convolution with summation fusion, and the Xception block are special cases of interleaved group convolutions. Empirical results over standard benchmarks, CIFAR-10, CIFAR-100, SVHN and ImageNet demonstrate that our networks are more efficient in using parameters and computation complexity with similar or higher accuracy.
Ting Zhang 0002, Guo-Jun Qi, Bin Xiao 0001, Jingdong Wang 0001
ICCV4
2017 Deeply-Learned Part-Aligned Representations for Person Re-identification
abstract
In this paper, we address the problem of person re-identification, which refers to associating the persons captured from different cameras. We propose a simple yet effective human part-aligned representation for handling the body part misalignment problem. Our approach decomposes the human body into regions (parts) which are discriminative for person matching, accordingly computes the representations over the regions, and aggregates the similarities computed between the corresponding regions of a pair of probe and gallery images as the overall matching score. Our formulation, inspired by attention models, is a deep neural network modeling the three steps together, which is learnt through minimizing the triplet loss function without requiring body part labeling information. Unlike most existing deep learning algorithms that learn a global or spatial partition-based local representation, our approach performs human body partition, and thus is more robust to pose changes and various human spatial distributions in the person bounding box. Our approach shows state-of-the-art results over standard datasets, Market-1501, CUHK03, CUHK01 and VIPeR.
Xi Li 0001, Yueting Zhuang, Jingdong Wang 0001
ICCV4
2017 Random Shifting for CNN: a Solution to Reduce Information Loss in Down-Sampling Layers
abstract
Down-sampling is widely adopted in deep convolutional neural networks (DCNN) for reducing the number of network parameters while preserving the transformation invariance. However, it cannot utilize information effectively because it only adopts a fixed stride strategy, which may result in poor generalization ability and information loss. In this paper, we propose a novel random strategy to alleviate these problems by embedding random shifting in the down-sampling layers during the training process. Random shifting can be universally applied to diverse DCNN models to dynamically adjust receptive fields by shifting kernel centers on feature maps in different directions. Thus, it can generate more robust features in networks and further enhance the transformation invariance of down-sampling operators. In addition, random shifting cannot only be integrated in all down-sampling layers including strided convolutional layers and pooling layers, but also improve performance of DCNN with negligible additional computational cost. We evaluate our method in different tasks (e.g., image classification and segmentation) with various network architectures (i.e., AlexNet, FCN and DFN-MR). Experimental results demonstrate the effectiveness of our proposed method.
Gangming Zhao, Jingdong Wang 0001, Zhaoxiang Zhang 0001
IJCAI2
2017 Mixture Factorized Ornstein-Uhlenbeck Processes for Time-Series Forecasting
abstract
Forecasting the future observations of time-series data can be performed by modeling the trend and fluctuations from the observed data. Many classical time-series analysis models like Autoregressive model (AR) and its variants have been developed to achieve such forecasting ability. While they are often based on the white noise assumption to model the data fluctuations, a more general Brownian motion has been adopted that results in Ornstein-Uhlenbeck (OU) process. The OU process has gained huge successes in predicting the future observations over many genres of time series, however, it is still limited in modeling simple diffusion dynamics driven by a single persistent factor that never evolves over time. However, in many real problems, a mixture of hidden factors are usually present, and when and how frequently they appear or disappear are unknown ahead of time. This imposes a challenge that inspires us to develop a Mixture Factorized OU process (MFOUP) to model evolving factors. The new model is able to capture the changing states of multiple mixed hidden factors, from which we can infer their roles in driving the movements of time series. We conduct experiments on three forecasting problems, covering sensor and market data streams. The results show its competitive performance on predicting future observations and capturing evolution patterns of hidden factors as compared with the other algorithms.
Guo-Jun Qi, Jiliang Tang, Jingdong Wang 0001, Jiebo Luo 0001
KDD3
2017 Finding the Secret of CNN Parameter Layout under Strict Size Constraint
abstract
Although deep convolutional neural networks (CNNs) have significantly boosted the performance of many computer vision tasks, their complexities~(the size or the number of parameters) are also dramatically increased even with slight performance improvement. However, the larger network leads to more computation requirements, which are unfavorable to resource-constrained scenarios, such as the widely used embedded systems. In this paper, we tentatively explore the essential effect of CNN parameter layout, ıe, the allocation of parameters in the convolution layers, on the discriminative capability of CNN. Instead of enlarging the breadth or depth of networks, we attempt to improve the discriminative ability of CNN by changing its parameter layout under strict size constraint. Toward this end, a novel energy function is proposed to represent the CNN parameter layout, which makes it possible to model the relationship between the allocation of parameters in the convolution layers and the discriminative ability of CNN. According to extensive experimental results with plain CNN models and Residual Nets, we find that the higher the energy of a specific CNN parameter layout is, the better its discriminative ability is. Following this finding, we propose a novel approach to learn the better parameter layout. Experimental results on two public image classification datasets show that the CNN models with the learned parameter layouts achieve the better image classification results under strict size constraint.
Lixin Liao, Yao Zhao 0001, Shikui Wei, Jingdong Wang 0001, Ruoyu Liu
ACM Multimedia4
2017 Exemplar-Guided Similarity Learning on Polynomial Kernel Feature Map for Person Re-identification
Dapeng Chen, Zejian Yuan, Jingdong Wang 0001, Badong Chen, Gang Hua 0001, Nanning Zheng 0001
Int. J. Comput. Vis.3
2017 Salient Object Detection: A Discriminative Regional Feature Integration Approach
Jingdong Wang 0001, Huaizu Jiang, Zejian Yuan, Ming-Ming Cheng, Xiaowei Hu 0003, Nanning Zheng 0001
Int. J. Comput. Vis.1
2017 Towards Reversal-Invariant Image Representation
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001
Int. J. Comput. Vis.2
2017 Special issue on intelligent urban computing with big data
Wu Liu 0005, Peng Cui 0001, Jukka K. Nurminen, Jingdong Wang 0001
Mach. Vis. Appl.4
2017 Multi-Timescale Collaborative Tracking
abstract
We present the multi-timescale collaborative tracker for single object tracking. The tracker simultaneously utilizes different types of "forces", namely attraction, repulsion and support, to take advantage of their complementary strengths. We model the three forces via three components that are learned from the sample sets with different timescales. The long-term descriptive component attracts the target sample, while the medium-term discriminative component repulses the target from the background. They are collaborated in the appearance model to benefit each other. The short-term regressive component combines the votes of the auxiliary samples to predict the target's position, forming the context-aware motion model. The appearance model and the motion model collaboratively determine the target state, and the optimal state is estimated by a novel coarse-to-fine search strategy. We have conducted an extensive set of experiments on the standard 50 video benchmark. The results confirm the effectiveness of each component and their collaboration, outperforming current state-of-the-art methods.
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Jingdong Wang 0001, Nanning Zheng 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Guest Editorial: Big Media Data: Understanding, Search, and Mining
abstract
The paper in this special section focus on big data in the media industry. Big media data is a new research area, and has been attracting a lot of research interests in both industry and academia. This editorial, as the fourth part of this special issue, introduces four papers on fine-grained image classification, object retrieval, and applications of social media.
Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal
IEEE Trans. Big Data1
2017 Learning Correspondence Structures for Person Re-Identification
abstract
This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure, which indicates the patchwise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global constraint-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Finally, we also extend our approach by introducing a multi-structure scheme, which learns a set of local correspondence structures to capture the spatial correspondence sub-patterns between a camera pair, so as to handle the spatial misalignments between individual images in a more precise way. Experimental results on various data sets demonstrate the effectiveness of our approach.
Weiyao Lin, Junchi Yan, Mingliang Xu 0001, Jianxin Wu 0001, Jingdong Wang 0001, Ke Lu 0002
IEEE Trans. Image Process.6
2016 Supervised Quantization for Similarity Search
abstract
In this paper, we address the problem of searching for semantically similar images from a large database. We present a compact coding approach, supervised quantization. Our approach simultaneously learns feature selection that linearly transforms the database points into a low-dimensional discriminative subspace, and quantizes the data points in the transformed space. The optimization criterion is that the quantized points not only approximate the transformed points accurately, but also are semantically separable: the points belonging to a class lie in a cluster that is not overlapped with other clusters corresponding to other classes, which is formulated as a classification problem. The experiments on several standard datasets show the superiority of our approach over the state-of-the art supervised hashing and unsupervised quantization algorithms.
Ting Zhang 0002, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001
CVPR5
2016 DisturbLabel: Regularizing CNN on the Loss Layer
abstract
During a long period of time we are combating overfitting in the CNN training process with model regularization, including weight decay, model averaging, data augmentation, etc. In this paper, we present DisturbLabel, an extremely simple algorithm which randomly replaces a part of labels as incorrect values in each iteration. Although it seems weird to intentionally generate incorrect training labels, we show that DisturbLabel prevents the network training from over-fitting by implicitly averaging over exponentially many networks which are trained with different label sets. To the best of our knowledge, DisturbLabel serves as the first work which adds noises on the loss layer. Meanwhile, DisturbLabel cooperates well with Dropout to provide complementary regularization functions. Experiments demonstrate competitive recognition results on several popular image recognition datasets.
Lingxi Xie, Jingdong Wang 0001, Meng Wang 0001, Qi Tian 0001
CVPR2
2016 InterActive: Inter-Layer Activeness Propagation
abstract
An increasing number of computer vision tasks can be tackled with deep features, which are the intermediate outputs of a pre-trained Convolutional Neural Network. Despite the astonishing performance, deep features extracted from low-level neurons are still below satisfaction, arguably because they cannot access the spatial context contained in the higher layers. In this paper, we present InterActive, a novel algorithm which computes the activeness of neurons and network connections. Activeness is propagated through a neural network in a top-down manner, carrying highlevel context and improving the descriptive power of lowlevel and mid-level neurons. Visualization indicates that neuron activeness can be interpreted as spatial-weighted neuron responses. We achieve state-of-the-art classification performance on a wide range of image datasets.
Lingxi Xie, Liang Zheng 0001, Jingdong Wang 0001, Alan L. Yuille, Qi Tian 0001
CVPR3
2016 Collaborative Quantization for Cross-Modal Similarity Search
abstract
Cross-modal similarity search is a problem about designing a search system supporting querying across content modalities, e.g., using an image to search for texts or using a text to search for images. This paper presents a compact coding solution for efficient search, with a focus on the quantization approach which has already shown the superior performance over the hashing solutions in the single-modal similarity search. We propose a cross-modal quantization approach, which is among the early attempts to introduce quantization into cross-modal search. The major contribution lies in jointly learning the quantizers for both modalities through aligning the quantized representations for each pair of image and text belonging to a document. In addition, our approach simultaneously learns the common space for both modalities in which quantization is conducted to enable efficient and effective search using the Euclidean distance computed in the common space with fast distance table lookup. Experimental results compared with several competitive algorithms over three benchmark datasets demonstrate that the proposed approach achieves the state-of-the-art performance.
Ting Zhang 0002, Jingdong Wang 0001
CVPR2
2016 Geometric Neural Phrase Pooling: Modeling the Spatial Co-occurrence of Neurons
Lingxi Xie, Qi Tian 0001, John Flynn, Jingdong Wang 0001, Alan L. Yuille
ECCV (1)4
2016 MARS: A Video Benchmark for Large-Scale Person Re-Identification
Liang Zheng 0001, Zhi Bie, Yifan Sun 0003, Jingdong Wang 0001, Chi Su, Shengjin Wang, Qi Tian 0001
ECCV (6)4
2016 Binary Optimized Hashing
abstract
This paper studies the problem of learning to hash, which is essentially a mixed integer optimization problem, containing both the binary hash code output and the (continuous) parameters forming the hash functions. Different from existing relaxation methods in hashing, which have no theoretical guarantees for the error bound of the relaxations, we propose binary optimized hashing (BOH), in which we prove that if the loss function is Lipschitz continuous, the binary optimization problem can be relaxed to a bound-constrained continuous optimization problem. Then we introduce a surrogate objective function, which only depends on unbinarized hash functions and does not need the slack variables transforming unbinarized hash functions to discrete functions, to approximate the relaxed objective function. We show that the approximation error is bounded and the bound is small when the problem is optimized. We apply the proposed approach to learn hash codes from either handcraft feature inputs or raw image inputs. Extensive experiments are carried out on three benchmarks, demonstrating that our approach outperforms state-of-the-arts with a significant margin on search accuracies.
Qi Dai 0001, Jingdong Wang 0001, Yu-Gang Jiang 0001
ACM Multimedia3
2016 Fast Nearest Neighbor Search in the Hamming Space
Zhansheng Jiang, Lingxi Xie, Xiaotie Deng, Weiwei Xu 0003, Jingdong Wang 0001
MMM (1)5
2016 Self-Paced Cross-Modal Subspace Matching
abstract
Cross-modal matching methods match data from different modalities according to their similarities. Most existing methods utilize label information to reduce the semantic gap between different modalities. However, it is usually time-consuming to manually label large-scale data. This paper proposes a Self-Paced Cross-Modal Subspace Matching (SCSM) method for unsupervised multimodal data. We assume that multimodal data are pair-wised and from several semantic groups, which form hard pair-wised constraints and soft semantic group constraints respectively. Then, we formulate the unsupervised cross-modal matching problem as a non-convex joint feature learning and data grouping problem. Self-paced learning, which learns samples from 'easy' to 'complex', is further introduced to refine the grouping result. Moreover, a multimodal graph is constructed to preserve the relationship of both inter- and intra-modality similarity. An alternating minimization method is employed to minimize the non-convex optimization problem, followed by the discussion on its convergence analysis and computational complexity. Experimental results on four multimodal databases show that SCSM outperforms state-of-the-art cross-modal subspace learning methods.
Jian Liang 0001, Zhihang Li, Ran He 0001, Jingdong Wang 0001
SIGIR5
2016 Detection of Co-salient Objects by Looking Deep and Wide
Dingwen Zhang, Junwei Han 0001, Chao Li 0028, Jingdong Wang 0001, Xuelong Li 0001
Int. J. Comput. Vis.4
2016 Accurate Image Search with Multi-Scale Contextual Evidences
Liang Zheng 0001, Shengjin Wang, Jingdong Wang 0001, Qi Tian 0001
Int. J. Comput. Vis.3
2016 Incorporating visual adjectives for image classification
Lingxi Xie, Jingdong Wang 0001, Bo Zhang 0010, Qi Tian 0001
Neurocomputing2
2016 Guest Editorial: Big Media Data: Understanding, Search, and Mining
abstract
The two papers in this special section focus on big media data. This is a new research area, and has been attracting a lot of research interests in both industry and academia. This editorial, as the third part of this special issue, introduces two papers on video copy detection data and personalized travel sequence recommendation from social media.
Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal
IEEE Trans. Big Data1
2016 DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection
abstract
A key problem in salient object detection is how to effectively model the semantic properties of salient objects in a data-driven manner. In this paper, we propose a multi-task deep saliency model based on a fully convolutional neural network with global input (whole raw images) and global output (whole saliency maps). In principle, the proposed saliency model takes a data-driven strategy for encoding the underlying saliency prior information, and then sets up a multi-task learning scheme for exploring the intrinsic correlations between saliency detection and semantic image segmentation. Through collaborative feature learning from such two correlated tasks, the shared fully convolutional layers produce effective features for object perception. Moreover, it is capable of capturing the semantic information on salient objects across different levels using the fully convolutional layers, which investigate the feature-sharing properties of salient object detection with a great reduction of feature redundancy. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Xi Li 0001, Lina Wei, Ming-Hsuan Yang 0001, Fei Wu 0001, Yueting Zhuang, Haibin Ling, Jingdong Wang 0001
IEEE Trans. Image Process.8
2016 Joint Multilabel Classification With Community-Aware Label Graph Learning
abstract
As an important and challenging problem in machine learning and computer vision, multilabel classification is typically implemented in a max-margin multilabel learning framework, where the inter-label separability is characterized by the sample-specific classification margins between labels. However, the conventional multilabel classification approaches are usually incapable of effectively exploring the intrinsic inter-label correlations as well as jointly modeling the interactions between inter-label correlations and multilabel classification. To address this issue, we propose a multilabel classification framework based on a joint learning approach called label graph learning (LGL) driven weighted Support Vector Machine (SVM). In principle, the joint learning approach explicitly models the inter-label correlations by LGL, which is jointly optimized with multilabel classification in a unified learning scheme. As a result, the learned label correlation graph well fits the multilabel classification task while effectively reflecting the underlying topological structures among labels. Moreover, the inter-label interactions are also influenced by label-specific sample communities (each community for the samples sharing a common label). Namely, if two labels have similar label-specific sample communities, they are likely to be correlated. Based on this observation, LGL is further regularized by the label Hypergraph Laplacian. Experimental results have demonstrated the effectiveness of our approach over several benchmark data sets.
Xi Li 0001, Xueyi Zhao, Zhongfei Zhang, Fei Wu 0001, Yueting Zhuang, Jingdong Wang 0001, Xuelong Li 0001
IEEE Trans. Image Process.6
2016 A Diffusion and Clustering-Based Approach for Finding Coherent Motions and Understanding Crowd Scenes
abstract
This paper addresses the problem of detecting coherent motions in crowd scenes and presents its two applications in crowd scene understanding: semantic region detection and recurrent activity mining. It processes input motion fields (e.g., optical flow fields) and produces a coherent motion field named thermal energy field. The thermal energy field is able to capture both motion correlation among particles and the motion trends of individual particles, which are helpful to discover coherency among them. We further introduce a two-step clustering process to construct stable semantic regions from the extracted time-varying coherent motions. These semantic regions can be used to recognize pre-defined activities in crowd scenes. Finally, we introduce a cluster-and-merge process, which automatically discovers recurrent activities in crowd scenes by clustering and merging the extracted coherent motions. Experiments on various videos demonstrate the effectiveness of our approach.
Weiyao Lin, Yang Mi, Weiyue Wang 0002, Jianxin Wu 0001, Jingdong Wang 0001, Tao Mei 0001
IEEE Trans. Image Process.5
2016 Weakly Supervised Metric Learning for Traffic Sign Recognition in a LIDAR-Equipped Vehicle
abstract
We address the problem of traffic sign recognition in a light detection and ranging (LIDAR)-equipped vehicle. With the help of 3-D LIDAR points, the 2-D multiview sign images will be easily detected from the captured images of street signs. After detection, the sign recognition problem is formulated as a multiview object recognition task. We develop a metric-learning-based template matching approach for this task and learn a distance metric between the captured images and the corresponding sign templates. For each sign, recognition is done via soft voting by the recognition results of its corresponding multiview images. We propose a latent structural support vector machine (SVM)-based weakly supervised metric learning (WSMLR) method to learn the metric and a reliability classifier. The reliability classifier is used to determine each image's reliability, which serves as each image's weight in both the learning and soft voting procedure. We evaluate the proposed method for multiview traffic sign recognition on a multiview traffic sign data set with 112 categories and observe very encouraging results compared with other state-of-the-art methods. In addition, the method can be customized to solve the single-view sign recognition. The performance of our method for single-view sign recognition is tested on two public data sets, showing that our method is comparable with other competitive ones.
Baoyuan Wang, Zhaohui Wu 0001, Jingdong Wang 0001, Gang Pan 0001
IEEE Trans. Intell. Transp. Syst.4
2016 A Distance-Computation-Free Search Scheme for Binary Code Databases
abstract
Recently, binary codes have been widely used in many multimedia applications to approximate high-dimensional multimedia features for practical similarity search due to the highly compact data representation and efficient distance computation. While the majority of the hashing methods aim at learning more accurate hash codes, only a few of them focus on indexing methods to accelerate the search for binary code databases. Among these indexing methods, most of them suffer from extremely high memory cost or extensive Hamming distance computations. In this paper, we propose a new Hamming distance search scheme for large scale binary code databases, which is free of Hamming distance computations to return the exact results. Without the necessity to compare database binary codes with queries, the search performance can be improved and databases can be externally maintained. More specifically, we adopt the inverted multi-index data structure to index binary codes. Importantly, the Hamming distance information embedded in the structure is utilized in the designed search scheme such that the verification of exact results no longer relies on Hamming distance computations. As a step further, we optimize the performance of the inverted multi-index structure by taking the code distributions among different bits into account for index construction. Empirical results on large-scale binary code databases demonstrate the superiority of our method over existing approaches in terms of both memory usage and search efficiency.
Jingkuan Song, Heng Tao Shen, Zi Huang, Nicu Sebe, Jingdong Wang 0001
IEEE Trans. Multim.6
2016 Dual Low-Rank Pursuit: Learning Salient Features for Saliency Detection
abstract
Saliency detection is an important procedure for machines to understand visual world as humans do. In this paper, we consider a specific saliency detection problem of predicting human eye fixations when they freely view natural images, and propose a novel dual low-rank pursuit (DLRP) method. DLRP learns saliency-aware feature transformations by utilizing available supervision information and constructs discriminative bases for effectively detecting human fixation points under the popular low-rank and sparsity-pursuit framework. Benefiting from the embedded high-level information in the supervised learning process, DLRP is able to predict fixations accurately without performing the expensive object segmentation as in the previous works. Comprehensive experiments clearly show the superiority of the proposed DLRP method over the established state-of-the-art methods. We also empirically demonstrate that DLRP provides stronger generalization performance across different data sets and inherits the advantages of both the bottom-up- and top-down-based saliency detection methods.
Congyan Lang, Jiashi Feng, Songhe Feng, Jingdong Wang 0001, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.4
2016 Generalized Deep Transfer Networks for Knowledge Propagation in Heterogeneous Domains
abstract
In recent years, deep neural networks have been successfully applied to model visual concepts and have achieved competitive performance on many tasks. Despite their impressive performance, traditional deep networks are subjected to the decayed performance under the condition of lacking sufficient training data. This problem becomes extremely severe for deep networks trained on a very small dataset, making them overfitting by capturing nonessential or noisy information in the training set. Toward this end, we propose a novel generalized deep transfer networks (DTNs), capable of transferring label information across heterogeneous domains, textual domain to visual domain. The proposed framework has the ability to adequately mitigate the problem of insufficient training images by bringing in rich labels from the textual domain. Specifically, to share the labels between two domains, we build parameter- and representation-shared layers. They are able to generate domain-specific and shared interdomain features, making this architecture flexible and powerful in capturing complex information from different domains jointly. To evaluate the proposed method, we release a new dataset extended from NUS-WIDE at http://imag.njust.edu.cn/NUS-WIDE-128.html. Experimental results on this dataset show the superior performance of the proposed DTNs compared to existing state-of-the-art methods.
Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Guo-Jun Qi, Jingdong Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2015 Similarity learning on an explicit polynomial kernel feature map for person re-identification
abstract
In this paper, we address the person re-identification problem, discovering the correct matches for a probe person image from a set of gallery person images. We follow the learning-to-rank methodology and learn a similarity function to maximize the difference between the similarity scores of matched and unmatched images for a same person. We introduce at least three contributions to person re-identification. First, we present an explicit polynomial kernel feature map, which is capable of characterizing the similarity information of all pairs of patches between two images, called soft-patch-matching, instead of greedily keeping only the best matched patch, and thus more robust. Second, we introduce a mixture of linear similarity functions that is able to discover different soft-patch-matching patterns. Last, we introduce a negative semi-definite regularization over a subset of the weights in the similarity function, which is motivated by the connection between explicit polynomial kernel feature map and the Mahalanobis distance, as well as the sparsity constraint over the parameters to avoid over-fitting. Experimental results over three public benchmarks demonstrate the superiority of our approach.
Dapeng Chen, Zejian Yuan, Gang Hua 0001, Nanning Zheng 0001, Jingdong Wang 0001
CVPR5
2015 Co-saliency detection via looking deep and wide
abstract
With the goal of effectively identifying common and salient objects in a group of relevant images, co-saliency detection has become essential for many applications such as video foreground extraction, surveillance, image retrieval, and image annotation. In this paper, we propose a unified co-saliency detection framework by introducing two novel insights: 1) looking deep to transfer higher-level representations by using the convolutional neural network with additional adaptive layers could better reflect the properties of the co-salient objects, especially their consistency among the image group; 2) looking wide to take advantage of the visually similar neighbors beyond a certain image group could effectively suppress the influence of the common background regions when formulating the intra-group consistency. In the proposed framework, the wide and deep information are explored for the object proposal windows extracted in each image, and the co-saliency scores are calculated by integrating the intra-image contrast and intra-group consistency via a principled Bayesian formulation. Finally the window-level co-saliency scores are converted to the superpixel-level co-saliency maps through a foreground region agreement strategy. Comprehensive experiments on two benchmark datasets have demonstrated the consistent performance gain of the proposed approach.
Dingwen Zhang, Junwei Han 0001, Chao Li 0028, Jingdong Wang 0001
CVPR4
2015 Sparse composite quantization
abstract
The quantization techniques have shown competitive performance in approximate nearest neighbor search. The state-of-the-art algorithm, composite quantization, takes advantage of the compositionabity, i.e., the vector approximation accuracy, as opposed to product quantization and Cartesian k-means. However, we have observed that the runtime cost of computing the distance table in composite quantization, which is used as a lookup table for fast distance computation, becomes nonnegligible in real applications, e.g., reordering the candidates retrieved from the inverted index when handling very large scale databases. To address this problem, we develop a novel approach, called sparse composite quantization, which constructs sparse dictionaries. The benefit is that the distance evaluation between the query and the dictionary element (a sparse vector) is accelerated using the efficient sparse vector operation, and thus the cost of distance table computation is reduced a lot. Experiment results on large scale ANN retrieval tasks (1M SIFTs and 1B SIFTs) and applications to object retrieval show that the proposed approach yields competitive performance: superior search accuracy to product quantization and Cartesian k-means with almost the same computing cost, and much faster ANN search than composite quantization with the same level of accuracy.
Ting Zhang 0002, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001
CVPR4
2015 Person Re-Identification with Correspondence Structure Learning
abstract
This paper addresses the problem of handling spatial misalignments due to camera-view changes or human-pose variations in person re-identification. We first introduce a boosting-based approach to learn a correspondence structure which indicates the patch-wise matching probabilities between images from a target camera pair. The learned correspondence structure can not only capture the spatial correspondence pattern between cameras but also handle the viewpoint or human-pose variation in individual images. We further introduce a global-based matching process. It integrates a global matching constraint over the learned correspondence structure to exclude cross-view misalignments during the image patch matching process, hence achieving a more reliable matching score between images. Experimental results on various datasets demonstrate the effectiveness of our approach.
Weiyao Lin, Junchi Yan, Jianxin Wu 0001, Jingdong Wang 0001
ICCV6
2015 RIDE: Reversal Invariant Descriptor Enhancement
abstract
In many fine-grained object recognition datasets, image orientation (left/right) might vary from sample to sample. Since handcrafted descriptors such as SIFT are not reversal invariant, the stability of image representation based on them is consequently limited. A popular solution is to augment the datasets by adding a left-right reversed copy for each original image. This strategy improves recognition accuracy to some extent, but also brings the price of almost doubled time and memory consumptions. In this paper, we present RIDE (Reversal Invariant Descriptor Enhancement) for fine-grained object recognition. RIDE is a generalized algorithm which cancels out the impact of image reversal by estimating the orientation of local descriptors, and guarantees to produce the identical representation for an image and its left-right reversed copy. Experimental results reveal the consistent accuracy gain of RIDE with various types of descriptors. We also provide insightful discussions on the working mechanism of RIDE and its generalization to other applications.
Lingxi Xie, Jingdong Wang 0001, Weiyao Lin, Bo Zhang 0010, Qi Tian 0001
ICCV2
2015 Scalable Person Re-identification: A Benchmark
abstract
This paper contributes a new high quality dataset for person re-identification, named "Market-1501". Generally, current datasets: 1) are limited in scale, 2) consist of hand-drawn bboxes, which are unavailable under realistic settings, 3) have only one ground truth and one query image for each identity (close environment). To tackle these problems, the proposed Market-1501 dataset is featured in three aspects. First, it contains over 32,000 annotated bboxes, plus a distractor set of over 500K images, making it the largest person re-id dataset to date. Second, images in Market-1501 dataset are produced using the Deformable Part Model (DPM) as pedestrian detector. Third, our dataset is collected in an open system, where each identity has multiple images under each camera. As a minor contribution, inspired by recent advances in large-scale image search, this paper proposes an unsupervised Bag-of-Words descriptor. We view person re-identification as a special task of image search. In experiment, we show that the proposed descriptor yields competitive accuracy on VIPeR, CUHK03, and Market-1501 datasets, and is scalable on the large-scale 500k dataset.
Liang Zheng 0001, Liyue Shen, Shengjin Wang, Jingdong Wang 0001, Qi Tian 0001
ICCV5
2015 Quantized Correlation Hashing for Fast Cross-Modal Search
Botong Wu, Qiang Yang 0010, Wei-Shi Zheng 0001, Yizhou Wang 0001, Jingdong Wang 0001
IJCAI5
2015 Weakly-Shared Deep Transfer Networks for Heterogeneous-Domain Knowledge Propagation
abstract
In recent years, deep networks have been successfully applied to model image concepts and achieved competitive performance on many data sets. In spite of impressive performance, the conventional deep networks can be subjected to the decayed performance if we have insufficient training examples. This problem becomes extremely severe for deep networks with powerful representation structure, making them prone to over fitting by capturing nonessential or noisy information in a small data set. In this paper, to address this challenge, we will develop a novel deep network structure, capable of transferring labeling information across heterogeneous domains, especially from text domain to image domain. This weakly-shared Deep Transfer Networks (DTNs) can adequately mitigate the problem of insufficient image training data by bringing in rich labels from the text domain.
Xiangbo Shu, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001
ACM Multimedia4
2015 Deep kinship verification
abstract
To improve the performance of kinship verification, we propose a novel deep kinship verification (DKV) model by integrating excellent deep learning architecture into metric learning. Unlike most existing shallow models based on metric learning for kinship verification, we employ a deep learning model followed by a metric learning formulation to select nonlinear features, which can find the appropriate project space to ensure the margin of negative sample pairs (i.e. parent and child without kinship relation) as large as possible and the margin of positive sample pairs (i.e. parent and child with kinship relation) as small as possible. Experimental results show that our method achieves satisfactory performance on two widely-used benchmarks, i.e. KFW-I and KFW-II.
Mengyin Wang, Zechao Li, Xiangbo Shu, Jingdong Wang 0001, Jinhui Tang 0001
MMSP4
2015 Guest Editorial: Ad Hoc Web Multimedia Analysis with Limited Supervision
Yahong Han, Yi Yang 0001, Jingdong Wang 0001
Multim. Tools Appl.3
2015 Guest Editorial: Big Media Data: Understanding, Search, and Mining
abstract
The papers in this special section focus on the topic of Big Data. There has been an explosive increase in media data, such as images, videos and social media in the Internet, mobile devices, and desktops. Accordingly, the area of big media data has been attracting a lot of research interest from both academia and industry. Big media data offers a promising possibility for in-depth media understanding, as well as exploring the very big scale media data to bridge the well-known semantic gap between high-level semantics and low-level features. It provides richer information, ranging from social relations to context information associated to rich media data of diverse modalities. It also provides us the opportunity to mine reliable and helpful knowledge from big media data for a wide variety of applications.
Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal
IEEE Trans. Big Data1
2015 Guest Editorial: Big Media Data: Understanding, Search, and Mining (Part 2)
abstract
The articles in this special section focus on Big Data. In the first part of this special issue, we have introduced three papers on large scale similar image search, image search quality improvement, and semi-supervised multi-label image annotation. This second part of this special issue includes two examples on large scale visual retrieval. The contributions cover two themes: visual semantic relationship learning and large scale indexing.
Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal
IEEE Trans. Big Data1
2015 Exploratory Product Image Search With Circle-to-Search Interaction
abstract
Exploratory search is emerging as a new form of information-seeking activity in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this paper, we investigate the challenges of understanding users' search interests from the product images being browsed and inferring their actual search intentions. We propose a novel interactive image exploring system for allowing users to lightly switch between browse and search processes, and naturally complete visual-based exploratory search tasks in an effective and efficient way. This system enables users to specify their visual search interests in product images by circling any visual objects in web pages, and then the system automatically infers users' underlying intent by analyzing the browsing context and by analyzing the same or similar product images obtained by large-scale image search technology. Users can then utilize the recommended queries to complete intent-specific exploratory tasks. The proposed solution is one of the first attempts to understand users' interests for a visual-based exploratory product search task by integrating the browse and search activities. We have evaluated our system performance based on five million product images. The evaluation study demonstrates that the proposed system provides accurate intent-driven search results and fast response to exploratory search demands compared with the conventional image search methods, and also, provides users with robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.3
2015 Optimized Cartesian K-Means
abstract
Product quantization-based approaches are effective to encode high-dimensional data points for approximate nearest neighbor search. The space is decomposed into a Cartesian product of low-dimensional subspaces, each of which generates a sub codebook. Data points are encoded as compact binary codes using these sub codebooks, and the distance between two data points can be approximated efficiently from their codes by the precomputed lookup tables. Traditionally, to encode a subvector of a data point in a subspace, only one sub codeword in the corresponding sub codebook is selected, which may impose strict restrictions on the search accuracy. In this paper, we propose a novel approach, named optimized cartesian K-means (ock-means), to better encode the data points for more accurate approximate nearest neighbor search. In ock-means, multiple sub codewords are used to encode the subvector of a data point in a subspace. Each sub codeword stems from different sub codebooks in each subspace, which are optimally generated with regards to the minimization of the distortion errors. The high-dimensional data point is then encoded as the concatenation of the indices of multiple sub codewords from all the subspaces. This can provide more flexibility and lower distortion errors than traditional methods. Experimental results on the standard real-life data sets demonstrate the superiority over state-of-the-art approaches for approximate nearest neighbor search.
Jingdong Wang 0001, Jingkuan Song, Xin-Shun Xu, Heng Tao Shen, Shipeng Li 0001
IEEE Trans. Knowl. Data Eng.2
2015 Fine-Grained Image Search
abstract
Large-scale image search has been attracting lots of attention from both academic and commercial fields. The conventional bag-of-visual-words (BoVW) model with inverted index is verified efficient at retrieving near-duplicate images, but it is less capable of discovering fine-grained concepts in the query and returning semantically matched search results. In this paper, we suggest that instance search should return not only near-duplicate images, but also fine-grained results, which is usually the actual intention of a user. We propose a new and interesting problem named fine-grained image search, which means that we prefer those images containing the same fine-grained concept with the query. We formulate the problem by constructing a hierarchical database and defining an evaluation method. We thereafter introduce a baseline system using fine-grained classification scores to represent and co-index images so that the semantic attributes are better incorporated in the online querying stage. Large-scale experiments reveal that promising search results are achieved with reasonable time and memory consumption. We hope this paper will be the foundation for future work on image search. We also expect more follow-up efforts along this research topic and look forward to commercial fine-grained image search engines.
Lingxi Xie, Jingdong Wang 0001, Bo Zhang 0010, Qi Tian 0001
IEEE Trans. Multim.2
2014 How Fashion Talks: Clothing-Region-Based Gender Recognition
Shengnan Cai, Jingdong Wang 0001, Long Quan
CIARP2
2014 Orientational Pyramid Matching for Recognizing Indoor Scenes
abstract
Scene recognition is a basic task towards image understanding. Spatial Pyramid Matching (SPM) has been shown to be an efficient solution for spatial context modeling. In this paper, we introduce an alternative approach, Orientational Pyramid Matching (OPM), for orientational context modeling. Our approach is motivated by the observation that the 3D orientations of objects are a crucial factor to discriminate indoor scenes. The novelty lies in that OPM uses the 3D orientations to form the pyramid and produce the pooling regions, which is unlike SPM that uses the spatial positions to form the pyramid. Experimental results on challenging scene classification tasks show that OPM achieves the performance comparable with SPM and that OPM and SPM make complementary contributions so that their combination gives the state-of-the-art performance.
Lingxi Xie, Jingdong Wang 0001, Baining Guo, Bo Zhang 0010, Qi Tian 0001
CVPR2
2014 Finding Coherent Motions and Semantic Regions in Crowd Scenes: A Diffusion and Clustering Approach
Weiyue Wang 0002, Weiyao Lin, Yuanzhe Chen, Jianxin Wu 0001, Jingdong Wang 0001, Bin Sheng 0001
ECCV (1)5
2014 Low-rank SIFT: An affine invariant feature for place recognition
abstract
In this paper, we study the problem of recognizing man-made objects and present a novel affine-invariant feature, Low-rank SIFT, which exploits the regular appearance property in man-made objects. The proposed feature achieves full affine invariance without needing to simulate over affine parameter space. We rectify local patches by converting them to their low-rank forms to achieve skew invariance, and perform the way similar to conventional SIFT to resolve rotation, translation and scaling ambiguity. The main contributions lie in two-fold: our method seeks to leverage low-rank prior to estimate affine parameters for local patches directly and we propose a fast algorithm to compute such parameters by introducing the Low-rank Integral Map. Besides, we describe a pipeline of constructing a geotagged building database from the ground up. We demonstrate the effectiveness of our approach in the application to place recognition.
Harry Yang, Shengnan Cai, Jingdong Wang 0001, Long Quan
ICIP3
2014 Composite Quantization for Approximate Nearest Neighbor Search
abstract
This paper presents a novel compact coding approach, composite quantization, for approximate nearest neighbor search. The idea is to use the composition of several elements selected from the dictionaries to accurately approximate a vector and to represent the vector by a short code composed of the indices of the selected elements. To efficiently compute the approximate distance of a query to a database vector using the short code, we introduce an extra constraint, constant inter-dictionary-element-product, resulting in that approximating the distance only using the distance of the query to each selected element is enough for nearest neighbor search. Experimental comparison with state-of-the-art algorithms over several benchmark datasets demonstrates the efficacy of the proposed approach.
Ting Zhang 0002, Jingdong Wang 0001
ICML3
2014 Optimized Distances for Binary Code Ranking
abstract
Binary encoding on high-dimensional data points has attracted much attention due to its computational and storage efficiency. While numerous efforts have been made to encode data points into binary codes, how to calculate the effective distance on binary codes to approximate the original distance is rarely addressed. In this paper, we propose an effective distance measurement for binary code ranking. In our approach, the binary code is firstly decomposed into multiple sub codes, each of which generates a query-dependent distance lookup table. Then the distance between the query and the binary code is constructed as the aggregation of the distances from all sub codes by looking up their respective tables. The entries of the lookup tables are optimized by minimizing the misalignment between the approximate distance and the original distance. Such a scheme is applied to both the symmetric distance and the asymmetric distance. Extensive experimental results show superior performance of the proposed approach over state-of-the-art methods on three real-world high-dimensional datasets for binary code ranking.
Heng Tao Shen, Shuicheng Yan, Nenghai Yu, Shipeng Li 0001, Jingdong Wang 0001
ACM Multimedia6
2014 Transductive 3D Shape Segmentation using Sparse Reconstruction
abstract
Abstract We propose a transductive shape segmentation algorithm, which can transfer prior segmentation results in database to new shapes without explicitly specification of prior category information. Our method first partitions an input shape into a set of segmentations as a data preparation, and then a linear integer programming algorithm is used to select segments from them to form the final optimal segmentation. The key idea is to maximize the segment similarity between the segments in the input shape and the segments in database, where the segment similarity is computed through sparse reconstruction error. The segment‐level similarity enables to handle a large amount of shapes with significant topology or shape variations with a small set of segmented example shapes. Experimental results show that our algorithm can generate high quality segmentation and semantic labeling results in the Princeton segmentation benchmark.
Weiwei Xu 0003, Zhouxu Shi, Kun Zhou 0001, Jingdong Wang 0001, JinRong Wang 0002, Zhenming Yuan
Comput. Graph. Forum5
2014 Image tag refinement by regularized latent Dirichlet allocation
Jingdong Wang 0001, Jiazhen Zhou, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
Comput. Vis. Image Underst.1
2014 Trinary-Projection Trees for Approximate Nearest Neighbor Search
abstract
We address the problem of approximate nearest neighbor (ANN) search for visual descriptor indexing. Most spatial partition trees, such as KD trees, VP trees, and so on, follow the hierarchical binary space partitioning framework. The key effort is to design different partition functions (hyperplane or hypersphere) to divide the points so that 1) the data points can be well grouped to support effective NN candidate location and 2) the partition functions can be quickly evaluated to support efficient NN candidate location. We design a trinary-projection direction-based partition function. The trinary-projection direction is defined as a combination of a few coordinate axes with the weights being 1 or -1. We pursue the projection direction using the widely adopted maximum variance criterion to guarantee good space partitioning and find fewer coordinate axes to guarantee efficient partition function evaluation. We present a coordinate-wise enumeration algorithm to find the principal trinary-projection direction. In addition, we provide an extension using multiple randomized trees for improved performance. We justify our approach on large-scale local patch indexing and similar image search.
Jingdong Wang 0001, Naiyan Wang, You Jia, Jian Li 0015, Hongbin Zha, Xian-Sheng Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Regularized Tree Partitioning and Its Application to Unsupervised Image Segmentation
abstract
In this paper, we propose regularized tree partitioning approaches. We study normalized cut (NCut) and average cut (ACut) criteria over a tree, forming two approaches: 1) normalized tree partitioning (NTP) and 2) average tree partitioning (ATP). We give the properties that result in an efficient algorithm for NTP and ATP. In addition, we present the relations between the solutions of NTP and ATP over the maximum weight spanning tree of a graph and NCut and ACut over this graph. To demonstrate the effectiveness of the proposed approaches, we show its application to image segmentation over the Berkeley image segmentation data set and present qualitative and quantitative comparisons with state-of-the-art methods.
Jingdong Wang 0001, Huaizu Jiang, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan
IEEE Trans. Image Process.1
2014 Browse-to-Search: Interactive Exploratory Search with Visual Entities
abstract
With the development of image search technology, users are no longer satisfied with searching for images using just metadata and textual descriptions. Instead, more search demands are focused on retrieving images based on similarities in their contents (textures, colors, shapes etc.). Nevertheless, one image may deliver rich or complex content and multiple interests. Sometimes users do not sufficiently define or describe their seeking demands for images even when general search interests appear, owing to a lack of specific knowledge to express their intents. A new form of information seeking activity, referred to as exploratory search, is emerging in the research community, which generally combines browsing and searching content together to help users gain additional knowledge and form accurate queries, thereby assisting the users with their seeking and investigation activities. However, there have been few attempts at addressing integrated exploratory search solutions when image browsing is incorporated into the exploring loop. In this work, we investigate the challenges of understanding users' search interests from the images being browsed and infer their actual search intentions. We develop a novel system to explore an effective and efficient way for allowing users to seamlessly switch between browse and search processes, and naturally complete visual-based exploratory search tasks. The system, called Browse-to-Search enables users to specify their visual search interests by circling any visual objects in the webpages being browsed, and then the system automatically forms the visual entities to represent users' underlying intent. One visual entity is not limited by the original image content, but also encapsulated by the textual-based browsing context and the associated heterogeneous attributes. We use large-scale image search technology to find the associated textual attributes from the repository. Users can then utilize the encapsulated visual entities to complete search tasks. The Browse-to-Search system is one of the first attempts to integrate browse and search activities for a visual-based exploratory search, which is characterized by four unique properties: (1) in session—searching is performed during browsing session and search results naturally accompany with browsing content; (2) in context—the pages being browsed provide text-based contextual cues for searching; (3) in focus—users can focus on the visual content of interest without worrying about the difficulties of query formulation, and visual entities will be automatically formed; and (4) intuitiveness—a touch and visual search-based user interface provides a natural user experience. We deploy the Browse-to-Search system on tablet devices and evaluate the system performance using millions of images. We demonstrate that it is effective and efficient in facilitating the user's exploratory search compared to the conventional image search methods and, more importantly, provides users with more robust results to satisfy their exploring experience.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, Shipeng Li 0001
ACM Trans. Inf. Syst.3
2014 Personalized Video Recommendation through Graph Propagation
abstract
The rapid growth of the number of videos on the Internet provides enormous potential for users to find content of interest. However, the vast quantity of videos also turns the finding process into a difficult task. In this article, we address the problem of providing personalized video recommendation for users. Rather than only exploring the user-video bipartite graph that is formulated using click information, we first combine the clicks and queries information to build a tripartite graph. In the tripartite graph, the query nodes act as bridges to connect user nodes and video nodes. Then, to further enrich the connections between users and videos, three subgraphs between the same kinds of nodes are added to the tripartite graph by exploring content-based information (video tags and textual queries). We propose an iterative propagation algorithm over the enhanced graph to compute the preference information of each user. Experiments conducted on a dataset with 1,369 users, 8,765 queries, and 17,712 videos collected from a commercial video search engine demonstrate the effectiveness of the proposed method.
Qinghua Huang, Bisheng Chen, Jingdong Wang 0001, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2013 Salient Object Detection: A Discriminative Regional Feature Integration Approach
abstract
Salient object detection has been attracting a lot of interest, and recently various heuristic computational models have been designed. In this paper, we regard saliency map computation as a regression problem. Our method, which is based on multi-level image segmentation, uses the supervised learning approach to map the regional feature vector to a saliency score, and finally fuses the saliency scores across multiple levels, yielding the saliency map. The contributions lie in two-fold. One is that we show our approach, which integrates the regional contrast, regional property and regional background ness descriptors together to form the master saliency map, is able to produce superior saliency maps to existing algorithms most of which combine saliency maps heuristically computed from different types of features. The other is that we introduce a new regional feature vector, background ness, to characterize the background, which can be regarded as a counterpart of the objectness descriptor [2]. The performance evaluation on several popular benchmark data sets validates that our approach outperforms existing state-of-the-arts.
Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Yang Wu 0001, Nanning Zheng 0001, Shipeng Li 0001
CVPR2
2013 Supervised Kernel Descriptors for Visual Recognition
abstract
In visual recognition tasks, the design of low level image feature representation is fundamental. The advent of local patch features from pixel attributes such as SIFT and LBP, has precipitated dramatic progresses. Recently, a kernel view of these features, called kernel descriptors (KDES), generalizes the feature design in an unsupervised fashion and yields impressive results. In this paper, we present a supervised framework to embed the image level label information into the design of patch level kernel descriptors, which we call supervised kernel descriptors (SKDES). Specifically, we adopt the broadly applied bag-of-words (BOW) image classification pipeline and a large margin criterion to learn the low-level patch representation, which makes the patch features much more compact and achieve better discriminative ability than KDES. With this method, we achieve competitive results over several public datasets comparing with state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Weiwei Xu 0003, Hongbin Zha, Shipeng Li 0001
CVPR2
2013 Online Robust Non-negative Dictionary Learning for Visual Tracking
abstract
This paper studies the visual tracking problem in video sequences and presents a novel robust sparse tracker under the particle filter framework. In particular, we propose an online robust non-negative dictionary learning algorithm for updating the object templates so that each learned template can capture a distinctive aspect of the tracked object. Another appealing property of this approach is that it can automatically detect and reject the occlusion and cluttered background in a principled way. In addition, we propose a new particle representation formulation using the Huber loss function. The advantage is that it can yield robust estimation without using trivial templates adopted by previous sparse trackers, leading to faster computation. We also reveal the equivalence between this new formulation and the previous one which uses trivial templates. The proposed tracker is empirically compared with state-of-the-art trackers on some challenging video sequences. Both quantitative and qualitative comparisons show that our proposed tracker is superior and more stable.
Naiyan Wang, Jingdong Wang 0001, Dit-Yan Yeung
ICCV2
2013 Fast Neighborhood Graph Search Using Cartesian Concatenation
abstract
In this paper, we propose a new data structure for approximate nearest neighbor search. This structure augments the neighborhood graph with a bridge graph. We propose to exploit Cartesian concatenation to produce a large set of vectors, called bridge vectors, from several small sets of subvectors. Each bridge vector is connected with a few reference vectors near to it, forming a bridge graph. Our approach finds nearest neighbors by simultaneously traversing the neighborhood graph and the bridge graph in the best-first strategy. The success of our approach stems from two factors: the exact nearest neighbor search over a large number of bridge vectors can be done quickly, and the reference vectors connected to a bridge (reference) vector near the query are also likely to be near the query. Experimental results on searching over large scale datasets (SIFT, GISTand HOG) show that our approach outperforms state-of-the-art ANN search algorithms in terms of efficiency and accuracy. The combination of our approach with the IVFADC system [18] also shows superior performance over the BIGANN dataset of 1 billion SIFT features compared with the best previously published result.
Jing Wang 0068, Jingdong Wang 0001, Rui Gan, Shipeng Li 0001, Baining Guo
ICCV2
2013 Learning CRFs for Image Parsing with Adaptive Subgradient Descent
abstract
We propose an adaptive sub gradient descent method to efficiently learn the parameters of CRF models for image parsing. To balance the learning efficiency and performance of the learned CRF models, the parameter learning is iteratively carried out by solving a convex optimization problem in each iteration, which integrates a proximal term to preserve the previously learned information and the large margin preference to distinguish bad labeling and the ground truth labeling. A solution of sub gradient descent updating form is derived for the convex optimization problem, with an adaptively determined updating step-size. Besides, to deal with partially labeled training data, we propose a new objective constraint modeling both the labeled and unlabeled parts in the partially labeled training data for the parameter learning of CRF models. The superior learning efficiency of the proposed method is verified by the experiment results on two public datasets. We also demonstrate the powerfulness of our method for handling partially labeled training data.
Honghui Zhang, Jingdong Wang 0001, Ping Tan 0002, Jinglu Wang, Long Quan
ICCV2
2013 Two dimensional analysis sparse model
abstract
An analysis sparse model represents an image signal by multiplying it using an analysis dictionary, leading to a sparse outcome. It transforms an image (two dimensional signal) into a one-dimensional (1D) vector. However, this 1D model ignores the two dimensional property and breaks the local spatial correlation inside images. In this paper, we propose a two dimensional (2D) analysis sparse model. Our 2D model uses two analysis dictionaries to efficiently exploit the horizontal and vertical features simultaneously. The corresponding sparse coding and dictionary learning algorithm are also presented in this paper. The 2D sparse model is further evaluated for image denoising. Experimental results demonstrate our 2D analysis sparse model outperforms a state-of-the-art 1D analysis model in terms of both denoising ability and memory usage.
Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Jingdong Wang 0001, Wenpeng Ding
ICIP4
2013 Two dimensional synthesis sparse model
abstract
Sparse representation has been proved to be very efficient in machine learning and image processing. Traditional image sparse representation formulates an image into a one dimensional (1D) vector which is then represented by a sparse linear combination of the basis atoms from a dictionary. This 1D representation ignores the local spatial correlation inside one image. In this paper, we propose a two dimensional (2D) sparse model to much efficiently exploit the horizontal and vertical features which are represented by two dictionaries simultaneously. The corresponding sparse coding and dictionary learning algorithm are also presented in this paper. The 2D synthesis model is further evaluated in image denoising. Experimental results demonstrate our 2D synthesis sparse model outperforms the state-of-the-art 1D model in terms of both objective and subjective qualities.
Na Qi, Yunhui Shi, Xiaoyan Sun 0001, Jingdong Wang 0001
ICME4
2013 Fixed-Point Model For Structured Labeling
abstract
In this paper, we propose a simple but effective solution to the structured labeling problem: a fixed-point model. Recently, layered models with sequential classifiers/regressors have gained an increasing amount of interests for structural prediction. Here, we design an algorithm with a new perspective on layered models; we aim to find a fixed-point function with the structured labels being both the output and the input. Our approach alleviates the burden in learning multiple/different classifiers in different layers. We devise a training strategy for our method and provide justifications for the fixed-point function to be a contraction mapping. The learned function captures rich contextual information and is easy to train and test. On several widely used benchmark datasets, the proposed method observes significant improvement in both performance and efficiency over many state-of-the-art algorithms.
Quannan Li, Jingdong Wang 0001, David P. Wipf, Zhuowen Tu
ICML (1)2
2013 Clickage: towards bridging semantic and intent gaps via mining click logs of search engines
abstract
The semantic gap between low-level visual features and high-level semantics has been investigated for decades but still remains a big challenge in multimedia. When "search" became one of the most frequently used applications, "intent gap", the gap between query expressions and users' search intents, emerged. Researchers have been focusing on three approaches to bridge the semantic and intent gaps: 1) developing more representative features, 2) exploiting better learning approaches or statistical models to represent the semantics, and 3) collecting more training data with better quality. However, it remains a challenge to close the gaps. In this paper, we argue that the massive amount of click data from commercial search engines provides a data set that is unique in the bridging of the semantic and intent gap. Search engines generate millions of click data (a.k.a. image-query pairs), which provide almost "unlimited" yet strong connections between semantics and images, as well as connections between users' intents and queries. To study the intrinsic properties of click data and to investigate how to effectively leverage this huge amount of data to bridge semantic and intent gap is a promising direction to advance multimedia research. In the past, the primary obstacle is that there is no such dataset available to the public research community. This changes as Microsoft has released a new large-scale real-world image click data to public. This paper presents preliminary studies on the power of large-scale click data with a variety of experiments, such as building large-scale concept detectors, tag processing, search, definitive tag detection, intent analysis, etc., with the goal to inspire deeper researches based on this dataset.
Xian-Sheng Hua 0001, Linjun Yang, Jingdong Wang 0001, Jing Wang 0068, Kuansan Wang, Yong Rui, Jin Li 0001
ACM Multimedia3
2013 Image search by graph-based label propagation with image representation from DNN
abstract
Our objective is to estimate the relevance of an image to a query for image search purposes. We address two limitations of the existing image search engines in this paper. First, there is no straightforward way of bridging the gap between semantic textual queries as well as users' search intents and image visual content. Image search engines therefore primarily rely on static and textual features. Visual features are mainly used to identify potentially useful recurrent patterns or relevant training examples for complementing search by image reranking. Second, image rankers are trained on query-image pairs labeled by human experts, making the annotation intellectually expensive and time-consuming. Furthermore, the labels may be subjective when the queries are ambiguous, resulting in difficulty in predicting the search intention. We demonstrate that the aforementioned two problems can be mitigated by exploring the use of click-through data, which can be viewed as the footprints of user searching behavior, as an effective means of understanding query. The correspondences between an image and a query are determined by whether the image was searched and clicked by users under the query in a commercial image search engine. We therefore hypothesize that the image click counts in response to a query are as their relevance indications. For each new image, our proposed graph-based label propagation algorithm employs neighborhood graph search to find the nearest neighbors on an image similarity graph built up with visual representations from deep neural networks and further aggregates their clicked queries/click counts to get the labels of the new image. We conduct experiments on MSR-Bing Grand Challenge and the results show consistent performance gain over various baselines. In addition, the proposed approach is very efficient, completing annotation of each query-image pair within just 15 milliseconds on a regular PC.
Yingwei Pan, Ting Yao 0003, Kuiyuan Yang, Houqiang Li, Chong-Wah Ngo, Jingdong Wang 0001, Tao Mei 0001
ACM Multimedia6
2013 Order preserving hashing for approximate nearest neighbor search
abstract
In this paper, we propose a novel method to learn similarity-preserving hash functions for approximate nearest neighbor (NN) search. The key idea is to learn hash functions by maximizing the alignment between the similarity orders computed from the original space and the ones in the hamming space. The problem of mapping the NN points into different hash codes is taken as a classification problem in which the points are categorized into several groups according to the hamming distances to the query. The hash functions are optimized from the classifiers pooled over the training points. Experimental results demonstrate the superiority of our approach over existing state-of-the-art hashing techniques.
Jingdong Wang 0001, Nenghai Yu, Shipeng Li 0001
ACM Multimedia2
2013 Structure-Sensitive Superpixels via Geodesic Distance
Peng Wang 0001, Rui Gan, Jingdong Wang 0001, Hongbin Zha
Int. J. Comput. Vis.4
2013 Interactive Multimodal Visual Search on Mobile Device
abstract
This paper describes a novel multimodal interactive image search system on mobile devices. The system, the Joint search with ImaGe, Speech, And Word Plus(JIGSAW+), takes full advantage of the multimodal input and natural user interactions of mobile devices. It is designed for users who already have pictures in their minds but have no precise descriptions or names to address them. By describing it using speech and then refining the recognized query by interactively composing a visual query using exemplary images, the user can easily find the desired images through a few natural multimodal interactions with his/her mobile device. Compared with our previous work JIGSAW, the algorithm has been significantly improved in three aspects: 1) segmentation-based image representation is adopted to remove the artificial block partitions; 2) relative position checking replaces the fixed position penalty; and 3) inverted index is constructed instead of brute force matching. The proposed JIGSAW+ is able to achieve 5% gain in terms of search performance and is ten times faster.
Houqiang Li, Tao Mei 0001, Jingdong Wang 0001, Shipeng Li 0001
IEEE Trans. Multim.4
2012 Image search results refinement via outlier detection using deep contexts
abstract
Visual reranking has become a widely-accepted method to improve traditional text-based image search results. The main principle is to exploit the visual aggregation property of relevant images among top results so as to boost ranking scores of relevant images, by explicitly or implicitly detecting the confident relevant images, and propagating ranking scores among visually similar images. However, such a visual aggregation property does not always hold, and thus these schemes may fail. In this paper, we instead propose to filter out the most probable irrelevant images using deep contexts, which is the extra information that is not limited in the current search results. The deep contexts for each image consist of sets of images that are returned by searches using the queries formed by the textual context of this image. We compare the popularity of this image in the current search results and the deep contexts to check the irrelevance score. Then the irrelevance scores are propagated to the images whose useful textual context is missed. We formulate the two schemes together to reach a Markov random field, which is effectively solved by graph cuts. The key is that our scheme does not rely on the assumption that relevant images are visually aggregated among top results and is based on the observation that an outlier under the current query is likely to be more popular under some other query. After that, we perform graph reranking over filtered results to reorder them. Experimental results on the INRIA dataset show that our proposed method achieves significant improvements over previous approaches.
Junyang Lu, Jiazhen Zhou, Jingdong Wang 0001, Tao Mei 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
CVPR3
2012 Fast approximate k-means via cluster closures
abstract
K-means, a simple and effective clustering algorithm, is one of the most widely used algorithms in computer vision community. Traditional k-means is an iterative algorithm - in each iteration new cluster centers are computed and each data point is re-assigned to its nearest center. The cluster re-assignment step becomes prohibitively expensive when the number of data points and cluster centers are large. In this paper, we propose a novel approximate k-means algorithm to greatly reduce the computational complexity in the assignment step. Our approach is motivated by the observation that most active points changing their cluster assignments at each iteration are located on or near cluster boundaries. The idea is to efficiently identify those active points by pre-assembling the data into groups of neighboring points using multiple random spatial partition trees, and to use the neighborhood information to construct a closure for each cluster, in such a way only a small number of cluster candidates need to be considered when assigning a data point to its nearest cluster. Using complexity analysis, real data clustering, and applications to image retrieval, we show that our approach out-performs state-of-the-art approximate k-means algorithms in terms of clustering quality and efficiency.
Jing Wang 0068, Jingdong Wang 0001, Qifa Ke, Shipeng Li 0001
CVPR2
2012 Salient object detection for searched web images via global saliency
abstract
In this paper, we deal with the problem of detecting the existence and the location of salient objects for thumbnail images on which most search engines usually perform visual analysis in order to handle web-scale images. Different from previous techniques, such as sliding window-based or segmentation-based schemes for detecting salient objects, we propose to use a learning approach, random forest in our solution. Our algorithm exploits global features from multiple saliency indicators to directly predict the existence and the position of the salient object. To validate our algorithm, we constructed a large image database collected from Bing image search, that contains hundreds of thousands of manually labeled web images. The experimental results using this new database and the resized MSRA database [16] demonstrate that our algorithm outperforms previous state-of-the-art methods.
Peng Wang 0001, Jingdong Wang 0001, Jie Feng 0012, Hongbin Zha, Shipeng Li 0001
CVPR2
2012 Scalable k-NN graph construction for visual descriptors
abstract
The k-NN graph has played a central role in increasingly popular data-driven techniques for various learning and vision tasks; yet, finding an efficient and effective way to construct k-NN graphs remains a challenge, especially for large-scale high-dimensional data. In this paper, we propose a new approach to construct approximate k-NN graphs with emphasis in: efficiency and accuracy. We hierarchically and randomly divide the data points into subsets and build an exact neighborhood graph over each subset, achieving a base approximate neighborhood graph; we then repeat this process for several times to generate multiple neighborhood graphs, which are combined to yield a more accurate approximate neighborhood graph. Furthermore, we propose a neighborhood propagation scheme to further enhance the accuracy. We show both theoretical and empirical accuracy and efficiency of our approach to k-NN graph construction and demonstrate significant speed-up in dealing with large scale visual data.
Jing Wang 0068, Jingdong Wang 0001, Zhuowen Tu, Rui Gan, Shipeng Li 0001
CVPR2
2012 A Probabilistic Approach to Robust Matrix Factorization
Naiyan Wang, Tiansheng Yao, Jingdong Wang 0001, Dit-Yan Yeung
ECCV (7)3
2012 Personalized video recommendation through tripartite graph propagation
abstract
The rapid growth of the number of videos on the Internet provides enormous potential for users to find content of interest to them. Video search, such as Google, Youtube, Bing, is a popular way to help users to find desired videos. However, it is still very challenging to discover new video contents for users. In this paper, we address the problem of providing personalized video suggestions for users. Rather than only exploring the user-video graph that is formulated using the click-through information, we also investigate other two useful graphs, the user-query graph indicating if a user ever issues a query, and the query-video graph indicating if a video appears in the search result of a query. The two graphs act as a bridge to connect users and videos, and have a large potential to improve the recommendation as the queries issued by a user essentially imply his interest. As a result, we reach a tripartite graph over (user, video, query). We develop an iterative propagation scheme over the tripartite graph to compute the preference information of each user. Experimental results on a dataset of 2,893 users, 23,630 queries and 55,114 videos collected during Feb. 1-28, 2011 demonstrate that the proposed method outperforms existing state-of-the-art approaches, co-views and random walks on the user-video bipartite graph.
Bisheng Chen, Jingdong Wang 0001, Qinghua Huang, Tao Mei 0001
ACM Multimedia2
2012 Browse-to-search
abstract
This demonstration presents a novel interactive online shopping application based on visual search technologies. When users want to buy something on a shopping site, they usually have the requirement of looking for related information from other web sites. Therefore users need to switch between the web page being browsed and other websites that provide search results. The proposed application enables users to naturally search products of interest when they browse a web page, and make their even causal purchase intent easily satisfied. The interactive shopping experience is characterized by: 1) in session---it allows users to specify the purchase intent in the browsing session, instead of leaving the current page and navigating to other websites; 2) in context---the browsed web page provides implicit context information which helps infer user purchase preferences; 3) in focus---users easily specify their search interest using gesture on touch devices and do not need to formulate queries in search box; 4) natural-gesture inputs and visual-based search provides users a natural shopping experience. The system is evaluated against a data set consisting of several millions commercial product images.
Shiyang Lu, Tao Mei 0001, Jingdong Wang 0001, Jian Zhang 0002, Zhiyong Wang 0001, David Dagan Feng, Jian-Tao Sun, Shipeng Li 0001
ACM Multimedia3
2012 Similar image search with a tiny bag-of-delegates representation
abstract
Similar image search over a large image database has been attracting a lot of attention recently. The widely-used solution is to use a set of codes, which we call bag-of-delegates, to represent each image, and to build inverted indices to organize the image database. The search can be conducted through the inverted indices, which is the same to the way of using texts to index images for search and has been shown to be efficient and effective.
Weiwen Tu, Jingdong Wang 0001
ACM Multimedia3
2012 Query-driven iterated neighborhood graph search for large scale indexing
abstract
In this paper, we address the approximate nearest neighbor (ANN) search problem over large scale visual descriptors. We investigate a simple but very effective approach, neighborhood graph search, which constructs a neighborhood graph to index the data points and conducts a local search, expanding neighborhoods with a best-first manner, for ANN search. Our empirical analysis shows that neighborhood expansion is very efficient, with O(1) cost, for a new NN candidate location, and has high chances to locate true NNs and hence it usually performs well. However, it often gets sub-optimal solutions since local search only checks the neighborhood of the current solution, or conducts exhaustive and continuous neighborhood expansions to find better solutions, which deteriorates the query efficiency.
Jingdong Wang 0001, Shipeng Li 0001
ACM Multimedia1
2012 Scalable similar image search by joint indices
abstract
Text-based image search is able to return desired images for simple queries, but has limited capabilities in finding images with additional visual requirements. As a result, an image is usually used to help describe the appearance requirements. In this demonstration, we show a similar image search system that can support the joint textual and visual query. We present an efficient and effective indexing algorithm, neighborhood graph index, which is suitable for millions of images, and use it to organize joint inverted indices to search over billions of images.
Jing Wang 0068, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia2
2012 Color filter for image search
abstract
Image search relying on surrounding texts can return reliably relevant images to some extent. Most recent efforts are focusing on utilizing visual contents to help users find images with specific visual requirements. In this demonstration, we show a color filter scheme for image search, which enables users to find images containing objects or scenes with their interested color. Color is one of the most crucial cues in describing visual contents and has been frequently used in various applications. The key components in developing our demo include salient object detection and perceptual color naming. The color filter in Microsoft Bing image search is developed using these techniques.
Peng Wang 0001, Dongqing Zhang, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia3
2012 Recommending Flickr groups with social topic model
Jingdong Wang 0001, Zhe Zhao 0001, Jiazhen Zhou, Hao Wang 0005, Bin Cui 0001, Guo-Jun Qi
Inf. Retr.1
2012 Correction to "Bayesian Visual Reranking"
abstract
In the above titled paper (ibid., vol. 13, no. 4, pp. 639-652, Aug. 2011), the first author's name appears incorrectly in the byline as "Xinmie Tian" instead of "Xinmei Tian." The name appears correctly in the biography section.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
IEEE Trans. Multim.3
2012 An interactive approach to semantic modeling of indoor scenes with an RGBD camera
abstract
We present an interactive approach to semantic modeling of indoor scenes with a consumer-level RGBD camera. Using our approach, the user first takes an RGBD image of an indoor scene, which is automatically segmented into a set of regions with semantic labels. If the segmentation is not satisfactory, the user can draw some strokes to guide the algorithm to achieve better results. After the segmentation is finished, the depth data of each semantic region is used to retrieve a matching 3D model from a database. Each model is then transformed according to the image depth to yield the scene. For large scenes where a single image can only cover one part of the scene, the user can take multiple images to construct other parts of the scene. The 3D models built for all images are then transformed and unified into a complete scene. We demonstrate the efficiency and robustness of our approach by modeling several real-world scenes.
Tianjia Shao, Weiwei Xu 0003, Kun Zhou 0001, Jingdong Wang 0001, Dongping Li, Baining Guo
ACM Trans. Graph.4
2011 Automatic salient object segmentation based on context and shape prior
abstract
We propose a novel automatic salient object segmentation algorithm which integrates both bottom-up salient stimuli and object-level shape prior, i.e., a salient object has a well-defined closed boundary. Our approach is formalized as an iterative energy minimization framework, leading to binary segmentation of the salient object. Such energy minimization is initialized with a saliency map which is computed through context analysis based on multi-scale superpixels. Object-level shape prior is then extracted combining saliency with object boundary information. Both saliency map and shape prior update after each iteration. Experimental results on two public benchmark datasets show that our proposed approach outperforms state-of-the-art methods. 1
Huaizu Jiang, Jingdong Wang 0001, Zejian Yuan, Nanning Zheng 0001
BMVC2
2011 A non-convex relaxation approach to sparse dictionary learning
abstract
Dictionary learning is a challenging theme in computer vision. The basic goal is to learn a sparse representation from an overcomplete basis set. Most existing approaches employ a convex relaxation scheme to tackle this challenge due to the strong ability of convexity in computation and theoretical analysis. In this paper we propose a non-convex online approach for dictionary learning. To achieve the sparseness, our approach treats a so-called minimax concave (MC) penalty as a nonconvex relaxation of the ℓ0penalty. This treatment expects to obtain a more robust and sparse representation than existing convex approaches. In addition, we employ an online algorithm to adaptively learn the dictionary, which makes the non-convex formulation computationally feasible. Experimental results on the sparseness comparison and the applications in image denoising and image inpainting demonstrate that our approach is more effective and flexible.
Jianping Shi, Guang Dai, Jingdong Wang 0001
CVPR4
2011 Multi-task low-rank affinity pursuit for image segmentation
abstract
This paper investigates how to boost region-based image segmentation by pursuing a new solution to fuse multiple types of image features. A collaborative image segmentation framework, called multi-task low-rank affinity pursuit, is presented for such a purpose. Given an image described with multiple types of features, we aim at inferring a unified affinity matrix that implicitly encodes the segmentation of the image. This is achieved by seeking the sparsity-consistent low-rank affinities from the joint decompositions of multiple feature matrices into pairs of sparse and low-rank matrices, the latter of which is expressed as the production of the image feature matrix and its corresponding image affinity matrix. The inference process is formulated as a constrained nuclear norm and ℓ2;1-norm minimization problem, which is convex and can be solved efficiently with the Augmented Lagrange Multiplier method. Compared to previous methods, which are usually based on a single type of features, the proposed method seamlessly integrates multiple types of features to jointly produce the affinity matrix within a single inference step, and produces more accurate and reliable segmentation results. Experiments on the MSRC dataset and Berkeley segmentation dataset well validate the superiority of using multiple features over single feature and also the superiority of our method over conventional methods for feature fusion. Moreover, our method is shown to be very competitive while comparing to other state-of-the-art methods.
Bin Cheng 0001, Guangcan Liu, Jingdong Wang 0001, ZhongYang Huang, Shuicheng Yan
ICCV3
2011 Complementary hashing for approximate nearest neighbor search
abstract
Recently, hashing based Approximate Nearest Neighbor (ANN) techniques have been attracting lots of attention in computer vision. The data-dependent hashing methods, e.g., Spectral Hashing, expects better performance than the data-blind counterparts, e.g., Locality Sensitive Hashing (LSH). However, most data-dependent hashing methods only employ a single hash table. When higher recall is desired, they have to retrieve exponentially growing number of hash buckets around the bucket containing the query, which may drag down the precision rapidly. In this paper, we propose a so-called complementary hashing approach, which is able to balance the precision and recall in a more effective way. The key idea is to employ multiple complementary hash tables, which are learned sequentially in a boosting manner, so that, given a query, its true nearest neighbors missed from the active bucket of one hash table are more likely to be found in the active bucket of the next hash table. Compared with LSH that also can exploit multiple hash tables, our approach is more effective to find true NNs, thanks to the complementarity property of the hash tables from our approach. Experimental results on large scale ANN search show that the proposed method significantly improves the performance and outperforms the state-of-the-art.
Jingdong Wang 0001, Shipeng Li 0001, Nenghai Yu
ICCV2
2011 Structure-sensitive superpixels via geodesic distance
abstract
Over-segments (i.e. superpixels) have been commonly used as supporting regions for feature vectors and primitives to reduce computational complexity in various image analysis tasks. In this paper, we describe a structuresensitive over-segmentation technique by exploiting Lloyd's algorithm with a geodesic distance. It generates smaller superpixels to achieve lower under-segmentation in structure-dense regions with high intensity or color variation, and produces larger segments to increase computational efficiency in structure-sparse regions with homogeneous appearance. We adopt geometric flows to compute the geodesic distances amongst pixels, and in the segmentation procedure, the density of over-segments is automatically adjusted according to an energy functional that embeds color homogeneity, structure density and compactness constraints. Comparative experiments with the Berkeley database show that the proposed algorithm outperforms prior arts while offering a comparable computational efficiency with fast methods, such as TurboPixels.
Peng Wang 0001, Jingdong Wang 0001, Rui Gan, Hongbin Zha
ICCV3
2011 Contextual image search
abstract
In this paper, we propose a novel image search scheme, contextual image search. Different from conventional image search schemes that present a separate interface (e.g., text input box) to allow users to submit a query, the new search scheme enables users to search images by only masking a few words when they are reading through Web pages or other documents. Rather than merely making use of the explicit query input that is often not sufficient to express user's search intent, our approach explores the context information to better understand the search intent with two key steps: query augmenting and search results reranking using context, and expects to obtain better search results. Beyond contextual Web search, the context in our case is much richer and includes images besides texts. In addition to this type of search scheme, called contextual image search with text input, we also present another type of scheme, called contextual image search with image input, to allow users to select an image as the search query from Web pages or other documents they are reading. The key idea is to use the search-to-annotation technique and the contextual textual query mining scheme to determine the corresponding textual query, to finally get semantically similar search results. Experiments show that the proposed schemes make image search more convenient and the search results are more relevant to user intention.
Wenhao Lu, Jingdong Wang 0001, Xian-Sheng Hua 0001, Shengjin Wang, Shipeng Li 0001
ACM Multimedia2
2011 Robust visual reranking via sparsity and ranking constraints
abstract
Visual reranking has become a widely-accepted method to improve traditional text-based image search engines. Its basic principle is that visually similar images should have similar ranking scores. While existing methods are different in specifics, almost all of them are based on explicit or implicit pseudo-relevance feedback (PRF). Explicit PRF-based approaches, including classification-based and clustering-based reranking, suffer from the difficulty of selecting reliable positive and negative samples. Implicit PRF-based approaches, such as graph-based and Bayesian visual reranking, deal with such unreliability by making use of the initial ranking in a soft manner, but have limited capability of promoting relevant images and lowering down irrelevant images.
Nobuyuki Morioka, Jingdong Wang 0001
ACM Multimedia2
2011 Web-scale image search by color sketch
abstract
Most existing image search engines rely on the associated texts or tags with images to index and retrieval images, which results in limited ability on searching images with visual requirement. In this demonstration, we present an image search system, which enables consumers to find images on the requirement of how the colors are spatially distributed. It is a well-designed trade-off between scalability and feasibility. To the best knowledge, this system is the first one to scale up to Web-scale images. The interface is very intuitive and requires users to only scribble a few color strokes or drag an image and mask a few regions of interest, to express the search intent.
Jingdong Wang 0001, Xian-Sheng Hua 0001
ACM Multimedia1
2011 JIGSAW: interactive mobile visual search with multimodal queries
abstract
The traditional text-based visual search has not been sufficiently improved over the years to accommodate the new emerging demand of mobile users. While on the go, searching on one's phone is becoming pervasive. This paper presents an innovative application for mobile phone users to facilitate their visual search experience. By taking advantage of smart phone functionalities such as multi-modal and multi-touch interactions, users can more conveniently formulate their search intent, and thus search performance can be significantly improved. The system, called JIGSAW (Joint search with ImaGe, Speech, And Words), represents one of the first attempts to create an interactive and multi-modal mobile visual search application. The key of JIGSAW is the composition of an exemplary image query generated from the raw speech via multi-touch user interaction, as well as the visual search based on the exemplary image. Through JIGSAW, users can formulate their search intent in a natural way like playing a jigsaw puzzle on the phone screen: 1) a user speaks a natural sentence as the query, 2) the speech is recognized and transferred to text which is further decomposed to keywords through entity extraction, 3) the user selects preferred exemplary images that can visually represent his/her intent and composes a query image via multi-touch, and 4) the composite image is then used as a visual query to search similar images. We have deployed JIGSAW on a real-world phone system, evaluated the performance on one million images, and demonstrated that it is an effective complement to existing mobile visual search applications.
Tao Mei 0001, Jingdong Wang 0001, Houqiang Li, Shipeng Li 0001
ACM Multimedia3
2011 Hybrid image summarization
abstract
In this paper, we address a problem of managing tagged images with hybrid summarization. We formulate this problem as finding a few image exemplars to represent the image set semantically and visually and solve it in a hybrid way by exploiting both visual and textual information associated with images. We propose a novel approach, called Homogeneous and Heterogeneous Message Propagation (H2MP), which extends affinity propagation that only works over homogeneous relations to heterogeneous relations. The summary obtained by our approach is both visually and semantically satisfactory. The experimental results demonstrate the effectiveness and efficiency of the proposed approach.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia2
2011 Document clustering with universum
abstract
Document clustering is a popular research topic, which aims to partition documents into groups of similar objects (i.e., clusters), and has been widely used in many applications such as automatic topic extraction, document organization and filtering. As a recently proposed concept, Universum is a collection of "non-examples" that do not belong to any concept/cluster of interest. This paper proposes a novel document clustering technique -- Document Clustering with Universum, which utilizes the Universum examples to improve the clustering performance. The intuition is that the Universum examples can serve as supervised information and help improve the performance of clustering, since they are known not belonging to any meaningful concepts/clusters in the target domain. In particular, a maximum margin clustering method is proposed to model both target examples and Universum examples for clustering. An extensive set of experiments is conducted to demonstrate the effectiveness and efficiency of the proposed algorithm.
Dan Zhang 0007, Jingdong Wang 0001, Luo Si
SIGIR2
2011 Discriminative Sketch-based 3D Model Retrieval via Robust Shape Matching
abstract
Abstract We propose a sketch‐based 3D shape retrieval system that is substantially more discriminative and robust than existing systems, especially for complex models. The power of our system comes from a combination of a contour‐based 2D shape representation and a robust sampling‐based shape matching scheme. They are defined over discriminative local features and applicable for partial sketches; robust to noise and distortions in hand drawings; and consistent when strokes are added progressively. Our robust shape matching, however, requires dense sampling and registration and incurs a high computational cost. We thus devise critical acceleration methods to achieve interactive performance: precomputing kNN graphs that record transformations between neighboring contour images and enable fast online shape alignment; pruning sampling and shape registration strategically and hierarchically; and parallelizing shape matching on multi‐core platforms or GPUs. We demonstrate the effectiveness of our system through various experiments, comparisons, and user studies.
Tianjia Shao, Weiwei Xu 0003, KangKang Yin, Jingdong Wang 0001, Kun Zhou 0001, Baining Guo
Comput. Graph. Forum4
2011 Interactive browsing via diversified visual summarization for image search results
Jingdong Wang 0001, Liyan Jia, Xian-Sheng Hua 0001
Multim. Syst.1
2011 Learning to Detect a Salient Object
abstract
In this paper, we study the salient object detection problem for images. We formulate this problem as a binary labeling task where we separate the salient object from the background. We propose a set of novel features, including multiscale contrast, center-surround histogram, and color spatial distribution, to describe a salient object locally, regionally, and globally. A conditional random field is learned to effectively combine these features for salient object detection. Further, we extend the proposed approach to detect a salient object from sequential images by introducing the dynamic salient features. We collected a large image database containing tens of thousands of carefully labeled images by multiple users and a video segment database, and conducted a set of experiments over them to demonstrate the effectiveness of the proposed approach.
Zejian Yuan, Jian Sun 0001, Jingdong Wang 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.4
2011 A transductive multi-label learning approach for video concept detection
Jingdong Wang 0001, Yinghai Zhao, Xiuqing Wu, Xian-Sheng Hua 0001
Pattern Recognit.1
2011 Interactive Image Search by Color Map
abstract
The availability of large-scale images from the Internet has made the research on image search attract a lot of attention. Text-based image search engines, for example, Google/Microsoft Bing/Yahoo! image search engines using the surrounding text, have been developed and widely used. However, they suffer from an inability to search image content. In this article, we present an interactive image search system, image search by color map, which can be applied to, but not limited to, enhance text-based image search. This system enables users to indicate how the colors are spatially distributed in the desired images, by scribbling a few color strokes, or dragging an image and highlighting a few regions of interest in an intuitive way. In contrast to the conventional sketch-based image retrieval techniques, our system searches images based on colors rather than shapes, and we, technically, propose a simple but effective scheme to mine the latent search intention from the user’s input, and exploit the dominant color filter strategy to make our system more efficient. We integrate our system to existing Web image search engines to demonstrate its superior performance over text-based image search. The user study shows that our system can indeed help users conveniently find desired images.
Jingdong Wang 0001, Xian-Sheng Hua 0001
ACM Trans. Intell. Syst. Technol.1
2011 Bayesian Visual Reranking
abstract
Visual reranking has been proven effective to refine text-based video and image search results. It utilizes visual information to recover “true” ranking list from the noisy one generated by text-based search, by incorporating both textual and visual information. In this paper, we model the textual and visual information from the probabilistic perspective and formulate visual reranking as an optimization problem in the Bayesian framework, termed Bayesian visual reranking. In this method, the textual information is modeled as a likelihood, to reflect the disagreement between reranked results and text-based search results which is called ranking distance. The visual information is modeled as a conditional prior, to indicate the ranking score consistency among visually similar samples which is called visual consistency. Bayesian visual reranking derives the best reranking results by maximizing visual consistency while minimizing ranking distance. To model the ranking distance more precisely, we propose a novel pair-wise method which measure the ranking distance based on the disagreement in terms of pair-wise orders. For visual consistency, we study three different regularizers to mine the best way for its modeling. We conduct extensive experiments on both video and image search datasets. Experimental results demonstrate the effectiveness of our proposed Bayesian visual reranking.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
IEEE Trans. Multim.3
2010 Optimizing kd-trees for scalable visual descriptor indexing
abstract
In this paper, we attempt to scale up the kd-tree indexing methods for large-scale vision applications, e.g., indexing a large number of SIFT features and other types of visual descriptors. To this end, we propose an effective approach to generate near-optimal binary space partitioning and need low time cost to access the nodes in the query stage. First, we relax the coordinate-axis-alignment constraint in partition axis selection used in conventional kd-trees, and form a partition axis with the great variance by combining a few coordinate axes in a binary manner for each node, which yields better space partitioning and requires almost the same time cost to visit internal nodes during the query stage thanks to cheap projection operations. Then, we introduce a simple but very effective scheme to guarantee the partition axis of each internal node is orthogonal to or parallel with those of its ancestors, which leads to efficient distance computation between a query point and the cell associated with each node and yields fast priority search. Compared with the conventional kd-trees, our approach takes a little more tree construction time, but obtains much better nearest neighbor search performance. Experimental results on large scale local patch indexing and image search with tiny images show that our approach outperforms the state-of-the-art kd-tree based indexing methods.
You Jia, Jingdong Wang 0001, Hongbin Zha, Xian-Sheng Hua 0001
CVPR2
2010 Learning to combine multi-resolution spatially-weighted co-occurrence matrices for image representation
abstract
Bag-of-Words is widely used to describe images for image classification. However, this approach is limited because the spatial relation over visual words is not well exploited and also it is difficult to generate a single comprehensive vocabulary. In this paper, we propose novel effective schemes to handle these two issues. First, we propose a structure propagation technique to build more reasonable co-occurrence matrices of visual words to exploit the spatial information, which assigns a higher weight to the co-occurrence over two patches that lie in the same object part. Second, we build the multiple-histogram representation over hierarchical vocabularies to avoid the ambiguity of single vocabulary, and particularly present a learning approach to combine the multiple histograms to integrate both within-vocabulary and cross-vocabulary information. We evaluate our proposed method using the Princeton sports event dataset. Compared to the state-of-the-art results, our proposed approach has shown promising results.
Xiangang Cheng, Jingdong Wang 0001, Liang-Tien Chia, Xian-Sheng Hua 0001
ICME2
2010 Dynamic Video Collage
Yan Wang 0059, Tao Mei 0001, Jingdong Wang 0001, Xian-Sheng Hua 0001
MMM3
2010 Image search by concept map
abstract
In this paper, we present a novel image search system, image search by concept map. This system enables users to indicate not only what semantic concepts are expected to appear but also how these concepts are spatially distributed in the desired images. To this end, we propose a new image search interface to enable users to formulate a query, called concept map, by intuitively typing textual queries in a blank canvas to indicate the desired spatial positions of the concepts. In the ranking process, by interpreting each textual concept as a set of representative visual instances, the concept map query is translated into a visual instance map, which is then used to evaluate the relevance of the image in the database. Experimental results demonstrate the effectiveness of the proposed system.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
SIGIR2
2010 Interactive image search by 2D semantic map
abstract
In this demo, we present a novel interactive image search system, image search by 2D semantic map. This system enables users to indicate what semantic concepts are expected to appear and even how these concepts are spatially distributed in the desired images. To this end, we design an intuitive interface for users to formulate a query in the form of 2D semantic map, called concept map, by typing textual queries in a blank canvas. In the ranking process, by interpreting each textual concept as a set of representative visual instances, the concept map query is translated into a visual instance map, which is then used for comparison with the images in the database. Besides, in this demo, we also show an image search system with a simplest semantic map, a 2D color map, where the concepts are limited from the colors.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
WWW2
2009 Summarizing tagged image collections by cross-media representativeness voting
abstract
In this paper, we address the problem of generating both visual and textual summaries for tagged image collections simultaneously. The visual and textual summaries consist of representative images and tags of the collection, which are selected through a proposed cross-media voting scheme. In the voting scheme, the likelihood of an image to be a representative is voted by not only other images but also the tags, according to the intra-media and cross-media affinities. The likelihood of a tag to be a representative is obtained in similar manner at the same time. We demonstrate that the proposed scheme produces more informative textual and visual summaries than summarizing images and tags separately.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ICME2
2009 Tag refinement by regularized LDA
abstract
Tagging is nowadays the most prevalent and practical way to make images searchable. However, in reality many tags are irrelevant to image content. To refine the tags, previous solutions usually mine tag relevance relying on the tag similarity estimated right from the corpus to be refined. The calculation of tag similarity is affected by the noisy tags in the corpus, which is not conducive to estimate accurate tag relevance. In this paper, we propose to do tag refinement from the angle of topic modeling. In the proposed scheme, tag similarity and tag relevance are jointly estimated in an iterative manner, so that they can benefit from each other. Specifically, a novel graphical model, regularized Latent Dirichlet Allocation (rLDA), is presented. It facilitates the topic modeling by exploiting both the statistics of tags and visual affinities of images in the corpus. The experiments on tag ranking and image retrieval demonstrate the advantages of the proposed method.
Jingdong Wang 0001, Xian-Sheng Hua 0001, Shipeng Li 0001
ACM Multimedia2
2009 Graph-based semi-supervised learning with multiple labels
Zhengjun Zha, Tao Mei 0001, Jingdong Wang 0001, Zengfu Wang, Xian-Sheng Hua 0001
J. Vis. Commun. Image Represent.3
2009 Linear Neighborhood Propagation and Its Applications
abstract
In this paper, a novel graph-based transductive classification approach, called Linear Neighborhood Propagation, is proposed. The basic idea is to predict the label of a data point according to its neighbors in a linear way. This method can be cast into a second-order intrinsic Gaussian Markov random field framework. Its result corresponds to a solution to an approximate inhomogeneous biharmonic equation with Dirichlet boundary conditions. Different from existing approaches, our approach provides a novel graph structure construction method by introducing multiple-wise edges instead of pairwise edges, and presents an effective scheme to estimate the weights for such multiple-wise edges. To the best of our knowledge, these two contributions are novel for semi-supervised classification. The experimental results on image segmentation and transductive classification demonstrate the effectiveness and efficiency of the proposed approach.
Jingdong Wang 0001, Fei Wang 0001, Changshui Zhang, Helen C. Shen, Long Quan
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary-to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a conditional random field model, which integrates image salience, canvas constraint, natural preference, and user interaction. Each image is represented by a group of weighted rectangles, which indicate the salient regions. Then picture collage is resolved by minimizing the energy, guided by the constraints. A two-step optimization method is proposed. First, a quick initialization algorithm based on the proposed 1D collage method is presented. Second, a very efficient Markov chain Monte Carlo method is designed for the refined optimization. We also integrate user interaction in the formulation and optimization to obtain an interactive collage reflecting personalized preference. Visual and quantitative experimental evaluations indicate the efficiency of the proposed collage creation technique.
Jingdong Wang 0001, Jian Sun 0001, Nanning Zheng 0001, Xiaoou Tang, Harry Shum
IEEE Trans. Multim.2
2008 Normalized tree partitioning for image segmentation
abstract
In this paper, we propose a novel graph based clustering approach with satisfactory clustering performance and low computational cost. It consists of two main steps: tree fitting and partitioning. We first introduce a probabilistic method to fit a tree to a data graph under the sense of minimum entropy. Then, we propose a novel tree partitioning method under a normalized cut criterion, called Normalized Tree Partitioning (NTP), in which a fast combinatorial algorithm is designed for exact bipartitioning. Moreover, we extend it to k-way tree partitioning by proposing an efficient best-first recursive bipartitioning scheme. Compared with spectral clustering, NTP produces the exact global optimal bipartition, introduces fewer approximations for k-way partitioning and can intrinsically produce superior performance. Compared with bottom-up aggregation methods, NTP adopts a global criterion and hence performs better. Last, experimental results on image segmentation demonstrate that our approach is more powerful compared with existing graph-based approaches.
Jingdong Wang 0001, Yangqing Jia, Xian-Sheng Hua 0001, Changshui Zhang, Long Quan
CVPR1
2008 Joint multi-label multi-instance learning for image classification
abstract
In real world, an image is usually associated with multiple labels which are characterized by different regions in the image. Thus image classification is naturally posed as both a multi-label learning and multi-instance learning problem. Different from existing research which has considered these two problems separately, we propose an integrated multi-label multi-instance learning (MLMIL) approach based on hidden conditional random fields (HCRFs), which simultaneously captures both the connections between semantic labels and regions, and the correlations among the labels in a single formulation. We apply this MLMIL framework to image classification and report superior performance compared to key existing approaches over the MSR Cambridge (MSRC) and Corel data sets.
Zhengjun Zha, Xian-Sheng Hua 0001, Tao Mei 0001, Jingdong Wang 0001, Guo-Jun Qi, Zengfu Wang
CVPR4
2008 Maximum Margin Clustering with Pairwise Constraints
abstract
Maximum margin clustering (MMC), which extends the theory of support vector machine to unsupervised learning, has been attracting considerable attention recently. The existing approaches mainly focus on reducing the computational complexity of MMC. The accuracy of these methods, however, has not always been guaranteed. In this paper, we propose to incorporate additional side-information, which is in the form of pairwise constraints, into MMC to further improve its performance. A set of pairwise loss functions are introduced into the clustering objective function which effectively penalize the violation of the given constraints. We show that the resulting optimization problem can be easily solved via constrained concave-convex procedure (CCCP). Moreover, for constrained multi-class MMC, we present an efficient cutting-plane algorithm to solve the sub-problem in each iteration of CCCP. The experiments demonstrate that the pairwise constrained MMC algorithms considerably outperform the unconstrained MMC algorithms and two other clustering algorithms that exploit the same type of side-information.
Yang Hu 0006, Jingdong Wang 0001, Nenghai Yu, Xian-Sheng Hua 0001
ICDM2
2008 Augmented tree partitioning for interactive image segmentation
abstract
In this paper, we propose a new fast semi-supervised image segmentation method based on augmented tree partitioning. Unlike many existing methods that use a graph structure to model the image, we use a tree-based structure called the augmented tree, which is built up by augmenting several abstract label nodes to the minimum spanning tree of the original graph. We then model image segmentation as the partitioning problem on the augmented tree. Dynamic programming is used to efficiently solve the optimization problem. Experimental results show that our method gives competitive segmentation results, and the speed is much faster than graph-based methods.
Yangqing Jia, Jingdong Wang 0001, Changshui Zhang, Xian-Sheng Hua 0001
ICIP2
2008 Transductive video annotation via local learnable kernel classifier
abstract
One crucial problem in transductive video annotation is how to estimate the label from the neighboring samples. Existing methods such as graph-based Gaussian random filed only considered the pair-wise similarity and then propagated the labels based on it. In this paper, we propose a new method from the perspective of local learning, which formulate the prediction of labels from the neighbors into a learning problem. Our contributions lie in two-fold: (1) we propose a new transductive video annotation method based on local kernel classifier; (2) local learnable is proposed to measure whether a sample can be learned from the neighbors well and we employ this measure into the optimization objective. Experiments on TRECVID 2005 dataset prove that the proposed method is effective and the local learning perspective is promising for video annotation.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
ICME3
2008 Optimized video scene segmentation
abstract
In this paper, we propose an optimized video scene segmentation approach with considering both content coherence and temporally contextual dissimilarity. First, a chain structure is constructed by connecting temporally adjacent shots to represent a video. Then the chain is partitioned such that the content within a chain segment is coherent enough and the contextual similarity of temporally adjacent chain segments is small enough. This task is formulated as a ratio function of content coherence and contextual similarity. Finally, we present an effective and efficient hierarchical chain partitioning approach to find the optimal scene segmentation. Experimental results on a set of home videos and feature movies demonstrate the superiority of the proposed approach over several existing key approaches.
Jingdong Wang 0001, Xinmei Tian 0001, Linjun Yang, Zhengjun Zha, Xian-Sheng Hua 0001
ICME1
2008 Graph-based semi-supervised learning with multi-label
abstract
Conventional graph-based semi-supervised learning methods predominantly focus on single label problem. However, it is more popular in real-world applications that an example is associated with multiple labels simultaneously. In this paper, we propose a novel graph-based learning framework in the setting of semi-supervised learning with multi-label. The proposed approach is characterized by simultaneously exploiting the inherent correlations among multiple labels and the label consistency over the graph. We apply the proposed framework to video annotation and report superior performance compared to key existing approaches over the TRECVID 2006 corpus.
Zhengjun Zha, Tao Mei 0001, Jingdong Wang 0001, Zengfu Wang, Xian-Sheng Hua 0001
ICME3
2008 Finding image exemplars using fast sparse affinity propagation
abstract
In this paper, we propose a novel approach to organize image search results obtained from state-of-the-art image search engines in order to improve user experience. We aim to discover exemplars from search results and simultaneously group the images. The exemplars are delivered to the user as a summary of search results instead of the large amount of unorganized images. This gives the user a brief overview of search results with a small amount of images, and helps the user to further find the images of interest. We adopt the idea of affinity propagation and design a fast sparse affinity propagation algorithm to find exemplars that best represent the image search results. Experiments on real-world data demonstrate the effectiveness of our method both visually and quantitatively.
Yangqing Jia, Jingdong Wang 0001, Changshui Zhang, Xian-Sheng Hua 0001
ACM Multimedia2
2008 Bayesian video search reranking
abstract
Content-based video search reranking can be regarded as a process that uses visual content to recover the "true" ranking list from the noisy one generated based on textual information. This paper explicitly formulates this problem in the Bayesian framework, i.e., maximizing the ranking score consistency among visually similar video shots while minimizing the ranking distance, which represents the disagreement between the objective ranking list and the initial text-based. Different from existing point-wise ranking distance measures, which compute the distance in terms of the individual scores, two new methods are proposed in this paper to measure the ranking distance based on the disagreement in terms of pair-wise orders. Specifically, hinge distance penalizes the pairs with reversed order according to the degree of the reverse, while preference strength distance further considers the preference degree. By incorporating the proposed distances into the optimization objective, two reranking methods are developed which are solved using quadratic programming and matrix computation respectively. Evaluation on TRECVID video search benchmark shows that the performance improvement up to 21% on TRECVID 2006 and 61.11% on TRECVID 2007 are achieved relative to text search baseline.
Xinmei Tian 0001, Linjun Yang, Jingdong Wang 0001, Yichen Yang 0001, Xiuqing Wu, Xian-Sheng Hua 0001
ACM Multimedia3
2008 Semi-Supervised Classification with Universum
abstract
The Universum data, defined as a collection of “non-examples” that do not belong to any class of interest, have been shown to encode some prior knowledge by representing meaningful concepts in the same domain as the problem at hand. In this paper, we address a novel semi-supervised classification problem, called semi-supervised Universum, that can simultaneously utilize the labeled data, unlabeled data and the Universum data to improve the classification performance. We propose a graph based method to make use of the Universum data to help depict the prior information for possible classifiers. Like conventional graph based semi-supervised methods, the graph regularization is also utilized to favor the consistency between the labels. Furthermore, since the proposed method is a graph based one, it can be easily extended to the multiclass case. The empirical experiments on the USPS and MNIST datasets are presented to show that the proposed method can obtain superior performances over conventional supervised and semi-supervised methods.
Dan Zhang 0007, Jingdong Wang 0001, Fei Wang 0001, Changshui Zhang
SDM2
2007 Joint Affinity Propagation for Multiple View Segmentation
abstract
A joint segmentation is a simultaneous segmentation of registered 2D images and 3D points reconstructed from the multiple view images. It is fundamental in structuring the data for subsequent modeling applications. In this paper, we treat this joint segmentation as a weighted graph labeling problem. First, we construct a 3D graph for the joint 3D and 2D points using a joint similarity measure. Then, we propose a hierarchical sparse affinity propagation algorithm to automatically and jointly segment 2D images and group 3D points. Third, a semi-supervised affinity propagation algorithm is proposed to refine the automatic results with the user assistance. Finally, intensive experiments demonstrate the effectiveness of the proposed approaches.
Jianxiong Xiao, Jingdong Wang 0001, Ping Tan 0002, Long Quan
ICCV2
2007 Image-Based Modeling by Joint Segmentation
Long Quan, Jingdong Wang 0001, Ping Tan 0002, Lu Yuan 0001
Int. J. Comput. Vis.2
2007 Face recognition using spectral features
Fei Wang 0001, Jingdong Wang 0001, Changshui Zhang, James T. Kwok
Pattern Recognit.2
2007 Image-based tree modeling
abstract
In this paper, we propose an approach for generating 3D models of natural-looking trees from images that has the additional benefit of requiring little user intervention. While our approach is primarily image-based, we do not model each leaf directly from images due to the large leaf count, small image footprint, and widespread occlusions. Instead, we populate the tree with leaf replicas from segmented source images to reconstruct the overall tree shape. In addition, we use the shape patterns of visible branches to predict those of obscured branches. We demonstrate our approach on a variety of trees.
Ping Tan 0002, Jingdong Wang 0001, Sing Bing Kang, Long Quan
ACM Trans. Graph.3
2006 Picture Collage
abstract
In this paper, we address a novel problem of automatically creating a picture collage from a group of images. Picture collage is a kind of visual image summary - to arrange all input images on a given canvas, allowing overlay, to maximize visible visual information. We formulate the picture collage creation problem in a Bayesian framework. The salient regions of each image are firstly extracted and represented as a set of weighted rectangles. Then, the image arrangement is formulated as a Maximum a Posterior (MAP) problem such that the output picture collage shows as many visible salient regions (without being overlaid by others) from all images as possible. Moreover, a very efficientMarkov chain Monte Carlo (MCMC) method is designed for the optimization. Applications to desktop image browsing and image search result summarization demonstrate the effectiveness of our approach.
Jingdong Wang 0001, Long Quan, Jian Sun 0001, Xiaoou Tang, Harry Shum
CVPR (1)1
2006 Semi-Supervised Classification Using Linear Neighborhood Propagation
abstract
In this paper, we address the general problem of learning from both labeled and unlabeled data. Based on the reasonable assumption that the label of each data can be linearly reconstructed from its neighbors’ labels, we develop a novel approach, called Linear Neighborhood Propagation (LNP), to learn the linear construction weights and predict the labels of the unlabeled data from the labeled data. The main procedure is composed of two steps: (1) it computes the local neighborhood geometry according to the image information by solving a constrained least squares problem, (2) the obtained geometry will be used to predict the labels of unlabeled data by solving a linearly constrained quadratic optimization problem. LNP has at least three characteristics. Firstly, it is much practical in that only one free parameter, the size of neighborhood, requires user-tuning. Secondly it can be easily extended to out-of-sample examples. Finally the labels achieved from LNP can be sufficiently smooth with respect to the intrinsic structure collectively revealed by the labeled and unlabeled points. Many promising experimental results, such as object recognition, data ranking and interactive image segmentation, demonstrate the effectiveness and efficiency of the proposed approach.
Fei Wang 0001, Changshui Zhang, Helen C. Shen, Jingdong Wang 0001
CVPR (1)4
2006 Image-based plant modeling
abstract
In this paper, we propose a semi-automatic technique for modeling plants directly from images. Our image-based approach has the distinct advantage that the resulting model inherits the realistic shape and complexity of a real plant. We designed our modeling system to be interactive, automating the process of shape recovery while relying on the user to provide simple hints on segmentation. Segmentation is performed in both image and 3D spaces, allowing the user to easily visualize its effect immediately. Using the segmented image and 3D data, the geometry of each leaf is then automatically recovered from the multiple views by fitting a deformable leaf model. Our system also allows the user to easily reconstruct branches in a similar manner. We show realistic reconstructions of a variety of plants, and demonstrate examples of plant editing.
Long Quan, Ping Tan 0002, Lu Yuan 0001, Jingdong Wang 0001, Sing Bing Kang
ACM Trans. Graph.5
2005 Data-dependent kernels for high-dimensional data classification
abstract
For high-dimensional data classification problems such as face recognition, one of the most efficient classifiers is the nearest neighbor (NN) classifier. What mostly affects the NN classification performance is the feature extracted by some methods. And the kernel method is one of the efficient methods for extracting features. However, the selection of kernel parameters is still difficult. In this paper, we propose a so-called data dependent kernel (DDK) which is defined by generalizing the Gaussian kernel. Also an efficient and practical method is presented to calculate the DDK parameters. Moreover, one DDK based on subspaces is given to improve the recognition performance. Experiments show that the proposed DDK can achieve promising classification performance in face recognition and SPECT heart diagnosis.
Jingdong Wang 0001, James T. Kwok, Helen C. Shen, Long Quan
IJCNN1
2005 Spectral feature analysis
abstract
We have seen a surge of interest in spectral-based methods and kernel-based methods for machine learning and data mining. Despite the significant research, these methods remain only loosely related. In this paper, we give theoretically an explicit relation between spectral clustering and weighted kernel principal component analysis (WKPCA). We show that spectral clustering is not only a method for data clustering, but also for feature extraction. We are then able to reinterpret the spectral clustering algorithm in terms of WKPCA and propose our spectral feature analysis (SFA) method. The spectral features extracted by SFA can capture the distinguishing information of data from different classes effectively. Finally some experimental results are presented to show the effectiveness of our method.
Fei Wang 0001, Jingdong Wang 0001, Changshui Zhang
IJCNN2
2005 Visual object recognition using probabilistic kernel subspace similarity
Jianguo Lee, Jingdong Wang 0001, Changshui Zhang, Zhaoqi Bian
Pattern Recognit.2
2004 Multi-view EM algorithm and its application to color image segmentation
abstract
We propose a new algorithm, the multi-view expectation and maximization algorithm (multi-view EM), to deal with real-world learning problems where there are some natural split of features. Multi-view EM does feature split in the same manner as co-training and co-EM, two successful semi-supervised learning algorithms in text learning tasks, but it considers the multi-view learning problem in the framework of the EM algorithm. The multi-view EM algorithm has impressive advantages compared with co-training and co-EM: its convergence is theoretically guaranteed; and it can deal with multiple views instead of only two views. We utilize it for color image segmentation and discuss the phenomenon that different weights for the color view and coordinate view lead to different segmentation results.
Xing Yi, Changshui Zhang, Jingdong Wang 0001
ICME3
2004 Probabilistic tangent subspace: a unified view
abstract
Tangent Distance (TD) is one classical method for invariant pattern classification. However, conventional TD need pre-obtain tangent vectors, which is difficult except for image objects. This paper extends TD to more general pattern classification tasks. The basic assumption is that tangent vectors can be approximately represented by the pattern variations. We propose three probabilistic subspace models to encode the variations: the linear subspace, nonlinear subspace, and manifold subspace models. These three models are addressed in a unified view, namely Probabilistic Tangent Subspace (PTS). Experiments show that PTS can achieve promising classification performance in non-image data sets.
Jianguo Lee, Jingdong Wang 0001, Changshui Zhang, Zhaoqi Bian
ICML2
2003 Kernel Trick Embedded Gaussian Mixture Model
Jingdong Wang 0001, Jianguo Lee, Changshui Zhang
ALT1
2003 Color Image Segmentation: Kernel Do the Feature Space
Jianguo Lee, Jingdong Wang 0001, Changshui Zhang
ECML2
2003 Kernel GMM and its application to image binarization
abstract
Gaussian mixture model (GMM) is an efficient method for parametric clustering. However, traditional GMM can't perform clustering well on data set with complex structure such as images. In this paper, kernel trick, successfully used by SVM and kernel PCA, is introduced into EM algorithm for solving parameter estimation of GMM, which is so called kernel GMM (kGMM). The basic idea of kernel GMM is to apply kernel based GMM in feature space instead of in input data space. In order to avoid the curse of dimension, the proposed kGMM also embeds a step to automatically select discriminative features in feature space. kGMM is employed for the task of image binarization. Result shows that the proposed approach is feasible.
Jingdong Wang 0001, Jianguo Lee, Changshui Zhang
ICME1