VLDB 2026 Research / reviewers in the wild / expert
Lei Zhang 0001
dblp:z/LeiZhang
· DBLP profile ↗
240ranked-venue papers
9as first author
86since 2021 · last 2026
0000-0001-6926-0538ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 165 · 9 first-author · 57 since 2021Artificial intelligence and machine learning · 136 · 77 since 2021Databases, data management, data science and information retrieval · 29Applied, interdisciplinary, general and emerging computing · 17 · 2 since 2021Systems, architecture and hardware · 3Computer networks · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Better Code Understanding in Decoder-Only Models with Contrastive LearningabstractRecent advances in large-scale code generation models have led to remarkable progress in producing high-quality code. These models are trained in a self-supervised manner on extensive unlabeled code corpora using a decoder-only architecture. However, despite their generative strength, decoder-only models often exhibit limited performance on code understanding tasks such as code search and clone detection, primarily due to their generation-oriented training objectives. While training large encoder-only models from scratch on massive code datasets can improve understanding ability but remains computationally expensive and time-consuming. In this paper, we explore a more efficient alternative by transferring knowledge from pre-trained decoder-only code generation models to code understanding tasks. We investigate how decoder-only architectures can be effectively adapted to learn discriminative and semantically meaningful code representations. To this end, we propose CL4D, a contrastive learning framework tailored to strengthen the representation capabilities of decoder-only models. Extensive experiments on multiple benchmark datasets demonstrate that CL4D achieves competitive or superior performance compared to existing methods on representative code understanding tasks, including code search and clone detection. Further analysis reveals that CL4D substantially improves the semantic alignment of code representations by reducing the distance between semantically similar code snippets. These findings highlight the feasibility of leveraging decoder-only models as a unified backbone for both code generation and understanding. Jiayi Lin 0004, Yibiao Yang, Lei Zhang 0001 |
AAAI | 4 |
| 2026 | SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D FeaturesabstractIn this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training images, SegDINO3D is designed to fully leverage 2D representation from a pre-trained 2D detection model, including both image-level and object-level features, for improving 3D representation. SegDINO3D takes both a point cloud and its associated 2D images as input. In the encoder stage, it first enriches each 3D point by retrieving 2D image features from its corresponding image views and then leverages a 3D encoder for 3D context fusion. In the decoder stage, it formulates 3D object queries as 3D anchor boxes and performs cross-attention from 3D queries to 2D object queries obtained from 2D images using the 2D detection model. These 2D object queries serve as a compact object-level representation of 2D images, effectively avoiding the challenge of keeping thousands of image feature maps in the memory while faithfully preserving the knowledge of the pre-trained 2D model. The introducing of 3D box queries also enables the model to modulate cross-attention using the predicted boxes for more precise querying. SegDINO3D achieves the state-of-the-art performance on the ScanNetV2 and ScanNet200 3D instance segmentation benchmarks. Notably, on the challenging ScanNet200 dataset, SegDINO3D significantly outperforms prior methods by +8.7 and +6.8 mAP on the validation and hidden test sets, respectively, demonstrating its superiority. Jinyuan Qu, Hongyang Li 0003, Xingyu Chen 0002, Shilong Liu 0004, Yukai Shi, Tianhe Ren, Ruitao Jing, Lei Zhang 0001 |
AAAI | 8 |
| 2026 | Robust-R1: Degradation-Aware Reasoning for Robust Visual UnderstandingabstractMultimodal Large Language Models struggle to maintain reliable performance under extreme real-world visual degradations, which impede their practical robustness. Existing robust MLLMs predominantly rely on implicit training/adaptation that focuses solely on visual encoder generalization, suffering from limited interpretability and isolated optimization. To overcome these limitations, we propose Robust-R1, a novel framework that explicitly models visual degradations through structured reasoning chains. Our approach integrates: (i) supervised fine-tuning for degradation-aware reasoning foundations, (ii) reward-driven alignment for accurately perceiving degradation parameters, and (iii) dynamic reasoning depth scaling adapted to degradation intensity. To facilitate this approach, we introduce a specialized 11K dataset featuring realistic degradations synthesized across four critical real-world visual processing stages, each annotated with structured chains connecting degradation parameters, perceptual influence, pristine semantic reasoning chain, and conclusion. Comprehensive evaluations demonstrate state-of-theart robustness: Robust-R1 outperforms all general and robust baselines on the real-world degradation benchmark R-Bench, while maintaining superior anti-degradation performance under multi-intensity adversarial degradations on MMMB, MMStar, and RealWorldQA. Jiaqi Tang 0005, Jianmin Chen, Wei Wei 0008, Xiaogang Xu 0002, Runtao Liu, Qipeng Xie, Jiafei Wu, Lei Zhang 0001, Qifeng Chen 0001 |
AAAI | 9 |
| 2026 | T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object DetectionabstractObject detection methods have evolved from closed-set to open-set paradigms over the years. Current open-set object detectors, however, remain constrained by their exclusive reliance on positive indicators based on given prompts like text descriptions or visual exemplars. This positive-only paradigm experiences consistent vulnerability to visually similar but semantically different distractors. We propose T-Rex-Omni, a novel framework that addresses this limitation by incorporating negative visual prompts to negate hard negative distractors. Specifically, we first introduce a unified visual prompt encoder that jointly processes positive and negative visual prompts. Next, a training-free Negating Negative Computing (NNC) module is proposed to dynamically suppress negative responses during the probability computing stage. To further boost performance through fine-tuning, our Negating Negative Hinge (NNH) loss enforces discriminative margins between positive and negative embeddings. T-Rex-Omni supports flexible deployment in both positive-only and joint positive-negative inference modes, accommodating either user-specified or automatically generated negative examples. Extensive experiments demonstrate remarkable zero-shot detection performance, significantly narrowing the performance gap between visual-prompted and text-prompted methods while showing particular strength in long-tailed scenarios (51.2 AP_r on LVIS-minival). This work establishes negative prompts as a crucial new dimension for advancing open-set visual recognition systems. Jiazhou Zhou, Kanghao Chen, Lutao Jiang, Yuanhuiyi Lyu, Ying-Cong Chen, Lei Zhang 0001 |
AAAI | 7 |
| 2026 | T-Rex2++: Toward Generic Object Perception via Text-Visual Prompt SynergyabstractWe present T-Rex2++, a unified and highly practical framework for generic open-set object perception, encompassing both object detection and instance segmentation. Previous methods relying on text prompts effectively encapsulate the abstract concept of common objects, but struggle with rare or complex object representation due to data scarcity and descriptive limitations. Conversely, visual prompts excel in depicting novel objects through concrete visual examples, but fall short in conveying the abstract concept of objects as effectively as text prompts. Recognizing these complementary strengths, we introduce a text-visual synergy mechanism that aligns both modalities within a single feature space via contrastive learning. Crucially, T-Rex2++ advances beyond the passive perception paradigm of its predecessor by introducing a novel Universal Prompt. This learnable component models generic objectness, empowering the system to autonomously discover and localize arbitrary objects without any user-provided cues, thereby closing the loop between human-guided interaction and fully automatic perception. Furthermore, we extend the synergy verification to the pixel level by integrating a zero-shot instance segmentation module, demonstrating that our contrastive alignment generalizes robustly to fine-grained masks. Comprehensive experiments demonstrate that T-Rex2++ exhibits strong zero-shot object perception capabilities across a wide spectrum of scenarios, validating T-Rex2++ as a versatile foundation for generic object perception. Feng Li 0040, Zhaoyang Zeng, Tianhe Ren, Shilong Liu 0004, Lei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape EstimationabstractExpressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods focus on training innovative architectural designs on confined datasets. In this work, we investigate the impact of scaling up EHPS towards a family of generalist foundation models. 1) For data scaling, we perform a systematic investigation on 40 EHPS datasets, encompassing a wide range of scenarios that a model trained on any single dataset cannot handle. More importantly, capitalizing on insights obtained from the extensive benchmarking process, we optimize our training scheme and select datasets that lead to a significant leap in EHPS capabilities. Ultimately, we achieve diminishing returns at 10 M training instances from diverse data sources. 2) For model scaling, we take advantage of vision transformers (up to ViT-Huge as the backbone) to study the scaling law of model sizes in EHPS. To exclude the influence of algorithmic design, we base our experiments on two minimalist architectures: SMPLer-X, which consists of an intermediate step for hand and face localization, and SMPLest-X, an even simpler version that reduces the network to its bare essentials and highlights significant advances in the capture of articulated hands. With Big Data and the large model, the foundation models exhibit strong performance across diverse test benchmarks and excellent transferability to even unseen environments. Moreover, our finetuning strategy turns the generalist into specialist models, allowing them to achieve further performance boosts. Notably, our foundation models consistently deliver state-of-the-art results on seven benchmarks such as AGORA, UBody, EgoBody, and our proposed SynHand dataset for comprehensive hand evaluation. Wanqi Yin, Zhongang Cai, Ruisi Wang, Ailing Zeng, Qingping Sun, Haiyi Mei, Hui En Pang, Lei Zhang 0001, Chen Change Loy, Atsushi Yamashita, Lei Yang 0045, Ziwei Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 11 |
| 2026 | Good Performance Estimation Strategies are All You Need in Neural Architecture SearchabstractRecent advances in Neural Architecture Search (NAS) are essentially attributed to Performance Estimation (PE), i.e., a method aims to effectively estimate an architecture. Meanwhile, Kendall's $\tau$τ is well recognized as the principled evaluation criteria for PE strategies in the literature. We argue that Kendall's $\tau$τ is not the optimal solution. Through extensive experiments and theoretical analysis, we take the initiative to reveal the problem behind the Kendall's $\tau$τ and propose a novel criterion named Minimum Keeping Ratio (MKR), which is closely connected to the final performance of NAS. It allows us to compare different PE approaches in a unified perspective, and use effective ablation studies to verify common beliefs and key differences of PE strategies. Based on the findings from MKR, we are able to derive a simple NAS method by integrating different PE strategies with random sampling. Such a method shows very strong performance in efficiency and effectiveness through extensive experiments on different challenging benchmarks. In particular, our simple random sampling NAS finds the optimal architecture in NASbenchMacro, NASbench201, and NASbench301. It is also well generalized to different search spaces (MobileNet) and tasks (semantic segmentation), finding an architecture surpasses the previous state-of-the-art architectures by 4.25 mIoU under $ 600M$600M FLOPs on ADE20K. Codes are available at https://anonymous.4open.science/r/Anonymization11264. Xiawu Zheng, Lei Zhang 0001, Binghan Chen, Fei Chao 0001, Chenglin Wu 0001, Shanshan Wang 0002, Rongrong Ji, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | CoCoCo: Improving Text-Guided Video Inpainting for Better Consistency, Controllability and CompatibilityabstractVideo inpainting is a crucial task with diverse applications, including fine-grained video editing, video recovery, and video dewatermarking. However, most existing video inpainting methods primarily focus on visual content completion while neglecting text information. There are only a limited number of text-guided video inpainting techniques, and these techniques struggle with maintaining visual quality and exhibit poor semantic representation capabilities. In this paper, we introduce CoCoCo, a text-guided video inpainting diffusion framework. To address the aforementioned challenges, we enhance both the training data and model structure. Specifically, we devise an instance-aware region selection strategy for masked area sampling and develop a novel motion block that incorporates efficient 3D full attention and textual cross attention. Additionally, our CoCoCo framework can be seamlessly integrated with various personalized text-to-image diffusion models through a delicate training-free transfer mechanism. Comprehensive experiments demonstrate that CoCoCo can create high-quality visual content with enhanced temporal consistency, improved text controllability, and better compatibility with personalized image models. Bojia Zi, Xianbiao Qi, Yukai Shi, Bin Liang 0004, Rong Xiao 0003, Kam-Fai Wong, Lei Zhang 0001 |
AAAI | 10 |
| 2025 | Adversarial Diffusion Compression for Real-World Image Super-ResolutionabstractReal-world image super-resolution (Real-ISR) aims to reconstruct high-resolution images from low-resolution inputs degraded by complex, unknown processes. While many Stable Diffusion (SD)-based Real-ISR methods have achieved remarkable success, their slow, multi-step inference hinders practical deployment. Recent SD-based one-step networks like OSEDiff and S3Diff alleviate this issue but still incur high computational costs due to their reliance on large pretrained SD models. This paper proposes a novel Real-ISR method, AdcSR, by distilling the one-step diffusion network OSEDiff into a streamlined diffusion-GAN model under our Adversarial Diffusion Compression (ADC) framework. We meticulously examine the modules of OSEDiff, categorizing them into two types: (1) Removable (VAE encoder, prompt extractor, text encoder, etc.) and (2) Prunable (denoising UNet and VAE decoder). Since direct removal and pruning can degrade the model’s generation capability, we pretrain our pruned VAE decoder to restore its ability to decode images and employ adversarial distillation to compensate for performance loss. This ADC-based diffusion-GAN hybrid design effectively reduces complexity by 73% in inference time, 78% in computation, and 74% in parameters, while preserving the model’s generation capability. Experiments manifest that our proposed AdcSR achieves competitive recovery quality on both synthetic and real-world datasets, offering up to 9.3× speedup over previous one-step diffusion-based methods. Code and models are available at https://github.com/Guaishou74851/AdcSR. Bin Chen 0006, Gehui Li, Rongyuan Wu, Jie Chen 0001, Jian Zhang 0018, Lei Zhang 0001 |
CVPR | 7 |
| 2025 | HandOS: 3D Hand Reconstruction in One StageabstractExisting approaches of hand reconstruction predominantly adhere to a multi-stage framework, encompassing detection, left-right classification, and pose estimation. This paradigm induces redundant computation and cumulative errors. In this work, we propose HandOS, an end-to-end framework for 3D hand reconstruction. Our central motivation lies in leveraging a frozen detector as the foundation while incorporating auxiliary modules for 2D and 3D keypoint estimation. In this manner, we integrate the pose estimation capacity into the detection framework, while at the same time obviating the necessity of using the left-right category as a prerequisite. Specifically, we propose an interactive 2D-3D decoder, where 2D joint semantics is derived from detection cues while 3D representation is lifted from those of 2D joints. Furthermore, hierarchical attention is designed to enable the concurrent modeling of 2D joints, 3D vertices, and camera translation. Consequently, we achieve an end-to-end integration of hand detection, 2D pose estimation, and 3D mesh reconstruction within a one-stage framework, so that the above multi-stage drawbacks are overcome. Meanwhile, the HandOS reaches state-of-the-art performances on public benchmarks, e.g., 5.0 PA-MPJPE on FreiHand and 64.6% [email protected] on HInt-Ego4D. Xingyu Chen 0002, Zhuheng Song, Xiaoke Jiang, Yaoqing Hu, Junzhi Yu 0001, Lei Zhang 0001 |
CVPR | 6 |
| 2025 | OSMamba: Omnidirectional Spectral Mamba with Dual-Domain Prior Generator for Exposure CorrectionabstractExposure correction is a fundamental problem in computer vision and image processing. Recently, frequency domainbased methods have achieved impressive improvement, yet they still struggle with complex real-world scenarios under extreme exposure conditions. This is due to the local convolutional receptive fields failing to model long-range dependencies in the spectrum, and the non-generative learning paradigm being inadequate for retrieving lost details from severely degraded regions. In this paper, we propose Omnidirectional Spectral Mamba (OSMamba), a novel exposure correction network that incorporates the advantages of state space models and generative diffusion models to address these limitations. Specifically, OSMamba introduces an omnidirectional spectral scanning mechanism that adapts Mamba to the frequency domain to capture comprehensive long-range dependencies in both the amplitude and phase spectra of deep image features, hence enhancing illumination correction and structure recovery. Furthermore, we develop a dual-domain prior generator that learns from well-exposed images to generate a degradation-free diffusion prior containing correct information about severely under- and over-exposed regions for better detail restoration. Extensive experiments on multiple-exposure and mixed-exposure datasets demonstrate that the proposed OSMamba achieves state-of-the-art performance both quantitatively and qualitatively. Our code and models can be found at https://github.com/cvsym/OSMamba. Gehui Li, Bin Chen 0006, Chen Zhao 0002, Lei Zhang 0001, Jian Zhang 0018 |
CVPR | 4 |
| 2025 | SkillMimic: Learning Basketball Interaction Skills from DemonstrationsabstractTraditional reinforcement learning methods for human- object interaction (HOI) rely on labor-intensive, manually designed skill rewards that do not generalize well across different interactions. We introduce SkillMimic, a unified data-driven framework that fundamentally changes how agents learn interaction skills by eliminating the need for skill-specific rewards. Our key insight is that a unified HOI imitation reward can effectively capture the essence of diverse interaction patterns from HOI datasets. This enables SkillMimic to learn a single policy that not only masters multiple interaction skills but also facilitates skill transitions, with both diversity and generalization improving as the HOI dataset grows. For evaluation, we collect and introduce two basketball datasets containing approximately 35 minutes of diverse basketball skills. Extensive experiments show that SkillMimic successfully masters a wide range of basketball skills including stylistic variations in dribbling, layup, and shooting. Moreover, these learned skills can be effectively composed by a high-level controller to accomplish complex and long-horizon tasks such as consecutive scoring, opening new possibilities for scalable and generalizable interaction skill learning. Project page: https://ingrid789.github.io/SkillMimic/ Yinhuai Wang, Qihan Zhao 0001, Runyi Yu 0003, Hok Wai Tsui, Ailing Zeng, Jiwen Yu, Xiu Li 0001, Qifeng Chen 0001, Jian Zhang 0018, Lei Zhang 0001, Ping Tan 0002 |
CVPR | 12 |
| 2025 | LeanGaussian: Breaking Pixel or Point Cloud Correspondence in Modeling 3D GaussiansabstractRencently, Gaussian splatting has demonstrated significant success in novel view synthesis. Current methods often regress Gaussians with pixel or point cloud correspondence, linking each Gaussian with a pixel or a 3D point. This leads to the redundancy of Gaussians being used to overfit the correspondence rather than the objects represented by the 3D Gaussians themselves, consequently wasting resources and lacking accurate geometries or textures. In this paper, we introduce LeanGaussian, a novel approach that treats each query in deformable Transformer as one 3D Gaussian ellipsoid, breaking the pixel or point cloud correspondence constraints. We leverage deformable decoder to iteratively refine the Gaussians layer-by-layer with the image features as keys and values. Notably, the center of each 3D Gaussian is defined as 3D reference points, which are then projected onto the image for deformable attention in 2D space. On both the ShapeNet SRN dataset (category level) and the Google Scanned Objects dataset (open-category level, trained with the Objaverse dataset), our approach, outperforms prior methods by approximately 6.1%, achieving a PSNR of 25.44 and 22.36, respectively. Additionally, our method achieves a 3D reconstruction speed of 7.2 FPS and rendering speed 500 FPS. Codes are available at https://github.com/jwubz123/LeanGaussian. Kenkun Liu, Xiaoke Jiang, Yuan Yao 0011, Lei Zhang 0001 |
CVPR | 6 |
| 2025 | HumanMM: Global Human Motion Recovery from Multi-shot VideosabstractIn this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions. Such long-sequence in-the-wild motions are highly valuable to applications such as motion generation and motion understanding, but are of great challenge to be recovered due to abrupt shot transitions, partial occlusions, and dynamic backgrounds presented in such videos. Existing methods primarily focus on single-shot videos, where continuity is maintained within a single camera view, or simplify multi-shot alignment in camera space only. In this work, we tackle the challenges by integrating an enhanced camera pose estimation with Human Motion Recovery (HMR) by incorporating a shot transition detector and a robust alignment module for accurate pose and orientation continuity across shots. By leveraging a custom motion integrator, we effectively mitigate the problem of foot sliding and ensure temporal consistency in human pose. Extensive evaluations on our created multi-shot dataset from public 3D human datasets demonstrate the robustness of our method in reconstructing realistic human motion in world coordinates. Guanlin Wu, Zhuokai Zhao, Xiaoke Jiang, Zhuoheng Li, Hao (Frank) Yang, Haoqian Wang, Lei Zhang 0001 |
CVPR | 11 |
| 2025 | Co-Painter: Fine-Grained Controllable Image Stylization via Implicit Decoupling and Adaptive Injection
Wei Wei 0008, Jiaqi Tang 0005, Jiangtao Nie, Yanyu Ye, Xiaogang Xu 0002, Ying-Cong Chen, Lei Zhang 0001 |
ICCV | 8 |
| 2025 | UniGS: Modeling Unitary 3D Gaussians for Novel View Synthesis from Sparse-View Images
Kenkun Liu, Xiaoke Jiang, Yuan Yao 0011, Lei Zhang 0001 |
ICCV | 5 |
| 2025 | Scaling Speech-Text Pre-training with Synthetic Interleaved DataabstractSpeech language models (SpeechLMs) accept speech input and produce speech output, allowing for more natural human-computer interaction compared to text-based large language models (LLMs).
Traditional approaches for developing SpeechLMs are constrained by the limited availability of unsupervised speech data and parallel speech-text data, which are significantly less abundant compared to text pre-training data, thereby limiting their scalability as LLMs.
We propose a novel approach to scaling speech-text pre-training by leveraging large-scale synthetic interleaved data derived from text corpora, eliminating the need for parallel speech-text datasets.
Our method efficiently constructs speech-text interleaved data by sampling text spans from existing text corpora and synthesizing corresponding speech spans using a text-to-token model, bypassing the need to generate actual speech.
We also employ a supervised speech tokenizer derived from an automatic speech recognition (ASR) model by incorporating a vector-quantized bottleneck into the encoder. This supervised training approach results in discrete speech tokens with strong semantic preservation even at lower sampling rates (e.g. 12.5Hz), while still maintaining speech reconstruction quality.
Starting from a pre-trained language model and scaling our pre-training to 1 trillion tokens (with 600B synthetic interleaved speech-text data), we achieve state-of-the-art performance in both speech language modeling and spoken question answering, improving performance on spoken questions tasks from the previous SOTA of 13\% (Moshi) to 31\%.
We further demonstrate that by fine-tuning the pre-trained model with speech dialogue data, we can develop an end-to-end spoken chatbot that achieves competitive performance comparable to existing baselines in both conversational abilities and speech quality, even operating exclusively in the speech domain. Aohan Zeng, Zhengxiao Du, Mingdao Liu, Lei Zhang 0001, Shengmin Jiang, Yuxiao Dong, Jie Tang 0001 |
ICLR | 4 |
| 2025 | Open-Set Image Tagging with Multi-Grained Text SupervisionabstractThis paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark. Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001 |
ACM Multimedia | 9 |
| 2025 | Motion2Motion: Cross-topology Motion Transfer with Sparse CorrespondenceabstractThis work studies the challenge of transfer animations between characters whose skeletal topologies differ substantially. While many techniques have advanced retargeting techniques in decades, transfer motions across diverse topologies remains less-explored. The primary obstacle lies in the inherent topological inconsistency between source and target skeletons, which restricts the establishment of straightforward one-to-one bone correspondences. Besides, the current lack of large-scale paired motion datasets spanning different topological structures severely constrains the development of data-driven approaches. To address these limitations, we introduce Motion2Motion, a novel, training-free framework. Simply yet effectively, Motion2Motion works with only one or a few example motions on the target skeleton, by accessing a sparse set of bone correspondences between the source and target skeletons. Through comprehensive qualitative and quantitative evaluations, we demonstrate that Motion2Motion achieves efficient and reliable performance in both similar-skeleton and cross-species skeleton transfer scenarios. The practical utility of our approach is further evidenced by its successful integration in downstream applications and user interfaces, highlighting its potential for industrial applications. Code and data are available at https://lhchen.top/Motion2Motion. Zixin Yin, Zhiyang Dou, Xin Chen 0040, Jingbo Wang 0003, Taku Komura, Lei Zhang 0001 |
SIGGRAPH Asia | 8 |
| 2025 | A Mutual Supervision Framework for Referring Expression Segmentation and Generation
Shijia Huang, Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Lei Zhang 0001, Liwei Wang 0009 |
Int. J. Comput. Vis. | 5 |
| 2025 | ED-Pose++: Enhanced Explicit Box Detection for Conventional and Interactive Multi-Object Keypoint DetectionabstractDetecting keypoints on diverse objects is essential for fine-grained visual understanding and analysis. This paper introduces Enhanced Explicit Box Detection (ED-Pose++), an end-to-end framework that leverages cascade box regression to realize both conventional and interactive multi-object keypoint detection. Unlike traditional one-stage methods, ED-Pose++ innovatively redefines multi-object keypoint detection as a dual-phase explicit box detection, achieving a unified representation and regression optimization process. Specifically, an object detection decoder first extracts each object's position and global features, establishing a good initialization for subsequent keypoint detection. To bring in contextual information near keypoints, we also regard each keypoint as a small box to learn both positions and their related local contents. In practice, an object-to-keypoint detection decoder adopts a collaborative learning strategy between object and keypoint features, facilitating efficient information propagation between global and local perspectives. Rooted on the architecture, we further equip dual-phase box detection with an interactive mechanism that enables the model to refine its predictions based on limited user feedback. During training, we incorporate an error correction scheme to equip the model with an adept self-correction capability for use during inference. The comprehensive experiments demonstrate ED-Pose++'s superior performance in conventional multi-object keypoint detection tasks. For the first time, ED-Pose++ outperforms heatmap-based top-down approaches across various benchmarks, despite operating within a fully end-to-end architecture. The interactive variant also dramatically reduces more than 10 times the labeling effort of 2D keypoint annotation compared with manual-only annotation. Ailing Zeng, Tianhe Ren, Shilong Liu 0004, Feng Li 0040, Ruimao Zhang, Lei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | TAPTR3D: Decoupled 3D Point Tracking Boosts 2D and Further Enhances 3D Tracking Accuracy
Hongyang Li 0003, Jinyuan Qu, Zhaoyang Zeng, Lei Zhang 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Diffusion Self-Distillation for Remote Sensing Scene ClassificationabstractRemote sensing scene classification, a fundamental task in remote image analysis, has obtained rapid progress due to the powerful capabilities of Convolutional Neural Networks (CNNs). Achieving precise classification performance heavily relies on the feature extraction capacity of the network. However, due to the large variation and severe distortion within the images, extracting robust feature representations is necessary but challenging. Self-distillation could enhance the shallow layers by providing stronger gradients and more accurate supervision from deeper layers, thereby promoting the extraction of spatially detailed features. Nonetheless, due to the limited capacity of shallow layers to learn truly valuable knowledge, shallow layer features can be viewed as the noisy version of deep layer features and contain more disruptive factors, which significantly impedes the effectiveness of self-distillation. To address this issue, in this paper, we establish the Diffusion Self-Distillation Network (DSDNet), which incorporates the conditional diffusion denoising model into the self-distillation framework. Specifically, DSDNet filters noise from shallow features through the diffusion denoising process, enabling more precise and accurate distillation between the refined student features and the teacher features. Extensive experiments on four challenging remote sensing datasets emonstrate that the proposed DSDNet achieves significant performance improvements over various backbone networks with negligible increases in parameters, delivering state-of-the-art classification performance. Our code and dataset are available on https://github.com/toggle1995/DSDNet. Yutao Hu 0002, Lei Zhang 0001, Xiaoyan Luo, Xianbin Cao 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Efficient Antibody Structure Refinement Using Energy-Guided SE(3) Flow MatchingabstractAntibodies are proteins produced by the immune system that recognize and bind to specific antigens, and their 3D structures are crucial for understanding their binding mechanism and designing therapeutic interventions. The specificity of antibody-antigen binding predominantly depends on the complementarity-determining regions (CDR) within antibodies.Despite recent advancements in antibody structure prediction, the quality of predicted CDRs remains suboptimal.In this paper, we develop a novel antibody structure refinement method termed FlowAB based on energy-guided flow matching. FlowAB adopts the powerful deep generative method SE(3) flow matching and simultaneously incorporates important physical prior knowledge into the flow model to guide the generation process.The extensive experiments demonstrate that FlowAB can significantly improve the antibody CDR structures. It achieves new state-of-the-art performance on the antibody structure prediction task when used in conjunction with an appropriate prior model while incurring only marginal computational overhead. This advantage makes FlowAB a practical tool in antibody engineering. Jiying Zhang, Zijing Liu, Shengyuan Bai, He Cao, Yu Li 0003, Lei Zhang 0001 |
BIBM | 6 |
| 2024 | Visual in-Context PromptingabstractIn-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existing visual prompting methods focus on referring segmentation to segment the most relevant object, falling short of addressing many generic vision tasks like open-set segmentation and detection. In this paper, we introduce a universal visual in-context prompting framework for both tasks, as shown in Fig. 1. In particular, we build on top of an encoder-decoder architecture, and develop a versatile prompt encoder to support a variety of prompts like strokes, boxes, and points. We further enhance it to take an arbitrary number of reference image segments as the context. Our extensive explorations show that the proposed visual in-context prompting elicits extraordinary referring and generic segmentation capabilities to refer and detect, yielding competitive performance to close-set in-domain datasets and showing promising results on many open-set segmentation datasets. By joint training on COCO and SA-1B, DINOv achieves 57.7 PQ on COCO and 23.2 PQ on ADE20K. Code will be available at https://github.com/UX-Decoder/DINOv Feng Li 0040, Hao Zhang 0097, Tianhe Ren, Shilong Liu 0004, Xueyan Zou, Huaizhe Xu, Hongyang Li 0003, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001 |
CVPR | 11 |
| 2024 | T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy
Feng Li 0040, Zhaoyang Zeng, Tianhe Ren, Shilong Liu 0004, Lei Zhang 0001 |
ECCV (33) | 6 |
| 2024 | TAPTR: Tracking Any Point with Transformers as Detection
Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Lei Zhang 0001 |
ECCV (16) | 7 |
| 2024 | Segment and Recognize Anything at Any Granularity
Feng Li 0040, Hao Zhang 0097, Peize Sun, Xueyan Zou, Shilong Liu 0004, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001 |
ECCV (48) | 8 |
| 2024 | LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Shilong Liu 0004, Hao Cheng 0002, Hao Zhang 0097, Feng Li 0040, Tianhe Ren, Xueyan Zou, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001, Jianfeng Gao 0001, Chunyuan Li |
ECCV (47) | 11 |
| 2024 | Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection
Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Hao Zhang 0097, Chunyuan Li, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ECCV (47) | 12 |
| 2024 | Compress3D: A Compressed Latent Space for 3D Generation from a Single Image
Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Xi Zhao 0002 |
ECCV (18) | 4 |
| 2024 | Progressive3D: Progressively Local Editing for Text-to-3D Content Creation with Complex Semantic PromptsabstractRecent text-to-3D generation methods achieve impressive 3D content creation capacity thanks to the advances in image diffusion models and optimizing strategies. However, current methods struggle to generate correct 3D content for a complex prompt in semantics, i.e., a prompt describing multiple interacted objects binding with different attributes. In this work, we propose a general framework named Progressive3D, which decomposes the entire generation into a series of locally progressive editing steps to create precise 3D content for complex prompts, and we constrain the content change to only occur in regions determined by user-defined region prompts in each editing step. Furthermore, we propose an overlapped semantic component suppression technique to encourage the optimization process to focus more on the semantic differences between prompts. Extensive experiments demonstrate that the proposed Progressive3D framework generates precise 3D content for prompts with complex semantics through progressive editing steps and is general for various text-to-3D methods driven by different 3D representations. Xinhua Cheng, Tianyu Yang 0003, Yu Li 0003, Lei Zhang 0001, Jian Zhang 0018, Li Yuan 0007 |
ICLR | 5 |
| 2024 | DreamTime: An Improved Optimization Strategy for Diffusion-Guided 3D GenerationabstractText-to-image diffusion models pre-trained on billions of image-text pairs have recently enabled 3D content creation by optimizing a randomly initialized differentiable 3D representation with score distillation. However, the optimization process suffers slow convergence and the resultant 3D models often exhibit two limitations: (a) quality concerns such as missing attributes and distorted shape and texture; (b) extremely low diversity comparing to text-guided image synthesis. In this paper, we show that the conflict between the 3D optimization process and uniform timestep sampling in score distillation is the main reason for these limitations. To resolve this conflict, we propose to prioritize timestep sampling with monotonically non-increasing functions, which aligns the 3D optimization process with the sampling process of diffusion model. Extensive experiments show that our simple redesign significantly improves 3D content creation with faster convergence, better quality and diversity. Yukai Shi, Boshi Tang, Xianbiao Qi, Lei Zhang 0001 |
ICLR | 6 |
| 2024 | Tag2Text: Guiding Vision-Language Model via Image TaggingabstractThis paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance. Youcai Zhang, Jinyu Ma, Rui Feng 0001, Yuejie Zhang, Yandong Guo, Lei Zhang 0001 |
ICLR | 9 |
| 2024 | TOSS: High-quality Text-guided Novel View Synthesis from a Single ImageabstractIn this paper, we present TOSS, which introduces text to the task of novel view synthesis (NVS) from just a single RGB image.
While Zero123 has demonstrated impressive zero-shot open-set NVS capabilities, it treats NVS as a pure image-to-image translation problem. This approach suffers from the challengingly under-constrained nature of single-view NVS: the process lacks means of explicit user control and often result in implausible NVS generations.
To address this limitation, TOSS uses text as high-level semantic information to constrain the NVS solution space.
TOSS fine-tunes text-to-image Stable Diffusion pre-trained on large-scale text-image pairs and introduces modules specifically tailored to image and camera pose conditioning, as well as dedicated training for pose correctness and preservation of fine details.
Comprehensive experiments are conducted with results showing that our proposed TOSS outperforms Zero123 with higher-quality NVS results and faster convergence. We further support these results with comprehensive ablations that underscore the effectiveness and potential of
the introduced semantic guidance and architecture design. Yukai Shi, He Cao, Boshi Tang, Xianbiao Qi, Tianyu Yang 0003, Shilong Liu 0004, Lei Zhang 0001, Harry Shum |
ICLR | 9 |
| 2024 | HumanTOMATO: Text-aligned Whole-body Motion GenerationabstractThis work targets a novel text-driven **whole-body** motion generation task, which takes a given textual description as input and aims at generating high-quality, diverse, and coherent facial expressions, hand gestures, and body motions simultaneously. Previous works on text-driven motion generation tasks mainly have two limitations: they ignore the key role of fine-grained hand and face controlling in vivid whole-body motion generation, and lack a good alignment between text and motion. To address such limitations, we propose a Text-aligned whOle-body Motion generATiOn framework, named HumanTOMATO, which is the first attempt to our knowledge towards applicable holistic motion generation in this research area. To tackle this challenging task, our solution includes two key designs: (1) a Holistic Hierarchical VQ-VAE (aka H${}^{2}$VQ) and a Hierarchical-GPT for fine-grained body and hand motion reconstruction and generation with two structured codebooks; and (2) a pre-trained text-motion-alignment model to help generated motion align with the input textual description explicitly. Comprehensive experiments verify that our model has significant advantages in both the quality of generated motions and their alignment with text. Shunlin Lu, Ailing Zeng, Ruimao Zhang, Lei Zhang 0001, Harry Shum |
ICML | 6 |
| 2024 | A small object detection network for remote sensing based on CS-PANet and DSAN
Lei Zhang 0001, Fengxian Wang, Yibin Chen |
Multim. Tools Appl. | 4 |
| 2024 | DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingabstractWe present in this paper a novel denoising training method to speed up DETR (DEtection TRansformer) training and offer a deepened understanding of the slow convergence issue of DETR-like methods. We show that the slow convergence results from the instability of bipartite graph matching which causes inconsistent optimization goals in early training stages. To address this issue, except for the Hungarian loss, our method additionally feeds GT bounding boxes with noises into the Transformer decoder and trains the model to reconstruct the original boxes, which effectively reduces the bipartite graph matching difficulty and leads to faster convergence. Our method is universal and can be easily plugged into any DETR-like method by adding dozens of lines of code to achieve a remarkable improvement. As a result, our DN-DETR results in a remarkable improvement ( +1.9AP) under the same setting and achieves 46.0 AP and 49.5 AP trained for 12 and 50 epochs with the ResNet-50 backbone. Compared with the baseline under the same setting, DN-DETR achieves comparable performance with 50% training epochs. We also demonstrate the effectiveness of denoising training in CNN-based detectors (Faster R-CNN), segmentation models (Mask2Former, Mask DINO), and more DETR-based models (DETR, Anchor DETR, Deformable DETR). Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Jian Guo 0016, Lionel M. Ni, Lei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | TLDW: Extreme Multimodal Summarization of News VideosabstractMultimodal summarisation with multimodal output is drawing increasing attention due to the rapid growth of multimedia data. While several methods have been proposed to summarise visual-text contents, their multimodal outputs are not succinct enough at an extreme level to address the information overload issue. To the end of extreme multimodal summarisation, we introduce a new task, eXtreme Multimodal Summarisation with Multimodal Output (XMSMO) for the scenario of TL;DW - Too Long; Didn’t Watch, akin to TL;DR. XMSMO aims to summarise a video-document pair into a summary with an extremely short length, which consists of one cover frame as the visual summary and one sentence as the textual summary. We propose a novel unsupervised Hierarchical Optimal Transport Network (HOT-Net) consisting of three components: hierarchical multimodal encoder, hierarchical multimodal fusion decoder, and optimal transport solver. Our method is trained, without using reference summaries, by optimising the visual and textual coverage from the perspectives of the distance between the semantic distributions under optimal transport plans. To facilitate the study on this task, we constructed a large-scale dataset, XMSMO-News, by harvesting 4,891 video-document pairs. The experimental results show that our method achieves promising performance in terms of ROUGE and IoU metrics. Our dataset and source code will be publicly available in GitHub. Peggy Tang, Kun Hu 0008, Lei Zhang 0001, Jiebo Luo 0001, Zhiyong Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Survey of Code Search Based on Deep LearningabstractCode writing is repetitive and predictable, inspiring us to develop various code intelligence techniques. This survey focuses on code search, that is, to retrieve code that matches a given natural language query by effectively capturing the semantic similarity between the query and code. Deep learning, being able to extract complex semantics information, has achieved great success in this field. Recently, various deep learning methods, such as graph neural networks and pretraining models, have been applied to code search with significant progress. Deep learning is now the leading paradigm for code search. In this survey, we provide a comprehensive overview of deep learning-based code search. We review the existing deep learning-based code search framework that maps query/code to vectors and measures their similarity. Furthermore, we propose a new taxonomy to illustrate the state-of-the-art deep learning-based code search in a three-step process: query semantics modeling, code semantics modeling, and matching modeling, which involves the deep learning model training. Finally, we suggest potential avenues for future research in this promising field. Jiayi Lin 0004, Hande Dong, Lei Zhang 0001, Zhonghai Wu |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and GroundingabstractIn this paper, we study the problem of visual grounding by considering both phrase extraction and grounding (PEG). In contrast to the previous phrase-known-at-test setting, PEG requires a model to extract phrases from text and locate objects from image simultaneously, which is a more practical setting in real applications. As phrase extraction can be regarded as a 1D text segmentation problem, we formulate PEG as a dual detection problem and propose a novel DQ-DETR model, which introduces dual queries to probe different features from image and text for object prediction and phrase mask prediction. Each pair of dual queries are designed to have shared positional parts but different content parts. Such a design effectively alleviates the difficulty of modality alignment between image and text (in contrast to a single query design) and empowers Transformer decoder to leverage phrase mask-guided attention to improve the performance. To evaluate the performance of PEG, we also propose a new metric CMAP (cross-modal average precision), analogous to the AP metric in object detection. The new metric overcomes the ambiguity of Recall@1 in many-box-to-one-phrase cases in phrase grounding. As a result, our PEG pre-trained DQ-DETR establishes new state-of-the-art results on all visual grounding benchmarks with a ResNet-101 backbone. For example, it achieves 91.04% and 83.51% in terms of recall rate on RefCOCO testA and testB with a ResNet-101 backbone. Shilong Liu 0004, Shijia Huang, Feng Li 0040, Hao Zhang 0097, Yaoyuan Liang, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
AAAI | 8 |
| 2023 | Are Transformers Effective for Time Series Forecasting?abstractRecently, there has been a surge of Transformer-based solutions for the long-term time series forecasting (LTSF) task. Despite the growing performance over the past few years, we question the validity of this line of research in this work. Specifically, Transformers is arguably the most successful solution to extract the semantic correlations among the elements in a long sequence. However, in time series modeling, we are to extract the temporal relations in an ordered set of continuous points. While employing positional encoding and using tokens to embed sub-series in Transformers facilitate preserving some ordering information, the nature of the permutation-invariant self-attention mechanism inevitably results in temporal information loss. To validate our claim, we introduce a set of embarrassingly simple one-layer linear models named LTSF-Linear for comparison. Experimental results on nine real-life datasets show that LTSF-Linear surprisingly outperforms existing sophisticated Transformer-based LTSF models in all cases, and often by a large margin. Moreover, we conduct comprehensive empirical studies to explore the impacts of various design elements of LTSF models on their temporal relation extraction capability. We hope this surprising finding opens up new research directions for the LTSF task. We also advocate revisiting the validity of Transformer-based solutions for other time series analysis tasks (e.g., anomaly detection) in the future. Ailing Zeng, Muxi Chen, Lei Zhang 0001, Qiang Xu 0001 |
AAAI | 3 |
| 2023 | DisCo-CLIP: A Distributed Contrastive Loss for Memory Efficient CLIP TrainingabstractWe propose DisCo-CLIP, a distributed memory-efficient CLIP training approach, to reduce the memory consumption of contrastive loss when training contrastive learning models. Our approach decomposes the contrastive loss and its gradient computation into two parts, one to calculate the intra-GPU gradients and the other to compute the inter-GPU gradients. According to our decomposition, only the intra-GPU gradients are computed on the current GPU, while the inter-GPU gradients are collected via all_reduce from other GPUs instead of being repeatedly computed on every GPU. In this way, we can reduce the GPU memory consumption of contrastive loss computation from$\mathcal{O}(B^{2})$to$\mathcal{O}(\frac{B^{2}}{N})$, where$B$and$N$are the batch size and the number of GPUs used for training. Such a distributed solution is mathematically equivalent to the original non-distributed contrastive loss computation, without sacrificing any computation accuracy. It is particularly efficient for large-batch CLIP training. For instance, DisCo-CLIP can enable contrastive training of a ViT-B/32 model with a batch size of 32K or 196K using 8 or 64 A100 40GB GPUs, compared with the original CLIP solution which requires 128 A100 40GB GPUs to train a ViT-B/32 model with a batch size of 32K. Xianbiao Qi, Lei Zhang 0001 |
CVPR | 4 |
| 2023 | Human-Art: A Versatile Human-Centric Dataset Bridging Natural and Artificial ScenesabstractHumans have long been recorded in a variety of forms since antiquity. For example, sculptures and paintings were the primary media for depicting human beings before the invention of cameras. However, most current human-centric computer vision tasks like human pose estimation and human image generation focus exclusively on natural images in the real world. Artificial humans, such as those in sculptures, paintings, and cartoons, are commonly neglected, making existing models fail in these scenarios. As an abstraction of life, art incorporates humans in both natural and artificial scenes. We take advantage of it and introduce the Human-Art dataset to bridge related tasks in natural and artificial scenarios. Specifically, Human-Art contains 50k high-quality images with over 123k person instances from 5 natural and 15 artificial scenarios, which are annotated with bounding boxes, keypoints, self-contact points, and text information for humans represented in both 2D and 3D. It is, therefore, comprehensive and versatile for various downstream tasks. We also provide a rich set of baseline results and detailed analyses for related tasks, including human detection, 2D and 3D human pose estimation, image generation, and motion transfer. As a challenging dataset, we hope Human-Art can provide insights for relevant research and open up new research questions. Xuan Ju, Ailing Zeng, Qiang Xu 0001, Lei Zhang 0001 |
CVPR | 5 |
| 2023 | Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and SegmentationabstractIn this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks (instance, panoptic, and semantic). It makes use of the query embeddings from DINO to dot-product a high-resolution pixel embedding map to predict a set of binary masks. Some key components in DINO are extended for segmentation through a shared architecture and training process. Mask DINO is simple, efficient, and scalable, and it can benefit from joint large-scale detection and segmentation datasets. Our experiments show that Mask DINO significantly outperforms all existing specialized segmentation methods, both on a ResNet-50 backbone and a pre-trained model with SwinL backbone. Notably, Mask DINO establishes the best results to date on instance segmentation (54.5 AP on COCO), panoptic segmentation (59.4 PQ on COCO), and semantic segmentation (60.8 mIoU on ADE20K) among models under one billion parameters. Code is available at https://github.com/IDEA-Research/MaskDINO. Feng Li 0040, Hao Zhang 0097, Huaizhe Xu, Shilong Liu 0004, Lei Zhang 0001, Lionel M. Ni, Harry Shum |
CVPR | 5 |
| 2023 | Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETRabstractRecent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of low-level features, are quite computationally inefficient, which hinders real applications of DETR models. In this paper, we present Lite DETR, a simple yet efficient end-to-end object detection framework that can effectively reduce the GFLOPs of the detection head by 60% while keeping 99% of the original performance. Specifically, we design an efficient encoder block to update high-level features (corresponding to small-resolution feature maps) and low-level features (corresponding to large-resolution feature maps) in an interleaved way. In addition, to better fuse cross-scale features, we develop a key-aware deformable attention to predict more reliable attention weights. Comprehensive experiments validate the effectiveness and efficiency of the proposed Lite DETR, and the efficient encoder strategy can generalize well across existing DETR-based models. The code will be available in https://github.com/IDEA-Research/Lite-DETR. Feng Li 0040, Ailing Zeng, Shilong Liu 0004, Hao Zhang 0097, Hongyang Li 0003, Lei Zhang 0001, Lionel M. Ni |
CVPR | 6 |
| 2023 | One-Stage 3D Whole-Body Mesh Recovery with Component Aware TransformerabstractWhole-body mesh recovery aims to estimate the 3D human body, face, and hands parameters from a single image. It is challenging to perform this task with a single network due to resolution issues, i.e., the face and hands are usually located in extremely small regions. Existing works usually detect hands and faces, enlarge their resolution to feed in a specific network to predict the parameter, and finally fuse the results. While this copy-paste pipeline can capture the fine-grained details of the face and hands, the connections between different parts cannot be easily recovered in late fusion, leading to implausible 3D rotation and unnatural pose. In this work, we propose a one-stage pipeline for expressive whole-body mesh recovery, named OSX, without separate networks for each part. Specifically, we design a Component Aware Transformer (CAT) composed of a global body encoder and a local face/hand decoder. The encoder predicts the body parameters and provides a high-quality feature map for the decoder, which performs a feature-level upsample-crop scheme to extract highresolution part-specific features and adopt keypointguided deformable attention to estimate hand and face precisely. The whole pipeline is simple yet effective without any manual post-processing and naturally avoids implausible prediction. Comprehensive experiments demonstrate the effectiveness of OSX. Lastly, we build a large-scale Upper-Body dataset (UBody) with high-quality 2D and 3D whole-body annotations. It contains persons with partially visible bodies in diverse real-life scenarios to bridge the gap between the basic task and downstream applications. Ailing Zeng, Haoqian Wang, Lei Zhang 0001, Yu Li 0003 |
CVPR | 4 |
| 2023 | MP-Former: Mask-Piloted Transformer for Image SegmentationabstractWe present a mask-piloted Transformer which improves masked-attention in Mask2Former for image segmentation. The improvement is based on our observation that Mask2Former suffers from inconsistent mask predictions between consecutive decoder layers, which leads to inconsistent optimization goals and low utilization of decoder queries. To address this problem, we propose a mask-piloted training approach, which additionally feeds noised ground-truth masks in masked-attention and trains the model to reconstruct the original ones. Compared with the predicted masks used in mask-attention, the ground-truth masks serve as a pilot and effectively alleviate the negative impact of inaccurate mask predictions in Mask2Former. Based on this technique, our MP-Former achieves a remarkable performance improvement on all three image segmentation tasks (instance, panoptic, and semantic), yielding +2.3AP and +1.6mIoU on the Cityscapes instance and semantic segmentation tasks with a ResNet-50 backbone. Our method also significantly speeds up the training, outperforming Mask2Former with half of the number of training epochs on ADE20K with both a ResNet-50 and a Swin-L backbones. Moreover, our method only introduces little computation during training and no extra computation during inference. Our code will be released at https://github.com/IDEA-Research/MP-Former. Hao Zhang 0097, Feng Li 0040, Huaizhe Xu, Shijia Huang, Shilong Liu 0004, Lionel M. Ni, Lei Zhang 0001 |
CVPR | 7 |
| 2023 | Automatic Network Pruning via Hilbert-Schmidt Independence Criterion Lasso under Information Bottleneck PrincipleabstractMost existing neural network pruning methods hand-crafted their importance criteria and structures to prune. This constructs heavy and unintended dependencies on heuristics and expert experience for both the objective and the parameters of the pruning approach. In this paper, we try to solve this problem by introducing a principled and unified framework based on Information Bottleneck (IB) theory, which further guides us to an automatic pruning approach. Specifically, we first formulate the channel pruning problem from an IB perspective, and then implement the IB principle by solving a Hilbert-Schmidt Independence Criterion (HSIC) Lasso problem under certain conditions. Based on the theoretical guidance, we then provide an automatic pruning scheme by searching for global penalty coefficients. Verified by extensive experiments, our method yields state-of-the-art performance on various benchmark networks and datasets. For example, with VGG-16, we achieve a 60%-FLOPs reduction by removing 76% of the parameters, with an improvement of 0.40% in top-1 accuracy on CIFAR-10. With ResNet-50, we achieve a 56%-FLOPs reduction by removing 50% of the parameters, with a small loss of 0.08% in the top-1 accuracy on ImageNet. The code is available at https://github.com/sunggo/APIB. Song Guo 0001, Lei Zhang 0001, Xiawu Zheng, Yan Wang 0059, Fei Chao 0001, Chenglin Wu 0001, Shengchuan Zhang, Rongrong Ji |
ICCV | 2 |
| 2023 | HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image GenerationabstractControllable human image generation (HIG) has numerous real-life applications. State-of-the-art solutions, such as ControlNet and T2I-Adapter, introduce an additional learnable branch on top of the frozen pre-trained stable diffusion (SD) model, which can enforce various conditions, including skeleton guidance of HIG. While such a plug-and-play approach is appealing, the inevitable and uncertain conflicts between the original images produced from the frozen SD branch and the given condition incur significant challenges for the learnable branch, which essentially conducts image feature editing for condition enforcement.In this work, we propose a native skeleton-guided diffusion model for controllable HIG called HumanSD. Instead of performing image editing with dual-branch diffusion, we fine-tune the original SD model using a novel heatmap-guided denoising loss. This strategy effectively and efficiently strengthens the given skeleton condition during model training while mitigating the catastrophic forgetting effects. HumanSD is fine-tuned on the assembly of three large-scale human-centric datasets with text-image-pose information, two of which are established in this work. Experimental results show that HumanSD outperforms ControlNet in terms of pose control and image quality, particularly when the given skeleton guidance is sophisticated. Code and data are available at: https://idea-research.github.io/HumanSD/. Xuan Ju, Ailing Zeng, Chenchen Zhao 0001, Lei Zhang 0001, Qiang Xu 0001 |
ICCV | 5 |
| 2023 | DFA3D: 3D Deformable Attention For 2D-to-3D Feature LiftingabstractIn this paper, we propose a new operator, called 3D DeFormable Attention (DFA3D), for 2D-to-3D feature lifting, which transforms multi-view 2D image features into a unified 3D space for 3D object detection. Existing feature lifting approaches, such as Lift-Splat-based and 2D attention-based, either use estimated depth to get pseudo LiDAR features and then splat them to a 3D space, which is a one-pass operation without feature refinement, or ignore depth and lift features by 2D attention mechanisms, which achieve finer semantics while suffering from a depth ambiguity problem. In contrast, our DFA3D-based method first leverages the estimated depth to expand each view’s 2D feature map to 3D and then utilizes DFA3D to aggregate features from the expanded 3D feature maps. With the help of DFA3D, the depth ambiguity problem can be effectively alleviated from the root, and the lifted features can be progressively refined layer by layer, thanks to the Transformerlike architecture. In addition, we propose a mathematically equivalent implementation of DFA3D which can significantly improve its memory efficiency and computational speed. We integrate DFA3D into several methods that use 2D attention-based feature lifting with only a few modifications in code and evaluate on the nuScenes dataset. The experiment results show a consistent improvement of +1.41% mAP on average, and up to +15.1% mAP improvement when high-quality depth information is available, demonstrating the superiority, applicability, and huge potential of DFA3D. The code is available at https://github.com/IDEAResearch/3D-deformable-attention.git. Hongyang Li 0003, Hao Zhang 0097, Zhaoyang Zeng, Shilong Liu 0004, Feng Li 0040, Tianhe Ren, Lei Zhang 0001 |
ICCV | 7 |
| 2023 | Detection Transformer with Stable MatchingabstractThis paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is caused by a multi-optimization path problem, which is highlighted by the one-to-one matching design in DETR. To address this problem, we show that the most important design is to use and only use positional metrics (like IOU) to supervise classification scores of positive examples. Under the principle, we propose two simple yet effective modifications by integrating positional metrics to DETR’s classification loss and matching cost, named position-supervised loss and position-modulated cost. We verify our methods on several DETR variants. Our methods show consistent improvements over baselines. By integrating our methods with DINO, we achieve 50.4 and 51.5 AP on the COCO detection benchmark using ResNet-50 backbones under 1× (12 epochs) and 2× (24 epochs) training settings, achieving a new record under the same setting. We achieve 63.8 AP on COCO detection test-dev with a Swin-Large backbone. Our code will be made available at https://github.com/IDEA-Research/Stable-DINO. Shilong Liu 0004, Tianhe Ren, Zhaoyang Zeng, Hao Zhang 0097, Feng Li 0040, Hongyang Li 0003, Jun Huang 0007, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ICCV | 11 |
| 2023 | Neural Interactive Keypoint DetectionabstractThis work proposes an end-to-end neural interactive keypoint detection framework named Click-Pose, which can significantly reduce more than 10 times labeling costs of 2D keypoint annotation compared with manual-only annotation. Click-Pose explores how user feedback can cooperate with a neural keypoint detector to correct the predicted keypoints in an interactive way for a faster and more effective annotation process. Specifically, we design the pose error modeling strategy that inputs the ground truth pose combined with four typical pose errors into the decoder and trains the model to reconstruct the correct poses, which enhances the self-correction ability of the model. Then, we attach an interactive human-feedback loop that allows receiving users’ clicks to correct one or several predicted keypoints and iteratively utilizes the decoder to update all other keypoints with a minimum number of clicks (NoC) for efficient annotation. We validate Click-Pose in in-domain, out-of-domain scenes, and a new task of keypoint adaptation. For annotation, Click-Pose only needs 1.97 and 6.45 NoC@95 (at precision 95%) on COCO and Human-Art, reducing 31.4% and 36.3% efforts than the SOTA model (ViTPose) with manual correction, respectively. Besides, without user clicks, Click-Pose surpasses the previous end-to-end model by 1.4 AP on COCO and 3.0 AP on Human-Art. Ailing Zeng, Feng Li 0040, Shilong Liu 0004, Ruimao Zhang, Lei Zhang 0001 |
ICCV | 6 |
| 2023 | A Simple Framework for Open-Vocabulary Segmentation and DetectionabstractWe present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets. To bridge the gap of vocabulary and annotation granularity, we first introduce a pre-trained text encoder to encode all the visual concepts in two tasks and learn a common semantic space for them. This gives us reasonably good results compared with the counterparts trained on segmentation task only. To further reconcile them, we identify two discrepancies: i) task discrepancy – segmentation requires extracting masks for both foreground objects and background stuff, while detection merely cares about the former; ii) data discrepancy – box and mask annotations are with different spatial granularity, and thus not directly interchangeable. To address these issues, we propose a decoupled decoding to reduce the interference between foreground/background and a conditioned mask decoding to assist in generating masks for given boxes. To this end, we develop a simple encoder-decoder model encompassing all three techniques and train it jointly on COCO and Objects365. After pre-training, our model exhibits competitive or stronger zero-shot transferability for both segmentation and detection. Specifically, OpenSeeD beats the state-of-the-art method for open-vocabulary instance and panoptic segmentation across 5 datasets, and outperforms previous work for open-vocabulary detection on LVIS and ODinW under similar settings. When transferred to specific tasks, our model achieves new SoTA for panoptic segmentation on COCO and ADE20K, and instance segmentation on ADE20K and Cityscapes (The bottom row in Fig. 1 shows a comparison of the performance of OpenSeeD and previous SoTA methods). Finally, we note that OpenSeeD is the first to explore the potential of joint training on segmentation and detection, and hope it can be received as a strong baseline for developing a single model for both tasks in the open world. Code will be released at https://github.com/IDEA-Research/OpenSeeD. Hao Zhang 0097, Feng Li 0040, Xueyan Zou, Shilong Liu 0004, Chunyuan Li, Lei Zhang 0001 |
ICCV | 7 |
| 2023 | DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection
Hao Zhang 0097, Feng Li 0040, Shilong Liu 0004, Lei Zhang 0001, Hang Su 0006, Jun Zhu 0001, Lionel M. Ni, Harry Shum |
ICLR | 4 |
| 2023 | LipsFormer: Introducing Lipschitz Continuity to Vision Transformers
Xianbiao Qi, Yukai Shi, Lei Zhang 0001 |
ICLR | 5 |
| 2023 | Explicit Box Detection Unifies End-to-End Multi-Person Pose Estimation
Ailing Zeng, Shilong Liu 0004, Feng Li 0040, Ruimao Zhang, Lei Zhang 0001 |
ICLR | 6 |
| 2023 | TopicCAT: Unsupervised Topic-Guided Co-Attention Transformer for Extreme Multimodal SummarisationabstractThe exponential growth of multimedia data has sparked a surge of interest in multimodal summarisation with multimodal output (MSMO). A relatively unexplored but essential task within this field is extreme multimodal summarisation, a process that involves creating extremely concise multimodal summaries to further address the issue of multimedia information overload. In this study, we propose a novel Unsupervised Topic-guided Co-Attention Transformer (TopicCAT) neural network to produce extreme multimodal summaries for video-document pairs. The approach consists of two learning stages for a comprehensive multimodal understanding, guided by topic-based insights: a unimodal learning stage and a cross-modal learning stage, in which a cross-modal topic model is devised to capture the overarching themes present in both documents and videos. To achieve unsupervised learning, eliminating the need for resource-expensive collection of ground-truth multimodal summaries, we propose an optimal transport-based optimisation scheme to evaluate summary coverage from a semantic distribution perspective at the topic-level. Comprehensive experiments demonstrate the effectiveness of our proposed TopicCAT method on a multimodal news dataset, achieving a BERTScore of 84.46 and an accuracy of 0.60. Peggy Tang, Kun Hu 0008, Lei Zhang 0001, Junbin Gao, Jiebo Luo 0001, Zhiyong Wang 0001 |
ACM Multimedia | 3 |
| 2023 | SMPLer-X: Scaling Up Expressive Human Pose and Shape EstimationabstractExpressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined set of training datasets. In this work, we investigate scaling up EHPS towards the first generalist foundation model (dubbed SMPLer-X), with up to ViT-Huge as the backbone and training with up to 4.5M instances from diverse data sources. With big data and the large model, SMPLer-X exhibits strong performance across diverse test benchmarks and excellent transferability to even unseen environments. 1) For the data scaling, we perform a systematic investigation on 32 EHPS datasets, including a wide range of scenarios that a model trained on any single dataset cannot handle. More importantly, capitalizing on insights obtained from the extensive benchmarking process, we optimize our training scheme and select datasets that lead to a significant leap in EHPS capabilities. 2) For the model scaling, we take advantage of vision transformers to study the scaling law of model sizes in EHPS. Moreover, our finetuning strategy turn SMPLer-X into specialist models, allowing them to achieve further performance boosts. Notably, our foundation model SMPLer-X consistently delivers state-of-the-art results on seven benchmarks such as AGORA (107.2 mm NMVE), UBody (57.4 mm PVE), EgoBody (63.6 mm PVE), and EHF (62.3 mm PVE without finetuning). Zhongang Cai, Wanqi Yin, Ailing Zeng, Qingping Sun, Wang Yanjun, Hui En Pang, Haiyi Mei, Lei Zhang 0001, Chen Change Loy, Lei Yang 0059, Ziwei Liu 0002 |
NeurIPS | 10 |
| 2023 | DreamWaltz: Make a Scene with Complex 3D Animatable AvatarsabstractWe present DreamWaltz, a novel framework for generating and animating complex 3D avatars given text guidance and parametric human body prior. While recent methods have shown encouraging results for text-to-3D generation of common objects, creating high-quality and animatable 3D avatars remains challenging. To create high-quality 3D avatars, DreamWaltz proposes 3D-consistent occlusion-aware Score Distillation Sampling (SDS) to optimize implicit neural representations with canonical poses. It provides view-aligned supervision via 3D-aware skeleton conditioning which enables complex avatar generation without artifacts and multiple faces. For animation, our method learns an animatable 3D avatar representation from abundant image priors of diffusion model conditioned on various poses, which could animate complex non-rigged avatars given arbitrary poses without retraining. Extensive evaluations demonstrate that DreamWaltz is an effective and robust approach for creating 3D avatars that can take on complex shapes and appearances as well as novel poses for animation. The proposed framework further enables the creation of complex scenes with diverse compositions, including avatar-avatar, avatar-object and avatar-scene interactions. See https://dreamwaltz3d.github.io/ for more vivid 3D avatar and animation results. Ailing Zeng, He Cao, Xianbiao Qi, Yukai Shi, Zhengjun Zha, Lei Zhang 0001 |
NeurIPS | 8 |
| 2023 | Motion-X: A Large-scale 3D Expressive Whole-body Human Motion DatasetabstractIn this paper, we present Motion-X, a large-scale 3D expressive whole-body motion dataset. Existing motion datasets predominantly contain body-only poses, lacking facial expressions, hand gestures, and fine-grained pose descriptions. Moreover, they are primarily collected from limited laboratory scenes with textual descriptions manually labeled, which greatly limits their scalability. To overcome these limitations, we develop a whole-body motion and text annotation pipeline, which can automatically annotate motion from either single- or multi-view videos and provide comprehensive semantic labels for each video and fine-grained whole-body pose descriptions for each frame. This pipeline is of high precision, cost-effective, and scalable for further research. Based on it, we construct Motion-X, which comprises 15.6M precise 3D whole-body pose annotations (i.e., SMPL-X) covering 81.1K motion sequences from massive scenes. Besides, Motion-X provides 15.6M frame-level whole-body pose descriptions and 81.1K sequence-level semantic labels. Comprehensive experiments demonstrate the accuracy of the annotation pipeline and the significant benefit of Motion-X in enhancing expressive, diverse, and natural motion generation, as well as 3D whole-body human mesh recovery. Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, Lei Zhang 0001 |
NeurIPS | 7 |
| 2023 | A Comprehensive Benchmark for Neural Human Radiance FieldsabstractThe past two years have witnessed a significant increase in interest concerning NeRF-based human body rendering. While this surge has propelled considerable advancements, it has also led to an influx of methods and datasets. This explosion complicates experimental settings and makes fair comparisons challenging. In this work, we design and execute thorough studies into unified evaluation settings and metrics to establish a fair and reasonable benchmark for human NeRF models. To reveal the effects of extant models, we benchmark them against diverse and hard scenes. Additionally, we construct a cross-subject benchmark pre-trained on large-scale datasets to assess generalizable methods. Finally, we analyze the essential components for animatability and generalizability, and make HumanNeRF from monocular videos generalizable, as the inaugural baseline. We hope these benchmarks and analyses could serve the community. Kenkun Liu, Derong Jin, Ailing Zeng, Xiaoguang Han 0001, Lei Zhang 0001 |
NeurIPS | 5 |
| 2022 | Large-Scale Pre-training for Person Re-identification with Noisy LabelsabstractThis paper aims to address the problem of pretraining for person re-identification (Re-ID) with noisy labels. To setup the pretraining task, we apply a simple online multi-object tracking system on raw videos of an existing un-labeled Re-ID dataset “LUPerson” and build the Noisy Labeled variant called “LUPerson-NL”. Since theses ID labels automatically derived from tracklets inevitably con-tain noises, we develop a large-scale Pre-training frame-work utilizing Noisy Labels (PNL), which consists of three learning modules: supervised Re-ID learning, prototype-based contrastive learning, and label-guided contrastive learning. In principle, joint learning of these three mod-ules not only clusters similar examples to one prototype, but also rectifies noisy labels based on the prototype as-signment. We demonstrate that learning directly from raw videos is a promising alternative for pre-training, which utilizes spatial and temporal correlations as weak super-vision. This simple pre-training task provides a scalable way to learn SOTA Re-ID representations from scratch on “LUPerson-NL” without bells and whistles. For example, by applying on the same supervised Re-ID method MGN, our pre-trained model improves the mAP over the unsu-pervised pre-training counterpart by 5.7%, 2.2%, 2.3% on CUHK03, DukeMTMC, and MSMT17 respectively. Under the small-scale or few-shot setting, the performance gain is even more significant, suggesting a better transferability of the learned representation. Code is available at https://github.com/DengpanFu/LUPerson-NL. Dengpan Fu, Dongdong Chen 0001, Hao Yang 0036, Jianmin Bao, Lu Yuan 0001, Lei Zhang 0001, Houqiang Li, Fang Wen 0001, Dong Chen 0003 |
CVPR | 6 |
| 2022 | DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingabstractWe present in this paper a novel denoising training method to speedup DETR (DEtection TRansformer) training and offer a deepened understanding of the slow convergence issue of DETR-like methods. We show that the slow convergence results from the instability of bipartite graph matching which causes inconsistent optimization goals in early training stages. To address this issue, except for the Hungarian loss, our method additionally feeds ground-truth bounding boxes with noises into Transformer decoder and trains the model to reconstruct the original boxes, which effectively reduces the bipartite graph matching difficulty and leads to a faster convergence. Our method is universal and can be easily plugged into any DETR-like methods by adding dozens of lines of code to achieve a remarkable improvement. As a result, our DN-DETR results in a remarkable improvement (+1.9AP) under the same setting and achieves the best result (AP 43.4 and 48.6 with 12 and 50 epochs of training respectively) among DETR-like methods with ResNet-50 backbone. Compared with the baseline under the same setting, DN-DETR achieves comparable performance with 50% training epochs. Code is available at https://github.com/FengLi-ust/DN-DETR. Feng Li 0040, Hao Zhang 0097, Shilong Liu 0004, Jian Guo 0016, Lionel M. Ni, Lei Zhang 0001 |
CVPR | 6 |
| 2022 | Grounded Language-Image Pre-trainingabstractThis paper presents a grounded language-image pretraining (GLIP) model for learning object-level, language-aware, and semantic-rich visual representations. GLIP unifies object detection and phrase grounding for pre-training. The unification brings two benefits: 1) it allows GLIP to learn from both detection and grounding data to improve both tasks and bootstrap a good grounding model; 2) GLIP can leverage massive image-text pairs by generating grounding boxes in a self-training fashion, making the learned representations semantic-rich. In our experiments, we pre-train GLIP on 27M grounding data, including 3M human-annotated and 24M web-crawled image-text pairs. The learned representations demonstrate strong zero-shot and few-shot transferability to various object-level recognition tasks. 1) When directly evaluated on COCO and LVIS (without seeing any images in COCO during pre-training), GLIP achieves 49.8 AP and 26.9 AP, respectively, surpassing many supervised baselines.11Supervised baselines on COCO object detection: Faster-RCNN w/ ResNet50 (40.2) or ResNet101 (42.0), and DyHead w/ Swin-Tiny (49.7). 2) After fine-tuned on COCO, GLIP achieves 60.8 AP on val and 61.5 AP on test-dev, surpassing prior SoTA. 3) When transferred to 13 downstream object detection tasks, a 1-shot GLIP rivals with a fully-supervised Dynamic Head. Code will be released at https://github.com/microsoft/GLIP. Liunian Harold Li, Pengchuan Zhang, Haotian Zhang 0005, Chunyuan Li, Yiwu Zhong, Lu Yuan 0001, Lei Zhang 0001, Jenq-Neng Hwang, Kai-Wei Chang 0001, Jianfeng Gao 0001 |
CVPR | 9 |
| 2022 | Neural Architecture Search with Representation Mutual InformationabstractPerformance evaluation strategy is one of the most important factors that determine the effectiveness and efficiency in Neural Architecture Search (NAS). Existing strategies, such as employing standard training or performance predictor, often suffer from high computational complexity and low generality. To address this issue, we propose to rank architectures by Representation Mutual Information (RMI). Specifically, given an arbitrary architecture that has decent accuracy, architectures that have high RMI with it always yield good accuracies. As an accurate performance indicator to facilitate NAS, RMI not only generalizes well to different search spaces, but is also efficient enough to evaluate architectures using only one batch of data. Building upon RMI, we further propose a new search algorithm termed RMI-NAS, facilitating with a theorem to guarantee the global optimal of the searched architecture. In particular, RMI-NAS first randomly samples architectures from the search space, which are then effectively classified as positive or negative samples by RMI. We then use these samples to train a random forest to explore new regions, while keeping track of the distribution of positive architectures. When the sample size is sufficient, the architecture with the largest probability from the aforementioned distribution is selected, which is theoretically proved to be the optimal solution. The architectures searched by our method achieve remarkable top-1 accuracies with the magnitude times faster search process. Besides, RMI-NAS also generalizes to different datasets and search spaces. Our code has been made available at https://git.openi.org.cn/PCL_AutoML/XNAS. Xiawu Zheng, Lei Zhang 0001, Chenglin Wu 0001, Fei Chao 0001, Jianzhuang Liu, Wei Zeng 0006, Yonghong Tian 0001, Rongrong Ji |
CVPR | 3 |
| 2022 | Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning
Mark Hamilton, Scott M. Lundberg, Stephanie Fu, Lei Zhang 0001, William T. Freeman |
ICLR | 4 |
| 2022 | DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR
Shilong Liu 0004, Feng Li 0040, Hao Zhang 0097, Xiao Yang 0028, Xianbiao Qi, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ICLR | 8 |
| 2022 | Online multi-object tracking with unsupervised re-identification learning and occlusion estimation
Qiankun Liu 0001, Dongdong Chen 0001, Qi Chu 0001, Lu Yuan 0001, Bin Liu 0016, Lei Zhang 0001, Nenghai Yu |
Neurocomputing | 6 |
| 2022 | Towards generalizable detection of face forgery via self-guided model-agnostic learning
Xiao Yang 0028, Shilong Liu 0004, Yinpeng Dong, Hang Su 0006, Lei Zhang 0001, Jun Zhu 0001 |
Pattern Recognit. Lett. | 5 |
| 2021 | VIVO: Visual Vocabulary Pre-Training for Novel Object CaptioningabstractIt is highly desirable yet challenging to generate image captions that can describe novel objects which are unseen in caption-labeled training data, a capability that is evaluated in the novel object captioning challenge (nocaps). In this challenge, no additional image-caption training data, other than COCO Captions, is allowed for model training. Thus, conventional Vision-Language Pre-training (VLP) methods cannot be applied. This paper presents VIsual VOcabulary pre-training (VIVO) that performs pre-training in the absence of caption annotations. By breaking the dependency of paired image-caption training data in VLP, VIVO can leverage large amounts of paired image-tag data to learn a visual vocabulary. This is done by pre-training a multi-layer Transformer model that learns to align image-level tags with their corresponding image region features. To address the unordered nature of image tags, VIVO uses a Hungarian matching loss with masked tag prediction to conduct pre-training. We validate the effectiveness of VIVO by fine-tuning the pre-trained model for image captioning. In addition, we perform an analysis of the visual-text alignment inferred by our model. The results show that our model can not only generate fluent image captions that describe novel objects, but also identify the locations of these objects. Our single model has achieved new state-of-the-art results on nocaps and surpassed the human CIDEr score. Xiaowei Hu 0006, Xi Yin 0006, Lei Zhang 0001, Jianfeng Gao 0001, Zicheng Liu 0001 |
AAAI | 4 |
| 2021 | Dynamic Head: Unifying Object Detection Heads With AttentionsabstractThe complex nature of combining localization and classification in object detection has resulted in the flourished development of methods. Previous works tried to improve the performance in various object detection heads but failed to present a unified view. In this paper, we present a novel dynamic head framework to unify object detection heads with attentions. By coherently combining multiple self-attention mechanisms between feature levels for scale-awareness, among spatial locations for spatial-awareness, and within output channels for task-awareness, the proposed approach significantly improves the representation ability of object detection heads without any computational overhead. Further experiments demonstrate that the effectiveness and efficiency of the proposed dynamic head on the COCO benchmark. With a standard ResNeXt-101-DCN backbone, we largely improve the performance over popular object detectors and achieve a new state-of-the-art at 54.0 AP. The code will be released at https://github.com/microsoft/DynamicHead. Xiyang Dai, Yinpeng Chen, Bin Xiao 0004, Dongdong Chen 0001, Mengchen Liu, Lu Yuan 0001, Lei Zhang 0001 |
CVPR | 7 |
| 2021 | Unsupervised Pre-Training for Person Re-IdentificationabstractIn this paper, we present a large scale unlabeled person re-identification (Re-ID) dataset "LUPerson" and make the first attempt of performing unsupervised pre-training for improving the generalization ability of the learned person Re-ID feature representation. This is to address the problem that all existing person Re-ID datasets are all of limited scale due to the costly effort required for data annotation. Previous research tries to leverage models pre-trained on ImageNet to mitigate the shortage of person Re-ID data but suffers from the large domain gap between ImageNet and person Re-ID data. LUPerson is an unlabeled dataset of 4M images of over 200K identities, which is 30× larger than the largest existing Re-ID dataset. It also covers a much diverse range of capturing environments (e.g., camera settings, scenes, etc.). Based on this dataset, we systematically study the key factors for learning Re-ID features from two perspectives: data augmentation and contrastive loss. Unsupervised pre-training performed on this large-scale dataset effectively leads to a generic Re-ID feature that can benefit all existing person Re-ID methods. Using our pre-trained model in some basic frameworks, our methods achieve state-of-the-art results without bells and whistles on four widely used Re-ID datasets: CUHK03, Market1501, DukeMTMC, and MSMT17. Our results also show that the performance improvement is more significant on small-scale target datasets or under few-shot setting. Dengpan Fu, Dongdong Chen 0001, Jianmin Bao, Hao Yang 0036, Lu Yuan 0001, Lei Zhang 0001, Houqiang Li, Dong Chen 0003 |
CVPR | 6 |
| 2021 | Unsupervised Part Segmentation Through Disentangling Appearance and ShapeabstractWe study the problem, of unsupervised discovery and segmentation of object parts, which, as an intermediate local representation, are capable of finding intrinsic object structure and providing more explainable recognition results. Recent unsupervised methods have greatly relaxed the dependency on annotated data which are costly to obtain, but still rely on additional information such as object segmentation mask or saliency map. To remove such a dependency and further improve the part segmentation performance, we develop a novel approach by disentangling the appearance and shape representations of object parts followed with reconstruction losses without using additional object mask information. To avoid degenerated solutions, a bottleneck block is designed to squeeze and expand the appearance representation, leading to a more effective disentanglement between geometry and appearance. Combined with a self-supervised part classification loss and an improved geometry concentration constraint, we can segment more consistent parts with semantic meanings. Comprehensive experiments on a wide variety of objects such as face, bird, and PASCAL VOC objects demonstrate the effectiveness of the proposed method. Shilong Liu 0004, Lei Zhang 0001, Xiao Yang 0028, Hang Su 0006, Jun Zhu 0001 |
CVPR | 2 |
| 2021 | TAP: Text-Aware Pre-Training for Text-VQA and Text-CaptionabstractIn this paper, we propose Text-Aware Pre-training (TAP) for Text-VQA and Text-Caption tasks. These two tasks aim at reading and understanding scene text in images for question answering and image caption generation, respectively. In contrast to conventional vision-language pretraining that fails to capture scene text and its relationship with the visual and text modalities, TAP explicitly incorporates scene text (generated from OCR engines) during pretraining. With three pre-training tasks, including masked language modeling (MLM), image-text (contrastive) matching (ITM), and relative (spatial) position prediction (RPP), pre-training with scene text effectively helps the model learn a better aligned representation among the three modalities: text word, visual object, and scene text. Due to this aligned representation learning, even pre-trained on the same downstream task dataset, TAP already boosts the absolute accuracy on the TextVQA dataset by +5:4%, compared with a non-TAP baseline. To further improve the performance, we build a large-scale scene text-related imagetext dataset based on the Conceptual Caption dataset, named OCR-CC, which contains 1:4 million images with scene text. Pre-trained on this OCR-CC dataset, our approach outperforms the state of the art by large margins on multiple tasks, i.e., +8:3% accuracy on TextVQA, +8:6% accuracy on ST-VQA, and +10:2 CIDEr score on TextCaps. Zhengyuan Yang, Yijuan Lu, Xi Yin 0006, Dinei A. F. Florêncio, Cha Zhang, Lei Zhang 0001, Jiebo Luo 0001 |
CVPR | 8 |
| 2021 | Lite-HRNet: A Lightweight High-Resolution NetworkabstractWe present an efficient high-resolution network, Lite-HRNet, for human pose estimation. We start by simply applying the efficient shuffle block in ShuffleNet to HRNet (high-resolution network), yielding stronger performance over popular lightweight networks, such as MobileNet, ShuffleNet, and Small HRNet. We find that the heavily-used pointwise (1 × 1) convolutions in shuffle blocks become the computational bottleneck. We introduce a lightweight unit, conditional channel weighting, to replace costly pointwise (1 × 1) convolutions in shuffle blocks. The complexity of channel weighting is linear w.r.t the number of channels and lower than the quadratic time complexity for pointwise convolutions. Our solution learns the weights from all the channels and over multiple resolutions that are readily available in the parallel branches in HRNet. It uses the weights as the bridge to exchange information across channels and resolutions, compensating the role played by the pointwise (1 × 1) convolution. Lite-HRNet demonstrates superior results on human pose estimation over popular lightweight networks. Moreover, Lite-HRNet can be easily applied to semantic segmentation task in the same lightweight manner. The code and models have been publicly available at https://github.com/HRNet/Lite-HRNet. Changqian Yu, Bin Xiao 0004, Changxin Gao, Lu Yuan 0001, Lei Zhang 0001, Nong Sang, Jingdong Wang 0001 |
CVPR | 5 |
| 2021 | VinVL: Revisiting Visual Representations in Vision-Language ModelsabstractThis paper presents a detailed study of improving visual representations for vision language (VL) tasks and develops an improved object detection model to provide object-centric representations of images. Compared to the most widely used bottom-up and top-down model [2], the new model is bigger, better-designed for VL tasks, and pre-trained on much larger training corpora that combine multiple public annotated object detection datasets. Therefore, it can generate representations of a richer collection of visual objects and concepts. While previous VL research focuses mainly on improving the vision-language fusion model and leaves the object detection model improvement untouched, we show that visual features matter significantly in VL models. In our experiments we feed the visual features generated by the new object detection model into a Transformer-based VL fusion model OSCAR [20], and utilize an improved approach OSCAR+ to pre-train the VL model and fine-tune it on a wide range of downstream VL tasks. Our results show that the new visual features significantly improve the performance across all VL tasks, creating new state-of-the-art results on seven public benchmarks. Code, models and pre-extracted features are released at https://github.com/pzzhang/VinVL. Pengchuan Zhang, Xiujun Li, Xiaowei Hu 0006, Lei Zhang 0001, Yejin Choi 0001, Jianfeng Gao 0001 |
CVPR | 5 |
| 2021 | DAP: Detection-Aware Pre-Training With Weak SupervisionabstractThis paper presents a detection-aware pre-training (DAP) approach, which leverages only weakly-labeled classification-style datasets (e.g., ImageNet) for pretraining, but is specifically tailored to benefit object detection tasks. In contrast to the widely used image classification-based pre-training (e.g., on ImageNet), which does not include any location-related training tasks, we transform a classification dataset into a detection dataset through a weakly supervised object localization method based on Class Activation Maps to directly pre-train a detector, making the pre-trained model location-aware and capable of predicting bounding boxes. We show that DAP can outperform the traditional classification pre-training in terms of both sample efficiency and convergence speed in downstream detection tasks including VOC and COCO. In particular, DAP boosts the detection accuracy by a large margin when the number of examples in the downstream task is small. Yuanyi Zhong, Jian Peng 0001, Yu-Xiong Wang, Lei Zhang 0001 |
CVPR | 6 |
| 2021 | Dynamic DETR: End-to-End Object Detection with Dynamic AttentionabstractIn this paper, we present a novel Dynamic DETR (Detection with Transformers) approach by introducing dynamic attentions into both the encoder and decoder stages of DETR to break its two limitations on small feature resolution and slow training convergence. To address the first limitation, which is due to the quadratic computational complexity of the self-attention module in Transformer encoders, we propose a dynamic encoder to approximate the Transformer encoder’s attention mechanism using a convolution-based dynamic encoder with various attention types. Such an encoder can dynamically adjust attentions based on multiple factors such as scale importance, spatial importance, and representation (i.e., feature dimension) importance. To mitigate the second limitation of learning difficulty, we introduce a dynamic decoder by replacing the cross-attention module with a ROI-based dynamic attention in the Transformer decoder. Such a decoder effectively assists Transformers to focus on region of interests from a coarse-to-fine manner and dramatically lowers the learning difficulty, leading to a much faster convergence with fewer training epochs. We conduct a series of experiments to demonstrate our advantages. Our Dynamic DETR significantly reduces the training epochs (by 14×), yet results in a much better performance (by 3.6 on mAP). Meanwhile, in the standard 1× setup with ResNet-50 backbone, we archive a new state-of-the-art performance that further proves the learning effectiveness of the proposed approach. Xiyang Dai, Yinpeng Chen, Pengchuan Zhang, Lu Yuan 0001, Lei Zhang 0001 |
ICCV | 6 |
| 2021 | Improve Unsupervised Pretraining for Few-label TransferabstractUnsupervised pretraining has achieved great success and many recent works have shown unsupervised pretraining can achieve comparable or even slightly better transfer performance than supervised pretraining on downstream target datasets. But in this paper, we find this conclusion may not hold when the target dataset has very few labeled samples for finetuning, i.e., few-label transfer. We analyze the possible reason from the clustering perspective: 1) The clustering quality of target samples is of great importance to few-label transfer; 2) Though contrastive learning is essential to learn how to cluster, its clustering quality is still inferior to supervised pretraining due to lack of label supervision. Based on the analysis, we interestingly discover that only involving some unlabeled target domain into the unsupervised pretraining can improve the clustering quality, subsequently reducing the transfer performance gap with supervised pretraining. This finding also motivates us to propose a new progressive few-label transfer algorithm for real applications, which aims to maximize the transfer performance under a limited annotation budget. To support our analysis and proposed method, we conduct extensive experiments on nine different target datasets. Experimental results show our proposed method can significantly boost the few-label transfer performance of unsupervised pretraining. Suichan Li, Dongdong Chen 0001, Yinpeng Chen, Lu Yuan 0001, Lei Zhang 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
ICCV | 5 |
| 2021 | MicroNet: Improving Image Recognition with Extremely Low FLOPsabstractThis paper aims at addressing the problem of substantial performance degradation at extremely low computational cost (e.g. 5M FLOPs on ImageNet classification). We found that two factors, sparse connectivity and dynamic activation function, are effective to improve the accuracy. The former avoids the significant reduction of network width, while the latter mitigates the detriment of reduction in network depth. Technically, we propose micro-factorized convolution, which factorizes a convolution matrix into low rank matrices, to integrate sparse connectivity into convolution. We also present a new dynamic activation function, named Dynamic Shift Max, to improve the non-linearity via maxing out multiple dynamic fusions between an input feature map and its circular channel shift. Building upon these two new operators, we arrive at a family of networks, named MicroNet, that achieves significant performance gains over the state of the art in the low FLOP regime. For instance, under the constraint of 12M FLOPs, MicroNet achieves 59.4% top-1 accuracy on ImageNet classification, outperforming MobileNetV3 by 9.6%. Source code is at https://github.com/liyunsheng13/micronet. Yunsheng Li, Yinpeng Chen, Xiyang Dai, Dongdong Chen 0001, Mengchen Liu, Lu Yuan 0001, Zicheng Liu 0001, Lei Zhang 0001, Nuno Vasconcelos |
ICCV | 8 |
| 2021 | CvT: Introducing Convolutions to Vision TransformersabstractWe present in this paper a new architecture, named Convolutional vision Transformer (CvT), that improves Vision Transformer (ViT) in performance and efficiency by introducing convolutions into ViT to yield the best of both de-signs. This is accomplished through two primary modifications: a hierarchy of Transformers containing a new convolutional token embedding, and a convolutional Transformer block leveraging a convolutional projection. These changes introduce desirable properties of convolutional neural networks (CNNs) to the ViT architecture (i.e. shift, scale, and distortion invariance) while maintaining the merits of Transformers (i.e. dynamic attention, global context, and better generalization). We validate CvT by conducting extensive experiments, showing that this approach achieves state-of-the-art performance over other Vision Transformers and ResNets on ImageNet-1k, with fewer parameters and lower FLOPs. In addition, performance gains are maintained when pretrained on larger datasets (e.g. ImageNet-22k) and fine-tuned to downstream tasks. Pretrained on ImageNet-22k, our CvT-W24 obtains a top-1 accuracy of 87.7% on the ImageNet-1k val set. Finally, our results show that the positional encoding, a crucial component in existing Vision Transformers, can be safely re-moved in our model, simplifying the design for higher resolution vision tasks. Code will be released at https://github.com/microsoft/CvT. Haiping Wu, Bin Xiao 0004, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan 0001, Lei Zhang 0001 |
ICCV | 7 |
| 2021 | Multi-Scale Vision Longformer: A New Vision Transformer for High-Resolution Image EncodingabstractThis paper presents a new Vision Transformer (ViT) architecture Multi-Scale Vision Longformer, which significantly enhances the ViT of [12] for encoding high-resolution images using two techniques. The first is the multi-scale model structure, which provides image encodings at multiple scales with manageable computational cost. The second is the attention mechanism of Vision Long-former, which is a variant of Longformer [3], originally developed for natural language processing, and achieves a linear complexity w.r.t. the number of input tokens. A comprehensive empirical study shows that the new ViT significantly outperforms several strong baselines, including the existing ViT models and their ResNet counterparts, and the Pyramid Vision Transformer from a concurrent work [47], on a range of vision tasks, including image classification, object detection, and segmentation. The models and source code are released at https://github.com/microsoft/vision-longformer. Pengchuan Zhang, Xiyang Dai, Bin Xiao 0004, Lu Yuan 0001, Lei Zhang 0001, Jianfeng Gao 0001 |
ICCV | 6 |
| 2021 | SEED: Self-supervised Distillation For Visual Representation
Zhiyuan Fang, Lei Zhang 0001, Yezhou Yang, Zicheng Liu 0001 |
ICLR | 4 |
| 2021 | Chasing Sparsity in Vision Transformers: An End-to-End ExplorationabstractVision transformers (ViTs) have recently received explosive popularity, but their enormous model sizes and training costs remain daunting. Conventional post-training pruning often incurs higher training budgets. In contrast, this paper aims to trim down both the training memory overhead and the inference complexity, without sacrificing the achievable accuracy. We carry out the first-of-its-kind comprehensive exploration, on taking a unified approach of integrating sparsity in ViTs "from end to end''. Specifically, instead of training full ViTs, we dynamically extract and train sparse subnetworks, while sticking to a fixed small parameter budget. Our approach jointly optimizes model parameters and explores connectivity throughout training, ending up with one sparse network as the final output. The approach is seamlessly extended from unstructured to structured sparsity, the latter by considering to guide the prune-and-grow of self-attention heads inside ViTs. We further co-explore data and architecture sparsity for additional efficiency gains by plugging in a novel learnable token selector to adaptively determine the currently most vital patches. Extensive results on ImageNet with diverse ViT backbones validate the effectiveness of our proposals which obtain significantly reduced computational cost and almost unimpaired generalization. Perhaps most surprisingly, we find that the proposed sparse (co-)training can sometimes \textit{improve the ViT accuracy} rather than compromising it, making sparsity a tantalizing "free lunch''. For example, our sparsified DeiT-Small at ($5\%$, $50\%$) sparsity for (data, architecture), improves $\mathbf{0.28\%}$ top-1 accuracy, and meanwhile enjoys $\mathbf{49.32\%}$ FLOPs and $\mathbf{4.40\%}$ running time savings. Our codes are available at https://github.com/VITA-Group/SViTE. Tianlong Chen 0001, Yu Cheng 0001, Zhe Gan, Lu Yuan 0001, Lei Zhang 0001, Zhangyang Wang |
NeurIPS | 5 |
| 2021 | Vision-Language Navigation Policy Learning and AdaptationabstractVision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms baseline methods by 10 percent on Success Rate weighted by Path Length (SPL) and achieves the state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore and adapt to unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7 to 11.7 percent). Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | Unified Vision-Language Pre-Training for Image Captioning and VQAabstractThis paper presents a unified Vision-Language Pre-training (VLP) model. The model is unified in that (1) it can be fine-tuned for either vision-language generation (e.g., image captioning) or understanding (e.g., visual question answering) tasks, and (2) it uses a shared multi-layer transformer network for both encoding and decoding, which differs from many existing methods where the encoder and decoder are implemented using separate models. The unified VLP model is pre-trained on a large amount of image-text pairs using the unsupervised learning objectives of two tasks: bidirectional and sequence-to-sequence (seq2seq) masked vision-language prediction. The two tasks differ solely in what context the prediction conditions on. This is controlled by utilizing specific self-attention masks for the shared transformer network. To the best of our knowledge, VLP is the first reported model that achieves state-of-the-art results on both vision-language generation and understanding tasks, as disparate as image captioning and visual question answering, across three challenging benchmark datasets: COCO Captions, Flickr30k Captions, and VQA 2.0. The code and the pre-trained models are available at https://github.com/LuoweiZhou/VLP. Luowei Zhou, Hamid Palangi, Lei Zhang 0001, Houdong Hu, Jason J. Corso, Jianfeng Gao 0001 |
AAAI | 3 |
| 2020 | MagGAN: High-Resolution Face Attribute Editing with Mask-Guided Generative Adversarial Network
Yi Wei 0006, Zhe Gan, Wenbo Li 0001, Siwei Lyu, Ming-Ching Chang, Lei Zhang 0001, Jianfeng Gao 0001, Pengchuan Zhang |
ACCV (4) | 6 |
| 2020 | Large-Scale Intelligent MicroservicesabstractDeploying Machine Learning (ML) algorithms within databases is a challenge due to the varied computational footprints of modern ML algorithms and the myriad of database technologies each with their own restrictive syntax. We introduce an Apache Spark-based micro-service orchestration framework that extends database operations to include web service primitives. Our system can orchestrate web services across hundreds of machines and takes full advantage of cluster, thread, and asynchronous parallelism. Using this framework, we provide large scale clients for intelligent services such as speech, vision, search, anomaly detection, and text analysis. This allows users to integrate ready-to-use intelligence into any datastore with an Apache Spark connector. To eliminate the majority of overhead from network communication, we also introduce a low-latency containerized version of our architecture. Finally, we demonstrate that the services we investigate are competitive on a variety of benchmarks, and present two applications of this framework to create intelligent search engines, and real time auto race analytics systems. Mark Hamilton, Nick Gonsalves, Anand Raman, Brendan Walsh, Siddhartha Prasad, Dalitso Banda, Lucy Zhang, Lei Zhang 0001, William T. Freeman |
IEEE BigData | 9 |
| 2020 | HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationabstractBottom-up human pose estimation methods have difficulties in predicting the correct pose for small persons due to challenges in scale variation. In this paper, we present HigherHRNet: a novel bottom-up human pose estimation method for learning scale-aware representations using high-resolution feature pyramids. Equipped with multi-resolution supervision for training and multi-resolution aggregation for inference, the proposed approach is able to solve the scale variation challenge in bottom-up multi-person pose estimation and localize keypoints more precisely, especially for small person. The feature pyramid in HigherHRNet consists of feature map outputs from HRNet and upsampled higher-resolution outputs through a transposed convolution. HigherHRNet outperforms the previous best bottom-up method by 2.5% AP for medium person on COCO test-dev, showing its effectiveness in handling scale variation. Furthermore, HigherHRNet achieves new state-of-the-art result on COCO test-dev (70.5% AP) without using refinement or other post-processing techniques, surpassing all existing bottom-up methods. HigherHRNet even surpasses all top-down methods on CrowdPose test (67.6% AP), suggesting its robustness in crowded scene. Bowen Cheng, Bin Xiao 0004, Jingdong Wang 0001, Humphrey Shi, Thomas S. Huang, Lei Zhang 0001 |
CVPR | 6 |
| 2020 | Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
Xiujun Li, Xi Yin 0006, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu 0006, Lei Zhang 0001, Houdong Hu, Li Dong 0004, Furu Wei, Yejin Choi 0001, Jianfeng Gao 0001 |
ECCV (30) | 6 |
| 2020 | Boosting Weakly Supervised Object Detection with Progressive Knowledge Transfer
Yuanyi Zhong, Jian Peng 0001, Lei Zhang 0001 |
ECCV (26) | 4 |
| 2020 | Anchor Box Optimization for Object DetectionabstractIn this paper, we propose a general approach to optimize anchor boxes for object detection. Nowadays, anchor boxes are widely adopted in state-of-the-art detection frameworks. However, these frameworks usually pre-define anchor box shapes in heuristic ways and fix the sizes during training. To improve the accuracy and reduce the effort of designing anchor boxes, we propose to dynamically learn the anchor shapes, which allows the anchors to automatically adapt to the data distribution and the network learning capability. The learning approach can be easily implemented with stochastic gradient descent and can be plugged into any anchor box-based detection framework. The extra training cost is almost negligible and it has no impact on the inference time or memory cost. Exhaustive experiments demonstrate that the proposed anchor optimization method consistently achieves significant improvement (≥ 1% mAP absolute gain) over the baseline methods on several benchmark datasets including Pascal VOC 07+12, MS COCO and Brainwash. Meanwhile, the robustness is also verified towards different anchor initialization methods and the number of anchor shapes, which greatly simplifies the problem of anchor box design. Yuanyi Zhong, Jian Peng 0001, Lei Zhang 0001 |
WACV | 4 |
| 2020 | Real-Time Burst Photo Selection Using a Light-Head Adversarial NetworkabstractWe present an automatic moment capture system that runs in real-time on mobile cameras. The system is designed to run in the viewfinder mode and capture a burst sequence of frames before and after the shutter is pressed. For each frame, the system predicts in real-time a goodness score, based on which the best moment in the burst can be selected immediately after the shutter is released. We develop a highly efficient deep neural network ranking model, which implicitly learns a latent relative attribute space to capture subtle visual differences within a sequence of burst images. The overall goodness is computed as a linear aggregation of the goodnesses of all the latent attributes. To obtain a compact model which can run on mobile devices in real-time, we have explored and evaluated a wide range of network design choices, taking into account the constraints of model size, computational cost, and accuracy. Extensive studies show that the best frame predicted by our model hit users' top-1 (out of 11 on average) choice for 64.1% cases and top-3 choices for 86.2% cases. Moreover, the model (only 0.47M Bytes) can run in real time on mobile devices, e.g. 13ms on iPhone 7. Baoyuan Wang, Noranart Vesdapunt, Utkarsh Sinha, Lei Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | High Frequency Residual Learning for Multi-Scale Image Classification
Bowen Cheng, Rong Xiao 0003, Thomas S. Huang, Lei Zhang 0001 |
BMVC | 5 |
| 2019 | Object-Driven Text-To-Image Synthesis via Adversarial TrainingabstractIn this paper, we propose Object-driven Attentive Generative Adversarial Newtorks (Obj-GANs) that allow attention-driven, multi-stage refinement for synthesizing complex images from text descriptions. With a novel object-driven attentive generative network, the Obj-GAN can synthesize salient objects by paying attention to their most relevant words in the text descriptions and their pre-generated class label. In addition, a novel object-wise discriminator based on the Fast R-CNN model is proposed to provide rich object-wise discrimination signals on whether the synthesized object matches the text description and the pre-generated class label. The proposed Obj-GAN significantly outperforms the previous state of the art in various metrics on the large-scale MS-COCO benchmark, increasing the inception score by 27% and decreasing the FID score by 11%. A thorough comparison between the classic grid attention and the new object-driven attention is provided through analyzing their mechanisms and visualizing their attention layers, showing insights of how the proposed model generates complex scenes in high quality. Wenbo Li 0001, Pengchuan Zhang, Lei Zhang 0001, Qiuyuan Huang, Xiaodong He 0001, Siwei Lyu, Jianfeng Gao 0001 |
CVPR | 3 |
| 2019 | Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language NavigationabstractVision-language navigation (VLN) is the task of navigating an embodied agent to carry out natural language instructions inside real 3D environments. In this paper, we study how to address three critical challenges for this task: the cross-modal grounding, the ill-posed feedback, and the generalization problems. First, we propose a novel Reinforced Cross-Modal Matching (RCM) approach that enforces cross-modal grounding both locally and globally via reinforcement learning (RL). Particularly, a matching critic is used to provide an intrinsic reward to encourage global matching between instructions and trajectories, and a reasoning navigator is employed to perform cross-modal grounding in the local visual scene. Evaluation on a VLN benchmark dataset shows that our RCM model significantly outperforms previous methods by 10% on SPL and achieves the new state-of-the-art performance. To improve the generalizability of the learned policy, we further introduce a Self-Supervised Imitation Learning (SIL) method to explore unseen environments by imitating its own past, good decisions. We demonstrate that SIL can approximate a better and more efficient policy, which tremendously minimizes the success rate performance gap between seen and unseen environments (from 30.7% to 11.7%). Xin Wang 0061, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao 0001, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, Lei Zhang 0001 |
CVPR | 8 |
| 2019 | REO-Relevance, Extraness, Omission: A Fine-grained Evaluation for Image CaptioningabstractMing Jiang, Junjie Hu, Qiuyuan Huang, Lei Zhang, Jana Diesner, Jianfeng Gao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ming Jiang 0018, Junjie Hu 0001, Qiuyuan Huang, Lei Zhang 0001, Jana Diesner, Jianfeng Gao 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | TIGEr: Text-to-Image Grounding for Image Caption EvaluationabstractMing Jiang, Qiuyuan Huang, Lei Zhang, Xin Wang, Pengchuan Zhang, Zhe Gan, Jana Diesner, Jianfeng Gao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Ming Jiang 0018, Qiuyuan Huang, Lei Zhang 0001, Xin Wang 0061, Pengchuan Zhang, Zhe Gan, Jana Diesner, Jianfeng Gao 0001 |
EMNLP/IJCNLP (1) | 3 |
| 2019 | WSOD2: Learning Bottom-Up and Top-Down Objectness Distillation for Weakly-Supervised Object DetectionabstractWe study on weakly-supervised object detection (WSOD) which plays a vital role in relieving human involvement from object-level annotations. Predominant works integrate region proposal mechanisms with convolutional neural networks (CNN). Although CNN is proficient in extracting discriminative local features, grand challenges still exist to measure the likelihood of a bounding box containing a complete object (i.e., “objectness”). In this paper, we propose a novel WSOD framework with Objectness Distillation (i.e., WSOD2) by designing a tailored training mechanism for weakly-supervised object detection. Multiple regression targets are specifically determined by jointly considering bottom-up (BU) and top-down (TD) objectness from low-level measurement and CNN confidences with an adaptive linear combination. As bounding box regression can facilitate a region proposal learning to approach its regression target with high objectness during training, deep objectness representation learned from bottom-up evidences can be gradually distilled into CNN by optimization. We explore different adaptive training curves for BU/TD objectness, and show that the proposed WSOD2 can achieve state-of-the-art results. Zhaoyang Zeng, Bei Liu 0001, Jianlong Fu, Hongyang Chao, Lei Zhang 0001 |
ICCV | 5 |
| 2019 | Improving 3D Human Pose Estimation Via 3D Part Affinity Fieldsabstract3D human pose estimation from monocular images has become a heated area in computer vision recently. For years, most deep neural network based practices have adopted either an end-to-end approach, or a two-stage approach. An end-to-end network typically estimates 3D human poses directly from 2D input images, but it suffers from the shortage of 3D human pose data. It is also obscure to know if the inaccuracy stems from limited visual under-standing or 2D-to-3D mapping. Whereas a two-stage directly lifts those 2D keypoint outputs to the 3D space, after utilizing an existing network for 2D keypoint detections. However, they tend to ignore some useful contextual hints from the 2D raw image pixels. In this paper, we introduce a two-stage architecture that can eliminate the main disadvantages of both these approaches. During the first stage we use an existing state-of-the-art detector to estimate 2D poses. To add more con-textual information to help lifting 2D poses to 3D poses, we propose 3D Part Affinity Fields (3D-PAFs). We use 3D-PAFs to infer 3D limb vectors, and combine them with 2D poses to regress the 3D coordinates. We trained and tested our proposed framework on Human3.6M, the most popular 3D human pose benchmark dataset. Our approach achieves the state-of-the-art performance, which proves that with right selections of contextual information, a simple regression model can be very powerful in estimating 3D poses. Ding Liu 0001, Xinchao Wang, Yuxiao Hu 0001, Lei Zhang 0001, Thomas S. Huang |
WACV | 5 |
| 2018 | Bottom-Up and Top-Down Attention for Image Captioning and Visual Question AnsweringabstractTop-down visual attention mechanisms have been used extensively in image captioning and visual question answering (VQA) to enable deeper image understanding through fine-grained analysis and even multiple steps of reasoning. In this work, we propose a combined bottom-up and top-down attention mechanism that enables attention to be calculated at the level of objects and other salient image regions. This is the natural basis for attention to be considered. Within our approach, the bottom-up mechanism (based on Faster R-CNN) proposes image regions, each with an associated feature vector, while the top-down mechanism determines feature weightings. Applying this approach to image captioning, our results on the MSCOCO test server establish a new state-of-the-art for the task, achieving CIDEr / SPICE / BLEU-4 scores of 117.9, 21.5 and 36.9, respectively. Demonstrating the broad applicability of the method, applying the same approach to VQA we obtain first place in the 2017 VQA Challenge. Peter Anderson 0001, Xiaodong He 0001, Chris Buehler, Damien Teney, Mark Johnson 0001, Stephen Gould, Lei Zhang 0001 |
CVPR | 7 |
| 2018 | CleanNet: Transfer Learning for Scalable Image Classifier Training With Label NoiseabstractIn this paper, we study the problem of learning image classification models with label noise. Existing approaches depending on human supervision are generally not scalable as manually identifying correct or incorrect labels is time-consuming, whereas approaches not relying on human supervision are scalable but less effective. To reduce the amount of human supervision for label noise cleaning, we introduce CleanNet, a joint neural embedding network, which only requires a fraction of the classes being manually verified to provide the knowledge of label noise that can be transferred to other classes. We further integrate CleanNet and conventional convolutional neural network classifier into one framework for image classification learning. We demonstrate the effectiveness of the proposed algorithm on both of the label noise detection task and the image classification on noisy data task on several large-scale datasets. Experimental results show that CleanNet can reduce label noise detection error rate on held-out classes where no human supervision available by 41.5% compared to current weakly supervised methods. It also achieves 47% of the performance gain of verifying all images with only 3.2% images verified on an image classification task. Source code and dataset will be available at kuanghuei.github.io/CleanNetProject. Kuang-Huei Lee, Xiaodong He 0001, Lei Zhang 0001, Linjun Yang |
CVPR | 3 |
| 2018 | AutoLoc: Weakly-Supervised Temporal Action Localization in Untrimmed Videos
Zheng Shou 0001, Lei Zhang 0001, Kazuyuki Miyazawa, Shih-Fu Chang |
ECCV (16) | 3 |
| 2018 | PatternNet: Visual Pattern Mining with Deep Neural NetworkabstractVisual patterns represent the discernible regularity in the visual world. They capture the essential nature of visual objects or scenes. Understanding and modeling visual patterns is a fundamental problem in visual recognition that has wide ranging applications. In this paper, we study the problem of visual pattern mining and propose a novel deep neural network architecture called PatternNet for discovering these patterns that are both discriminative and representative. The proposed PatternNet leverages the filters in the last convolution layer of a convolutional neural network to find locally consistent visual patches, and by combining these filters we can effectively discover unique visual patterns. In addition, PatternNet can discover visual patterns efficiently without performing expensive image patch sampling, and this advantage provides an order of magnitude speedup compared to most other approaches. We evaluate the proposed PatternNet subjectively by showing randomly selected visual patterns which are discovered by our method and quantitatively by performing image classification with the identified visual patterns and comparing our performance with the current state-of-the-art. We also directly evaluate the quality of the discovered visual patterns by leveraging the identified patterns as proposed objects in an image and compare with other relevant methods. Our proposed network and procedure, PatterNet, is able to outperform competing methods for the tasks described. Hongzhi Li 0001, Joseph G. Ellis, Lei Zhang 0001, Shih-Fu Chang |
ICMR | 3 |
| 2018 | Turbo Learning for CaptionBot and DrawingBotabstractWe study in this paper the problems of both image captioning and text-to-image generation, and present a novel turbo learning approach to jointly training an image-to-text generator (a.k.a. CaptionBot) and a text-to-image generator (a.k.a. DrawingBot). The key idea behind the joint training is that image-to-text generation and text-to-image generation as dual problems can form a closed loop to provide informative feedback to each other. Based on such feedback, we introduce a new loss metric by comparing the original input with the output produced by the closed loop. In addition to the old loss metrics used in CaptionBot and DrawingBot, this extra loss metric makes the jointly trained CaptionBot and DrawingBot better than the separately trained CaptionBot and DrawingBot. Furthermore, the turbo-learning approach enables semi-supervised learning since the closed loop can provide peudo-labels for unlabeled samples. Experimental results on the COCO dataset demonstrate that the proposed turbo learning can significantly improve the performance of both CaptionBot and DrawingBot by a large margin. Qiuyuan Huang, Pengchuan Zhang, Dapeng Oliver Wu, Lei Zhang 0001 |
NeurIPS | 4 |
| 2018 | CLOTHO: A Large-Scale Internet of Things-Based Crowd Evacuation Planning System for Disaster ManagementabstractIn recent years, different kinds of natural hazards or man-made disasters happened that were diversified and difficult to control with heavy casualties. In this paper, we focus on the rapid and systematic evacuation of large-scale densities of people after disasters to reduce loss in an effective manner. The optimal evacuation planning is a key challenge and becomes a hotspot of research and development. We design our system based on an Internet of Things (IoT) scenario that utilizes a mobile cloud computing platform in order to develop the crowd lives oriented track and help optimization system (CLOTHO). CLOTHO is an evacuation planning system for large-scale densities of people in disasters. It includes the mobile terminal (IoT side) for data collection and the cloud backend system for storage and analytics. We build our solution upon a typical IoT/fog disaster management scenario and we propose an IoT application based on an evacuation planning algorithm that uses the artificial potential field (APF), which is the core of CLOTHO. APF is conceptualized as an IoT service, and can determine the direction of evacuation automatically according to the gradient direction of the potential field, suitable for rapid evacuation of large population. Based on APF, we propose an evacuation planning algorithm names as APF with relationship attraction (APF-RA). APF-RA guides the evacuees with relationship to move to the same shelter as much as possible, to calm evacuees and realize a more humanitarian evacuation. The experimental results show that CLOTHO (using APF and APF-RA) can effectively improve convergence rate, shorten the evacuation route length and evacuation time, and make the remaining capacity of the surrounding shelters well balanced. Xiaolong Xu 0002, Lei Zhang 0001, Stelios Sotiriadis, Eleana Asimakopoulou, Maozhen Li 0001, Nik Bessis |
IEEE Internet Things J. | 2 |
| 2018 | Visual instance mining from the graph perspective
Wei Li 0152, Jianmin Li 0001, Changhu Wang, Lei Zhang 0001, Bo Zhang 0010 |
Multim. Syst. | 4 |
| 2016 | MS-Celeb-1M: A Dataset and Benchmark for Large-Scale Face Recognition
Yandong Guo, Lei Zhang 0001, Yuxiao Hu 0001, Xiaodong He 0001, Jianfeng Gao 0001 |
ECCV (3) | 2 |
| 2015 | Scalable Visual Instance Mining with Instance GraphabstractIn this paper we address the problem of visual instance mining, which is to automatically discover frequently appearing visual instances from a large collection of images. We propose a scalable mining method by leveraging the graph structure with images as vertices. Different from most existing work that focused on either instance-level similarities or image-level context properties, our graph captures both information. The instance-level information is integrated during the construction of a weighted and undirected instance graph based on the similarity between augmented local features, while the image-level context is explored with a greedy breadth-first search algorithm to discover clusters of visual instances from the graph. This method is capable of mining challenging small visual instances with diverse variations. We evaluated our method on two fully annotated datasets and outperformed the state of the arts on both datasets with higher recalls. We also applied our method on a one-million Flickr dataset and proved its scalability. Wei Li 0152, Changhu Wang, Lei Zhang 0001, Yong Rui, Bo Zhang 0010 |
BMVC | 3 |
| 2015 | Sketch-based Image Retrieval via Shape WordsabstractThe explosive growth of touch screens has provided a good platform for sketch-based image retrieval. However, most previous works focused on low level descriptors of shapes and sketches. In this paper, we try to step forward and propose to leverage shape words descriptor for sketch-based image retrieval. First, the shape words are defined and an efficient algorithm is designed for shape words extraction. Then we generalize the classic Chamfer Matching algorithm to address the shape words matching problem. Finally, a novel inverted index structure is proposed to make shape words representation scalable to large scale image databases. Experimental results show that our method achieves competitive accuracy but requires much less memory, e.g., less than 3% of memory storage of MindFinder. Due to its competitive accuracy and low memory cost, our method can scale up to much larger database. Changcheng Xiao, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ICMR | 4 |
| 2015 | IdeaPanel: A Large Scale Interactive Sketch-based Image Search SystemabstractIn this work, we introduce the IdeaPanel system, an interactive sketch-based image search engine with millions of images. IdeaPanel enables users to sketch the target image in their minds and also supports tagging to describe their intentions. After a search is triggered, similar images will be returned in real time, based on which users can interactively refine their query sketches until ideal images are returned. Different from existing work, most of which requires a huge amount of memory for indexing and matching, IdeaPanel can achieve very competitive performance but requires much less memory storage. IdeaPanel needs only about 240MB memory to index 1.3M images (less than 3% of previous MindFinder system). Due to its high accuracy and low memory cost, IdealPanel can scale up to much larger database and thus has larger potential to return the most desired images for users. Changcheng Xiao, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ICMR | 4 |
| 2015 | Near Duplicate Image Discovery on One Billion ImagesabstractNear-duplicate image discovery is the task of detecting all clusters of images which duplicate at a significant region. Previous work generally take divide and conquer approaches composed of two steps: generating cluster seeds using min-hashing, and growing the seeds by searching the entire image space with the seeds as queries. Since the computational complexity of the seed growing step is generally O (NL) where N and L are the number of images and seeds respectively, existing work can hardly be scaled up to a billion-scale dataset because L is typically millions. In this paper, we study a feasible solution of near-duplicate image discovery on one billion images, which is easily implemented on MapReduce framework. The major contribution of this work is to introduce the seed growing step designed to efficiently reduce the number of false positives among cluster seeds with O (cNL) time complexity, where c is small enough for a billion-scale dataset. The basis component of the seed growing step is a bottom-k min-hash, which generates different signatures in a sketch to remove all candidate images that share only one common visual word with a cluster seed. Our evaluations suggest that the proposed method can discover near-duplicate clusters with high precision and recall, and represent some interesting properties of our 1 billion dataset. Saehoon Kim, Xin-Jing Wang, Lei Zhang 0001, Seungjin Choi 0001 |
WACV | 3 |
| 2015 | Regularized discriminant embedding for visual descriptor learning
Kye-Hyeon Kim, Rui Cai 0002, Lei Zhang 0001, Seungjin Choi 0001 |
Neurocomputing | 3 |
| 2015 | Partial-Duplicate Clustering and Visual Pattern Discovery on Web Scale Image DatabaseabstractIn this paper, we study the problem of discovering visual patterns and partial-duplicate images, which is fundamental to visual concept representation and image parsing, but very challenging when the database is extremely large, such as billions of images indexed by a commercial search engine. Although extensive research with sophisticated algorithms has been conducted for either partial-duplicate clustering or visual pattern discovery, most of them can not be easily extended to this scale, since both are clustering problems in nature and require pairwise comparisons. To tackle this computational challenge, we introduce a novel and highly parallelizable framework to discover partial-duplicate images and visual patterns in a unified way in distributed computing systems. We emphasize the nested property of local features, and propose the generalized nested feature (GNF) as a mid-level representation for regions and local patterns. Initial coarse clusters are then discovered by GNFs, upon which$n$-gram GNF is defined to represent co-occurrent visual patterns. After that, efficient merging and refining algorithms are used to get the partial-duplicate clusters, and logical combinations of probabilistic GNF models are leveraged to represent the visual patterns of partially duplicate images. Extensive experiments show the parallelizable property and effectiveness of the algorithms on both partial-duplicate clustering and visual pattern discovery. With 2000 machines, it costs about eight and 400 minutes to process one million and 40 million images respectively, which is quite efficient compared to previous methods. Wei Li 0152, Changhu Wang, Lei Zhang 0001, Yong Rui, Bo Zhang 0010 |
IEEE Trans. Multim. | 3 |
| 2014 | Mining text snippets for images on the webabstractImages are often used to convey many different concepts or illustrate many different stories. We propose an algorithm to mine multiple diverse, relevant, and interesting text snippets for images on the web. Our algorithm scales to all images on the web. For each image, all webpages that contain it are considered. The top-K text snippet selection problem is posed as combinatorial subset selection with the goal of choosing an optimal set of snippets that maximizes a combination of relevancy, interestingness, and diversity. The relevancy and interestingness are scored by machine learned models. Our algorithm is run at scale on the entire image index of a major search engine resulting in the construction of a database of images with their corresponding text snippets. We validate the quality of the database through a large-scale comparative study. We showcase the utility of the database through two web-scale applications: (a) augmentation of images on the web as webpages are browsed and (b)~an image browsing experience (similar in spirit to web browsing) that is enabled by interconnecting semantically related images (which may not be visually related) through shared concepts in their corresponding text snippets. Anitha Kannan, Simon Baker, Krishnan Ramnath, Juliet Fiss, Dahua Lin, Lucy Vanderwende, Rizwan Ansary, Ashish Kapoor, Qifa Ke, Matthew Uyttendaele, Xin-Jing Wang, Lei Zhang 0001 |
KDD | 12 |
| 2014 | Viewpoint-Aware Representation for Sketch-Based 3D Model RetrievalabstractWe study the problem of sketch-based 3D model retrieval, and propose a solution powered by a new query-to-model distance metric and a powerful feature descriptor based on the bag-of-features framework. The main idea of the proposed query-to-model distance metric is to represent a query sketch using a compact set of sample views (called basic views) of each model, and to rank the models in ascending order of the representation errors. To better differentiate between relevant and irrelevant models, the representation is constrained to be essentially a combination of basic views with similar viewpoints. In another aspect, we propose a mid-level descriptor (called BOF-JESC) which robustly characterizes the edge information within junction-centered patches, to extract the salient shape features from sketches or model views. The combination of the query-to-model distance metric and the BOF-JESC descriptor achieves effective results on two latest benchmark datasets. Changqing Zou, Changhu Wang, Yafei Wen, Lei Zhang 0001, Jianzhuang Liu |
IEEE Signal Process. Lett. | 4 |
| 2014 | Introduction to the Special Issue Best Papers of ACM Multimedia 2013abstractNo abstract available. Zhengjun Zha, Lei Zhang 0001, Max Mühlhäuser, Alan F. Smeaton |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Efficient 2D-to-3D Correspondence Filtering for Scalable 3D Object Recognitionabstract3D model-based object recognition has been a noticeable research trend in recent years. Common methods find 2D-to-3D correspondences and make recognition decisions by pose estimation, whose efficiency usually suffers from noisy correspondences caused by the increasing number of target objects. To overcome this scalability bottleneck, we propose an efficient 2D-to-3D correspondence filtering approach, which combines a light-weight neighborhood-based step with a finer-grained pairwise step to remove spurious correspondences based on 2D/3D geometric cues. On a dataset of 300 3D objects, our solution achieves ~10 times speed improvement over the baseline, with a comparable recognition accuracy. A parallel implementation on a quad-core CPU can run at ~3fps for 1280×720 images. Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001, Yong Rui |
CVPR | 4 |
| 2013 | Exploring Implicit Image Statistics for Visual Representativeness ModelingabstractIn this paper, we propose a computational model of visual representative ness by integrating cognitive theories of representative ness heuristics with computer vision and machine learning techniques. Unlike previous models that build their representative ness measure based on the visible data, our model takes the initial inputs as explicit positive reference and extend the measure by exploring the implic it negatives. Given a group of images that contains obvious visual concepts, we create a customized image ontology consisting of both positive and negative instances by mining the most related and confusable neighbors of the positive concept in ontological semantic knowledge bases. The representative ness of a new item is then determined by its likelihoods for both the positive and negative references. To ensure the effectiveness of probability inference as well as the cognitive plausibility, we discover the potential prototypes and treat them as an intermediate representation of semantic concepts. In the experiment, we evaluate the performance of representative ness models based on both human judgements and user-click logs of commercial image search engine. Experimental results on both Image Net and image sets of general concepts demonstrate the superior performance of our model against the state-of-the-arts. Xiaoshuai Sun, Xin-Jing Wang, Hongxun Yao, Lei Zhang 0001 |
CVPR | 4 |
| 2013 | The shortest warping path based multiple images alignmentabstractIn this paper, we propose a method to align multiple images of the same category. Images with large variations are aligned via a smooth transition formed by some intermediate images. These intermediate images are found by shortest warping path algorithm on a directed complete graph. Moreover, the common regions in the images are discovered to further improve alignment performance. The experimental results show that our method is effective to map and align images of the same category but with large variations of appearance, shape and view. Wei Yu 0004, Hongxun Yao, Kuiyuan Yang, Lei Zhang 0001 |
ICIP | 4 |
| 2013 | Mobile multimedia travelogue generation by exploring geo-locations and image tagsabstractTraveling experience sharing has become pervasive. In this paper, we present a system which automatically generates a multimedia travelogue for mobile users. Multimedia travelogue shows user footprint with photos on the corresponding location on the map, which could offer the space distribution inner or between the locations. Meanwhile, location overviews with representative tags and images show comprehensive knowledge of the landmarks. The key challenges are: 1) when mapping footprint, some photos are without geo-tags; 2)user-contributed photos may be not clear and graceful; 3) to detect which landmark is actually on the photo to offer location overview. We solve these challenges by combining both geo-tags and tags to estimate the location. First we use the relationship of timestamp and geo-tag in the same trip group to estimate the footprints. Second we localize the landmark on the photo by tag match method which matching user tags with representative tags of the landmark. Experimental results on a Flickr image collection of nearly 2 million images of 6,581 users demonstrate the effectiveness of our approach. Shuhui Jiang, Xueming Qian, Ke Lan, Lei Zhang 0001, Tao Mei 0001 |
ISCAS | 4 |
| 2013 | MagicBrush: image search by color sketchabstractIn this paper, we showcase the MagicBrush system, a novel painting-based image search engine. This system enables users to draw a color sketch as a query to find images. Different from existing works on sketch-based image retrieval, most of which focus on matching the shape structure without carefully considering other important visual modalities, MagicBrush takes into account the indispensable value of "color" related to "shape", and explores to make use of both the shape and color expectations that users usually have when they're imaging or searching for an image. To achieve this, we 1) develop a user-friendly interface to allow users to easily "paint out" their colorful visual expectations; 2) design a compact feature "color-edge word" to encode both shape and color information in a organic way; and 3) develop a novel matching and index structure to support a real-time response in 6.4 million images. By taking into account both shape and color information, the MagicBrush system helps users to vividly present what they are imagining, and retrieve images in a more natural way. Xinghai Sun, Changhu Wang, Avneesh Sud, Chao Xu 0006, Lei Zhang 0001 |
ACM Multimedia | 5 |
| 2013 | Indexing billions of images for sketch-based retrievalabstractBecause of the popularity of touch-screen devices, it has become a highly desirable feature to retrieve images from a huge repository by matching with a hand-drawn sketch. Although searching images via keywords or an example image has been successfully launched in some commercial search engines of billions of images, it is still very challenging for both academia and industry to develop a sketch-based image retrieval system on a billion-level database. In this work, we systematically study this problem and try to build a system to support query-by-sketch for two billion images. The raw edge pixel and Chamfer matching are selected as the basic representation and matching in this system, owning to the superior performance compared with other methods in extensive experiments. To get a more compact feature and a faster matching, a vector-like Chamfer feature pair is introduced, based on which the complex matching is reformulated as the crossover dot-product of feature pairs. Based on this new formulation, a compact shape code is developed to represent each image/sketch by projecting the Chamfer features to a linear subspace followed by a non-linear source coding. Finally, the multi-probe Kmedoids-LSH is leveraged to index database images, and the compact shape codes are further used for fast reranking. Extensive experiments show the effectiveness of the proposed features and algorithms in building such a sketch-based image search system. Xinghai Sun, Changhu Wang, Chao Xu 0006, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2013 | Special issue on image feature detection and description
Yanwei Pang, Xianbin Cao 0001, Lei Zhang 0001, Amir Hussein |
Neurocomputing | 3 |
| 2013 | Image search - from thousands to billions in 20 yearsabstractThis article presents a comprehensive review and analysis on image search in the past 20 years, emphasizing the challenges and opportunities brought by the astonishing increase of dataset scales from thousands to billions in the same time period, which was witnessed first-hand by the authors as active participants in this research area. Starting with a retrospective review of three stages of image search in the history, the article highlights major breakthroughs around the year 2000 in image search features, indexing methods, and commercial systems, which marked the transition from stage two to stage three. Subsequent sections describe the image search research from four important aspects: system framework, feature extraction and image representation, indexing, and big data's potential. Based on the review, the concluding section discusses open research challenges and suggests future research directions in effective visual representation, image knowledge base construction, implicit user feedback and crowdsourcing, mobile image search, and creative multimedia interfaces. Lei Zhang 0001, Yong Rui |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2012 | Hierarchical Object Representations for Visual Recognition via Weakly Supervised Learning
Tianzhu Zhang 0001, Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Hanqing Lu |
ACCV (1) | 4 |
| 2012 | 3D visual phrases for landmark recognitionabstractIn this paper, we study the problem of landmark recognition and propose to leverage 3D visual phrases to improve the performance. A 3D visual phrase is a triangular facet on the surface of a reconstructed 3D landmark model. In contrast to existing 2D visual phrases which are mainly based on co-occurrence statistics in 2D image planes, such 3D visual phrases explicitly characterize the spatial structure of a 3D object (landmark), and are highly robust to projective transformations due to viewpoint changes. We present an effective solution to discover, describe, and detect 3D visual phrases. The experiments on 10 landmarks have achieved promising results, which demonstrate that our approach provides a good balance between precision and recall of landmark recognition while reducing the dependence on post-verification to reject false positives. Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001, Yanwei Pang, Feng Wu 0001 |
CVPR | 4 |
| 2012 | The scale of edgesabstractAlthough the scale of isotropic visual elements such as blobs and interest points, e.g. SIFT[12], has been well studied and adopted in various applications, how to determine the scale of anisotropic elements such as edges is still an open problem. In this paper, we study the scale of edges, and try to answer two questions: 1) what is the scale of edges, and 2) how to calculate it. From the points of human cognition and physical interpretation, we illustrate the existence of the scale of edges and provide a quantitative definition. Then, an automatic edge scale selection approach is proposed. Finally, a cognitive experiment is conducted to validate the rationality of the detected scales. Moreover, the importance of identifying the scale of edges is also shown in applications such as boundary detection and hierarchical edge parsing. Xianming Liu 0005, Changhu Wang, Hongxun Yao, Lei Zhang 0001 |
CVPR | 4 |
| 2012 | QsRank: Query-sensitive hash code ranking for efficient ∊-neighbor searchabstractAlthough binary hash code-based image indexing methods have been recently developed for large-scale applications, the problem of ranking such hash codes has been barely studied. In this paper, we propose a query sensitive ranking algorithm (QsRank) to rank PCA-based hash codes for the ∊-neighbor search problem. The QsRank algorithm takes the target neighborhood radius ∊ and the raw feature of a given query as input, and models the statistical properties of the target ∊-neighbors in the space of hash codes. Unlike the Hamming distance, the proposed algorithm does not compress query points to hash codes. Therefore, it suffers less information loss and is more effective than Hamming distance-based approaches. Based on the QsRank method, we developed an efficient indexing structure and retrieval algorithm for large-scale ∊-neighbor search. Evaluations on two datasets of 10 million web images and 10 million SIFT descriptors demonstrate that the proposed retrieval system achieves higher accuracy with less memory cost and faster speed. Lei Zhang 0001, Harry Shum |
CVPR | 2 |
| 2012 | Pairwise Rotation Invariant Co-occurrence Local Binary Pattern
Xianbiao Qi, Rong Xiao 0003, Jun Guo 0002, Lei Zhang 0001 |
ECCV (6) | 4 |
| 2012 | Free Hand-Drawn Sketch Segmentation
Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ECCV (1) | 4 |
| 2012 | Efficient Tag Mining via Mixture Modeling for Real-Time Search-Based Image AnnotationabstractAlthough it has been extensively studied for many years, automatic image annotation is still a challenging problem. Recently, data-driven approaches have demonstrated their great success to image auto-annotation. Such approaches leverage abundant partially annotated web images to annotate an uncaptioned image. Specifically, they first retrieve a group of visually closely similar images given an uncaptioned image as a query, then figure out meaningful phrases from the surrounding texts of the image search results. Since the surrounding texts are generally noisy, how to effectively mine meaningful phrases is crucial for the success of such approaches. We propose a mixture modeling approach which assumes that a tag is generated from a convex combination of topics. Different from a typical topic modeling approach like LDA, topics in our approach are explicitly learnt from a definitive catalog of the Web, i.e. the Open Directory Project (ODP). Compared with previous works, it has two advantages: Firstly, it uses an open vocabulary rather than a limited one defined by a training set. Secondly, it is efficient for real-time annotation. Experimental results conducted on two billion web images show the efficiency and effectiveness of the proposed approach. Lican Dai, Xin-Jing Wang, Lei Zhang 0001, Nenghai Yu |
ICME | 3 |
| 2012 | A rapid flower/leaf recognition systemabstractIn this work, we introduce a rapid and accurate flower/leaf recognition system. The system could process one query in less than 0.35s with users' simple interaction. Meanwhile, high accuracy and recall is achieved. Furthermore, low computational resource and memory cost are required by the system. Now, the system is demonstrated on 172 categories of flowers, the largest flower dataset until now, and 220 categories of leaves. Xianbiao Qi, Rong Xiao 0003, Lei Zhang 0001, Chun-Guang Li, Jun Guo 0002 |
ACM Multimedia | 3 |
| 2012 | Query-adaptive shape topic mining for hand-drawn sketch recognitionabstractIn this work, we study the problem of hand-drawn sketch recognition. Due to large intra-class variations presented in hand-drawn sketches, most of existing work was limited to a particular domain or limited pre-defined classes. Different from existing work, we target at developing a general sketch recognition system, to recognize any semantically meaningful object that a child can recognize. To increase the recognition coverage, a web-scale clipart image collection is leveraged as the knowledge base of the recognition system. To alleviate the problems of intra-class shape variation and inter-class shape ambiguity in this unconstrained situation, a query-adaptive shape topic model is proposed to mine object topics and shape topics related to the sketch, in which, multiple layers of information such as sketch, object, shape, image, and semantic labels are modeled in a generative process. Besides sketch recognition, the proposed topic model can also be used for related applications such as sketch tagging, image tagging, and sketch-based image search. Extensive experiments on different applications show the effectiveness of the proposed topic model and the recognition system. Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2012 | Sketch2Tag: automatic hand-drawn sketch recognitionabstractIn this work, we introduce the Sketch2Tag system for hand-drawn sketch recognition. Due to large variations presented in hand-drawn sketches, most of existing work was limited to a particular domain or limited predefined classes. Different from existing work, Sketch2Tag is a general sketch recognition system, towards recognizing any semantically meaningful object that a child can recognize. This system enables a user to draw a sketch on the query panel, and then provides real-time recognition results. To increase the recognition coverage, a web-scale clipart image collection is leveraged as the knowledge base of the recognition system. Better understanding a user's drawing will be of great value to a variety of applications, such as, improving the sketch-based image search by combining the recognition results as textual queries. Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2012 | Towards indexing representative images on the webabstractEven after 20 years of research on real-world image retrieval, there is still a big gap between what search engines can provide and what users expect to see. To bridge this gap, we present an image knowledge base, ImageKB, a graph representation of structured entities, categories, and representative images, as a new basis for practical image indexing and search. ImageKB is automatically constructed via a both bottom-up and top-down, scalable approach that efficiently matches 2 billion web images onto an ontology with millions of nodes. Our approach consists of identifying duplicate image clusters from billions of images, obtaining a candidate set of entities and their images, discovering definitive texts to represent an image and identifying representative images for an entity. To date, ImageKB contains 235.3M representative images corresponding to 0.52M entities, much larger than the state-of-the-art alternative ImageNet that contains 14.2M images for 0.02M synsets. Compared to existing image databases, ImageKB reflects the distributions of both images on the web and users' interests, contains rich semantic descriptions for images and entities, and can be widely used for both text to image search and image to text understanding. Xin-Jing Wang, Zheng Xu 0002, Lei Zhang 0001, Ce Liu 0001, Yong Rui |
ACM Multimedia | 3 |
| 2012 | A probabilistic graphical model for topic and preference discovery on social media
Lu Liu 0005, Feida Zhu 0001, Lei Zhang 0001, Shiqiang Yang |
Neurocomputing | 3 |
| 2012 | Preface
Qiang Yang 0001, Jie Tang 0001, Lei Zhang 0001, Bin Cao 0001 |
J. Comput. Sci. Technol. | 3 |
| 2012 | Duplicate-Search-Based Image Annotation Using Web-Scale DataabstractEasy photo-taking and photo-sharing today make image an increasingly important type of media in people's everyday life, which arouses a growing demand for a practical image understanding technique. Traditional computer vision or machine learning methods which learn models based on a set of training data are still in the stage of tackling hundreds of object categories. Such a scale is far from practical usage. In recent years, the technique of search-based image annotation on a large-scale data set has demonstrated great success. Rather than directly mapping visual features to texts which is inevitably hindered by the semantic gap, it understands the content of an image by propagating labels of its similar images in a large-scale data set. Since similarity search is performed among homogenous data, the difficulty is greatly reduced. This paper summarizes the extensive work on web image annotation using the large-scale metadata and social information available on the Web, and introduces the Arista system, which is a nonparametric image annotation platform built upon two billion web images. We propose a highly efficient and scalable duplicate-search technique so that the Arista system can be deployed on a few servers. A few interesting applications such as building large-scale celebrity face database and text-to-image translation are also presented in this paper. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
Proc. IEEE | 2 |
| 2012 | Finding Celebrities in Billions of Web ImagesabstractIn this paper, we present a face annotation system to automatically collect and label celebrity faces from the web. With the proposed system, we have constructed a large-scale dataset called “Celebrities on the Web,” which contains 2.45 million distinct images of 421 436 celebrities and is orders of magnitude larger than previous datasets. Lei Zhang 0001, Xin-Jing Wang, Harry Shum |
IEEE Trans. Multim. | 2 |
| 2011 | User browsing behavior-driven web crawlingabstractTo optimize the performance of web crawlers, various page importance measures have been studied to select and order URLs in crawling. Most sophisticated measures (e.g. breadth-first and PageRank) are based on link structure. In this paper, we treat the problem from another perspective and propose to measure page importance through mining user interest and behaviors from web browse logs. Unlike most existing approaches which work on single URL, in this paper, both the log mining and the crawl ordering are performed at the granularity of URL pattern. The proposed URL pattern-based crawl orderings are capable to properly predict the importance of newly created (unseen) URLs. Promising experimental results proved the feasibility of our approach. Minghai Liu, Rui Cai 0002, Ming Zhang 0004, Lei Zhang 0001 |
CIKM | 4 |
| 2011 | Edgel index for large-scale sketch-based image searchabstractRetrieving images to match with a hand-drawn sketch query is a highly desired feature, especially with the popularity of devices with touch screens. Although query-by-sketch has been extensively studied since 1990s, it is still very challenging to build a real-time sketch-based image search engine on a large-scale database due to the lack of effective and efficient matching/indexing solutions. The explosive growth of web images and the phenomenal success of search techniques have encouraged us to revisit this problem and target at solving the problem of web-scale sketch-based image retrieval. In this work, a novel index structure and the corresponding raw contour-based matching algorithm are proposed to calculate the similarity between a sketch query and natural images, and make sketch-based image retrieval scalable to millions of images. The proposed solution simultaneously considers storage cost, retrieval accuracy, and efficiency, based on which we have developed a real-time sketch-based image search engine by indexing more than 2 million images. Extensive experiments on various retrieval tasks (basic shape search, specific image search, and similar image search) show better accuracy and efficiency than state-of-the-art methods. Yang Cao 0008, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001 |
CVPR | 4 |
| 2011 | Rank-SIFT: Learning to rank repeatable local interest pointsabstractScale-invariant feature transform (SIFT) has been well studied in recent years. Most related research efforts focused on designing and learning effective descriptors to characterize a local interest point. However, how to identify stable local interest points is still a very challenging problem. In this paper, we propose a set of differential features, and based on them we adopt a data-driven approach to learn a ranking function to sort local interest points according to their stabilities across images containing the same visual objects. Compared with the handcrafted rule-based method used by the standard SIFT algorithm, our algorithm substantially improves the stability of detected local interest point on a very challenging benchmark dataset, in which images were generated under very different imaging conditions. Experimental results on the Oxford and PASCAL databases further demonstrate the superior performance of the proposed algorithm on both object image retrieval and category recognition. Rong Xiao 0003, Zhiwei Li 0006, Rui Cai 0002, Bao-Liang Lu, Lei Zhang 0001 |
CVPR | 6 |
| 2011 | Grassmann Hashing for approximate nearest neighbor search in high dimensional spaceabstractLocality-Sensitive Hashing (LSH) approximates nearest neighbors in high dimensions by projecting original data into low-dimensional subspaces. The basic idea is to hash data samples to ensure that the probability of collision is much higher for samples that are close to each other than for those that are far apart. However, by applying k random hashing functions on original data, LSH fails to find the most discriminant hashing-subspaces, so the nearest neighbor approximation is inefficient. To alleviate this problem, we propose the Grassmann Hashing (GRASH) for approximate nearest neighbor search in high dimensions. GRASH first introduces a set of subspace candidates from Linear Discriminant Analysis (LDA). Then it applies Grassmann metric to select the optimal subspaces for hashing. Finally, it generates hashing codes based on non-uniform bucket size design motivated by Lloyd-Max quantization. The proposed GRASH model enjoys a number of merits: 1) GRASH introduces the Grassmann metric to measure the similarity between different hashing subspaces, so the hashing function can better capture the data diversity; 2) GRASH obtains the subspace candidates from LDA, so it incorporates the discriminant information into the hashing functions; 3) GRASH extends LSH's 1-d hashing subspaces to m-d, i.e. it is a multidimensional extension of hashing approximation; 4) motivated by Lloyd-Max quantization, GRASH applies non-uniform size bucket to generate hashing codes, so the distortion can be minimized. Experimental results on a number of datasets confirm the validity of our proposed model. Xinchao Wang, Zhu Li 0001, Lei Zhang 0001, Junsong Yuan 0001 |
ICME | 3 |
| 2011 | Contextual synonym dictionary for visual object retrievalabstractIn this paper, we study the problem of visual object retrieval by introducing a dictionary of contextual synonyms to narrow down the semantic gap in visual word quantization. The basic idea is to expand a visual word in the query image with its synonyms to boost the retrieval recall. Unlike the existing work such as soft-quantization, which only focuses on the Euclidean (l2) distance in descriptor space, we utilize the visual words which are more likely to describe visual objects with the same semantic meaning by identifying the words with similar contextual distributions (i.e. contextual synonyms). We describe the contextual distribution of a visual word using the statistics of both co-occurrence and spatial information averaged over all the image patches having this visual word, and propose an efficient system implementation to construct the contextual synonym dictionary for a large visual vocabulary. The whole construction process is unsupervised and the synonym dictionary can be naturally integrated into a standard bag-of-feature image retrieval system. Experimental results on several benchmark datasets are quite promising. The contextual synonym dictionary-based expansion consistently outperforms the l2 distance-based soft-quantization, and advances the state-of-the-art performance remarkably. Rui Cai 0002, Zhiwei Li 0006, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2011 | Sketch2Cartoon: composing cartoon images by sketchingabstractIn this paper, we introduce the Sketch2Cartoon system, which is an automatic cartoon making system by leveraging a novel sketch-based clipart image search engine. Different from existing work, most of which either limited users to the pre-prepared characters or only used keyword queries to search materials, Sketch2Cartoon enables users to sketch major curves of characters and props in their mind, and real-time search results from millions of clipart images could be selected to compose the cartoon images. The selected components are vectorized and thus could be further edited. By enabling sketch-based input, the cartoon image making process becomes more natural, and even a child who is too young to read or write can draw whatever he/she imagines and get interesting cartoon images. Changhu Wang, Bruce Yang, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2011 | Semantic point detectorabstractLocal features are the building blocks of many visual systems, and local point detector is usually the first component for local feature extraction. Existing local point detector are designed with target for matching and it may not perform well when applied in image content representation. Actually many existing studies demonstrate that the simple dense sampling strategy can achieve better performance than many local point detection methods in image classification tasks. In this paper, we propose a novel point detector named semantic point detector, which detects a set of semantically meaningful patches from each image and yields more compact and complete image representation. It is learned from an set of images with concepts from a large ontology. We conduct extensive experiments based on the proposed detector, and the experimental results demonstrate the effectiveness of our approach. Kuiyuan Yang, Lei Zhang 0001, Meng Wang 0001, HongJiang Zhang |
ACM Multimedia | 2 |
| 2011 | Multi-feature pLSA for combining visual features in image annotationabstractWe study in this paper the problem of combining low-level visual features for image region annotation. The problem is tackled with a novel method that combines texture and color features via a mixture model of their joint distribution. The structure of the presented model can be considered as an extension of the probabilistic latent semantic analysis (pLSA) in that it handles data from two different visual feature domains by attaching one more leaf node to the graphical structure of the original pLSA. Therefore, the proposed approach is referred to as multi-feature pLSA (MF-pLSA). The supervised paradigm is adopted to classify a new image region into one of a few pre-defined object categories using the MF-pLSA. To evaluate the performance, the VOC2009 and LabelMe databases were employed in our experiments, along with various experimental settings in terms of the number of visual words and mixture components. Evaluated based on the average recall and precision, the MF-pLSA is demonstrated superior to seven other approaches, including other schemes for visual feature combination. Rui Zhang 0010, Lei Zhang 0001, Xin-Jing Wang, Ling Guan |
ACM Multimedia | 2 |
| 2011 | From one tree to a forest: a unified solution for structured web data extractionabstractStructured data, in the form of entities and associated attributes, has been a rich web resource for search engines and knowledge databases. To efficiently extract structured data from enormous websites in various verticals (e.g., books, restaurants), much research effort has been attracted, but most existing approaches either require considerable human effort or rely on strong features that lack of flexibility. We consider an ambitious scenario -- can we build a system that (1) is general enough to handle any vertical without re-implementation and (2) requires only one labeled example site from each vertical for training to automatically deal with other sites in the same vertical? In this paper, we propose a unified solution to demonstrate the feasibility of this scenario. Specifically, we design a set of weak but general features to characterize vertical knowledge (including attribute-specific semantics and inter-attribute layout relationships). Such features can be adopted in various verticals without redesign; meanwhile, they are weak enough to avoid overfitting of the learnt knowledge to seed sites. Given a new unseen site, the learnt knowledge is first applied to identify page-level candidate attribute values, while inevitably involve false positives. To remove noise, site-level information of the new site is then exploited to boost up the true values. The site-level information is derived in an unsupervised manner, without harm to the applicability of the solution. Promising experimental performance on 80 websites in 8 distinct verticals demonstrated the feasibility and flexibility of the proposed solution. Rui Cai 0002, Yanwei Pang, Lei Zhang 0001 |
SIGIR | 4 |
| 2011 | Query by document via a decomposition-based two-level retrieval approachabstractRetrieving similar documents from a large-scale text corpus according to a given document is a fundamental technique for many applications. However, most of existing indexing techniques have difficulties to address this problem due to special properties of a document query, e.g. high dimensionality, sparse representation and semantic issue. Towards addressing this problem, we propose a two-level retrieval solution based on a document decomposition idea. A document is decomposed to a compact vector and a few document specific keywords by a dimension reduction approach. The compact vector embodies the major semantics of a document, and the document specific keywords complement the discriminative power lost in dimension reduction process. We adopt locality sensitive hashing (LSH) to index the compact vectors, which guarantees to quickly find a set of related documents according to the vector of a query document. Then we re-rank documents in this set by their document Linkai Weng, Zhiwei Li 0006, Rui Cai 0002, Yaoxue Zhang, Yue-Zhi Zhou, Laurence T. Yang, Lei Zhang 0001 |
SIGIR | 7 |
| 2011 | Summarizing tourist destinations by mining user-generated travelogues and photos
Yanwei Pang, Yuan Yuan 0001, Tanji Hu, Rui Cai 0002, Lei Zhang 0001 |
Comput. Vis. Image Underst. | 6 |
| 2010 | Spatial-bag-of-featuresabstractIn this paper, we study the problem of large scale image retrieval by developing a new class of bag-of-features to encode geometric information of objects within an image. Beyond existing orderless bag-of-features, local features of an image are first projected to different directions or points to generate a series of ordered bag-of-features, based on which different families of spatial bag-of-features are designed to capture the invariance of object translation, rotation, and scaling. Then the most representative features are selected based on a boosting-like method to generate a new bag-of-features-like vector representation of an image. The proposed retrieval framework works well in image retrieval task owing to the following three properties: 1) the encoding of geometric information of objects for capturing objects' spatial transformation, 2) the supervised feature selection and combination strategy for enhancing the discriminative power, and 3) the representation of bag-of-features for effective image matching and indexing for large scale image retrieval. Extensive experiments on 5000 Oxford building images and 1 million Panoramio images show the effectiveness and efficiency of the proposed features as well as the retrieval framework. Yang Cao 0008, Changhu Wang, Zhiwei Li 0006, Liqing Zhang 0001, Lei Zhang 0001 |
CVPR | 5 |
| 2010 | Probabilistic models for supervised dictionary learningabstractDictionary generation is a core technique of the bag-of-visual-words (BOV) models when applied to image categorization. Most of previous approaches generate dictionaries by unsupervised clustering techniques, e.g. k-means. However, the features obtained by such kind of dictionaries may not be optimal for image classification. In this paper, we propose a probabilistic model for supervised dictionary learning (SDLM) which seamlessly combines an unsuper-vised model (a Gaussian Mixture Model) and a supervised model (a logistic regression model) in a probabilistic framework. In the model, image category information directly affects the generation of a dictionary. A dictionary obtained by this approach is a trade-off between minimization of distortions of clusters and maximization of discriminative power of image-wise representations, i.e. histogram representations of images. We further extend the model to incorporate spatial information during the dictionary learning process in a spatial pyramid matching like manner. We extensively evaluated the two models on various benchmark dataset and obtained promising results. Xiao-Chen Lian, Zhiwei Li 0006, Changhu Wang, Bao-Liang Lu, Lei Zhang 0001 |
CVPR | 5 |
| 2010 | ARISTA - image search to annotation on billions of web photosabstractThough it has cost great research efforts for decades, object recognition is still a challenging problem. Traditional methods based on machine learning or computer vision are still in the stage of tackling hundreds of object categories. In recent years, non-parametric approaches have demonstrated great success, which understand the content of an image by propagating labels of its similar images in a large-scale dataset. However, due to the limited dataset size and imperfect image crawling strategy, previous work can only address a biased small subset of image concepts. Here we introduce the Arista project, which aims to build a practical image annotation engine targeting at popular concepts in the real world. In this project, we are particularly interested in understanding how many image concepts can be addressed by the data-driven annotation approach (coverage) and how good the performance is (precision). This paper reports the first stage of the work. Two billions web images were indexed, and based on simple yet effective near-duplicate detection, the system is capable of automatically generating accurate tags for popular web images having near-duplicates in the database. We found that about 8.1% web images have more than ten near duplicate and the number increases to 28.5% for top images in search results. Further, based on random samples in the latter case, we observed the precision of 57.9% at the point of the highest recall of 28% on ground truth tags. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
CVPR | 2 |
| 2010 | Interest seam imageabstractWe propose interest seam image, an efficient visual synopsis for video. To extract an interest seam image, a spatiotemporal energy map is constructed for the target video shot. Then an optimal seam which encompasses the highest energy is identified by an efficient dynamic programming algorithm. The optimal seam is used to extract a seam of pixels from each video frame to form one column of an image, based on which an interest seam image is finally composited. The interest seam image is efficient both in terms of computation and memory cost. Therefore it is able to power a wide variety of web-scale video content analysis applications, such as near duplicate video clip search, video genre recognition and classification, as well as video clustering, etc. The representation capacity of the proposed interest seam image is demonstrated in a large scale video retrieval task. Its advantages are clearly exhibited when compared with previous works, as reported in our experiments. Gang Hua 0001, Lei Zhang 0001, Harry Shum |
CVPR | 3 |
| 2010 | Max-Margin Dictionary Learning for Multiclass Image Categorization
Xiao-Chen Lian, Zhiwei Li 0006, Bao-Liang Lu, Lei Zhang 0001 |
ECCV (4) | 4 |
| 2010 | An efficient location extraction algorithm by leveraging web contextual informationabstractA typical location extraction approach consists of two steps, location name detection and location entity disambiguation. Promising results have been obtained in the last decade based on natural language processing technologies. However, there are still two challenges which requires further investigation: 1)How to leverage the prior and contextual evidence to improve the location extraction performance, and 2) How to utilize the interdependence information between the named entity recognition step and disambiguation step. In this paper, we propose an iterative detection-ranking framework to address these problems as well as a set of novel features to mine contextual information from web resources. Experimental results show that our solution outperforms the state-of-the-art approaches, including Metacarta GeoTagger and Yahoo Placemaker. Teng Qin, Rong Xiao 0003, Lei Fang 0004, Xing Xie 0001, Lei Zhang 0001 |
GIS | 5 |
| 2010 | Robust semantic sketch based specific image retrievalabstractSpecific images refer to images one has a certain episodic memory about, e.g. a picture one has ever seen before. Specific image retrieval is a frequent daily information need and the episodic memory is the key to find a specific image. In this paper, we propose a novel semantic sketch-based interface to incorporate the episodic memory for specific image retrieval. The interface allows a user to specify the semantic category and rough area/color of the objects in his memory. To bridge the semantic gap between the query sketch and database images, in the back end, a sampling method selects exemplars from a reference dataset which contains many object instances with user-provided tags and bounding boxes. After that, an exemplar matching algorithm ranks images to retrieve the target image to match the user's memory. In practice, we have observed that query sketches are usually error prone. That is, the position or the color of an object may not be accurate. Meanwhile, the annotations in the reference dataset are also noisy. Thus, the search algorithm has to handle two kinds of errors: 1) reference dataset label noise; 2) user sketch error such as position or scale. For the former, we propose a robust sampling method. For the latter, we derive an efficient spatial reranking algorithm to tolerate inaccurate user sketches. Detailed experimental results on the LabelMe dataset show that the proposed approach is robust to both kinds of errors. Cailiang Liu, Dong Wang 0022, Changhu Wang, Lei Zhang 0001, Bo Zhang 0010 |
ICME | 5 |
| 2010 | MindFinder: interactive sketch-based image search on millions of imagesabstractIn this paper, we showcase the MindFinder system, which is an interactive sketch-based image search engine. Different from existing work, most of which is limited to a small scale database or only enables single modality input, MindFinder is a sketch-based multimodal search engine for million-level database. It enables users to sketch major curves of the target image in their mind, and also supports tagging and coloring operations to better express their search intentions. Owning to a friendly interface, our system supports multiple actions, which help users to flexibly design their queries. After each operation, top returned images are updated in real time, based on which users could interactively refine their initial thoughts until ideal images are returned. The novelty of the MindFinder system includes the following two aspects: 1) A multimodal searching scheme is proposed to retrieve images which meet users' requirements not only in structure, but also in semantic meaning and color tone. 2) An indexing framework is designed to make MindFinder scalable in terms of database size, memory cost, and response time. By scaling up the database to more than two million images, MindFinder not only helps users to easily present whatever they are imagining, but also has the potential to retrieve the most desired images in their mind. Yang Cao 0008, Changhu Wang, Zhiwei Li 0006, Liqing Zhang 0001, Lei Zhang 0001 |
ACM Multimedia | 6 |
| 2010 | Photo2Trip: generating travel routes from geo-tagged photos for trip planningabstractTravel route planning is an important step for a tourist to prepare his/her trip. As a common scenario, a tourist usually asks the following questions when he/she is planning his/her trip in an unfamiliar place: 1) Are there any travel route suggestions for a one-day or three-day trip in Beijing? 2) What is the most popular travel path within the Forbidden City? To facilitate a tourist's trip planning, in this paper, we target at solving the problem of automatic travel route planning. We propose to leverage existing travel clues recovered from 20 million geo-tagged photos collected from www.panoramio.com to suggest customized travel route plans according to users' preferences. As the footprints of tourists at memorable destinations, the geo-tagged photos could be naturally used to discover the travel paths within a destination (attractions/landmarks) and travel routes between destinations. Based on the information discovered from geo-tagged photos, we can provide a customized trip plan for a tourist, i.e., the popular destinations to visit, the visiting order of destinations, the time arrangement in each destination, and the typical travel path within each destination. Users are also enabled to specify personal preference such as visiting location, visiting time/season, travel duration, and destination style in an interactive manner to guide the system. Owning to 20 million geo-tagged photos and 200,000 travelogues, an online system has been developed to help users plan travel routes for over 30,000 attractions/landmarks in more than 100 countries and territories. Experimental results show the intelligence and effectiveness of the proposed framework. Changhu Wang, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001 |
ACM Multimedia | 5 |
| 2010 | Understanding multimedia content using web scale social media dataabstractNowadays, increasingly rich and massive social media data (such as texts, images, audios, videos, blogs, and so on) are being posted to the web, including social networking websites (e.g., MySpace, Facebook), photo and video sharing websites (e.g., Flickr, YouTube), and photo forums (e.g., Photosig.com and Photo.net). Recently, researchers from multidisciplinary areas have proposed to use data-driven approaches for multimedia content understanding by leveraging such unlimited web images and videos as well as their associated rich contextual information (e.g., tag, comments, category, title and metadata). In this three hour tutorial, we plan to introduce the important general concepts and themes of this timely topic. We will also review and summarize the recent multimedia content analysis methods using web-scale social media data as well as present insight into the challenges and future directions in this area. Moreover, we will also show extensive demos on image annotation and retrieval by using rich social media data. Dong Xu 0001, Lei Zhang 0001, Jiebo Luo 0001 |
ACM Multimedia | 2 |
| 2010 | Photo2Trip: an interactive trip planning system based on geo-tagged photosabstractIn this technical demonstration, we present a novel interactive trip planning system, i.e. Photo2Trip, by leveraging existing travel clues recovered from 20 million geo-tagged photos. Compared with the most common ways of trip planning, such as surveying travelogues and resorting to travel forums, Photo2Trip enables users to plan their trips in a more effective way. To meet users' diverse travel requirements, the system considers the following preferences: travel location (e.g. Beijing, Paris, or New York), travel duration (e.g. a two-day trip or a five-day trip), visiting time (e.g. summer, winter, March, or October), and travel style preference (e.g. prefer historic or prefer scenery sites). According to user requirements, Photo2Trip can automatically recommend popular travel routes among multiple destinations (attractions/ landmarks), and suggest typical internal paths within each destination. Moreover, users are allowed to interactively adjust the suggested plans by adding or removing destinations to get more customized travel routes from the system. Owning to 20 million geo-tagged photos and 200,000 travelogues, Photo2Trip is capable of supporting users plan travel routes for over 30,000 attractions/landmarks in more than 100 countries and territories. Huagang Yin, Changhu Wang, Nenghai Yu, Lei Zhang 0001 |
ACM Multimedia | 5 |
| 2010 | Mining adjacent markets from a large-scale ads video collection for image advertisingabstractThe research on image advertising is still in its infancy. Most previous approaches suggest ads by directly matching an ad to a query image, which lacks the power to identify ads from adjacent market. In this paper, we tackle the problem by mining knowledge on adjacent markets from ads videos with a novel Multi-Modal Dirichlet Process Mixture Sets model, which is a unified model of (video frames) clustering and (ads) ranking. Our approach is not only capable of discovering relevant ads (e.g. car ads for a query car image), but also suggesting ads from adjacent markets (e.g. tyre ads). Experimental results show that our proposed approach is fairly effective. Guwen Feng, Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
SIGIR | 3 |
| 2010 | Equip tourists with knowledge mined from traveloguesabstractWith the prosperity of tourism and Web 2.0 technologies, more and more people have willingness to share their travel experiences on the Web (e.g., weblogs, forums, or Web 2.0 communities). These so-called travelogues contain rich information, particularly including location-representative knowledge such as attractions (e.g., Golden Gate Bridge), styles (e.g., beach, history), and activities (e.g., diving, surfing). The location-representative information in travelogues can greatly facilitate other tourists' trip planning, if it can be correctly extracted and summarized. However, since most travelogues are unstructured and contain much noise, it is difficult for common users to utilize such knowledge effectively. In this paper, to mine location-representative knowledge from a large collection of travelogues, we propose a probabilistic topic model, named as Location-Topic model. This model has the advantages of (1) differentiability between two kinds of topics, i.e., local topics which characterize locations and global topics which represent other common themes shared by various locations, and (2) representation of locations in the local topic space to encode both location-representative knowledge and similarities between locations. Some novel applications are developed based on the proposed model, including (1) destination recommendation for on flexible queries, (2) characteristic summarization for a given destination with representative tags and snippets, and (3) identification of informative parts of a travelogue and enriching such highlights with related images. Based on a large collection of travelogues, the proposed framework is evaluated using both objective and subjective evaluation methods and shows promising results. Rui Cai 0002, Changhu Wang, Rong Xiao 0003, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001 |
WWW | 7 |
| 2010 | A pattern tree-based approach to learning URL normalization rulesabstractDuplicate URLs have brought serious troubles to the whole pipeline of a search engine, from crawling, indexing, to result serving. URL normalization is to transform duplicate URLs to a canonical form using a set of rewrite rules. Nowadays URL normalization has attracted significant attention as it is lightweight and can be flexibly integrated into both the online (e.g. crawling) and the offline (e.g. index compression) parts of a search engine. To deal with a large scale of websites, automatic approaches are highly desired to learn rewrite rules for various kinds of duplicate URLs. In this paper, we rethink the problem of URL normalization from a global perspective and propose a pattern tree-based approach, which is remarkably different from existing approaches. Most current approaches learn rewrite rules by iteratively inducing local duplicate pairs to more general forms, and inevitably suffer from noisy training data and are practically inefficient. Given a training set of URLs partitioned into duplicate clusters for a targeted website, we develop a simple yet efficient algorithm to automatically construct a URL pattern tree. With the pattern tree, the statistical information from all the training samples is leveraged to make the learning process more robust and reliable. The learning process is also accelerated as rules are directly summarized based on pattern tree nodes. In addition, from an engineering perspective, the pattern tree helps select deployable rules by removing conflicts and redundancies. An evaluation on more than 70 million duplicate URLs from 200 websites showed that the proposed approach achieves very promising performance, in terms of both de-duping effectiveness and computational efficiency. Rui Cai 0002, Jiang-Ming Yang, Yan Ke, Xiaodong Fan, Lei Zhang 0001 |
WWW | 6 |
| 2010 | Diversifying landmark image search results by learning interested views from community photosabstractIn this paper, we demonstrate a novel landmark photo search and browsing system: Agate, which ranks landmark image search results considering their relevance, diversity and quality. Agate learns from community photos the most interested aspects and related activities of a landmark, and generates adaptively a Table of Content (TOC) as a summary of the attractions to facilitate the user browsing. Image search results are thus re-ranked with the TOC so as to ensure a quick overview of the attractions of the landmarks. A novel non-parametric TOC generation and set-based ranking algorithm, MoM-DPM Sets, is proposed as the key technology of Agate. Experimental results based on human evaluation show the effectiveness of our model and users' preference for Agate. Yuheng Ren, Mo Yu, Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
WWW | 4 |
| 2010 | MindFinder: image search by interactive sketching and taggingabstractIn this technical demonstration, we showcase the MindFinder system $-$ a novel image search engine. Different from existing interactive image search engines, most of which only provide image-level relevance feedback, MindFinder enables users to sketch and tag query images at object level. By considering the image database as a huge repository, MindFinder is able to help users present and refine their initial thoughts in their mind, and finally turn thoughts to a beautiful image(s). Multiple actions are enabled for users to flexibly design their queries in a bilateral interactive manner by leveraging the whole image database, including tagging, refining query by dragging and dropping objects from search results, as well as editing objects. After each action, the search results will be updated in real time to provide users up-to-date materials to further formulate the query. By the deliberate but easy design of the query, MindFinder not only tries to enable users to present on the query panel whatever they are imagining, but also returns to users the most similar images to the picture in users' mind. By scaling up the image database to 10 million, MindFinder has the potential to reveal whatever in users' mind, that is where the name MindFinder comes from. Changhu Wang, Zhiwei Li 0006, Lei Zhang 0001 |
WWW | 3 |
| 2010 | Constructing Concept Lexica With Small Semantic GapsabstractIn recent years, constructing mathematical models for visual concepts by using content features, i.e., color, texture, shape, or local features, has led to the fast development of concept-based multimedia retrieval. In concept-based multimedia retrieval, defining a good lexicon of high-level concepts is the first and important step. However, which concepts should be used for data collection and model construction is still an open question. People agree that concepts that can be easily described by low-level visual features can construct a good lexicon. These concepts are called concepts with small semantic gaps. Unfortunately, there is very little research found on semantic gap analysis and on automatically choosing multimedia concepts with small semantic gaps, even though differences of semantic gaps among concepts are well worth investigating. Yijuan Lu, Lei Zhang 0001, Jiemin Liu, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2009 | Multiplicative nonnegative graph embeddingabstractIn this paper, we study the problem of nonnegative graph embedding, originally investigated in [14] for reaping the benefits from both nonnegative data factorization and the specific purpose characterized by the intrinsic and penalty graphs [13]. Our contributions are two-fold. On the one hand, we present a multiplicative iterative procedure for nonnegative graph embedding, which significantly reduces the computational cost compared with the iterative procedure in [14] involving the matrix inverse calculation of an M-matrix. On the other hand, the nonnegative graph embedding framework is expressed in a more general way by encoding each datum as a tensor of arbitrary order, which brings a group of byproducts, e.g., nonnegative discriminative tensor factorization algorithm, with admissible time and memory cost. Extensive experiments compared with the state-of-the-art algorithms on nonnegative data factorization, graph embedding, and tensor representation demonstrate the algorithmic properties in computation speed, sparsity, discriminating power, and robustness to realistic image occlusions. Changhu Wang, Shuicheng Yan, Lei Zhang 0001, HongJiang Zhang |
CVPR | 4 |
| 2009 | Multi-label sparse coding for automatic image annotationabstractIn this paper, we present a multi-label sparse coding framework for feature extraction and classification within the context of automatic image annotation. First, each image is encoded into a so-called supervector, derived from the universal Gaussian Mixture Models on orderless image patches. Then, a label sparse coding based subspace learning algorithm is derived to effectively harness multi-label information for dimensionality reduction. Finally, the sparse coding method for multi-label data is proposed to propagate the multi-labels of the training images to the query image with the sparse ℓ1reconstruction coefficients. Extensive image annotation experiments on the Corel5k and Corel30k databases both show the superior performance of the proposed multi-label sparse coding framework over the state-of-the-art algorithms. Changhu Wang, Shuicheng Yan, Lei Zhang 0001, HongJiang Zhang |
CVPR | 3 |
| 2009 | The data deluge: Challenges and opportunities of unlimited data in statistical signal processingabstractRecently, there has been a dramatic increase of the amount of audio, video, and images created and shared on the internet by users around the world. Much of this content is publicly available and free of cost. When viewed through the lens of pattern classification, this content can be seen as a virtually unlimited supply of training data for various statistical modeling and labeling tasks such as speech recognition and computer vision. In order to effectively exploit this data resource, significant research challenges must be addressed. In this paper, we present three significant challenges that must be solved to harness the potential of this “data deluge”. We then describe recent work in spoken language processing and image processing that has begun to address these challenges in order to tackle large-scale classification tasks. By bringing together the work of these two communities, we hope to stimulate the cross-pollination of ideas and methods among different signal processing communities. Michael L. Seltzer, Lei Zhang 0001 |
ICASSP | 2 |
| 2009 | Efficient indexing for large scale visual searchabstractWith the popularity of “bag of visual terms” representations of images, many text indexing techniques have been applied in large-scale image retrieval systems. However, due to a fundamental difference between an image query (e.g. 1500 visual terms) and a text query (e.g. 3-5 terms), the usages of some text indexing techniques, e.g. inverted list, are misleading. In this work, we develop a novel indexing technique for this problem. The basic idea is to decompose a document-like representation of an image into two components, one for dimension reduction and the other for residual information preservation. The computing of similarity of two images can be transferred to measuring similarities of their components. The decomposition has two major merits: (1) these components have good properties which enable them to be efficiently indexed and retrieved; (2) The decomposition has better generalization ability than other dimension reduction algorithms. The decomposition can be achieved by either a graphical model or a matrix factorization approach. Theoretic analysis and extensive experiments over a 2.3 million image database show that this framework is scalable to index large scale image database to support fast and accurate visual search. Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma, Harry Shum |
ICCV | 3 |
| 2009 | A lexica family with small semantic gapabstractDefining a lexicon of high-level concepts is the first step for data collection and model construction in concept-based image retrieval. Differences of semantic gaps among concepts are well worth considering. By measuring consistency in visual space and textual space, concepts with small semantic gap can be obtained. Considering so many diverse concepts in large-scale image dataset, we construct a lexica family of high-level concepts with small semantic gap based on different low-level features and different consistency measurements. In this lexica family, the lexica are independent to each other and mutually complementary. It provides helpful suggestions about data collection, feature selection and search model construction for large-scale image retrieval. Jiemin Liu, Qi Tian 0001, Yijuan Lu, Changhu Wang, Lei Zhang 0001, Xiaokang Yang 0001, Shipeng Li 0001 |
ICME | 5 |
| 2009 | Advertising based on users' photosabstractIn this paper, we tackle the problem of learning a user's interest from his photo collections and suggesting relevant ads. We address two key challenges in this work: 1) understanding a user's photos to detect his interest, and 2) bridging the lexical and semantic gap between the vocabulary of ads and that of general users' photos. We solve the first problem by employing a data-driven image annotation approach to annotate each photo and modeling a group of photos, and tackle the second problem by learning and matching the topics of users' photos and ads. The experiments based on real flicker data showed the effectiveness of the approach. Xin-Jing Wang, Mo Yu, Lei Zhang 0001, Wei-Ying Ma |
ICME | 3 |
| 2009 | User grouping behavior in online forumsabstractOnline forums represent one type of social media that is particularly rich for studying human behavior in information seeking and diffusing. The way users join communities is a reflection of the changing and expanding of their interests toward information. In this paper, we study the patterns of user participation behavior, and the feature factors that influence such behavior on different forum datasets. We find that, despite the relative randomness and lesser commitment of structural relationships in online forums, users’ community joining behaviors display some strong regularities. One particularly interesting observation is that the very weak relationships between users defined by online replies have similar diffusion curves as those of real friendships or co-authorships. We build social selection models, Bipartite Markov Random Field (BiMRF), to quantitatively evaluate the prediction performance of those feature factors and their relationships. Using these models, we show that some features carry supplementary information, and the effectiveness of different features vary in different types of forums. Moreover, the results of BiMRF with two-star configurations suggest that the feature of user similarity defined by frequency of communication or number of common friends is inadequate to predict grouping behavior, but adding node-level features can improve the fit of the model. Jun Zhu 0001, Rui Cai 0002, Lei Zhang 0001 |
KDD | 4 |
| 2009 | Incorporating site-level knowledge for incremental crawling of web forums: a list-wise strategyabstractWe study in this paper the problem of incremental crawling of web forums, which is a very fundamental yet challenging step in many web applications. Traditional approaches mainly focus on scheduling the revisiting strategy of each individual page. However, simply assigning different weights for different individual pages is usually inefficient in crawling forum sites because of the different characteristics between forum sites and general websites. Instead of treating each individual page independently, we propose a list-wise strategy by taking into account the site-level knowledge. Such site-level knowledge is mined through reconstructing the linking structure, called sitemap, for a given forum site. With the sitemap, posts from the same thread but distributed on various pages can be concatenated according to their timestamps. After that, for each thread, we employ a regression model to predict the time when the next post arrives. Based on this model, we develop an efficient crawler which is 260% faster than some state-of-the-art methods in terms of fetching new generated content; and meanwhile our crawler also ensure a high coverage ratio. Experimental results show promising performance of Coverage, Bandwidth utilization, and Timeliness of our crawler on 18 various forums. Jiang-Ming Yang, Rui Cai 0002, Chunsong Wang, Hua Huang 0002, Lei Zhang 0001, Wei-Ying Ma |
KDD | 5 |
| 2009 | Generating location overviews with images and tags by mining user-generated traveloguesabstractAutomatically generating location overviews in the form of both visual and textual descriptions is highly desired for online services such as travel planning, to provide attractive and comprehensive outlines of travel destinations. Actually, user-generated content (e.g., travelogues) on the Web provides abundant information to various aspects (e.g., landmarks, styles, activities) of most locations in the world. To leverage the experience shared by Web users, in this paper we propose a location overview generation approach, which first mines location-representative tags from travelogues and then uses such tags to retrieve web images. The learnt tags and retrieved images are finally presented via a novel user interface which provides an informative overview for a given location. Experimental results based on 23,756 travelogues and evaluation over 20 locations show promising results on both travelogue mining and location overview generation. Rui Cai 0002, Xin-Jing Wang, Jiang-Ming Yang, Yanwei Pang, Lei Zhang 0001 |
ACM Multimedia | 6 |
| 2009 | TravelScope: standing on the shoulders of dedicated travelersabstractIn this paper, we propose a system called TravelScope that helps users experience virtual tours by presenting information mined from user-generated travelogues and photos. The system can (1) recommend popular places for a given region; (2) characterize comprehensive aspects (e.g., landmarks, styles, activities) for a location; and (3) show representative images for a landmark. A novel user interface is designed to provide a better user experience, by organizing both the textual and visual information generated by dedicated travelers in an attractive way. Rui Cai 0002, Jiang-Ming Yang, Rong Xiao 0003, Like Liu, Lei Zhang 0001 |
ACM Multimedia | 7 |
| 2009 | Argo: intelligent advertising made possible from users' photosabstractThough monetizing user-generated photos has a great potential in image business, this topic is seldom touched due to the difficulties of both image understanding and ads-to-images vocabulary matching. In this technical demonstration, we show case the Argo system, which attempts to monetize UGC (user-generated content) photos by mining a user's interest from a group of his photos and advertising the photos accordingly. Given a page of photos, it first auto-tags each photo by a large-scale search-based image annotation method, then maps both image annotations and the textual descriptions of ads onto an ODP-based topic hierarchy. The mapping produces semantic features which are statistical distributions on ODP topics. Ads are ranked by their similarities to such topic distributions of the photos and the top-ranked ones are output. Xin-Jing Wang, Mo Yu, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2009 | Multimedia content analysis: model-based approaches vs. data-driven approachesabstractThe explosive growth of multimedia data on the Web has had a great impact on research, and opened a potentially controversial topic in the multimedia community about model-based approaches and data-driven approaches. In this panel, we expect intelligent discussions between the panelists and the audience on this topic. For example, which approach is more advantageous: model-based or data-driven? How can these two approaches learn from the latest developments of the other to further improve existing limitations? The debate in this panel and the insight shared by the panelists are highly anticipated and expected to inspire and drive researchers to answer these questions and push them into working on other, related problems. Lei Zhang 0001, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2009 | LogisticLDA: Regularizing Latent Dirichlet Allocation by Logistic Regression
Jia-Cheng Guo, Bao-Liang Lu, Zhiwei Li 0006, Lei Zhang 0001 |
PACLIC | 4 |
| 2009 | Simultaneously modeling semantics and structure of threaded discussions: a sparse coding approach and its applicationsabstractThe huge amount of knowledge in web communities has motivated the research interests in threaded discussions. The dynamic nature of threaded discussions poses lots of challenging problems for computer scientists. Although techniques such as semantic models and structural models have been shown to be useful in a number of areas, they are inefficient in understanding threaded discussions due to three reasons: (I) as most of users read existing messages before posting, posts in a discussion thread are temporally dependent on the previous ones; It causes the semantics and structure to be coupled with each other in threaded discussions; (II) in online discussion threads, there are a lot of junk posts which are useless and may disturb content analysis; and (III) it is very hard to judge the quality of a post. In this paper, we propose a sparse coding-based model named SMSS to Simultaneously Model Semantics and Structure of threaded discussions. The model projects each post into a topic space, and approximates each post by a linear combination of previous posts in the same discussion thread. Meanwhile, the model also imposes two sparse constraints to force a sparse post reconstruction in the topic space and a sparse post approximation from previous posts. The sparse properties effectively take into account the characteristics of threaded discussions. Towards the above three problems, we demonstrate the competency of our model in three applications: reconstructing reply structure of threaded discussions, identifying junk posts, and finding experts in a given board/sub-board in web communities. Experimental results show encouraging performance of the proposed SMSS model in all these applications. Chen Lin 0001, Jiang-Ming Yang, Rui Cai 0002, Xin-Jing Wang, Wei Wang 0009, Lei Zhang 0001 |
SIGIR | 6 |
| 2009 | Ranking community answers by modeling question-answer relationships via analogical reasoningabstractThe method of finding high-quality answers has significant impact on user satisfaction in community question answering systems. However, due to the lexical gap between questions and answers as well as spam typically existing in usergenerated content, filtering and ranking answers is very challenging. Previous solutions mainly focus on generating redundant features, or finding textual clues using machine learning techniques; none of them ever consider questions and their answers as relational data but instead model them as independent information. Moreover, they only consider the answers of the current question, and ignore any previous knowledge that would be helpful to bridge the lexical and semantic gap. We assume that answers are connected to their questions with various types of latent links, i.e. positive links indicating high-quality answers, negative links indicating incorrect answers or user-generated spam, and propose an analogical reasoning-based approach which measures the analogy between the new question-answer linkages and those of previous relevant knowledge which contains only positive links; the candidate answer which has the most analogous link is assumed to be the best answer. We conducted experiments based on 29.8 million Yahoo!Answer question-answer threads and showed the effectiveness of our approach. Xin-Jing Wang, Xudong Tu, Dan Feng 0001, Lei Zhang 0001 |
SIGIR | 4 |
| 2009 | Modeling semantics and structure of discussion threadsabstractThe abundant knowledge in web communities has motivated the research interests in discussion threads. The dynamic nature of discussion threads poses interesting and challenging problems for computer scientists. Although techniques such as semantic models or structural models have been shown to be useful in a number of areas, they are inefficient in understanding discussion threads due to the temporal dependence among posts in a discussion thread. Such dependence causes that semantics and structure coupled with each other in discussion threads. In this paper, we propose a sparse coding-based model named SMSS to Simultaneously Model Semantic and Structure of discussion threads. Chen Lin 0001, Jiang-Ming Yang, Rui Cai 0002, Xin-Jing Wang, Wei Wang 0009, Lei Zhang 0001 |
WWW | 6 |
| 2009 | Ranking community answers via analogical reasoningabstractDue to the lexical gap between questions and answers, automatically detecting right answers becomes very challenging for community question-answering sites. In this paper, we propose an analogical reasoning-based method. It treats questions and answers as relational data and ranks an answer by measuring the analogy of its link to a query with the links embedded in previous relevant knowledge; the answer that links in the most analogous way to the new question is assumed to be the best answer. We based our experiments on 29.8 million Yahoo!Answer question-answer threads and showed the effectiveness of the approach. Xudong Tu, Xin-Jing Wang, Dan Feng 0001, Lei Zhang 0001 |
WWW | 4 |
| 2009 | Incorporating site-level knowledge to extract structured data from web forumsabstractWeb forums have become an important data resource for many web applications, but extracting structured data from unstructured web forum pages is still a challenging task due to both complex page layout designs and unrestricted user created posts. In this paper, we study the problem of structured data extraction from various web forum sites. Our target is to find a solution as general as possible to extract structured data, such as post title, post author, post time, and post content from any forum site. In contrast to most existing information extraction methods, which only leverage the knowledge inside an individual page, we incorporate both page-level and site-level knowledge and employ Markov logic networks (MLNs) to effectively integrate all useful evidence by learning their importance automatically. Site-level knowledge includes (1) the linkages among different object pages, such as list pages and post pages, and (2) the interrelationships of pages belonging to the same object. The experimental results on 20 forums show a very encouraging information extraction performance, and demonstrate the ability of the proposed approach on various forums. We also show that the performance is limited if only page-level knowledge is used, while when incorporating the site-level knowledge both precision and recall can be significantly improved. Jiang-Ming Yang, Rui Cai 0002, Yida Wang 0008, Jun Zhu 0001, Lei Zhang 0001, Wei-Ying Ma |
WWW | 5 |
| 2009 | A Unified Relevance Feedback Framework for Web Image RetrievalabstractAlthough relevance feedback (RF) has been extensively studied in the content-based image retrieval community, no commercial Web image search engines support RF because of scalability, efficiency, and effectiveness issues. In this paper, we propose a unified relevance feedback framework for Web image retrieval. Our framework shows advantage over traditional RF mechanisms in the following three aspects. First, during the RF process, both textual feature and visual feature are used in a sequential way. To seamlessly combine textual feature-based RF and visual feature-based RF, a query concept-dependent fusion strategy is automatically learned. Second, the textual feature-based RF mechanism employs an effective search result clustering (SRC) algorithm to obtain salient phrases, based on which we could construct an accurate and low-dimensional textual space for the resulting Web images. Thus, we could integrate RF into Web image retrieval in a practical way. Last, a new user interface (UI) is proposed to support implicit RF. On the one hand, unlike traditional RF UI which enforces users to make explicit judgment on the results, the new UI regards the users' click-through data as implicit relevance feedback in order to release burden from the users. On the other hand, unlike traditional RF UI which hardily substitutes subsequent results for previous ones, a recommendation scheme is used to help the users better understand the feedback process and to mitigate the possible waiting caused by RF. Experimental results on a database consisting of nearly three million Web images show that the proposed framework is wieldy, scalable, and effective. En Cheng, Lei Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2009 | FPGA Acceleration of RankBoost in Web Search EnginesabstractSearch relevance is a key measurement for the usefulness of search engines. Shift of search relevance among search engines can easily change a search company's market cap by tens of billions of dollars. With the ever-increasing scale of the Web, machine learning technologies have become important tools to improve search relevance ranking. RankBoost is a promising algorithm in this area, but it is not widely used due to its long training time. To reduce the computation time for RankBoost, we designed a FPGA-based accelerator system and its upgraded version. The accelerator, plugged into a commodity PC, increased the training speed on MSN search engine data up to 1800x compared to the original software implementation on a server. The proposed accelerator has been successfully used by researchers in the search relevance ranking. Ningyi Xu, Xiongfei Cai, Lei Zhang 0001, Feng-Hsiung Hsu |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2008 | Search-based query suggestionabstractIn this paper, we proposed a unified strategy to combine query log and search results for query suggestion. In this way, we leverage both the users' search intentions for popular queries and the power of search engines for unpopular queries. The suggested queries are also ranked according to their relevance and qualities; and each suggestion is described with a rich snippet including a photo and related description. Jiang-Ming Yang, Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
CIKM | 5 |
| 2008 | What are the high-level concepts with small semantic gaps?abstractConcept-based multimedia search has become more and more popular in Multimedia Information Retrieval (MIR). However, which semantic concepts should be used for data collection and model construction is still an open question. Currently, there is very little research found on automatically choosing multimedia concepts with small semantic gaps. In this paper, we propose a novel framework to develop a lexicon of high-level concepts with small semantic gaps (LCSS) from a large-scale web image dataset. By defining a confidence map and content-context similarity matrix, images with small semantic gaps are selected and clustered. The final concept lexicon is mined from the surrounding descriptions (titles, categories and comments) of these images. This lexicon offers a set of high-level concepts with small semantic gaps, which is very helpful for people to focus for data collection, annotation and modeling. It also shows a promising application potential for image annotation refinement and rejection. The experimental results demonstrate the validity of the developed concepts lexicon. Yijuan Lu, Lei Zhang 0001, Qi Tian 0001, Wei-Ying Ma |
CVPR | 2 |
| 2008 | Delivering online advertisements inside imagesabstractWe present in this paper a new channel to deliver online advertisements along with Web images and show a new business model to monetize billions of Web images. The idea is intuitively inspired by image displaying processes on the Web, which typically require people to wait a few seconds before they see full resolution images. This is due to large file sizes and limited network bandwidth. To utilize idle time and the display area, we propose an innovative method for non-intrusively embedding ads into images in a visually pleasant manner. To maintain a smooth user experience, we utilize the thumbnail of the full-resolution image because it is small and visually similar to the full-resolution image. At the client side, a rendering engine first enlarges and blurs the thumbnail, and then blends the pre-chosen ads information into the enlarged image. Based on this idea, we propose three typical scenarios that can adopt the proposed image-advertising mode. More importantly, we can encourage providers of images or other users to participate in our online image ads service by tagging or annotating images. We envision revenue sharing with the providers participating in our service, and we expect that a large number of users will actively submit, tag and annotate images using the system. We have implemented a prototype image ads system, and conducted a series of experiments and user studies to evaluate such a new advertisement channel. The experimental results and user studies show that the proposed online image ad delivery is a non-intrusive ads mode, and the proposed solution is practical. This work also opens multiple new research directions ranging from multimedia to web data mining Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 2 |
| 2008 | Exploring traversal strategy for web forum crawlingabstractIn this paper, we study the problem of Web forum crawling. Web forum has now become an important data source of many Web applications; while forum crawling is still a challenging task due to complex in-site link structures and login controls of most forum sites. Without carefully selecting the traversal path, a generic crawler usually downloads many duplicate and invalid pages from forums, and thus wastes both the precious bandwidth and the limited storage space. To crawl forum data more effectively and efficiently, in this paper, we propose an automatic approach to exploring an appropriate traversal strategy to direct the crawling of a given target forum. In detail, the traversal strategy consists of the identification of the skeleton links and the detection of the page-flipping links. The skeleton links instruct the crawler to only crawl valuable pages and meanwhile avoid duplicate and uninformative ones; and the page-flipping links tell the crawler how to completely download a long discussion thread which is usually shown in multiple pages in Web forums. The extensive experimental results on several forums show encouraging performance of our approach. Following the discovered traversal strategy, our forum crawler can archive more informative pages in comparison with previous related work and a commercial generic crawler. Yida Wang 0008, Jiang-Ming Yang, Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
SIGIR | 5 |
| 2008 | Learning to reduce the semantic gap in web image retrieval and annotationabstractWe study in this paper the problem of bridging the semantic gap between low-level image features and high-level semantic concepts, which is the key hindrance in content-based image retrieval. Piloted by the rich textual information of Web images, the proposed framework tries to learn a new distance measure in the visual space, which can be used to retrieve more semantically relevant images for any unseen query image. The framework differentiates with traditional distance metric learning methods in the following ways. 1) A ranking-based distance metric learning method is proposed for image retrieval problem, by optimizing the leave-one-out retrieval performance on the training data. 2) To be scalable, millions of images together with rich textual information have been crawled from the Web to learn the similarity measure, and the learning framework particularly considers the indexing problem to ensure the retrieval efficiency. 3) To alleviate the noises in the unbalanced labels of images and fully utilize the textual information, a Latent Dirichlet Allocation based topic-level text model is introduced to define pairwise semantic similarity between any two images. The learnt distance measure can be directly applied to applications such as content-based image retrieval and search-based image annotation. Experimental results on the two applications in a two million Web image database show both the effectiveness and efficiency of the proposed framework. Changhu Wang, Lei Zhang 0001, HongJiang Zhang |
SIGIR | 2 |
| 2008 | iRobot: an intelligent crawler for web forumsabstractWe study in this paper the Web forum crawling problem, which is a very fundamental step in many Web applications, such as search engine and Web data mining. As a typical user-created content (UCC), Web forum has become an important resource on the Web due to its rich information contributed by millions of Internet users every day. However, Web forum crawling is not a trivial problem due to the in-depth link structures, the large amount of duplicate pages, as well as many invalid pages caused by login failure issues. In this paper, we propose and build a prototype of an intelligent forum crawler, iRobot, which has intelligence to understand the content and the structure of a forum site, and then decide how to choose traversal paths among different kinds of pages. To do this, we first randomly sample (download) a few pages from the target forum site, and introduce the page content layout as the characteristics to group those pre-sampled pages and re-construct the forum's sitemap. After that, we select an optimal crawling path which only traverses informative pages and skips invalid and duplicate ones. The extensive experimental results on several forums show the performance of our system in the following aspects: 1) Effectiveness - Compared to a generic crawler, iRobot significantly decreases the duplicate and invalid pages; 2) Efficiency - With a small cost of pre-sampling a few pages for learning the necessary knowledge, iRobot saves substantial network bandwidth and storage as it only fetches informative pages from a forum site; and 3) Long threads that are divided into multiple pages can be re-concatenated and archived as a whole thread, which is of great help for further indexing and data mining. Rui Cai 0002, Jiang-Ming Yang, Yida Wang 0008, Lei Zhang 0001 |
WWW | 5 |
| 2008 | Improving relevance judgment of web search results with image excerptsabstractCurrent web search engines return result pages containing mostly text summary even though the matched web pages may contain informative pictures. A text excerpt (i.e. snippet) is generated by selecting keywords around the matched query terms for each returned page to provide context for user's relevance judgment. However, in many scenarios, we found that the pictures in web pages, if selected properly, could be added into search result pages and provide richer contextual description because a picture is worth a thousand words. Such new summary is named as image excerpts. By well designed user study, we demonstrate image excerpts can help users make much quicker relevance judgment of search results for a wide range of query types. To implement this idea, we propose a practicable approach to automatically generate image excerpts in the result pages by considering the dominance of each picture in each web page and the relevance of the picture to the query. We also outline an efficient way to incorporate image excerpts in web search engines. Web search engines can adopt our approach by slightly modifying their index and inserting a few low cost operations in their workflow. Our experiments on a large web dataset indicate the performance of the proposed approach is very promising. Zhiwei Li 0006, Shuming Shi 0001, Lei Zhang 0001 |
WWW | 3 |
| 2008 | Scalable search-based image annotation
Changhu Wang, Lei Zhang 0001, HongJiang Zhang |
Multim. Syst. | 3 |
| 2008 | Annotating Images by Mining Image Search ResultsabstractAlthough it has been studied for years by the computer vision and machine learning communities, image annotation is still far from practical. In this paper, we propose a novel attempt at model-free image annotation, which is a data-driven approach that annotates images by mining their search results. Some 2.4 million images with their surrounding text are collected from a few photo forums to support this approach. The entire process is formulated in a divide-and-conquer framework where a query keyword is provided along with the uncaptioned image to improve both the effectiveness and efficiency. This is helpful when the collected data set is not dense everywhere. In this sense, our approach contains three steps: 1) the search process to discover visually and semantically similar search results, 2) the mining process to identify salient terms from textual descriptions of the search results, and 3) the annotation rejection process to filter out noisy terms yielded by Step 2. To ensure real-time annotation, two key techniques are leveraged-one is to map the high-dimensional image visual features into hash codes, the other is to implement it as a distributed system, of which the search and mining processes are provided as Web services. As a typical result, the entire process finishes in less than 1 second. Since no training data set is required, our approach enables annotating with unlimited vocabulary and is highly scalable and robust to outliers. Experimental results on both real Web images and a benchmark image data set show the effectiveness and efficiency of the proposed algorithm. It is also worth noting that, although the entire approach is illustrated within the divide-and conquer framework, a query keyword is not crucial to our current implementation. We provide experimental results to prove this. Xin-Jing Wang, Lei Zhang 0001, Xirong Li 0001, Wei-Ying Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Reconstruction and Recognition of Tensor-Based Objects With Concurrent Subspaces AnalysisabstractPrincipal Components Analysis (PCA) has traditionally been utilized with data expressed in the form of 1-D vectors, but there exists much data such as gray-level images, video sequences, Gabor-filtered images and so on, that are intrinsically in the form of second or higher order tensors. For representations of image objects in their intrinsic form and order rather than concatenating all the object data into a single vector, we propose in this paper a new optimal object reconstruction criterion with which the information of a high-dimensional tensor is represented as a much lower dimensional tensor computed from projections to multiple concurrent subspaces. In each of these subspaces, correlations with respect to one of the tensor dimensions are reduced, enabling better object reconstruction performance. Concurrent subspaces analysis (CSA) is presented to efficiently learn these subspaces in an iterative manner. In contrast to techniques such as PCA which vectorize tensor data, CSA's direct use of data in tensor form brings an enhanced ability to learn a representative subspace and an increased number of available projection directions. These properties enable CSA to outperform traditional algorithms in the common case of small sample sizes, where CSA can be effective even with only a single sample per class. Extensive experiments on images of faces and digital numbers encoded as second or third order tensors demonstrate that the proposed CSA outperforms PCA-based algorithms in object reconstruction and object recognition. Dong Xu 0001, Shuicheng Yan, Lei Zhang 0001, Stephen Lin 0001, HongJiang Zhang, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | IGroup: presenting web image search results in semantic clustersabstractCurrent web image search engines still rely on user typing textual description: query word(s) for visual targets. As the queries are often short, general or even ambiguous, the images in resulting pages vary in content and style. Thus, browsing with these results is likely to be tedious, frustrating and unpredictable. Jibo He, Qixing Du, Lei Zhang 0001 |
CHI | 5 |
| 2007 | Learning query-biased web page summarizationabstractQuery-biased Web page summarization is the summarization of a Web page reflecting the relevance of it to a specific query. It plays an important role in search results representation of Web search engines. In this paper, we propose a learning-based query-biased Web page summarization method. The summarization problem is solved within the typical sentence selection framework. Different from existing Web page summarization methods that use page content or link context alone, both of them are considered as the sources of sentences in this work. Most of existing learning-based summarization methods treat summarization as a sentence classification problem and train a classifier to discriminate between extracted sentences and non-extracted sentences of all training documents. The basic assumption of these methods is that sentences from different documents are comparable with respect to the class information. In contrast to the classification scheme, a ranking scheme is introduced to rank extracted sentences higher than non-extracted sentences of each training document. The underlying assumption that sentences within a document are comparable is weaker and more reasonable than the assumption of classification-based scheme. Extensive results using intrinsic evaluation metrics gauge many aspects of the proposed method. Changhu Wang, Lei Zhang 0001, HongJiang Zhang |
CIKM | 3 |
| 2007 | Content-Based Image Annotation RefinementabstractAutomatic image annotation has been an active research topic due to its great importance in image retrieval and management. However, results of the state-of-the-art image annotation methods are often unsatisfactory. Despite continuous efforts in inventing new annotation algorithms, it would be advantageous to develop a dedicated approach that could refine imprecise annotations. In this paper, a novel approach to automatically refining the original annotations of images is proposed. For a query image, an existing image annotation method is first employed to obtain a set of candidate annotations. Then, the candidate annotations are re-ranked and only the top ones are reserved as the final annotations. By formulating the annotation refinement process as a Markov process and defining the candidate annotations as the states of a Markov chain, a content-based image annotation refinement (CIAR) algorithm is proposed to re-rank the candidate annotations. It leverages both corpus information and the content feature of a query image. Experimental results on a typical Corel dataset show not only the validity of the refinement, but also the superiority of the proposed algorithm over existing ones. Changhu Wang, Lei Zhang 0001, HongJiang Zhang |
CVPR | 3 |
| 2007 | FPGA-based Accelerator Design for RankBoost in Web Search EnginesabstractSearch relevance is a key measurement for the usefulness of search engines. Shift of search relevance among search engines can easily change a search company's market cap by tens of billions of dollars. With the ever-increasing scale of the Web, machine learning technologies have become important tools to improve search relevance ranking. RankBoost is a promising algorithm in this area, but it is not widely used due to its long training time. To reduce the computation time for RankBoost, we designed a FPGA-based accelerator system. The accelerator, plugged into a commodity PC, increased the training speed on MSN search engine data by 2 orders of magnitude compared to the original software implementation on a server. The proposed accelerator has been successfully used by researchers in the search relevance ranking. Ningyi Xu, Xiongfei Cai, Lei Zhang 0001, Feng-Hsiung Hsu |
FPT | 4 |
| 2007 | Automated Music Video Generation using WEB Image ResourceabstractIn this paper, we proposed a novel prototype of automated music video generation using web image resource. In this prototype, the salient words/phrases of a song's lyrics are first automatically extracted and then used as queries to retrieve related high-quality images from web search engines. To guarantee the coherence among the chosen images' visual representation and the music song, the returned images are further re-ranked and filtered based on their content characteristics such as color, face, landscape, as well as the song's mood type. Finally, those selected images are concatenated to generate a music video using the Photo2Video technique, based on the rhythm information of the music. Preliminary evaluations of the proposed prototype have shown promising results. Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
ICASSP (2) | 2 |
| 2007 | Search Result Clustering Based Relevance Feedback for Web Image RetrivalabstractAlthough relevance feedback (RF) has been extensively studied in the information retrieval community, no commercial Web image search engines support RF because of usability, scalability, and efficiency issues. In this paper, we proposed a search result clustering (SRC) -based RF mechanism for Web image retrieval. The proposed SRC-based RF mechanism employs an effective Search Result Clustering (SRC) algorithm to obtain salient phrases, based on which we could construct an accurate and low-dimensional textual space for the resulting Web images. Given the textual space, we could integrate RF into Web image retrieval in a practical way. The proposed mechanism shows advantage over traditional relevance feedback methods in the following two aspects. On the one hand, our relevance feedback scheme could catch and reflect user's search intension precisely, for the noisy terms would be exempted from the term list with the aid of clustering, thus, the usability of RF in textual space for Web image retrieval is guaranteed. On the other hand, with the exemption of noisy term, the computation with regards to the low-dimensioned textual space is feasible; therefore, the issues of scalability and efficiency for Web image retrieval are addressed. Experimental results on a database consisting of nearly three million Web images show that the proposed mechanism is wieldy, scalable and effective. En Cheng, Lei Zhang 0001 |
ICASSP (1) | 4 |
| 2007 | Retrieving Web Images to Enrich Music RepresentationabstractAudiovisual media which integrates visual media with audio to enrich music representation, such as music video (MV) or music slideshow, is now more welcome than only audio. In this paper, we proposed a novel approach to automatically retrieve web images appropriately suitable to a given music song. In this approach, an imageability measurement is first proposed to select those meaningful words and phrases from lyrics, as queries for further image search. Then, considering the possible ambiguities of queries may cause web image search engines to return images with various semantic concepts, we also introduced a search result clustering (SRC)based strategy to select those images which are more likely to be relevant to the content of music, using a naive Bayesian inference. Preliminary evaluations of the proposed approach on around 100 popular English music songs have shown promising results. Rui Cai 0002, Lei Zhang 0001, Jianmin Li 0001 |
ICME | 3 |
| 2007 | MusicSense: contextual music recommendation using emotional allocation modelingabstractIn this paper, we present a novel contextual music recommendation approach, MusicSense, to automatically suggest music when users read Web documents such as Weblogs. MusicSense matches music to a document's content, in terms of the emotions expressed by both the document and the music songs. To achieve this, we propose a generative model - Emotional Allocation Modeling - in which a collection of word terms is considered as generated with a mixture of emotions. This model also integrates knowledge discovering from a Web-scale corpus and guidance from psychological studies of emotion. Music songs are also described using textual information extracted from their meta-data and relevant Web pages. Thus, both music songs and Web documents can be characterized as distributions over the emotion mixtures through the emotional allocation modeling. For a given document, the songs with the most matched emotion distributions are finally selected as the recommendations. Preliminary experiments on Weblogs show promising results on both emotion allocation and music recommendation. Rui Cai 0002, Chong Wang 0002, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 4 |
| 2007 | Scalable music recommendation by searchabstractThe growth of music resources on personal devices and Internet radio has increased the need for music recommendations. In this paper, aiming at providing an efficient and general solution, we present a search-based solution for scalable music recommendations. In this solution a music piece is first transformed to a music signature sequence in which each signature characterizes the timbre of a local music clip. Based on such signatures, a scale-sensitive method is then proposed to index the music pieces for similarity search, using the locality sensitive hashing (LSH). The scale-sensitive method can numerically find the appropriate parameters for indexing various scales of music collections, and thus can guarantee a proper number of nearest neighbors are found in search. In the recommendation stage, representative signatures from snippets of a seed piece are extracted as query terms, to retrieve pieces with similar melodies for suggestions. We also design a relevance-ranking function to sort the search results, based on the criteria that include matching ratio, temporal order, term weight, and matching confidence. Finally, with the search results, we propose a strategy to generate a dynamic playlist which can automatically expand with time. Evaluations of several music collections at various scales showed that our approach achieves encouraging results in terms of recommendation satisfaction and system scalability. Rui Cai 0002, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2007 | SBIA: search-based image annotation by leveraging web-scale imagesabstractIn this technical demonstration, we showcase the SBIA system - a search-based image annotation system. At the heart of the system lies a very large-scale image search engine which indexed three million Web images and supports both text and visual queries. Given an image (with initial annotations), SBIA first finds semantically/visually similar images via the search engine, and then mines representative keywords from the retrieved images. These keywords, after annotation rejection and relevance ranking, are finally used to annotate the query image. Xirong Li 0001, Xin-Jing Wang, Changhu Wang, Lei Zhang 0001 |
ACM Multimedia | 4 |
| 2007 | Multilinear Discriminant Analysis for Face RecognitionabstractThere is a growing interest in subspace learning techniques for face recognition; however, the excessive dimension of the data space often brings the algorithms into the curse of dimensionality dilemma. In this paper, we present a novel approach to solve the supervised dimensionality reduction problem by encoding an image object as a general tensor of second or even higher order. First, we propose a discriminant tensor criterion, whereby multiple interrelated lower dimensional discriminative subspaces are derived for feature extraction. Then, a novel approach, called k-mode optimization, is presented to iteratively learn these subspaces by unfolding the tensor along different tensor directions. We call this algorithm multilinear discriminant analysis (MDA), which has the following characteristics: 1) multiple interrelated subspaces can collaborate to discriminate different classes, 2) for classification problems involving higher order tensors, the MDA algorithm can avoid the curse of dimensionality dilemma and alleviate the small sample size problem, and 3) the computational cost in the learning stage is reduced to a large extent owing to the reduced data dimensions in k-mode optimization. We provide extensive experiments on ORL, CMU PIE, and FERET databases by encoding face images as second- or third-order tensors to demonstrate that the proposed MDA algorithm based on higher order tensors has the potential to outperform the traditional vector-based subspace learning algorithms, especially in the cases with small sample sizes. Shuicheng Yan, Dong Xu 0001, Qiang Yang 0001, Lei Zhang 0001, Xiaoou Tang, HongJiang Zhang |
IEEE Trans. Image Process. | 4 |
| 2006 | Ranking web objects from multiple communitiesabstractVertical search is a promising direction as it leverages domain-specific knowledge and can provide more precise information for users. In this paper, we study the Web object-ranking problem, one of the key issues in building a vertical search engine. More specifically, we focus on this problem in cases when objects lack relationships between different Web communities, and take high-quality photo search as the test bed for this investigation. We proposed two score fusion methods that can automatically integrate as many Web communities (Web forums) with rating information as possible. The proposed fusion methods leverage the hidden links discovered by a duplicate photo detection algorithm, and aims at minimizing score differences of duplicate photos in different forums. Both intermediate results and user studies show the proposed fusion methods are practical and efficient solutions to Web object ranking in cases we have described. Though the experiments were conducted on high-quality photo ranking, the proposed algorithms are also applicable to other ranking problems, such as movie ranking and music ranking. Lei Zhang 0001, Kefeng Deng, Wei-Ying Ma |
CIKM | 2 |
| 2006 | AnnoSearch: Image Auto-Annotation by SearchabstractAlthough it has been studied for several years by computer vision and machine learning communities, image annotation is still far from practical. In this paper, we present AnnoSearch, a novel way to annotate images using search and data mining technologies. Leveraging the Web-scale images, we solve this problem in two-steps: 1) searching for semantically and visually similar images on the Web, 2) and mining annotations from them. Firstly, at least one accurate keyword is required to enable text-based search for a set of semantically similar images. Then content-based search is performed on this set to retrieve visually similar images. At last, annotations are mined from the descriptions (titles, URLs and surrounding texts) of these images. It worth highlighting that to ensure the efficiency, high dimensional visual features are mapped to hash codes which significantly speed up the content-based search process. Our proposed approach enables annotating with unlimited vocabulary, which is impossible for all existing approaches. Experimental results on real web images show the effectiveness and efficiency of the proposed algorithm. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
CVPR (2) | 2 |
| 2006 | Scalable relevance feedback using click-through data for web image retrievalabstractRelevance feedback (RF) has been extensively studied in the content-based image retrieval community. However, no commercial Web image search engines support RF because of scalability, efficiency and effectiveness issues. In this paper we proposed a scalable relevance feedback mechanism using click-through data for web image retrieval. The proposed mechanism regards users' click-through data as implicit feedback which could be collected at lower cost, in larger quantities and without extra burden on the user. During RF process, both textual feature and visual feature are used in a sequential way. To seamlessly combine textual feature-based RF and visual feature-based RF, a query concept-dependent fusion strategy is automatically learned. Experimental results on a database consisting of nearly three million Web images show that the proposed mechanism is wieldy, scalable and effective. En Cheng, Lei Zhang 0001, Hai Jin 0001 |
ACM Multimedia | 3 |
| 2006 | IGroup: web image search results clusteringabstractIn this paper, we propose, IGroup, an efficient and effective algorithm that organizes Web image search results into clusters. IGroup is different from all existing Web image search results clustering algorithms that only cluster the top few images using visual or textual features. Our proposed algorithm first identifies several query-related semantic clusters based on a key phrases extraction algorithm originally proposed for clustering general Web search results. Then, all the resulting images are separated and assigned to corresponding clusters. As a result, all the resulting images are organized into a clustering structure with semantic level. To make the best use of the clustering results, a new user interface (UI) is proposed. Different from existing Web image search interfaces, which show only a limited number of suggested query terms or representative image thumbnails of some clusters, the proposed interface displays both representative thumbnails and appropriate titles of semantically coherent image clusters. Comprehensive user studies have been completed to evaluate both the clustering algorithm and the new UI. Changhu Wang, Yuhuan Yao, Kefeng Deng, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 5 |
| 2006 | IGroup: a web image search engine with semantic clustering of search resultsabstractIn this demo, we present IGroup, a Web image search engine that organizes the search results into semantic clusters. Different from all existing Web image search results clustering algorithms that only cluster the top few images using visual or textual features, IGroup first identifies several query-related semantic clusters based on a key phrases extraction algorithm originally proposed for clustering general Web search results. Then, all the resulting images are separated and assigned to corresponding clusters. To make the best use of the clustering results, a new user interface is proposed. Please go to http://igroup.msra.cn for real experience. Changhu Wang, Yuhuan Yao, Kefeng Deng, Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 5 |
| 2006 | VirtualTour: an online travel assistant based on high quality imagesabstractWith the popularity of both travel and Web, more and more people use online travel services to facilitate their travel activities or share their travel experiences. Considering that existing services emphasize more on the textual content with the pictorial content only as supplement, we propose the VirtualTour system. It is an online travel service dedicated on high quality images, which helps travelers plan their trip. The images of VirtualTour are from photo forum sites. They have rich and accurate metadata which could be used to extract geographic location information of them and assess the quality of them. A representative sights identification algorithm is also proposed to automatically identify the possible related sights of a region. Based on the map services that seamlessly integrated into the system, a wieldy UI is designed to support several useful features, e.g. query by map, location or path. Lei Zhang 0001, Wei-Ying Ma |
ACM Multimedia | 2 |
| 2006 | Image annotation by large-scale content-based image retrievalabstractImage annotation has been an active research topic in recent years due to its potentially large impact on both image understanding and Web image search. In this paper, we target at solving the automatic image annotation problem in a novel search and mining framework. Given an uncaptioned image, first in the search stage, we perform content-based image retrieval (CBIR) facilitated by high-dimensional indexing to find a set of visually similar images from a large-scale image database. The database consists of images crawled from the World Wide Web with rich annotations, e.g. titles and surrounding text. Then in the mining stage, a search result clustering technique is utilized to find most representative keywords from the annotations of the retrieved image subset. These keywords, after salience ranking, are finally used to annotate the uncaptioned image. Based on search technologies, this framework does not impose an explicit training stage, but efficiently leverages large-scale and well-annotated images, and is potentially capable of dealing with unlimited vocabulary. Based on 2.4 million real Web images, comprehensive evaluation of image annotation on Corel and U. Washington image databases show the effectiveness and efficiency of the proposed approach. Xirong Li 0001, Lei Zhang 0001, Fuzong Lin, Wei-Ying Ma |
ACM Multimedia | 3 |
| 2006 | Image annotation refinement using random walk with restartsabstractImage annotation plays an important role in image retrieval and management. However, the results of the state-of-the-art image annotation methods are often unsatisfactory. Therefore, it is necessary to refine the imprecise annotations obtained by existing annotation methods. In this paper, a novel approach to automatically refine the original annotations of images is proposed. On the one hand, for Web images, textual information, e.g. file name and surrounding text, is used to retrieve a set of candidate annotations. On the other hand, for non-Web images that are lack of textual information, a relevance model-based algorithm using visual information is used to decide the candidate annotations. Then, candidate annotations are re-ranked and only the top ones are reserved as the final annotations. To re-rank the annotations, an algorithm using Random Walk with Restarts (RWR) is proposed to leverage both the corpus information and the original confidence information of the annotations. Experimental results on both non-Web images of Corel dataset and Web images of photo forum sites demonstrate the effectiveness of the proposed method. Changhu Wang, Lei Zhang 0001, HongJiang Zhang |
ACM Multimedia | 3 |
| 2006 | EnjoyPhoto: a vertical image search engine for enjoying high-quality photosabstractIn this paper, we propose building a vertical image search engine called EnjoyPhoto that leverages rich metadata from various photo forum web sites to meet users' requirements for enjoying high-quality photos, which is virtually impossible in traditional image search engines. To solve the ranking problem when aggregating multiple photo forums, we propose a novel rank fusion algorithm that uses duplicate photos to normalize rating scores. To further improve user experiences in enjoying photos, we design an in-place image browsing interface, and compare it with several other interfaces in a user study. With rich metadata and rating information, more attractive user interfaces are enabled, including slideshow authoring and photo recommendations. We conducted experiments and user studies on a 2.5-million image database to evaluate the proposed rank fusion algorithm, investigate the rationale behind building a vertical image search engine, and study user interfaces and preferences for the purpose of enjoying high-quality photos. The experimental results demonstrate the effectiveness of the proposed ranking algorithm. The results also show that the 2.5-million high-quality image database in EnjoyPhoto performs comparably with Google's 1- billion image database for queries related to location, nature, and daily life categories. Finally, our results show that the in-place browsing interface-called Force-Transfer view-is much more convenient for users than traditional interfaces. Lei Zhang 0001, Kefeng Deng, Wei-Ying Ma |
ACM Multimedia | 1 |
| 2006 | Image annotation using search and mining technologiesabstractIn this paper, we present a novel solution to the image annotation problem which annotates images using search and data mining technologies. An accurate keyword is required to initialize this process, and then leveraging a large-scale image database, it 1) searches for semantically and visually similar images, 2) and mines annotations from them. A notable advantage of this approach is that it enables unlimited vocabulary, while it is not possible for all existing approaches. Experimental results on real web images show the effectiveness and efficiency of the proposed algorithm. Xin-Jing Wang, Lei Zhang 0001, Wei-Ying Ma |
WWW | 2 |
| 2006 | Human Gait Recognition With Matrix RepresentationabstractHuman gait is an important biometric feature. It can be perceived from a great distance and has recently attracted greater attention in video-surveillance-related applications, such as closed-circuit television. We explore gait recognition based on a matrix representation in this paper. First, binary silhouettes over one gait cycle are averaged. As a result, each gait video sequence, containing a number of gait cycles, is represented by a series of gray-level averaged images. Then, a matrix-based unsupervised algorithm, namely coupled subspace analysis (CSA), is employed as a preprocessing step to remove noise and retain the most representative information. Finally, a supervised algorithm, namely discriminant analysis with tensor representation, is applied to further improve classification ability. This matrix-based scheme demonstrates a much better gait recognition performance than state-of-the-art algorithms on the standard USF HumanID Gait database. Dong Xu 0001, Shuicheng Yan, Dacheng Tao, Lei Zhang 0001, Xuelong Li 0001, HongJiang Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2005 | Concurrent Subspaces AnalysisabstractA representative subspace is significant for image analysis, while the corresponding techniques often suffer from the curse of dimensionality dilemma. In this paper, we propose a new algorithm, called concurrent subspaces analysis (CSA), to derive representative subspaces by encoding image objects as 2/sup nd/ or even higher order tensors. In CSA, an original higher dimensional tensor is transformed into a lower dimensional one using multiple concurrent subspaces that characterize the most representative information of different dimensions, respectively. Moreover, an efficient procedure is provided to learn these subspaces in an iterative manner. As analyzed in this paper, each sub-step of CSA takes the column vectors of the matrices, which are acquired from the k-mode unfolding of the tensors, as the new objects to be analyzed, thus the curse of dimensionality dilemma can be effectively avoided. The extensive experiments on the 3/sup rd/ order tensor data, simulated video sequences and Gabor filtered digital number image database show that CSA outperforms principal component analysis in terms of both reconstruction and classification capability. Dong Xu 0001, Shuicheng Yan, Lei Zhang 0001, HongJiang Zhang, Zhengkai Liu, Harry Shum |
CVPR (2) | 3 |
| 2005 | Discriminant Analysis with Tensor RepresentationabstractIn this paper, we present a novel approach to solving the supervised dimensionality reduction problem by encoding an image object as a general tensor of 2nd or higher order. First, we propose a discriminant tensor criterion (DTC), whereby multiple interrelated lower-dimensional discriminative subspaces are derived for feature selection. Then, a novel approach called k-mode cluster-based discriminant analysis is presented to iteratively learn these subspaces by unfolding the tensor along different tensor dimensions. We call this algorithm discriminant analysis with tensor representation (DATER), which has the following characteristics: 1) multiple interrelated subspaces can collaborate to discriminate different classes; 2) for classification problems involving higher-order tensors, the DATER algorithm can avoid the curse of dimensionality dilemma and overcome the small sample size problem; and 3) the computational cost in the learning stage is reduced to a large extent owing to the reduced data dimensions in generalized eigenvalue decomposition. We provide extensive experiments by encoding face images as 2nd or 3rd order tensors to demonstrate that the proposed DATER algorithm based on higher order tensors has the potential to outperform the traditional subspace learning algorithms, especially in the small sample size cases. Shuicheng Yan, Dong Xu 0001, Qiang Yang 0001, Lei Zhang 0001, Xiaoou Tang, HongJiang Zhang |
CVPR (1) | 4 |
| 2005 | Coupled Kernel-Based Subspace LearningabstractIt was prescriptive that an image matrix was transformed into a vector before the kernel-based subspace learning. In this paper, we take the kernel discriminant analysis (KDA) algorithm as an example to perform kernel analysis on 2D image matrices directly. First, each image matrix is decomposed as the product of two orthogonal matrices and a diagonal one by using singular value decomposition; then an image matrix is expanded to be of higher or even infinite dimensions by applying the kernel trick on the column vectors of the two orthogonal matrices; finally, two coupled discriminative kernel subspaces are iteratively learned for dimensionality reduction by optimizing the Fisher criterion measured by Frobenius norm. The derived algorithm, called coupled kernel discriminant analysis (CKDA), effectively utilizes the underlying spatial structure of objects and the discriminating information is encoded in two coupled kernel subspaces respectively. The experiments on real face databases compared with KDA and Fisherface validate the effectiveness of CKDA. Shuicheng Yan, Dong Xu 0001, Lei Zhang 0001, Benyu Zhang, HongJiang Zhang |
CVPR (1) | 3 |
| 2005 | Neighborhood Preserving Projections (NPP): A Novel Linear Dimension Reduction Method
Yanwei Pang, Lei Zhang 0001, Zhengkai Liu, Nenghai Yu, Houqiang Li |
ICIC (1) | 2 |
| 2005 | Natural Image Retrieval with SketchesabstractIn this paper, we present a method to retrieve natural images by sketch query. To measure the similarity between the sketch and an image, relevant regions are first located in that image through a multi-resolution search, and a normalized local shape similarity is proposed for image retrieval. Efficiency and other implementation issues are discussed. Experimental results show that it is an effective approach for content-based image retrieval Jinyi Yao, Mingjing Li, Zhiwei Li 0006, Lei Zhang 0001, Wei-Ying Ma |
ICME | 4 |
| 2005 | Auto cropping for digital photographsabstractIn this paper, we propose an effective approach to the nearly untouched problem, still photograph auto cropping, which is one of the important features to automatically enhance photographs. To obtain an optimal result, we first formulate auto cropping as an optimization problem by defining an energy function, which consists of three sub models: composition sub model, conservative sub model, and penalty sub model. Then, particle swarm optimization (PSO) is employed to obtain the optimal solution by maximizing the objective function. Experimental results and user studies over hundreds of photographs show that the proposed approach is effective and accurate in most cases, and can be used in many practical multimedia applications. Mingju Zhang, Lei Zhang 0001, Wei-Ying Ma |
ICME | 2 |
| 2005 | Iteratively clustering web images based on link and attribute reinforcementsabstractImage clustering is an important research topic which contributes to a wide range of applications. Traditional image clustering approaches are based on image content features only, while content features alone can hardly describe the semantics of the images. In the context of Web, images are no longer assumed homogeneous and "flatdistributed but are richly structured. There are two kinds of reinforcements embedded in such data: 1) the reinforcement between attributes of different data types (intra-type links reinforcements); and 2) the reinforcement between object attributes and the inter-type links (inter-type links reinforcements). Unfortunately, most of the previous works addressing relational data failed to fully explore the reinforcements. In this paper, we propose a reinforcement clustering framework to tackle this problem. It reinforces images and texts' attributes via inter-type links and inversely uses these attributes to update these links. The iterative reinforcing nature of this framework promises the discovery of the semantic structure of images, which is the basis of image clustering. Experimental results show the effectiveness of our proposed framework. Xin-Jing Wang, Wei-Ying Ma, Lei Zhang 0001, Xing Li 0001 |
ACM Multimedia | 3 |
| 2005 | Parallel Image Matrix Compression for Face RecognitionabstractThe canonical face recognition algorithm Eigenface and Fisherface are both based on one dimensional vector representation. However, with the high feature dimensions and the small training data, face recognition often suffers from the curse of dimension and the small sample problem. Recent research [4] shows that face recognition based on direct 2D matrix representation, i.e. 2DPCA, obtains better performance than that based on traditional vector representation. However, there are three questions left unresolved in the 2DPCA algorithm: I ) what is the meaning of the eigenvalue and eigenvector of the covariance matrix in 2DPCA; 2) why 2DPCA can outperform Eigenface; and 3) how to reduce the dimension after 2DPCA directly. In this paper, we analyze 2DPCA in a different view and proof that is 2DPCA actually a "localized" PCA with each row vector of an image as object. With this explanation, we discover the intrinsic reason that 2DPCA can outperform Eigenface is because fewer feature dimensions and more samples are used in 2DPCA when compared with Eigenface. To further reduce the dimension after 2DPCA, a two-stage strategy, namely parallel image matrix compression (PIMC), is proposed to compress the image matrix redundancy, which exists among row vectors and column vectors. The exhaustive experiment results demonstrate that PIMC is superior to 2DPCA and Eigenface, and PIMC+LDA outperforms 2DPC+LDA and Fisherface. Dong Xu 0001, Shuicheng Yan, Lei Zhang 0001, Mingjing Li, Wei-Ying Ma, Zhengkai Liu, HongJiang Zhang |
MMM | 3 |
| 2005 | Efficient 3D reconstruction for face recognition
Dalong Jiang, Yuxiao Hu 0001, Shuicheng Yan, Lei Zhang 0001, HongJiang Zhang, Wen Gao 0001 |
Pattern Recognit. | 4 |
| 2005 | Boosting image classification with LDA-based feature combination for digital photograph management
Xuezheng Liu, Lei Zhang 0001, Mingjing Li, HongJiang Zhang, Dingxing Wang |
Pattern Recognit. | 2 |
| 2004 | Automated red-eye detection and correction in digital photographsabstractCaused by light reflected off the subject's retina, red-eye is a troublesome problem in consumer photography. Although most of the cameras have the red-eye reduction mode, the result reality is that no on-camera system is completely effective. In this paper, we propose a fully automatic approach to detecting and correcting red-eyes in digital images. In order to detect red-eyes in a picture, a heuristic yet efficient algorithm is first adopted to detect a group of candidate red regions and then an eye classifier is utilized to confirm whether each candidate region is a human eye. Thereafter, for each detected redeye, we can correct it by the correction algorithm. In case that a red-eye cannot be detected automatically, another algorithm is also provided to detect red-eyes manually with the user's interaction by clicking on an eye. Experimental results on about 300 images with various red-eye appearances demonstrate that the proposed solution is robust and effective. Lei Zhang 0001, Mingjing Li, HongJiang Zhang |
ICIP | 1 |
| 2004 | Efficient propagation for face annotation in family albumsabstractIn this paper, we propose and investigate a new user scenario for face annotation, in which users are allowed to multi-select a group of photographs and assign names to these photographs. The system will then attempt to propagate names from photograph level to face level, i.e. to infer the correspondence between name and face. Given the face similarity measure which combines methodologies from face recognition and content-based image retrieval, we formulate name propagation as an optimization problem. We define the objective function as the sum of similarities between each pair of faces of the same individual in different photographs, and propose an iterative optimization algorithm to infer the optimal correspondence. To make the propagation result reliable, a reject scheme is adopted to reject those with low confidence scores. Furthermore, we investigate the combination and alternation of browsing mode for propagation and viewer mode for annotation, so that each mode can benefit from additional inputs from the other mode. The experimental evaluation has been conducted within a typical family album of over one thousand photographs and the results show that the proposed approach is effective and efficient in automated face annotation in family albums. Lei Zhang 0001, Yuxiao Hu 0001, Mingjing Li, Wei-Ying Ma, HongJiang Zhang |
ACM Multimedia | 1 |
| 2003 | An efficient memorization scheme for relevance feedback in image retrievalabstractWe propose a novel approach to memorizing content relevance information accumulated in relevance feedback sessions (historical relevance feedback) in image retrieval. It is based on the idea that by storing identification numbers of several representative images for each positive image after a query session, it is possible to provide more precise and diverse retrieval results when this image is used later as a query in a new search session. To do so, a criterion for a "good" buddy image is proposed. Based on such criterion, an algorithm is designed to maximize the space spanned by the selected buddy images. Since the proposed approach requires only a constant storage space for each image, it has good scalability for a large size of image database. Experimental results on a database of 10,000 images show the high efficiency and good scalability of the proposed memorization scheme. Lei Zhang 0001, Fang Qian, Mingjing Li, HongJiang Zhang |
ICME | 1 |
| 2003 | Automated annotation of human faces in family albumsabstractAutomatic annotation of photographs is one of the most desirable needs in family photograph management systems. In this paper, we present a learning framework to automate the face annotation in family photograph albums. Firstly, methodologies of content-based image retrieval and face recognition are seamlessly integrated to achieve automated annotation. Secondly, face annotation is formulated in a Bayesian framework, in which the face similarity measure is defined as maximum a posteriori (MAP) estimation. Thirdly, to deal with the missing features, marginal probability is used so that samples which have missing features are compared with those having the full feature set to ensure a non-biased decision. The experimental evaluation has been conducted within a family album of few thousands of photographs and the results show that the proposed approach is effective and efficient in automated face annotation in family albums. Lei Zhang 0001, Longbin Chen, Mingjing Li, HongJiang Zhang |
ACM Multimedia | 1 |
| 2002 | Chinese Named Entity Identification Using Class-based Language Model
Jian Sun 0001, Jianfeng Gao 0001, Lei Zhang 0001, Ming Zhou 0001, Changning Huang |
COLING | 3 |
| 2002 | Gaussian mixture model for relevance feedback in image retrievalabstractRelevance feedback (RF) has become a powerful technique in content-based image retrieval. Most RF methods assume that positive images follow the single Gaussian distribution, which is not sufficient to model the actual distribution of images due to the gap between the semantic concept and low-level features. In this paper, the Gaussian mixture model (GMM) is applied to represent the distribution of positive images in relevance feedback, and a novel method is proposed to estimate the parameters of the GMM. Both positive and negative examples are used to estimate the number of Gaussian components. Furthermore, due to the lack of training samples, unlabeled data are also incorporated to estimate the covariance matrices. Experimental results show that our GMM-based RF method outperforms that based on a single Gaussian model. Fang Qian, Mingjing Li, Lei Zhang 0001, HongJiang Zhang, Bo Zhang 0010 |
ICME (1) | 3 |
| 2002 | MyPhotos: a system for home photo management and processingabstractMyPhotos is a prototype system for home photo management and processing. Several home user orientated image processing and analysis tools are provided. And several auto grouping methods can help user to organize photos. The system also provides a natural user interface and a workflow for easy browsing and searching. HongJiang Zhang, Lei Zhang 0001, Mingjing Li |
ACM Multimedia | 3 |
| 2002 | Boosting Image Orientation Detection with Indoor vs. Outdoor ClassificationabstractAutomatic detection of image orientation is a very important operation in photo image management. In this paper, we propose an automated method based on the boosting algorithm to estimate image orientations. The proposed method has the capability of rejecting images based on the confidence score of the orientation detection. Also, images are classified into indoor and outdoor, and this classification result is used to further refine the orientation detection. To select features more sensitive to the rotation, we combine the features by subtraction operation and select the most useful features by boosting algorithm. The proposed method has several advantages: small model size, fast classification speed, and effective rejection scheme. Lei Zhang 0001, Mingjing Li, HongJiang Zhang |
WACV | 1 |
| 2001 | Support vector machine learning for image retrievalabstractA novel method of relevance feedback is presented based on support vector machine learning in the content-based image retrieval system. A SVM classifier can be learned from training data of relevance images and irrelevance images marked by users. Using the classifier, the system can retrieve more images relevant to the query in the database efficiently. Experiments were carried out on a large-size database of 9918 images. It shows that the interactive learning and retrieval process can find correct images increasingly. It also shows the generalization ability of SVM under the condition of limited training samples. Lei Zhang 0001, Fuzong Lin, Bo Zhang 0010 |
ICIP (2) | 1 |