EDBT 2026 Demo / reviewers in the wild / expert
Ming Yang 0007
dblp:98/2604-7
· DBLP profile ↗
102ranked-venue papers
16as first author
44since 2021 · last 2026
0000-0003-1691-6817ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 81 · 14 first-author · 35 since 2021Artificial intelligence and machine learning · 70 · 10 first-author · 29 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAN: Self-Calibrated AutoregressioN for High-Quality Visual GenerationabstractHuman artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCAN) model capable of self-evaluating and refining generation quality without regenerating the entire image. We unify image token generation and quality evaluation into a single autoregressive model, formulating both tasks as categorical prediction problems. During inference, the model first generates a coarse initial image, then iteratively refines the lowest-quality patches until satisfactory image quality is achieved. Experimental results demonstrate that SCAN effectively handles diverse real-world generation errors and achieves a promising balance between image quality and speed. For example, SCAN-XL achieves an FID of 2.10 and an IS of 326.1, surpassing the LlamaGen-XL by 1.29 (+38%) in FID and 99.0 (+43.6%) in IS, with a 5.6× speedup (19.76s to 3.56s). Compared to recent works, SCAN improves FID and speed by +18.3% and +23% over VAR-d20, and by +7% and +46% over RandAR-XL. Zhanzhou Feng, Qingpei Guo, Jingdong Chen, Ming Yang 0007, Shiliang Zhang |
AAAI | 5 |
| 2026 | Zippo: RGB-Alpha Joint Modeling With a Unified Diffusion ModelabstractRecent advances in generative models have sparked growing interest in moving beyond pure image generation toward transparent image generation, i.e., joint generation of image and its alpha mask. However, most existing approaches adopt a two-stage pipeline, where a diffusion-based model first generates an RGB image and a subsequent matting head predicts the alpha mask. This separation not only leads to error accumulation and inaccurate predictions but also overlooks the intrinsic correlation between the cross-modal data. In this work, we introduce Zippo, a unified diffusion framework, zipping color and transparency distributions into a single diffusion model, by learning joint distribution of RGB image and alpha mask. Zippo not only generates high-fidelity images but also produces plausible and sharp alpha masks. In practice, Zippo inflates the latent space into a unified representation that encodes cross-modal data, and builds upon it with a modality-aware diffusion process that flexibly switches between RGB and alpha domains. In this process, conditioning on one modality while denoising the other allows the model to generate RGB images from alpha masks and predict transparency from input images. In addition to single-modality prediction, we further design a modality-aware noise reassignment strategy to empower Zippo with the joint generation capability of RGB images and their corresponding alpha masks under text guidance. With these techniques, Zippo supports a wide range of transparent image generation tasks, including image-alpha joint generation, image matting, and alpha mask conditioned image generation. Extensive experiments demonstrate that Zippo not only delivers superior visual fidelity but also achieves competitive performance in visual downstream prediction, highlighting joint image-alpha modeling as a powerful alternative to traditional paradigms. Kangyang Xie, Chenchen Jing, Cheng Peng 0011, Ming Yang 0007, Heqian Qiu, Hongliang Li 0001, Hao Chen 0041 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | HomoMatcher: Achieving Dense Feature Matching with Semi-Dense Efficiency by Homography EstimationabstractFeature matching between image pairs is a fundamental problem in computer vision that drives many applications, such as SLAM. Recently, semi-dense matching approaches have achieved substantial performance enhancements and established a widely-accepted coarse-to-fine paradigm. However, the majority of existing methods focus on improving coarse feature representation rather than the fine-matching module. Prior fine-matching techniques, which rely on point-to-patch matching probability expectation or direct regression, often lack precision and do not guarantee the continuity of feature points across sequential images. To address this limitation, this paper concentrates on enhancing the fine-matching module in the semi-dense matching framework. We employ a lightweight and efficient homography estimation network to generate the perspective mapping between patches obtained from coarse matching. This patch-to-patch approach achieves the overall alignment of two patches, resulting in a higher sub-pixel accuracy by incorporating additional constraints. By leveraging the homography estimation between patches, we can achieve a dense matching result with low computational cost. Extensive experiments demonstrate that our method achieves higher accuracy compared to previous semi-dense matchers. Meanwhile, our dense matching results exhibit similar end-point-error accuracy compared to previous dense matchers while maintaining semi-dense efficiency. Xiaolong Wang 0013, Lei Yu 0005, Jiangwei Lao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Yu Zhang 0018, Ming Yang 0007 |
AAAI | 9 |
| 2025 | VQAGuider: Guiding Multimodal Large Language Models to Answer Complex Video QuestionsabstractComplex video question-answering (VQA) requires in-depth understanding of video contents including object and action recognition as well as video classification and summarization, which exhibits great potential in emerging applications in education and entertainment, etc. Multimodal large language models (MLLMs) may accomplish this task by grasping the intention of a question and decomposing it to a series of visual recognition sub-tasks to find out the answer with the help of an agent. To tackle this task, we first collect a new dedicated Complex VQA dataset named CVQA and then propose VQAGuider, an innovative framework planning a few atomic visual recognition tools by video-related API matching. VQAGuider facilitates a deep engagement with video content and precise responses to complex video-related questions by MLLMs, which is beyond aligning visual and language features for simple VQA tasks. Our experiments demonstrate VQAGuider is capable of navigating the complex VQA tasks by MLLMs and improves the accuracy by 29.6% and 17.2% on CVQA and the existing VQA datasets, respectively, highlighting its potential in advancing MLLMs’s capabilities in video understanding. Yuyan Chen, Jiyuan Jia, Yu Guan 0001, Ming Yang 0007, Qingpei Guo |
ACL (1) | 6 |
| 2025 | DynFocus: Dynamic Cooperative Network Empowers LLMs with Video UnderstandingabstractThe challenge in LLM-based video understanding lies in preserving visual and semantic information in long videos while maintaining a memory-affordable token count. However, redundancy and correspondence in videos have hindered the performance potential of existing methods. Through statistical learning on current datasets, we observe that redundancy occurs in both repeated and answer-irrelevant frames, and the corresponding frames vary with different questions. This suggests the possibility of adopting dynamic encoding to balance detailed video information preservation with token budget reduction. To this end, we propose a dynamic cooperative network, DynFocus, for memory-efficient video encoding in this paper. Specifically, i) a Dynamic Event Prototype Estimation (DPE) module to dynamically select meaningful frames for question answering; (ii) a Compact Cooperative Encoding (CCE) module that encodes meaningful frames with detailed visual appearance and the remaining frames with sketchy perception separately. We evaluate our method on five publicly available benchmarks, and experimental results consistently demonstrate that our method achieves competitive performance. Code is available at https://github.com/Simon98-AI/ DynFocus Qingpei Guo, Liyuan Pan, Liu Liu 0009, Yu Guan 0001, Ming Yang 0007 |
CVPR | 6 |
| 2025 | Reversing Flow for Image RestorationabstractImage restoration aims to recover high-quality (HQ) images from degraded low-quality (LQ) ones by reversing the effects of degradation. Existing generative models for image restoration, including diffusion and score-based models, often treat the degradation process as a stochastic transformation, which introduces inefficiency and complexity. In this work, we propose ResFlow, a novel image restoration framework that models the degradation process as a deterministic path using continuous normalizing flows. ResFlow augments the degradation process with an auxiliary process that disambiguates the uncertainty in HQ prediction to enable reversible modeling of the degradation process. ResFlow adopts entropy-preserving flow paths and learns the augmented degradation flow by matching the velocity field. ResFlow significantly improves the performance and speed of image restoration, completing the task in fewer than four sampling steps. Extensive experiments demonstrate that ResFlow achieves state-of-the-art results across various image restoration benchmarks, offering a practical and efficient solution for real-world applications. Haina Qin, Wenyang Luo, Jingdong Chen, Ming Yang 0007, Bing Li 0001, Weiming Hu 0004 |
CVPR | 6 |
| 2025 | MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video GenerationabstractThe image-to-video (I2V) generation is conditioned on the static image, which has been enhanced recently by the motion intensity as an additional control signal. These motion-aware models are appealing to generate diverse motion patterns, yet there lacks a reliable motion estimator for training such models on large-scale video set in the wild. Traditional metrics, e.g., SSIM or optical flow, are hard to generalize to arbitrary videos, while, it is very tough for human annotators to label the abstract motion intensity neither. Furthermore, the motion intensity shall reveal both local object motion and global camera movement, which has not been studied before. This paper addresses the challenge with a new motion estimator, capable of measuring the decoupled motion intensities of objects and cameras in video. We leverage the contrastive learning on randomly paired videos and distinguish the video with greater motion intensity. Such a paradigm is friendly for annotation and easy to scale up to achieve stable performance on motion estimation. We then present a new I2V model, named MotionStone, developed with the decoupled motion estimator. Experimental results demonstrate the stability of the proposed motion estimator and the state-of-the-art performance of MotionStone on I2V generation. These advantages warrant the decoupled motion estimator to serve as a general plug-in enhancer for both data processing and video generation training. Shuwei Shi, Biao Gong, Zizheng Yang, Yuyuan Li 0001, Jingwen He, Kecheng Zheng, Jingdong Chen, Ming Yang 0007, Yinqiang Zheng |
CVPR | 11 |
| 2025 | Mimir: Improving Video Diffusion Models for Precise Text UnderstandingabstractText serves as the key control signal in video generation due to its narrative nature. To render text descriptions into video clips, current video diffusion models borrow features from text encoders yet struggle with limited text comprehension. The recent success of large language models (LLMs) showcases the power of decoder-only transformers, which offers three clear benefits for text-to-video (T2V) generation, namely, precise text understanding resulting from the superior scalability, imagination beyond the input text enabled by next token prediction, and flexibility to prioritize user interests through instruction tuning. Nevertheless, the feature distribution gap emerging from the two different text modeling paradigms hinders the direct use of LLMs in established T2V models. This work addresses this challenge with Mimir, an end-to-end training framework featuring a carefully tailored token fuser to harmonize the outputs from text encoders and LLMs. Such a design allows the T2V model to fully leverage learned video priors while capitalizing on the text-related capability of LLMs. Extensive quantitative and qualitative results demonstrate the effectiveness of Mimir in generating high-quality videos with excellent text comprehension, especially when processing short captions and managing shifting motions. Project page: https://lucaria-academy.github.io/Mimir/ Biao Gong, Yutong Feng, Kecheng Zheng, Shuwei Shi, Yujun Shen, Jingdong Chen, Ming Yang 0007 |
CVPR | 9 |
| 2025 | SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language ModelingabstractOpen-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two challenges. 1) Existing RS semantic categories are limited, particularly for pixel-level interpretation datasets. 2) Distinguishing among diverse RS spatial regions solely by language space is challenging due to the dense and intricate spatial distribution in open-world RS imagery. To address the first issue, we develop a fine-grained RS interpretation dataset, Sky-SA, which contains 183,375 high-quality local image-text pairs with full-pixel manual annotations, covering 1,763 category labels, exhibiting richer semantics and higher density than previous datasets. Afterwards, to solve the second issue, we introduce the vision-centric principle for vision-language modeling. Specifically, in the pre-training stage, the visual self-supervised paradigm is incorporated into image-text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across open-category texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarthOV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https://github.com/zqcrafts/SkySense-O. Qi Zhu 0010, Jiangwei Lao, Deyi Ji, Lixiang Ru, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Dong Liu 0002, Feng Zhao 0004 |
CVPR | 10 |
| 2025 | SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator TrajectoriesabstractWhile MLLMs have demonstrated impressive image understanding capabilities, they still struggle with pixel-level comprehension, limiting their practical applications. Current evaluation tasks such as VQA and visual grounding remain too coarse to assess fine-grained pixel comprehension accurately. Although segmentation is foundational for pixel-level understanding, existing methods often require MLLMs to generate implicit tokens, decoded through external pixel decoders. This approach disrupts the MLLM’s text output space, potentially compromising language capabilities and reducing flexibility and extensibility while failing to reflect the model’s intrinsic pixel-level understanding. Thus, we introduce the Human-Like Mask Annotation Task (HLMAT), a new paradigm where MLLMs mimic human annotators using interactive segmentation tools. Modelling segmentation as a multi-step Markov Decision Process, HLMAT enables MLLMs to iteratively generate text-based click points, achieving high-quality masks without architectural changes or implicit tokens. Through this setup, we develop SegAgent, a model fine-tuned on human-like annotation trajectories, which achieves performance comparable to SoTA methods and supports additional tasks like mask refinement and annotation filtering. HLMAT provides a protocol for assessing fine-grained pixel understanding in MLLMs and introduces a vision-centric, multi-step decision-making task that facilitates the exploration of MLLMs’ visual reasoning abilities. Our adaptations of policy improvement method StaR and PRM guided tree search further enhance model robustness in complex segmentation tasks, laying a foundation for future advancements in fine-grained visual perception and multi-step decision-making for MLLMs. Code can be found at https://github.com/aim-uofa/SegAgent. Muzhi Zhu, Yuzhuo Tian, Hao Chen 0041, Chunluan Zhou, Qingpei Guo, Yang Liu 0357, Ming Yang 0007, Chunhua Shen |
CVPR | 7 |
| 2025 | CasP: Improving Semi-Dense Feature Matching Pipeline Leveraging Cascaded Correspondence Priors for GuidanceabstractSemi-dense feature matching methods have shown strong performance in challenging scenarios. However, the existing pipeline relies on a global search across the entire feature map to establish coarse matches, limiting further improvements in accuracy and efficiency. Motivated by this limitation, we propose a novel pipeline, CasP, which leverages cascaded correspondence priors for guidance. Specifically, the matching stage is decomposed into two progressive phases, bridged by a region-based selective cross-attention mechanism designed to enhance feature discriminability. In the second phase, one-to-one matches are determined by restricting the search range to the one-to-many prior areas identified in the first phase. Additionally, this pipeline benefits from incorporating high-level features, which helps reduce the computational costs of low-level feature extraction. The acceleration gains of CasP increase with higher resolution, and our lite model achieves a speedup of $\sim2.2\times$ at a resolution of 1152 compared to the most efficient method, ELoFTR. Furthermore, extensive experiments demonstrate its superiority in geometric estimation, particularly with impressive cross-domain generalization. These advantages highlight its potential for latency-sensitive and high-robustness applications, such as SLAM and UAV systems. Code is available at https://github.com/pq-chen/CasP. Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yingying Pei, Xinyi Liu 0002, Yongxiang Yao, Lixiang Ru, Liheng Zhong, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002 |
ICCV | 11 |
| 2025 | Unified Visual Generation via Next-Set Prediction in Continuous Domain
Zhanzhou Feng, Qingpei Guo, Xinyu Xiao, Ruihan Xu 0002, Ming Yang 0007, Shiliang Zhang |
ICCV | 5 |
| 2025 | Animate-X: Universal Character Image Animation with Enhanced Motion RepresentationabstractCharacter image animation, which generates high-quality videos from a reference image and target pose sequence, has seen significant progress in recent years. However, most existing methods only apply to human figures, which usually do not generalize well on anthropomorphic characters commonly used in industries like gaming and entertainment. Our in-depth analysis suggests to attribute this limitation to their insufficient modeling of motion, which is unable to comprehend the movement pattern of the driving video, thus imposing a pose sequence rigidly onto the target character. To this end, this paper proposes $\texttt{Animate-X}$, a universal animation framework based on LDM for various character types (collectively named $\texttt{X}$), including anthropomorphic characters. To enhance motion representation, we introduce the Pose Indicator, which captures comprehensive motion pattern from the driving video through both implicit and explicit manner. The former leverages CLIP visual features of a driving video to extract its gist of motion, like the overall movement pattern and temporal relations among motions, while the latter strengthens the generalization of LDM by simulating possible inputs in advance that may arise during inference. Moreover, we introduce a new Animated Anthropomorphic Benchmark ($\texttt{$A^2$Bench}$) to evaluate the performance of $\texttt{Animate-X}$ on universal and widely applicable animation images. Extensive experiments demonstrate the superiority and effectiveness of $\texttt{Animate-X}$ compared to state-of-the-art methods. Biao Gong, Xiang Wang 0012, Shiwei Zhang 0001, Ruobing Zheng, Kecheng Zheng, Jingdong Chen, Ming Yang 0007 |
ICLR | 9 |
| 2025 | Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
Ruobing Zheng, Jingdong Chen, Ming Yang 0007 |
ACM Multimedia | 5 |
| 2025 | Versatile Multimodal Controls for Expressive Talking Human Animation
Ruobing Zheng, Zixin Zhu, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
ACM Multimedia | 7 |
| 2025 | Incremental Model Enhancement via Memory-based Contrastive Learning
Shiyu Xuan, Ming Yang 0007, Shiliang Zhang |
Int. J. Comput. Vis. | 2 |
| 2025 | SHE-Net: Syntax-Hierarchy-Enhanced Text-Video RetrievalabstractThe user base of short video apps has experienced unprecedented growth in recent years, resulting in a significant demand for video content analysis. In particular, text-video retrieval, which aims to find the top matching videos given text descriptions from a vast video corpus, is an essential function, the primary challenge of which is to bridge the modality gap. Nevertheless, most existing approaches treat texts merely as discrete tokens and neglect their syntax structures. Moreover, the abundant spatial and temporal clues in videos are often underutilized due to the lack of interaction with text. To address these issues, we argue that using texts as guidance to focus on relevant temporal frames and spatial regions within videos is beneficial. In this paper, we propose a novel Syntax-Hierarchy-Enhanced text-video retrieval method (SHE-Net) that exploits the inherent semantic and syntax hierarchy of texts to bridge the modality gap from two perspectives. First, to facilitate a more fine-grained integration of visual content, we employ the text syntax hierarchy, which reveals the grammatical structure of text descriptions, to guide the visual representations. Second, to further enhance the multi-modal interaction and alignment, we also utilize the syntax hierarchy to guide the similarity calculation. We evaluated our method on four public text-video retrieval datasets of MSR-VTT, MSVD, DiDeMo, and ActivityNet. The experimental results and ablation studies confirm the advantages of our proposed method. Xuzheng Yu, Chen Jiang 0006, Xingning Dong, Tian Gan 0002, Ming Yang 0007, Qingpei Guo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Towards Better Vision-Inspired Vision-Language ModelsabstractVision-language (VL) models have achieved unprece-dented success recently, in which the connection module is the key to bridge the modality gap. Nevertheless, the abun-dant visual clues are not sufficiently exploited in most existing methods. On the vision side, most existing approaches only use the last feature of the vision tower, without using the low-level features. On the language side, most existing meth-ods only introduce shallow vision-language interactions. In this paper, we present a vision-inspired vision-language con-nection module, dubbed as VIVL, which efficiently exploits the vision cue for VL models. To take advantage of the lower-level information from the vision tower, a feature pyramid extractor (FPE) is introduced to combine features from differ-ent intermediate layers, which enriches the visual cue with negligible parameters and computation overhead. To en-hance VL interactions, we propose deep vision-conditioned prompts (DVCP) that allows deep interactions of vision and language features efficiently. Our VIVL exceeds the previous state-of-the-art method by 18.1 CIDEr when training from scratch on the COCO caption task, which greatly improves the data efficiency. When used as a plug-in module, VIVL consistently improves the performance for various backbones and VL frameworks, delivering new state-of-the-art results on multiple benchmarks, e.g., NoCaps and VQAv2. Yun-Hao Cao, Kaixiang Ji, Chuanyang Zheng, Jiajia Liu 0002, Jian Wang 0108, Jingdong Chen, Ming Yang 0007 |
CVPR | 8 |
| 2024 | SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation ImageryabstractPrior studies on Remote Sensing Foundation Model (RSFM) reveal immense potential towards a generic model for Earth Observation. Nevertheless, these works primar-ily focus on a single modality without temporal and geo-context modeling, hampering their capabilities for diverse tasks. In this study, we present SkySense, a generic billion-scale model, pretrained on a curated multimodal Remote Sensing Imagery (RSI) dataset with 21.5 million temporal sequences. SkySense incorporates a factorized multimodal spatiotemporal encoder taking temporal sequences of opti-cal and Synthetic Aperture Radar (SAR) data as input. This encoder is pretrained by our proposed Multi-Granularity Contrastive Learning to learn representations across different modal and spatial granularities. To further enhance the RSI representations by the geo-context clue, we introduce Geo-Context Prototype Learning to learn region-aware prototypes upon RSI's multimodal spatiotemporal features. To our best knowledge, SkySense is the largest Multi-Modal RSFM to date, whose modules can be flexibly combined or used individually to accommodate various tasks. It demonstrates remarkable generalization capabilities on a thor-ough evaluation encompassing 16 datasets over 7 tasks, from single- to multimodal, static to temporal, and classification to localization. SkySense surpasses 18 recent RSFMs in all test scenarios. Specifically, it outperforms the latest models such as GFM, SatLas and Scale-MAE by a large margin, i.e., 2.76%, 3.67% and 3.61% on average respectively. We will release the pretrained weights to facilitate future research and Earth Observation applications. Xin Guo 0010, Jiangwei Lao, Bo Dang 0002, Lei Yu 0005, Lixiang Ru, Liheng Zhong, Dingxiang Hu, Huimei He, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Yongjun Zhang 0002, Yansheng Li 0001 |
CVPR | 14 |
| 2024 | Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMsabstractMulti-modal Large Language Models (MLLMs) have shown remarkable capabilities in various multi-modal tasks. Nevertheless, their performance in fine-grained image understanding tasks is still limited. To address this issue, this paper proposes a new framework to enhance the fine-grained image understanding abilities of MLLMs. Specifically, we present a new method for constructing the instruction tuning dataset at a low cost by leveraging annotations in existing datasets. A self-consistent bootstrapping method is also introduced to extend existing dense object annotations into high-quality referring-expression-bounding-box pairs. These methods enable the generation of high-quality instruction data which includes a wide range of fundamental abilities essential for fine-grained image perception. Moreover, we argue that the visual encoder should be tuned during instruction tuning to mitigate the gap between full image perception and fine-grained image perception. Experimental results demonstrate the superior performance of our method. For instance, our model exhibits a 5.2% accuracy improvement over Qwen-VL on GQA and surpasses the accuracy of Kosmos-2 by 24.7% on RefCOCO_val. We have also attained the top rank on the leaderboard of MM-Bench. This promising performance is achieved by training on only publicly available data, making it easily reproducible. The models, datasets, and codes are publicly available at https://github.com/SY-Xuan/Pink. Shiyu Xuan, Qingpei Guo, Ming Yang 0007, Shiliang Zhang |
CVPR | 3 |
| 2024 | Learning Dynamic Tetrahedra for High-Quality Talking Head SynthesisabstractRecent works in implicit representations, such as Neural Radiance Fields (NeRF), have advanced the generation of realistic and animatable head avatars from video sequences. These implicit methods are still confronted by visual artifacts and jitters, since the lack of explicit geometric constraints poses a fundamental challenge in accurately modeling complex facial deformations. In this paper, we introduce Dynamic Tetrahedra (DynTet), a novel hybrid representation that encodes explicit dynamic meshes by neural networks to ensure geometric consistency across various motions and viewpoints. DynTet is parameterized by the coordinate-based networks which learn signed distance, deformation, and material texture, anchoring the training data into a predefined tetrahedra grid. Leveraging Marching Tetrahedra, DynTet efficiently decodes textured meshes with a consistent topology, enabling fast rendering through a differentiable rasterizer and supervision via a pixel loss. To enhance training efficiency, we incorporate classical 3D Morphable Models to facilitate geometry learning and define a canonical space for simplifying texture learning. These advantages are readily achievable owing to the effective geometric representation employed in DynTet. Compared with prior works, DynTet demonstrates significant improvements in fidelity, lip synchronization, and real-time performance according to various metrics. Beyond producing stable and visually appealing synthesis videos, our method also outputs the dynamic meshes which is promising to enable many emerging applications. Code is available at https://github.com/zhangzc21/DynTet. Ruobing Zheng, Bonan Li, Congying Han, Tiande Guo, Jingdong Chen, Ziwen Liu 0001, Ming Yang 0007 |
CVPR | 10 |
| 2024 | EcoMatcher: Efficient Clustering Oriented Matcher for Detector-Free Image Matching
Peiqi Chen, Lei Yu 0005, Yi Wan 0001, Yongjun Zhang 0002, Jian Wang 0108, Liheng Zhong, Jingdong Chen, Ming Yang 0007 |
ECCV (68) | 8 |
| 2024 | StyleTokenizer: Defining Image Style by a Single Instance for Controlling Diffusion Models
Wen Li 0024, Muyuan Fang, Biao Gong, Ruobing Zheng, Jingdong Chen, Ming Yang 0007 |
ECCV (28) | 8 |
| 2024 | POA: Pre-training Once for Models of All Sizes
Xin Guo 0010, Jiangwei Lao, Lei Yu 0005, Lixiang Ru, Jian Wang 0108, Guo Ye, Huimei He, Jingdong Chen, Ming Yang 0007 |
ECCV (3) | 10 |
| 2024 | SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal AlignmentabstractMultimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and Image Caption (IC) into a unified framework, resulting in impressive results. CLIP imposes a bidirectional constraints on global representations of entire images and sentences. Although IC conducts an unidirectional image-to-text generation on local representation, it lacks any constraint on local text-to-image reconstruction, which limits the ability to understand images at a fine-grained level when aligned with texts. To achieve multimodal alignment from both global and local perspectives, this paper proposes Symmetrizing Contrastive Captioners (SyCoCa), which introduces bidirectional interactions on images and texts across the global and local representation levels. Specifically, we expand a Text-Guided Masked Image Modeling (TG-MIM) head based on ITC and IC heads. The improved SyCoCa further leverages textual cues to reconstruct contextual images and visual cues to predict textual contents. When implementing bidirectional local interactions, the local contents of images tend to be cluttered or unrelated to their textual descriptions. Thus, we employ an attentive masking strategy to select effective image patches for interaction. Extensive experiments on five vision-language tasks, including image-text retrieval, image-captioning, visual question answering, and zero-shot/finetuned image classification, validate the effectiveness of our proposed method. Ziping Ma 0002, Furong Xu, Ming Yang 0007, Qingpei Guo |
ICML | 4 |
| 2024 | EVE: Efficient Zero-Shot Text-Based Video Editing With Depth Map Guidance and Temporal Consistency Constraints
Xingning Dong, Tian Gan 0002, Chunluan Zhou, Ming Yang 0007, Qingpei Guo |
IJCAI | 5 |
| 2024 | Parameter-Efficient Complementary Expert Learning for Long-Tailed Visual RecognitionabstractLong-tailed recognition (LTR) aims to learn balanced models from extremely unbalanced training data. Fine-tuning pretrained foundation models has recently emerged as a promising research direction for LTR. However, we observe that the fine-tuning process tends to degrade the intrinsic representation capability of pretrained models and lead to model bias towards certain classes, thereby hindering the overall recognition performance. To unleash the intrinsic representation capability of pretrained foundation models, in this work, we propose a new Parameter-Efficient Complementary Expert Learning (PECEL) for LTR. Specifically, PECEL consists of multiple experts, where individual experts are trained via Parameter-Efficient Fine-Tuning (PEFT) and encouraged to learn different expertise on complementary sub-categories via the proposed sample-aware logit adjustment loss. By aggregating the predictions of different experts, PECEL effectively achieves a balanced performance on long-tailed classes. Nevertheless, learning multiple experts generally introduces extra trainable parameters. To ensure parameter efficiency, we further propose a parameter sharing strategy which decomposes and shares the parameters in each expert. Extensive experiments on 4 LTR benchmarks show that the proposed PECEL can effectively learn multiple complementary experts without increasing the trainable parameters and achieve new state-of-the-art performance. Lixiang Ru, Xin Guo 0010, Lei Yu 0005, Jiangwei Lao, Jian Wang 0108, Jingdong Chen, Yansheng Li 0001, Ming Yang 0007 |
ACM Multimedia | 9 |
| 2024 | Accelerating Pre-training of Multimodal LLMs via Chain-of-SightabstractThis paper introduces Chain-of-Sight, a vision-language bridge module that accelerates the pre-training of Multimodal Large Language Models (MLLMs).
Our approach employs a sequence of visual resamplers that capture visual details at various spacial scales.
This architecture not only leverages global and local visual contexts effectively, but also facilitates the flexible extension of visual tokens through a compound token scaling strategy, allowing up to a 16x increase in the token count post pre-training.
Consequently, Chain-of-Sight requires significantly fewer visual tokens in the pre-training phase compared to the fine-tuning phase.
This intentional reduction of visual tokens during pre-training notably accelerates the pre-training process, cutting down the wall-clock training time by $\sim$73\%.
Empirical results on a series of vision-language benchmarks reveal that the pre-train acceleration through Chain-of-Sight is achieved without sacrificing performance, matching or surpassing the standard pipeline of utilizing all visual tokens throughout the entire training process.
Further scaling up the number of visual tokens for pre-training leads to stronger performances, competitive to existing approaches in a series of benchmarks. Kaixiang Ji, Biao Gong, Zhiwu Qing, Kecheng Zheng, Jian Wang 0108, Jingdong Chen, Ming Yang 0007 |
NeurIPS | 9 |
| 2024 | Referencing Where to Focus: Improving Visual Grounding with Referential QueryabstractVisual Grounding aims to localize the referring object in an image given a natural language expression. Recent advancements in DETR-based visual grounding methods have attracted considerable attention, as they directly predict the coordinates of the target object without relying on additional efforts, such as pre-generated proposal candidates or pre-defined anchor boxes. However, existing research primarily focuses on designing stronger multi-modal decoder, which typically generates learnable queries by random initialization or by using linguistic embeddings. This vanilla query generation approach inevitably increases the learning difficulty for the model, as it does not involve any target-related information at the beginning of decoding. Furthermore, they only use the deepest image feature during the query learning process, overlooking the importance of features from other levels. To address these issues, we propose a novel approach, called RefFormer. It consists of the query adaption module that can be seamlessly integrated into CLIP and generate the referential query to provide the prior context for decoder, along with a task-specific decoder. By incorporating the referential query into the decoder, we can effectively mitigate the learning difficulty of the decoder, and accurately concentrate on the target object. Additionally, our proposed query adaption module can also act as an adapter, preserving the rich knowledge within CLIP without the need to tune the parameters of the backbone network. Extensive experiments demonstrate the effectiveness and efficiency of our proposed method, outperforming state-of-the-art approaches on five visual grounding benchmarks. Zhuotao Tian, Qingpei Guo, Sanping Zhou, Ming Yang 0007, Le Wang 0003 |
NeurIPS | 6 |
| 2024 | M2-RAAP: A Multi-Modal Recipe for Advancing Adaptation-based Pre-training towards Effective and Efficient Zero-shot Video-text RetrievalabstractWe present a Recipe for Effective and Efficient zero-shot video-text Retrieval, dubbed M2-RAAP. Upon popular image-text models like CLIP, most current adaptation-based video-text pre-training methods are confronted by three major issues, i.e., noisy data corpus, time-consuming pre-training, and limited performance gain. Towards this end, we conduct a comprehensive study including four critical steps in video-text pre-training. Specifically, we investigate 1) data filtering and refinement, 2) video input type selection, 3) temporal modeling, and 4) video feature enhancement. We then summarize this empirical study into the M2-RAAP recipe, where our technical contributions lie in 1) the data filtering and text re-writing pipeline resulting in 1M high-quality bilingual video-text pairs, 2) the promotion of video inputs with key-frames to accelerate pre-training, and 3) the Auxiliary-Caption-Guided (ACG) strategy to enhance video features. We conduct extensive experiments by adapting three image-text foundation models on two refined video-text datasets from different languages, validating the robustness and reproducibility of M2-RAAP for adaptation-based pre-training. Results demonstrate that M2-RAAP yields superior performance with significantly less data (-90%) and time consumption (-95%), establishing a new SOTA on four English zero-shot retrieval datasets and two Chinese ones. Codebase and refined bilingual data annotations are available at https://github.com/alipay/Ant-Multi-Modal-Framework/tree/main/prj/M2_RAAP. Xingning Dong, Zipeng Feng, Chunluan Zhou, Xuzheng Yu, Ming Yang 0007, Qingpei Guo |
SIGIR | 5 |
| 2024 | Adapting Vision-Language Models via Learning to Inject KnowledgeabstractPre-trained vision-language models (VLM) such as CLIP, have demonstrated impressive zero-shot performance on various vision tasks. Trained on millions or even billions of image-text pairs, the text encoder has memorized a substantial amount of appearance knowledge. Such knowledge in VLM is usually leveraged by learning specific task-oriented prompts, which may limit its performance in unseen tasks. This paper proposes a new knowledge injection framework to pursue a generalizable adaption of VLM to downstream vision tasks. Instead of learning task-specific prompts, we extract task-agnostic knowledge features, and insert them into features of input images or texts. The fused features hence gain better discriminative capability and robustness to intra-category variances. Those knowledge features are generated by inputting learnable prompt sentences into text encoder of VLM, and extracting its multi-layer features. A new knowledge injection module (KIM) is proposed to refine text features or visual features using knowledge features. This knowledge injection framework enables both modalities to benefit from the rich knowledge memorized in the text encoder. Experiments show that our method outperforms recently proposed methods under few-shot learning, base-to-new classes generalization, cross-dataset transfer, and domain generalization settings. For instance, it outperforms CoOp by 4.5% under the few-shot learning setting, and CoCoOp by 4.4% under the base-to-new classes generalization setting. Our code will be released. Shiyu Xuan, Ming Yang 0007, Shiliang Zhang |
IEEE Trans. Image Process. | 2 |
| 2023 | Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive LearningabstractIn recent years, the explosion of web videos makes text-video retrieval increasingly essential and popular for video filtering, recommendation, and search. Text-video retrieval aims to rank relevant text/video higher than irrelevant ones. The core of this task is to precisely measure the cross-modal similarity between texts and videos. Recently, contrastive learning methods have shown promising results for text-video retrieval, most of which focus on the construction of positive and negative pairs to learn text and video representations. Nevertheless, they do not pay enough attention to hard negative pairs and lack the ability to model different levels of semantic similarity. To address these two issues, this paper improves contrastive learning using two novel techniques. First, to exploit hard examples for robust discriminative power, we propose a novel Dual-Modal Attention-Enhanced Module (DMAE) to mine hard negative pairs from textual and visual clues. By further introducing a Negative-aware InfoNCE (NegNCE) loss, we are able to adaptively identify all these hard negatives and explicitly highlight their impacts in the training loss. Second, our work argues that triplet samples can better model fine-grained semantic similarity compared to pairwise samples. We thereby present a new Triplet Partial Margin Contrastive Learning (TPM-CL) module to construct partial order triplet samples by automatically generating fine-grained hard negatives for matched text-video pairs. The proposed TPM-CL designs an adaptive token masking strategy with cross-modal interaction to model subtle semantic differences. Extensive experiments demonstrate that the proposed approach outperforms existing methods on four widely-used text-video retrieval datasets, including MSR-VTT, MSVD, DiDeMo and ActivityNet. Chen Jiang 0006, Xuzheng Yu, Qing Wang 0068, Jia Xu 0013, Zhongyi Liu 0001, Qingpei Guo, Ming Yang 0007, Yuan Qi 0001 |
ACM Multimedia | 10 |
| 2023 | Fine-grained Pseudo Labels for Scene Text RecognitionabstractPseudo-Labeling based semi-supervised learning has shown promising advantages in Scene Text Recognition (STR). Most of them usually use a pre-trained model to generate sequence-level pseudo labels for text images and then re-train the model. Recently, conducting Pseudo-Labeling in a teacher-student framework (a student model is supervised by the pseudo labels from a teacher model) has become increasingly popular, which trains in an end-to-end manner and yields outstanding performance in semi-supervised learning. However, applying this framework directly to Pseudo-Labeling STR exhibits unstable convergence, as generating pseudo labels at the coarse-grained sequence-level leads to inefficient utilization of unlabelled data. Furthermore, the inherent domain shift between labeled and unlabeled data results in low quality of derived pseudo labels. To mitigate the above issues, we propose a novel Cross-domain Pseudo-Labeling (CPL) approach for scene text recognition, which makes better utilization of unlabeled data at the character-level and provides more accurate pseudo labels. Specifically, our proposed Pseudo-Labeled Curriculum Learning dynamically adjusts the thresholds for different character classes according to the model's learning status. Moreover, an Adaptive Distribution Regularizer is employed to bridge the domain gap and improve the quality of pseudo labels. Extensive experiments show that CPL boosts those representative STR models to achieve state-of-the-art results on six challenging STR benchmarks. Besides, it can be effectively generalized to handwritten text. Xiaoxue Chen, Zuming Huang, Lele Xie, Jingdong Chen, Ming Yang 0007 |
ACM Multimedia | 6 |
| 2023 | Complementary Coarse-to-Fine Matching for Video Object SegmentationabstractSemi-supervised Video Object Segmentation (VOS) needs to establish pixel-level correspondences between a video frame and preceding segmented frames to leverage their segmentation clues. Most works rely on features at a single scale to establish those correspondences, e.g., perform dense matching with Convolutional Neural Network (CNN) features from a deep layer. Differently, this work explores complementary features at different scales to pursue more robust feature matching. A coarse feature from a deep layer is first adopted to get coarse pixel-level correspondences. We hence evaluate the quality of those correspondences, and select pixels with low-quality correspondences for fine-scale feature matching. Segmentation clues of previous frames are propagated by both coarse and fine-scale correspondences, which are fused with appearance features for object segmentation. Compared with previous works, this coarse-to-fine matching scheme is more robust to distractions by similar objects and better preserves object details. The sparse fine-scale matching also ensures a fast inference speed. On popular VOS datasets including DAVIS and YouTube-VOS, the proposed method shows promising performance compared with recent works. Zhen Chen 0020, Ming Yang 0007, Shiliang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Joint Global-Local Alignment for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptation has shown promising results in leveraging synthetic (source) images for semantic segmentation of real (target) images. One key issue is how to align data distributions between the source and target domains. Adversarial learning has been applied to align these distributions. However, most existing approaches focus on aligning the output distributions related to image (global) segmentation. Such global alignment may not result in effective alignment due to the inherent high dimensionality feature space involved in the alignment. Moreover, global alignment might be hindered by the noisy outputs corresponding to background pixels in the source domain. To address this limitation, we propose a local output alignment. Such an approach can also mitigate the influences of noisy background pixels from the source domain when performing the local alignment. Our experiments show that by adding local output alignment into various global alignment based domain adaptation, our joint global-local alignment methods improves semantic segmentation. Code is available at https://github.com/skrya/globallocal. Sudhir Yarram, Ming Yang 0007, Junsong Yuan 0001, Chunming Qiao |
ICASSP | 2 |
| 2022 | Asymmetric Label Propagation for Video Object SegmentationabstractSemi-supervised video object segmentation aims to segment foreground objects across a video sequence based on their masks given at the first frame. The motion in adjacent frames tends to be smooth, yet object appearances could change substantially in subsequent frames due to clutters or occlusions. Most existing works segment a video frame by equally referring to segmentation masks of its previous frame and the first frame, and are prone to unreliable matching and accumulated segmentation errors. In order to alleviate this issue, this paper proposes to treat the first and previous frames differently to leverage the motion and appearance clues reliably, and presents an Asymmetric Label Propagation (ALP) method. ALP consists of a Confidence-guided Local Propagation (CLP) module and a Global Label Matching (GLM) module, respectively. CLP propagates labels from the previous frame to the current frame based on local affinity and appearance matching uncertainty. To further recover potential missing objects and alleviate error accumulation, GLM matches the current frame to both the foreground and background of the first frame, and adaptively fuses their matching results. The CLP and GLM outputs are fused to generate object-specific feature maps to perform multi-object segmentation. Extensive experiments on DAVIS and Youtube-VOS datasets demonstrate the effectiveness of the proposed method. Zhen Chen 0020, Ming Yang 0007, Shiliang Zhang |
MMAsia | 2 |
| 2022 | Adversarial structured prediction for domain-adaptive semantic segmentation
Sudhir Yarram, Junsong Yuan 0001, Ming Yang 0007 |
Mach. Vis. Appl. | 3 |
| 2022 | BDCN: Bi-Directional Cascade Network for Perceptual Edge DetectionabstractExploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a bi-directional cascade network (BDCN) architecture, where an individual layer is supervised by labeled edges at its specific scale, rather than directly applying the same supervision to different layers. Furthermore, to enrich multi-scale representations learned by each layer of BDCN, we introduce a scale enhancement module (SEM), which utilizes dilated convolution to generate multi-scale features, instead of using deeper CNNs. These new approaches encourage the learning of multi-scale representations in different layers and detect edges that are well delineated by their scales. Learning scale dedicated layers also results in a compact network with a fraction of parameters. We evaluate our method on three datasets, i.e., BSDS500, NYUDv2, and Multicue, and achieve ODS F-measure of 0.832, 2.7 percent higher than current state-of-the-art on the BSDS500 dataset. We also applied our edge detection result to other vision tasks. Experimental results show that, our method further boosts the performance of image segmentation, optical flow estimation, and object proposal generation. Shiliang Zhang, Ming Yang 0007, Yanhu Shan, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | ForestDet: Large-Vocabulary Long-Tailed Object Detection and Instance SegmentationabstractObject detection and instance segmentation with a large number of object categories and long-tailed data distribution are challenging for most existing deep learning models. As the number of classes increases, the outputs of a classifier become sensitive to likely noisy logits, which can easily result in an incorrect recognition. To alleviate the large-vocabulary problem, we cluster fine-grained classes into coarser parent classes and then build a classification tree to classify an object into a fine-grained class via its parent class. Because the number of parent class is much fewer, their logits are more stable to suppress the wrong/noisy logits existed in the fine-grained class nodes. Due to a variety of ways for clustering fine-grained classes into parent classes, we can further construct multiple trees to build a classification forest where each single tree contributes its vote to the fine-grained classification. Moreover, a simple yet effective resampling method, termed as NMS Resampling, is proposed aiming at solving the long tail (data imbalance) problem. Our method, coined as ForestDet, serves as a plug-and-play module, which can be readily employed in both one-stage and two-stage object recognition models for recognizing more than 1000 categories. Extensive experiments are conducted on the large vocabulary dataset LVIS. Compared to the Mask R-CNN baseline, our two-stage counterpart Forest R-CNN significantly boosts the performance by 11.5% and 3.9% AP improvements on the rare categories and overall categories, respectively. Compared to the RetinaNet baseline, our one-stage counterpart Forest RetinaNet improves 2.1% AP on overall categories. Moreover, we achieve state-of-the-art results on the LVIS dataset.Code and models are available athttps://github.com/JialianW/Forest_RCNN. Jialian Wu, Liangchen Song, Qian Zhang 0009, Ming Yang 0007, Junsong Yuan 0001 |
IEEE Trans. Multim. | 4 |
| 2021 | Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student ModelabstractWhen adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf models is not trivial and generally requires considerable try-and-error and parameter tuning. In this paper, we denote a well-trained model as a teacher network and a model for the new task as a student network. We aim to ease the efforts of transferring knowledge from the teacher to the student network, robust to the gaps between their network architectures, domain data, and task definitions. Specifically, we propose a hybrid forward scheme in training the teacher-student models, alternately updating layer weights of the student model. The key merit of our hybrid forward scheme is on the dynamical balance between the knowledge transfer loss and task specific loss in training. We demonstrate the effectiveness of our method on a variety of tasks, e.g., model compression, segmentation, and detection, under a variety of knowledge transfer settings. Liangchen Song, Jialian Wu, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001 |
AAAI | 3 |
| 2021 | Track To Detect and Segment: An Online Multi-Object TrackerabstractMost online multi-object trackers perform object detection stand-alone in a neural net without any input from tracking. In this paper, we present a new online joint detection and tracking model, TraDeS (TRAck to DEtect and Segment), exploiting tracking clues to assist detection end-to-end. TraDeS infers object tracking offset by a cost volume, which is used to propagate previous object features for improving current object detection and segmentation. Effectiveness and superiority of TraDeS are shown on 4 datasets, including MOT (2D tracking), nuScenes (3D tracking), MOTS and Youtube-VIS (instance segmentation tracking). Project page: https://jialianwu.com/projects/TraDeS.html. Jialian Wu, Jiale Cao, Liangchen Song, Yu Wang 0032, Ming Yang 0007, Junsong Yuan 0001 |
CVPR | 5 |
| 2021 | Stacked Homography Transformations for Multi-View Pedestrian DetectionabstractMulti-view pedestrian detection aims to predict a bird’s eye view (BEV) occupancy map from multiple camera views. This task is confronted with two challenges: how to establish the 3D correspondences from views to the BEV map and how to assemble occupancy information across views. In this paper, we propose a novel Stacked HOmography Transformations (SHOT) approach, which is motivated by approximating projections in 3D world coordinates via a stack of homographies. We first construct a stack of transformations for projecting views to the ground plane at different height levels. Then we design a soft selection module so that the network learns to predict the likelihood of the stack of transformations. Moreover, we provide an in-depth theoretical analysis on constructing SHOT and how well SHOT approximates projections in 3D world coordinates. SHOT is empirically verified to be capable of estimating accurate correspondences from individual views to the BEV map, leading to new state-of-the-art performance on standard evaluation benchmarks. Liangchen Song, Jialian Wu, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001 |
ICCV | 3 |
| 2021 | Handling Difficult Labels for Multi-label Image Classification via Uncertainty DistillationabstractMulti-label image classification aims to predict multiple labels for a single image. However, the difficulties of predicting different labels may vary dramatically due to semantic variations of the label as well as the image context. Direct learning of multi-label classification models has the risk of being biased and overfitting those difficult labels, e.g., deep network based classifiers are over-trained on the difficult labels, therefore, lead to false-positive errors of those difficult labels during testing. To handle difficult labels of multi-label image classification, we propose to calibrate the model, which not only predicts the labels but also estimates the uncertainty of the prediction. With the new calibration branch of the network, the classification model is trained with the pick-all-labels normalized loss and optimized pertaining to the number of positive labels. Moreover, to improve performance on difficult labels, instead of annotating them, we leverage the calibrated model as the teacher network and teach the student network about handling difficult labels via uncertainty distillation. Our proposed uncertainty distillation teaches the student network which labels are highly uncertain through prediction distribution distillation, and locates the image regions that cause such uncertain predictions through uncertainty attention distillation. Conducting extensive evaluations on benchmark datasets, we demonstrate that our proposed uncertainty distillation is valuable to handle difficult labels of multi-label image classification. Liangchen Song, Jialian Wu, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001 |
ACM Multimedia | 3 |
| 2021 | 3D Object Representation Learning: A Set-to-Set Matching PerspectiveabstractIn this paper, we tackle the 3D object representation learning from the perspective of set-to-set matching. Given two 3D objects, calculating their similarity is formulated as the problem of set-to-set similarity measurement between two set of local patches. As local convolutional features from convolutional feature maps are natural representations of local patches, the set-to-set matching between sets of local patches is further converted into a local features pooling problem. To highlight good matchings and suppress the bad ones, we exploit two pooling methods: 1) bilinear pooling and 2) VLAD pooling. We analyze their effectiveness in enhancing the set-to-set matching and meanwhile establish their connection. Moreover, to balance different components inherent in a bilinear-pooled feature, we propose the harmonized bilinear pooling operation, which follows the spirits of intra-normalization used in VLAD pooling. To achieve an end-to-end trainable framework, we implement the proposed harmonized bilinear pooling and intra-normalized VLAD as two layers to construct two types of neural network, multi-view harmonized bilinear network (MHBN) and multi-view VLAD network (MVLADN). Systematic experiments conducted on two public benchmark datasets demonstrate the efficacy of the proposed MHBN and MVLADN in 3D object recognition. Jingjing Meng, Ming Yang 0007, Junsong Yuan 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Temporal-Context Enhanced Detection of Heavily Occluded PedestriansabstractState-of-the-art pedestrian detectors have performed promisingly on non-occluded pedestrians, yet they are still confronted by heavy occlusions. Although many previous works have attempted to alleviate the pedestrian occlusion issue, most of them rest on still images. In this paper, we exploit the local temporal context of pedestrians in videos and propose a tube feature aggregation network (TFAN) aiming at enhancing pedestrian detectors against severe occlusions. Specifically, for an occluded pedestrian in the current frame, we iteratively search for its relevant counterparts along temporal axis to form a tube. Then, features from the tube are aggregated according to an adaptive weight to enhance the feature representations of the occluded pedestrian. Furthermore, we devise a temporally discriminative embedding module (TDEM) and a part-based relation module (PRM), respectively, which adapts our approach to better handle tube drifting and heavy occlusions. Extensive experiments are conducted on three datasets, Caltech, NightOwls and KAIST, showing that our proposed method is significantly effective for heavily occluded pedestrian detection. Moreover, we achieve the state-of-the-art performance on the Caltech and NightOwls datasets. Jialian Wu, Chunluan Zhou, Ming Yang 0007, Qian Zhang 0009, Junsong Yuan 0001 |
CVPR | 3 |
| 2020 | Self-Mimic Learning for Small-scale Pedestrian DetectionabstractDetecting small-scale pedestrians is one of the most challenging problems in pedestrian detection. Due to the lack of visual details, the representations of small-scale pedestrians tend to be weak to be distinguished from background clutters. In this paper, we conduct an in-depth analysis of the small-scale pedestrian detection problem, which reveals that weak representations of small-scale pedestrians are the main cause for a classifier to miss them. To address this issue, we propose a novel Self-Mimic Learning (SML) method to improve the detection performance on small-scale pedestrians. We enhance the representations of small-scale pedestrians by mimicking the rich representations from large-scale pedestrians. Specifically, we design a mimic loss to force the feature representations of small-scale pedestrians to approach those of large-scale pedestrians. The proposed SML is a general component that can be readily incorporated into both one-stage and two-stage detectors, with no additional network layers and incurring no extra computational cost during inference. Extensive experiments on both the CityPersons and Caltech datasets show that the detector trained with the mimic loss is significantly effective for small-scale pedestrian detection and achieves state-of-the-art results on CityPersons and Caltech, respectively. Jialian Wu, Chunluan Zhou, Qian Zhang 0009, Ming Yang 0007, Junsong Yuan 0001 |
ACM Multimedia | 4 |
| 2020 | Learning Recurrent 3D Attention for Video-Based Person Re-IdentificationabstractIn this paper, we propose to learn recurrent 3D attention (A3D) for video-based person re-identification. Attention model plays a key role in both spatial and temporal domains for video representation. Most existing methods apply spatial attention model to extract feature from a single image and aggregate image features with attentive temporal pooling or RNN. However, the inherent consistencies and correlations between spatial and temporal clues are not leveraged. Our A3D method aims to utilize the joint constraints of temporal and spatial attentions to enhance the robustness of attention model. Towards this goal, we treat the pedestrian video as a unified 3D bin where the temporal domain is denoted as an additional dimension. Then we develop an attention agent to iteratively select the locations of the salient spatial-temporal parts in the 3D bin. In addition, we formulate our sequential 3D attention learning as a Markov Decision Process and train the representation network and attention detector with the policy gradient method in an end-to-end manner. We evaluate the proposed method on three challenging datasets including iLIDS-VID, PRID-2011 and the large-scale MARS dataset, and consistently improve the performance in comparison with the state-of-the-art methods. Guangyi Chen 0002, Jiwen Lu, Ming Yang 0007, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Bi-Directional Cascade Network for Perceptual Edge DetectionabstractExploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a Bi-Directional Cascade Network (BDCN) structure, where an individual layer is supervised by labeled edges at its specific scale, rather than directly applying the same supervision to all CNN outputs. Furthermore, to enrich multi-scale representations learned by BDCN, we introduce a Scale Enhancement Module (SEM) which utilizes dilated convolution to generate multi-scale features, instead of using deeper CNNs or explicitly fusing multi-scale edge maps. These new approaches encourage the learning of multi-scale representations in different layers and detect edges that are well delineated by their scales. Learning scale dedicated layers also results in compact network with a fraction of parameters. We evaluate our method on three datasets, i.e., BSDS500, NYUDv2, and Multicue, and achieve ODS Fmeasure of 0.828, 1.3% higher than current state-of-the art on BSDS500. Shiliang Zhang, Ming Yang 0007, Yanhu Shan, Tiejun Huang 0001 |
CVPR | 3 |
| 2019 | SSAP: Single-Shot Instance Segmentation With Affinity PyramidabstractRecently, proposal-free instance segmentation has received increasing attention due to its concise and efficient pipeline. Generally, proposal-free methods generate instance-agnostic semantic segmentation labels and instance-aware features to group pixels into different object instances. However, previous methods mostly employ separate modules for these two sub-tasks and require multiple passes for inference. We argue that treating these two sub-tasks separately is suboptimal. In fact, employing multiple separate modules significantly reduces the potential for application. The mutual benefits between the two complementary sub-tasks are also unexplored. To this end, this work proposes a single-shot proposal-free instance segmentation method that requires only one single pass for prediction. Our method is based on a pixel-pair affinity pyramid, which computes the probability that two pixels belong to the same instance in a hierarchical manner. The affinity pyramid can also be jointly learned with the semantic class labeling and achieve mutual benefits. Moreover, incorporating with the learned affinity pyramid, a novel cascaded graph partition module is presented to sequentially generate instances from coarse to fine. Unlike previous time-consuming graph partition methods, this module achieves 5× speedup and 9% relative improvement on Average-Precision (AP). Our approach achieves new state of the art on the challenging Cityscapes dataset. Naiyu Gao, Yanhu Shan, Yupei Wang, Xin Zhao 0012, Yinan Yu, Ming Yang 0007, Kaiqi Huang |
ICCV | 6 |
| 2019 | Discriminative Feature Transformation for Occluded Pedestrian DetectionabstractDespite promising performance achieved by deep con- volutional neural networks for non-occluded pedestrian de- tection, it remains a great challenge to detect partially oc- cluded pedestrians. Compared with non-occluded pedes- trian examples, it is generally more difficult to distinguish occluded pedestrian examples from background in featue space due to the missing of occluded parts. In this paper, we propose a discriminative feature transformation which en- forces feature separability of pedestrian and non-pedestrian examples to handle occlusions for pedestrian detection. Specifically, in feature space it makes pedestrian exam- ples approach the centroid of easily classified non-occluded pedestrian examples and pushes non-pedestrian examples close to the centroid of easily classified non-pedestrian ex- amples. Such a feature transformation partially compen- sates the missing contribution of occluded parts in feature space, therefore improving the performance for occluded pedestrian detection. We implement our approach in the Fast R-CNN framework by adding one transformation net- work branch. We validate the proposed approach on two widely used pedestrian detection datasets: Caltech and CityPersons. Experimental results show that our approach achieves promising performance for both non-occluded and occluded pedestrian detection. Chunluan Zhou, Ming Yang 0007, Junsong Yuan 0001 |
ICCV | 2 |
| 2019 | Resolution-invariant Person Re-IdentificationabstractExploiting resolution invariant representation is critical for person Re-Identification (ReID) in real applications, where the resolutions of captured person images may vary dramatically. This paper learns person representations robust to resolution variance through jointly training a Foreground-Focus Super-Resolution (FFSR) module and a Resolution-Invariant Feature Extractor (RIFE) by end-to-end CNN learning. FFSR upscales the person foreground using a fully convolutional auto-encoder with skip connections learned with a foreground focus training loss. RIFE adopts two feature extraction streams weighted by a dual-attention block to learn features for low and high resolution images, respectively. These two complementary modules are jointly trained, leading to a strong resolution invariant representation. We evaluate our methods on five datasets containing person images at a large range of resolutions, where our methods show substantial superiority to existing solutions. For instance, we achieve Rank-1 accuracy of 36.4% and 73.3% on CAVIAR and MLR-CUHK03, outperforming the state-of-the art by 2.9% and 2.6%, respectively. Shunan Mao, Shiliang Zhang, Ming Yang 0007 |
IJCAI | 3 |
| 2019 | Spatial-Temporal Attention-Aware Learning for Video-Based Person Re-IdentificationabstractIn this paper, we present a spatial-temporal attention-aware learning (STAL) method for video-based person re-identification. Most existing person re-identification methods aggregate image features identically to represent persons, which are extracted from the same receptive field across video frames. However, the image quality may be varying for different spatial regions and changing over time, which shall contribute to person representation and matching adaptively. Our STAL method aims to attend to the salient parts of persons in videos jointly in both spatial and temporal domains. To achieve this, we slice the video into multiple spatial-temporal units which preserve the body structure of a person and develop a joint spatial-temporal attention model to learn the quality scores of these units. We evaluate the proposed method on three challenging datasets including iLIDS-VID, PRID-2011, and the large-scale MARS dataset, and consistently improve the rank-1 accuracy by a large margin of 5.7%, 0.9%, and 6.6% respectively, in comparison with the state-of-the-art methods. Guangyi Chen 0002, Jiwen Lu, Ming Yang 0007, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | Learning Semantics-Preserving Attention and Contextual Interaction for Group Activity RecognitionabstractIn this paper, we investigate the problem of group activity recognition by learning semantics-perserving attention and contextual interaction among different people. Conventional methods usually aggregate the features extracted from individual persons by pooling operations, which lack physical meaning and cannot fully explore the contextual information for group activity recognition. To address this, we develop a Semantics-Preserving Teacher-Student (SPTS) networks architecture. Our SPTS networks first learn a Teacher Network in the semantic domain that classifies the word of group activity based on the words of individual actions. Then we design a Student Network in the appearance domain that recognizes the group activity according to the input video. We enforce the Student Network to mimic the Teacher Network in the learning procedure. In this way, we allocate semantics-preserving attention to different people, which is more effective to seek the key people and discard the misleading people, while no extra labelled data are required. Moreover, a group of people inherently lie in a graphbased structure, where the people and their relationship can be regarded as the nodes and edges of a graph respectively. Based on this, we build two graph convolutional modules on both the Teacher Network and the Student Network to reason the dependency among different people. Furthermore, we extend our approach on action segmentation task based on its intermediate features. Experimental results on four datasets for group activity analysis clearly show the superior performance of our method in comparisons with the state-of-the-arts. Yansong Tang, Jiwen Lu, Ming Yang 0007, Jie Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Conditional Generative Adversarial Network for Structured Domain AdaptationabstractIn recent years, deep neural nets have triumphed over many computer vision problems, including semantic segmentation, which is a critical task in emerging autonomous driving and medical image diagnostics applications. In general, training deep neural nets requires a humongous amount of labeled data, which is laborious and costly to collect and annotate. Recent advances in computer graphics shed light on utilizing photo-realistic synthetic data with computer generated annotations to train neural nets. Nevertheless, the domain mismatch between real images and synthetic ones is the major challenge against harnessing the generated data and labels. In this paper, we propose a principled way to conduct structured domain adaption for semantic segmentation, i.e., integrating GAN into the FCN framework to mitigate the gap between source and target domains. Specifically, we learn a conditional generator to transform features of synthetic images to real-image like features, and a discriminator to distinguish them. For each training batch, the conditional generator and the discriminator compete against each other so that the generator learns to produce real-image like features to fool the discriminator; afterwards, the FCN parameters are updated to accommodate the changes of GAN. In experiments, without using labels of real image data, our method significantly outperforms the baselines as well as state-of-the-art methods by 12% ~ 20% mean IoU on the Cityscapes dataset. Weixiang Hong 0001, Ming Yang 0007, Junsong Yuan 0001 |
CVPR | 3 |
| 2018 | Deep Reinforcement Learning with Iterative Shift for Visual Tracking
Liangliang Ren, Xin Yuan 0006, Jiwen Lu, Ming Yang 0007, Jie Zhou 0001 |
ECCV (9) | 4 |
| 2018 | Mining Semantics-Preserving Attention for Group Activity RecognitionabstractIn this paper, we propose a Semantics-Preserving Teacher-Student (SPTS) model for group activity recognition in videos, which aims to mine the semantics-preserving attention to automatically seek the key people and discard the misleading people. Conventional methods usually aggregate the features extracted from individual persons by pooling operations, which cannot fully explore the contextual information for group activity recognition. To address this, our SPTS networks first learn a Teacher Network in semantic domain, which classifies the word of group activity based on the words of individual actions. Then we carefully design a Student Network in vision domain, which recognizes the group activity according to the input videos, and enforce the Student Network to mimic the Teacher Network during the learning process. In this way, we allocate semantics-preserving attention to different people, which adequately explores the contextual information of different people and requires no extra labelled data. Experimental results on two widely used benchmarks for group activity recognition clearly show the superior performance of our method in comparisons with the state-of-the-arts. Yansong Tang, Jiwen Lu, Ming Yang 0007, Jie Zhou 0001 |
ACM Multimedia | 5 |
| 2018 | Collaborative Active Visual Recognition from Crowds: A Distributed Ensemble ApproachabstractActive learning is an effective way of engaging users to interactively train models for visual recognition more efficiently. The vast majority of previous works focused on active learning with a single human oracle. The problem of active learning with multiple oracles in a collaborative setting has not been well explored. We present a collaborative computational model for active learning with multiple human oracles, the input from whom may possess different levels of noises. It leads to not only an ensemble kernel machine that is robust to label noises, but also a principled label quality measure to online detect irresponsible labelers. Instead of running independent active learning processes for each individual human oracle, our model captures the inherent correlations among the labelers through shared data among them. Our experiments with both simulated and real crowd-sourced noisy labels demonstrate the efficacy of our model. Gang Hua 0001, Chengjiang Long, Ming Yang 0007, Yan Gao 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Assessing Image Retrieval Quality at the First GlanceabstractImage retrieval has achieved remarkable improvements with the rapid progress on visual representation and indexing techniques. Given a query image, search engines are expected to retrieve relevant results in which the top-ranked short list is of most value to users. However, it is challenging to measure the retrieval quality on-the-fly without direct user feedbacks. In this paper, we aim at evaluating the quality of retrieval results at the first glance (i.e., with the top-ranked images). For each retrieval result, we compute a correlation based feature matrix that comprises of contextual information from the retrieval list, and then feed it into a convolutional neural network regression model for retrieval quality evaluation. In this proposed framework, multiple visual features are integrated together for robust representations. We optimize the output of this simpleyet- effective evaluation method to be consistent with Discounted Cumulative Gain (DCG), the intuitive measure for the quality of the top-ranked results. We evaluate our method in terms of prediction accuracy and consistency with the ground truth, and demonstrate its practicability in applications such as rank list selection and database image abundance analyses. Shaoyan Sun, Wengang Zhou 0001, Qi Tian 0001, Ming Yang 0007, Houqiang Li |
IEEE Trans. Image Process. | 4 |
| 2016 | Scalable Feature Matching by Dual Cascaded Scalar Quantization for Image RetrievalabstractIn this paper, we investigate the problem of scalable visual feature matching in large-scale image search and propose a novel cascaded scalar quantization scheme in dual resolution. We formulate the visual feature matching as a range-based neighbor search problem and approach it by identifying hyper-cubes with a dual-resolution scalar quantization strategy. Specifically, for each dimension of the PCA-transformed feature, scalar quantization is performed at both coarse and fine resolutions. The scalar quantization results at the coarse resolution are cascaded over multiple dimensions to index an image database. The scalar quantization results over multiple dimensions at the fine resolution are concatenated into a binary super-vector and stored into the index list for efficient verification. The proposed cascaded scalar quantization (CSQ) method is free of the costly visual codebook training and thus is independent of any image descriptor training set. The index structure of the CSQ is flexible enough to accommodate new image features and scalable to index large-scale image database. We evaluate our approach on the public benchmark datasets for large-scale image retrieval. Experimental results demonstrate the competitive retrieval performance of the proposed method compared with several recent retrieval algorithms on feature quantization. Wengang Zhou 0001, Ming Yang 0007, Xiaoyu Wang 0002, Houqiang Li, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Web-scale training for face identificationabstractScaling machine learning methods to very large datasets has attracted considerable attention in recent years, thanks to easy access to ubiquitous sensing and data from the web. We study face recognition and show that three distinct properties have surprising effects on the transferability of deep convolutional networks (CNN): (1) The bottleneck of the network serves as an important transfer learning regularizer, and (2) in contrast to the common wisdom, performance saturation may exist in CNN's (as the number of training samples grows); we propose a solution for alleviating this by replacing the naive random subsampling of the training set with a bootstrapping process. Moreover, (3) we find a link between the representation norm and the ability to discriminate in a target domain, which sheds lights on how such networks represent faces. Based on these discoveries, we are able to improve face recognition accuracy on the widely used LFW benchmark, both in the verification (1:1) and identification (1:N) protocols, and directly compare, for the first time, with the state of the art Commercially-Off-The-Shelf system and show a sizable leap in performance. Yaniv Taigman, Ming Yang 0007, Marc'Aurelio Ranzato, Lior Wolf |
CVPR | 2 |
| 2015 | Regionlets for Generic Object DetectionabstractGeneric object detection is confronted by dealing with different degrees of variations, caused by viewpoints or deformations in distinct object classes, with tractable computations. This demands for descriptive and flexible object representations which can be efficiently evaluated in many locations. We propose to model an object class with a cascaded boosting classifier which integrates various types of features from competing local regions, each of which may consist of a group of subregions, named as regionlets. A regionlet is a base feature extraction region defined proportionally to a detection window at an arbitrary resolution (i.e., size and aspect ratio). These regionlets are organized in small groups with stable relative positions to be descriptive to delineate fine-grained spatial layouts inside objects. Their features are aggregated into a one-dimensional feature within one group so as to be flexible to tolerate deformations. The most discriminative regionlets for each object class are selected through a boosting learning procedure. Our regionlet approach achieves very competitive performance on popular multi-class detection benchmark datasets with a single method, without any context. It achieves a detection mean average precision of 41.7 percent on the PASCAL VOC 2007 dataset, and 39.7 percent on the VOC 2010 for 20 object categories. We further develop support pixel integral images to efficiently augment regionlet features with the responses learned by deep convolutional neural networks. Our regionlet based method won second place in the ImageNet Large Scale Visual Object Recognition Challenge (ILSVRC 2013). Xiaoyu Wang 0002, Ming Yang 0007, Shenghuo Zhu, Yuanqing Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Query Specific Rank Fusion for Image RetrievalabstractRecently two lines of image retrieval algorithms demonstrate excellent scalability: 1) local features indexed by a vocabulary tree, and 2) holistic features indexed by compact hashing codes. Although both of them are able to search visually similar images effectively, their retrieval precision may vary dramatically among queries. Therefore, combining these two types of methods is expected to further enhance the retrieval precision. However, the feature characteristics and the algorithmic procedures of these methods are dramatically different, which is very challenging for the feature-level fusion. This motivates us to investigate how to fuse the ordered retrieval sets, i.e., the ranks of images, given by multiple retrieval methods, to boost the retrieval precision without sacrificing their scalability. In this paper, we model retrieval ranks as graphs of candidate images and propose a graph-based query specific fusion approach, where multiple graphs are merged and reranked by conducting a link analysis on a fused graph. The retrieval quality of an individual method is measured on-the-fly by assessing the consistency of the top candidates' nearest neighborhoods. Hence, it is capable of adaptively integrating the strengths of the retrieval methods using local or holistic features for different query images. This proposed method does not need any supervision, has few parameters, and is easy to implement. Extensive and thorough experiments have been conducted on four public datasets, i.e., the UKbench, Corel-5K, Holidays and the large-scale San Francisco Landmarks datasets. Our proposed method has achieved very competitive performance, including state-of-the-art results on several data sets, e.g., the N-S score 3.83 for UKbench. Shaoting Zhang 0001, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | Semantic-Aware Co-Indexing for Image RetrievalabstractIn content-based image retrieval, inverted indexes allow fast access to database images and summarize all knowledge about the database. Indexing multiple clues of image contents allows retrieval algorithms search for relevant images from different perspectives, which is appealing to deliver satisfactory user experiences. However, when incorporating diverse image features during online retrieval, it is challenging to ensure retrieval efficiency and scalability. In this paper, for large-scale image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. Specifically, for an initial set of inverted indexes of local features, we utilize semantic attributes to filter out isolated images and insert semantically similar images to this initial set. Encoding these two distinct and complementary cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online retrieval, because only local features but no semantic attributes are employed for the query. Hence, this co-indexing is different from existing image retrieval methods fusing multiple features or retrieval results. Extensive experiments and comparisons with recent retrieval methods manifest the competitive performance of our method. Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Accurate Object Detection with Location Relaxation and Regionlets Re-localization
Chengjiang Long, Xiaoyu Wang 0002, Gang Hua 0001, Ming Yang 0007, Yuanqing Lin |
ACCV (1) | 4 |
| 2014 | DeepFace: Closing the Gap to Human-Level Performance in Face VerificationabstractIn modern face recognition, the conventional pipeline consists of four stages: detect => align => represent => classify. We revisit both the alignment step and the representation step by employing explicit 3D face modeling in order to apply a piecewise affine transformation, and derive a face representation from a nine-layer deep neural network. This deep network involves more than 120 million parameters using several locally connected layers without weight sharing, rather than the standard convolutional layers. Thus we trained it on the largest facial dataset to-date, an identity labeled dataset of four million facial images belonging to more than 4, 000 identities. The learned representations coupling the accurate model-based alignment with the large facial database generalize remarkably well to faces in unconstrained environments, even with a simple classifier. Our method reaches an accuracy of 97.35% on the Labeled Faces in the Wild (LFW) dataset, reducing the error of the current state of the art by more than 27%, closely approaching human-level performance. Yaniv Taigman, Ming Yang 0007, Marc'Aurelio Ranzato, Lior Wolf |
CVPR | 2 |
| 2014 | Towards Codebook-Free: Scalable Cascaded Hashing for Mobile Image SearchabstractState-of-the-art image retrieval algorithms using local invariant features mostly rely on a large visual codebook to accelerate the feature quantization and matching. This codebook typically contains millions of visual words, which not only demands for considerable resources to train offline but also consumes large amount of memory at the online retrieval stage. This is hardly affordable in resource limited scenarios such as mobile image search applications. To address this issue, we propose a codebook-free algorithm for large scale mobile image search. In our method, we first employ a novel scalable cascaded hashing scheme to ensure the recall rate of local feature matching. Afterwards, we enhance the matching precision by an efficient verification with the binary signatures of these local features. Consequently, our method achieves fast and accurate feature matching free of a huge visual codebook. Moreover, the quantization and binarizing functions in the proposed scheme are independent of small collections of training images and generalize well for diverse image datasets. Evaluated on two public datasets with a million distractor images, the proposed algorithm demonstrates competitive retrieval accuracy and scalability against four recent retrieval methods in literature. Wengang Zhou 0001, Ming Yang 0007, Houqiang Li, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | Collaborative Active Learning of a Kernel Machine Ensemble for RecognitionabstractActive learning is an effective way of engaging users to interactively train models for visual recognition. The vast majority of previous works, if not all of them, focused on active learning with a single human oracle. The problem of active learning with multiple oracles in a collaborative setting has not been well explored. Moreover, most of the previous works assume that the labels provided by the human oracles are noise free, which may often be violated in reality. We present a collaborative computational model for active learning with multiple human oracles. It leads to not only an ensemble kernel machine that is robust to label noises, but also a principled label quality measure to online detect irresponsible labelers. Instead of running independent active learning processes for each individual human oracle, our model captures the inherent correlations among the labelers through shared data among them. Our simulation experiments and experiments with real crowd-sourced noisy labels demonstrated the efficacy of our model. Gang Hua 0001, Chengjiang Long, Ming Yang 0007, Yan Gao 0003 |
ICCV | 3 |
| 2013 | Regionlets for Generic Object DetectionabstractGeneric object detection is confronted by dealing with different degrees of variations in distinct object classes with tractable computations, which demands for descriptive and flexible object representations that are also efficient to evaluate for many locations. In view of this, we propose to model an object class by a cascaded boosting classifier which integrates various types of features from competing local regions, named as region lets. A region let is a base feature extraction region defined proportionally to a detection window at an arbitrary resolution (i.e. size and aspect ratio). These region lets are organized in small groups with stable relative positions to delineate fine grained spatial layouts inside objects. Their features are aggregated to a one-dimensional feature within one group so as to tolerate deformations. Then we evaluate the object bounding box proposal in selective search from segmentation cues, limiting the evaluation locations to thousands. Our approach significantly outperforms the state-of-the-art on popular multi-class detection benchmark datasets with a single method, without any contexts. It achieves the detection mean average precision of 41.7% on the PASCAL VOC 2007 dataset and 39.7% on the VOC 2010 for 20 object categories. It achieves 14.7% mean average precision on the Image Net dataset for 200 object categories, outperforming the latest deformable part-based model (DPM) by 4.7%. Xiaoyu Wang 0002, Ming Yang 0007, Shenghuo Zhu, Yuanqing Lin |
ICCV | 2 |
| 2013 | Semantic-Aware Co-indexing for Image RetrievalabstractInverted indexes in image retrieval not only allow fast access to database images but also summarize all knowledge about the database, so that their discriminative capacity largely determines the retrieval performance. In this paper, for vocabulary tree based image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. For an initial set of inverted indexes of local features, we utilize 1000 semantic attributes to filter out isolated images and insert semantically similar images to the initial set. Encoding these two distinct cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online query cause only local features but no semantic attributes are used for query. Experiments and comparisons with recent retrieval methods on 3 datasets, i.e., UKbench, Holidays, Oxford5K, and 1.3 million images from Flickr as distractors, manifest the competitive performance of our method. Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
ICCV | 2 |
| 2013 | 3D Convolutional Neural Networks for Human Action RecognitionabstractWe consider the automated recognition of human actions in surveillance videos. Most current methods build classifiers based on complex handcrafted features computed from the raw inputs. Convolutional neural networks (CNNs) are a type of deep model that can act directly on the raw inputs. However, such models are currently limited to handling 2D inputs. In this paper, we develop a novel 3D CNN model for action recognition. This model extracts features from both the spatial and the temporal dimensions by performing 3D convolutions, thereby capturing the motion information encoded in multiple adjacent frames. The developed model generates multiple channels of information from the input frames, and the final feature representation combines information from all channels. To further boost the performance, we propose regularizing the outputs with high-level features and combining the predictions of a variety of different models. We apply the developed models to recognize human actions in the real-world environment of airport surveillance videos, and they achieve superior performance in comparison to baseline methods. Shuiwang Ji, Wei Xu 0007, Ming Yang 0007, Kai Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Query Specific Fusion for Image Retrieval
Shaoting Zhang 0001, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Dimitris N. Metaxas |
ECCV (2) | 2 |
| 2011 | Large-scale image classification: Fast feature extraction and SVM trainingabstractMost research efforts on image classification so far have been focused on medium-scale datasets, which are often defined as datasets that can fit into the memory of a desktop (typically 4G~48G). There are two main reasons for the limited effort on large-scale image classification. First, until the emergence of ImageNet dataset, there was almost no publicly available large-scale benchmark data for image classification. This is mostly because class labels are expensive to obtain. Second, large-scale classification is hard because it poses more challenges than its medium-scale counterparts. A key challenge is how to achieve efficiency in both feature extraction and classifier training without compromising performance. This paper is to show how we address this challenge using ImageNet dataset as an example. For feature extraction, we develop a Hadoop scheme that performs feature extraction in parallel using hundreds of mappers. This allows us to extract fairly sophisticated features (with dimensions being hundreds of thousands) on 1.2 million images within one day. For SVM training, we develop a parallel averaging stochastic gradient descent (ASGD) algorithm for training one-against-all 1000-class SVM classifiers. The ASGD algorithm is capable of dealing with terabytes of training data and converges very fast-typically 5 epochs are sufficient. As a result, we achieve state-of-the-art performance on the ImageNet 1000-class classification, i.e., 52.9% in classification accuracy and 71.8% in top 5 hit rate. Yuanqing Lin, Fengjun Lv, Shenghuo Zhu, Ming Yang 0007, Timothée Cour, Kai Yu 0001, Liangliang Cao, Thomas S. Huang |
CVPR | 4 |
| 2011 | Correspondence driven adaptation for human profile recognitionabstractVisual recognition systems for videos using statistical learning models often show degraded performance when being deployed to a real-world environment, primarily due to the fact that training data can hardly cover sufficient variations in reality. To alleviate this issue, we propose to utilize the object correspondences in successive frames as weak supervision to adapt visual recognition models, which is particularly suitable for human profile recognition. Specifically, we substantialize this new strategy on an advanced convolutional neural network (CNN) based system to estimate human gender, age, and race. We enforce the system to output consistent and stable results on face images from the same trajectories in videos by using incremental stochastic training. Our baseline system already achieves competitive performance on gender and age estimation as compared to the state-of-the-art algorithms on the FG-NET database. Further, on two new video datasets containing about 900 persons, the proposed supervision of correspondences improves the estimation accuracy by a large margin over the baseline. Ming Yang 0007, Shenghuo Zhu, Fengjun Lv, Kai Yu 0001 |
CVPR | 1 |
| 2011 | Mining discriminative co-occurrence patterns for visual recognitionabstractThe co-occurrence pattern, a combination of binary or local features, is more discriminative than individual features and has shown its advantages in object, scene, and action recognition. We discuss two types of co-occurrence patterns that are complementary to each other, the conjunction (AND) and disjunction (OR) of binary features. The necessary condition of identifying discriminative co-occurrence patterns is firstly provided. Then we propose a novel data mining method to efficiently discover the optimal co-occurrence pattern with minimum empirical error, despite the noisy training dataset. This mining procedure of AND and OR patterns is readily integrated to boosting, which improves the generalization ability over the conventional boosting decision trees and boosting decision stumps. Our versatile experiments on object, scene, and action categorization validate the advantages of the discovered discriminative co-occurrence patterns. Junsong Yuan 0001, Ming Yang 0007, Ying Wu 0001 |
CVPR | 2 |
| 2011 | Contextual weighting for vocabulary tree based image retrievalabstractIn this paper we address the problem of image retrieval from millions of database images. We improve the vocabulary tree based approach by introducing contextual weighting of local features in both descriptor and spatial domains. Specifically, we propose to incorporate efficient statistics of neighbor descriptors both on the vocabulary tree and in the image spatial domain into the retrieval. These contextual cues substantially enhance the discriminative power of individual local features with very small computational overhead. We have conducted extensive experiments on benchmark datasets, i.e., the UKbench, Holidays, and our new Mobile dataset, which show that our method reaches state-of-the-art performance with much less computation. Furthermore, the proposed method demonstrates excellent scalability in terms of both retrieval accuracy and efficiency on large-scale experiments using 1.26 million images from the ImageNet database as distractors. Xiaoyu Wang 0002, Ming Yang 0007, Timothée Cour, Shenghuo Zhu, Kai Yu 0001, Tony X. Han |
ICCV | 2 |
| 2011 | Real-time clothing recognition in surveillance videosabstractRecognition of clothing categories from videos is appealing to emerging applications such as intelligent customer profile analysis and computer-aided fashion design. This paper presents a complete system to tag clothing categories in real-time, which addresses some practical complications in surveillance videos. Specifically, we take advantage of face detection and tracking to locate human figures and develop an efficient clothing segmentation method utilizing Voronoi images to select seeds for region growing. We compare clothing representations combining color histograms and 3 different texture descriptors. Evaluated on a video dataset with 937 persons and 25441 cloth instances, the system demonstrates promising results in recognizing 8 clothing categories. Ming Yang 0007, Kai Yu 0001 |
ICIP | 1 |
| 2010 | 3D Convolutional Neural Networks for Human Action Recognition
Shuiwang Ji, Wei Xu 0007, Ming Yang 0007, Kai Yu 0001 |
ICML | 3 |
| 2010 | AdaBoost-based face detection for embedded systems
Ming Yang 0007, Jim E. Crenshaw, Bruce Augustine, Russell Mareachen, Ying Wu 0001 |
Comput. Vis. Image Underst. | 1 |
| 2009 | Semi Supervised Image Spam Hunter: A Regularized Discriminant EM Approach
Yan Gao 0003, Ming Yang 0007, Alok N. Choudhary |
ADMA | 2 |
| 2009 | Detection driven adaptive multi-cue integration for multiple human trackingabstractIn video surveillance scenarios, appearances of both human and their nearby scenes may experience large variations due to scale and view angle changes, partial occlusions, or interactions of a crowd. These challenges may weaken the effectiveness of a dedicated target observation model even based on multiple cues, which demands for an agile framework to adjust target observation models dynamically to maintain their discriminative power. Towards this end, we propose a new adaptive way to integrate multi-cue in tracking multiple human driven by human detections. Given a human detection can be reliably associated with an existing trajectory, we adapt the way how to combine specifically devised models based on different cues in this tracker so as to enhance the discriminative power of the integrated observation model in its local neighborhood. This is achieved by solving a regression problem efficiently. Specifically, we employ 3 observation models for a single person tracker based on color models of part of torso regions, an elliptical head model, and bags of local features, respectively. Extensive experiments on 3 challenging surveillance datasets demonstrate long-term reliable tracking performance of this method. Ming Yang 0007, Fengjun Lv, Wei Xu 0007, Yihong Gong |
ICCV | 1 |
| 2009 | Detecting video events based on action recognition in complex scenes using spatio-temporal descriptorabstractEvent detection plays an essential role in video content analysis and remains a challenging open problem. In particular, the study on detecting human-related video events in complex scenes with both a crowd of people and dynamic motion is still limited. In this paper, we investigate detecting video events that involve elementary human actions, e.g. making cellphone call, putting an object down, and pointing to something, in complex scenes using a novel spatio-temporal descriptor based approach. A new spatio-temporal descriptor, which temporally integrates the statistics of a set of response maps of low-level features, e.g. image gradients and optical flows, in a space-time cube, is proposed to capture the characteristics of actions in terms of their appearance and motion patterns. Based on this kind of descriptors, the bag-of-words method is utilized to describe a human figure as a concise feature vector. Then, these features are employed to train SVM classifiers at multiple spatial pyramid levels to distinguish different actions. Finally, a Gaussian kernel based temporal filtering is conducted to segment the sequences of events from a video stream taking account of the temporal consistency of actions. The proposed approach is capable of tolerating spatial layout variations and local deformations of human actions due to diverse view angles and rough human figure alignment in complex scenes. Extensive experiments on the 50-hour video dataset of TRECVid 2008 event detection task demonstrate that our approach outperforms the well-known SIFT descriptor based methods and effectively detects video events in challenging real-world conditions. Guangyu Zhu 0002, Ming Yang 0007, Kai Yu 0001, Wei Xu 0007, Yihong Gong |
ACM Multimedia | 2 |
| 2009 | Context-Aware Visual TrackingabstractEnormous uncertainties in unconstrained environments lead to a fundamental dilemma that many tracking algorithms have to face in practice: Tracking has to be computationally efficient, but verifying whether or not the tracker is following the true target tends to be demanding, especially when the background is cluttered and/or when occlusion occurs. Due to the lack of a good solution to this problem, many existing methods tend to be either effective but computationally intensive by using sophisticated image observation models or efficient but vulnerable to false alarms. This greatly challenges long-duration robust tracking. This paper presents a novel solution to this dilemma by considering the context of the tracking scene. Specifically, we integrate into the tracking process a set of auxiliary objects that are automatically discovered in the video on the fly by data mining. Auxiliary objects have three properties, at least in a short time interval: 1) persistent co-occurrence with the target, 2) consistent motion correlation to the target, and 3) easy to track. Regarding these auxiliary objects as the context of the target, the collaborative tracking of these auxiliary objects leads to efficient computation as well as strong verification. Our extensive experiments have exhibited exciting performance in very challenging real-world testing cases. Ming Yang 0007, Ying Wu 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Tracking Nonstationary Visual Appearances by Data-Driven AdaptationabstractWithout any prior about the target, the appearance is usually the only cue available in visual tracking. However, in general, the appearances are often nonstationary which may ruin the predefined visual measurements and often lead to tracking failure in practice. Thus, a natural solution is to adapt the observation model to the nonstationary appearances. However, this idea is threatened by the risk of adaptation drift that originates in its ill-posed nature, unless good data-driven constraints are imposed. Different from most existing adaptation schemes, we enforce three novel constraints for the optimal adaptation: 1) negative data, 2) bottom-up pair-wise data constraints, and 3) adaptation dynamics. Substantializing the general adaptation problem as a subspace adaptation problem, this paper presents a closed-form solution as well as a practical iterative algorithm for subspace tracking. Extensive experiments have demonstrated that the proposed approach can largely alleviate adaptation drift and achieve better tracking results for a large variety of nonstationary scenes. Ming Yang 0007, Zhimin Fan 0002, Jialue Fan, Ying Wu 0001 |
IEEE Trans. Image Process. | 1 |
| 2008 | Vital sign estimation from passive thermal videoabstractConventional wired detection of vital signs limits the use of these important physiological parameters by many applications, such as airport health screening, elder care, and workplace preventive care. In this paper, we explore contact-free heart rate and respiratory rate detection through measuring infrared light modulation emitted near superficial blood vessels or a nasal area respectively. To deal with complications caused by subjects’ movements, facial expressions, and partial occlusions of the skin, we propose a novel algorithm based on contour segmentation and tracking, clustering of informative pixels, and dominant frequency component estimation. The proposed method achieves robust subject regions-of-interest alignment and motion compensation in infrared video with low SNR. It relaxes some strong assumptions used in previous work and substantially improves on previously reported performance. Preliminary experiments on heart rate estimation for 20 subjects and respiratory rate estimation for 8 subjects exhibit promising results. Ming Yang 0007, Qiong Liu 0003, Thea Turner, Ying Wu 0001 |
CVPR | 1 |
| 2008 | Granularity and elasticity adaptation in visual trackingabstractThe observation models in tracking algorithms are critical to both tracking performance and applicable scenarios but are often simplified to focus on fixed level of certain target properties such as appearances and structures. In this paper, we propose a unified tracking paradigm in which targets are represented by Markov random fields of interest regions and introduce a new way to adapt observation models by automatically tuning the feature granularity and model elasticity, i.e. the abstraction level of features and the model’s degree of flexibility to tolerate deformations. Specifically, we employ a multi-scale scheme to extract features from interest regions and adjust the parameters of the potential functions of the MRF model to maximize the likelihoods of tracking results. Experiments demonstrate the method can estimate translation, scaling and rotation and deal with deformation, partial occlusions, and camouflage objects within this unified framework. Ming Yang 0007, Ying Wu 0001 |
CVPR | 1 |
| 2008 | Image spam hunterabstractSpammers are constantly creating sophisticated new weapons in their arms race with anti-spam technology, the latest of which is image-based spam. The newest image-based spam uses simple image processing technologies to vary the content of individual messages, e.g. by changing foreground colors, backgrounds, font types, or even rotating and adding artifacts to the images. Thus, they pose great challenges to conventional spam filters. In this paper, we propose a system using a probabilistic boosting tree to determine whether an incoming image is a spam or not based on global image features, i.e. color and gradient orientation histograms. The system identifies spam without the need for OCR and is robust in the face of the kinds of variation found in current spam images. Evaluation results show the system correctly classifies 90% of spam images while mislabeling only 0.86% of non-spam images as spam. Yan Gao 0003, Ming Yang 0007, Xiaonan Zhao, Bryan Pardo, Ying Wu 0001, Thrasyvoulos N. Pappas, Alok N. Choudhary |
ICASSP | 2 |
| 2008 | A bi-subspace model for robust visual trackingabstractThe changes of the target’s visual appearance often lead to tracking failure in practice. Hence, trackers need to be adaptive to non-stationary appearances to achieve robust visual tracking. However, the risk of adaptation drift is common in most existing adaptation schemes. This paper describes a bi-subspace model that stipulates the interactions of two different visual cues. The visual appearance of the target is represented by two interactive subspaces, each of which corresponds to a particular cue. The adaption of the subspaces is through the interaction of the two cues, which leads to robust tracking performance. Extensive experiments show that the proposed approach can largely alleviate adaptation drift and obtain better tracking results. Jialue Fan, Ming Yang 0007, Ying Wu 0001 |
ICIP | 2 |
| 2007 | Detector EnsembleabstractComponent-based detection methods have demonstrated their promise by integrating a set of part-detectors to deal with large appearance variations of the target. However, an essential and critical issue, i.e., how to handle the imperfectness of part-detectors in the integration, is not well addressed in the literature. This paper proposes a detector ensemble model that consists of a set of substructure-detectors, each of which is composed of several part-detectors. Two important issues are studied both in theory and in practice, (1) finding an optimal detector ensemble, and (2) detecting targets based on an ensemble. Based on some theoretical analysis, a new model selection strategy is proposed to learn an optimal detector ensemble that has a minimum number of false positives and satisfies the design requirement on the capacity of tolerating missing parts. In addition, this paper also links ensemble-based detection to the inference in Markov random field, and shows that the target detection can be done by a max-product belief propagation algorithm. Shengyang Dai, Ming Yang 0007, Ying Wu 0001, Aggelos K. Katsaggelos |
CVPR | 2 |
| 2007 | Spatial selection for attentional visual trackingabstractLong-duration tracking of general targets is quite challenging for computer vision, because in practice target may undergo large uncertainties in its visual appearance and the unconstrained environments may be cluttered and distractive, although tracking has never been a challenge to the human visual system. Psychological and cognitive findings indicate that the human perception is attentional and selective, and both early attentional selection that may be innate and late attentional selection that may be learned are necessary for human visual tracking. This paper proposes a new visual tracking approach by reflecting some aspects of spatial selective attention, and presents a novel attentional visual tracking (AVT) algorithm. In AVT, the early selection process extracts a pool of attentional regions (ARs) that are defined as the salient image regions which have good localization properties, and the late selection process dynamically identifies a subset of discriminative attentional regions (D-ARs) through a discriminative learning on the historical data on the fly. The computationally demanding process of matching of the AR pool is done in an efficient and innovative way by using the idea in the locality-sensitive hashing (LSH) technique. The proposed AVT algorithm is general, robust and computationally efficient, as shown in extensive experiments on a large variety of real-world video. Ming Yang 0007, Junsong Yuan 0001, Ying Wu 0001 |
CVPR | 1 |
| 2007 | Discovery of Collocation Patterns: from Visual Words to Visual PhrasesabstractA visual word lexicon can be constructed by clustering primitive visual features, and a visual object can be described by a set of visual words. Such a "bag-of-words" representation has led to many significant results in various vision tasks including object recognition and categorization. However, in practice, the clustering of primitive visual features tends to result in synonymous visual words that over-represent visual patterns, as well as polysemous visual words that bring large uncertainties and ambiguities in the representation. This paper aims at generating a higher-level lexicon, i.e. visual phrase lexicon, where a visual phrase is a meaningful spatially co-occurrent pattern of visual words. This higher-level lexicon is much less ambiguous than the lower-level one. The contributions of this paper include: (1) a fast and principled solution to the discovery of significant spatial co-occurrent patterns using frequent itemset mining; (2) a pattern summarization method that deals with the compositional uncertainties in visual phrases; and (3) a top-down refinement scheme of the visual word lexicon by feeding back discovered phrases to tune the similarity measure through metric learning. Junsong Yuan 0001, Ying Wu 0001, Ming Yang 0007 |
CVPR | 3 |
| 2007 | False Positive Reduction in Lung GGO Nodule Detection with 3D Volume Shape DescriptorabstractLung nodule detection, especially ground glass opacity (GGO) detection, in helical computed tomography (CT) images is a challenging computer-aided detection (CAD) task due to the enormous variances in nodules' volumes, shapes, appearances, and the structures nearby. Most of the detection algorithms employ some efficient candidate generation (CG) algorithms to spot the suspicious volumes with high sensitivity at the cost of low specificity, e.g. tens even hundreds of false positives per volume. This paper proposes a learning based method to reduce the number of false positives given by CG based on a new general 3D volume shape descriptor. The 3D volume shape descriptor is constructed by concatenating spatial histograms of gradient orientations, which is robust to large variabilities in intensity levels, shapes, and appearances. The proposed method achieves promising performance on a difficult mixture lung nodule dataset with average 81% detection rate and 4.3 false positives per volume. Ming Yang 0007, Senthil Periaswamy, Ying Wu 0001 |
ICASSP (1) | 1 |
| 2007 | Game-Theoretic Multiple Target TrackingabstractVideo-based multiple target tracking (MTT) is a challenging task when similar targets are present in close vicinity. Because their visual observations are mixed and difficult to segment, their motions have to be estimated jointly. Most existing approaches perform this joint motion estimation in a centralized fashion and involve searching a rather high dimensional space, and thus leading to quite complicated joint trackers. This paper brings a new view to MTT from a game-theoretic perspective, bridging the joint motion estimation and the Nash equilibrium of a game. Instead of designing a centralized tracker, MTT is decentralized and a set of individual trackers is used, each of which tries to maximize its visual evidence for explaining its motion as well as generates interferences to others. Modelling this competition behavior, a special game is designed so that the difficult joint motion estimation is achieved at the Nash Equilibrium of this game where no individual tracker has incentives to change its motion estimate. This paper substantializes this novel idea in a solid case study where individual trackers are kernel-based trackers. An efficient best response updating procedure is designed to find the Nash equilibrium. The powerfulness of this game-theoretic MTT is shown by promising results on difficult real videos. Ming Yang 0007, Ting Yu 0003, Ying Wu 0001 |
ICCV | 1 |
| 2007 | Mining Auxiliary Objects for Tracking by Multibody GroupingabstractOn-line discovery of some auxiliary objects to verify the tracking results is a novel approach to achieving robust tracking by balancing the need for strong verification and computational efficiency. However, the applicability and effectiveness of this approach highly depend on how to reliably validate the motion correlation between the target and the auxiliary objects so as to estimate the motion model. In this paper, we extend the algorithm of mining auxiliary objects for tracking by incorporating multibody grouping to detect the motion correlation and estimate the motion model, which imposes more general motion correlation constraints. The proposed method discovers the auxiliary objects that exhibit strong affine motion correlation and estimates the closed-form affine models. The proposed tracking algorithm shows good performance in real-world test sequences. Ming Yang 0007, Ying Wu 0001, Shihong Lao |
ICIP (3) | 1 |
| 2007 | From frequent itemsets to semantically meaningful visual patternsabstractData mining techniques that are successful in transaction and text data may not be simply applied to image data that contain high-dimensional features and have spatial structures. It is not a trivial task to discover meaningful visual patterns in image databases, because the content variations and spatial dependency in the visual data greatly challenge most existing methods. This paper presents a novel approach to coping with these difficulties for mining meaningful visual patterns. Specifically, the novelty of this work lies in the following new contributions: (1) a principled solution to the discovery of meaningful itemsets based on frequent itemset mining; (2) a self-supervised clustering scheme of the high-dimensional visual features by feeding back discovered patterns to tune the similarity measure through metric learning; and (3) a pattern summarization method that deals with the measurement noises brought by the image data. The experimental results in the real images show that our method can discover semantically meaningful patterns efficiently and effectively. Junsong Yuan 0001, Ying Wu 0001, Ming Yang 0007 |
KDD | 3 |
| 2007 | Multiple Collaborative Kernel TrackingabstractThose motion parameters that cannot be recovered from image measurements are unobservable in the visual dynamic system. This paper studies this important issue of singularity in the context of kernel-based tracking and presents a novel approach that is based on a motion field representation which employs redundant but sparsely correlated local motion parameters instead of compact but uncorrelated global ones. This approach makes it easy to design fully observable kernel-based motion estimators. This paper shows that these high-dimensional motion fields can be estimated efficiently by the collaboration among a set of simpler local kernel-based motion estimators, which makes the new approach very practical. Zhimin Fan 0002, Ming Yang 0007, Ying Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Efficient Optimal Kernel Placement for Reliable Visual TrackingabstractThis paper describes a novel approach to optimal kernel placement in kernel-based tracking. If kernels are placed at arbitrary places, kernel-based methods are likely to be trapped in ill-conditioned locations, which prevents the reliable recovery of the motion parameters and jeopardizes the tracking performance. The theoretical analysis presented in this paper indicates that the optimal kernel placement can be evaluated based on a closed-form criterion, and achieved efficiently by a novel gradient-based algorithm. Based on that, new methods for temporal-stable multiple kernel placement and scale-invariant kernel placement are proposed. These new theoretical results and new algorithms greatly advance the study of kernel-based tracking in both theory and practice. Extensive real-time experimental results demonstrate the improved tracking reliability. Zhimin Fan 0002, Ming Yang 0007, Ying Wu 0001, Gang Hua 0001, Ting Yu 0003 |
CVPR (1) | 2 |
| 2006 | Intelligent Collaborative Tracking by Mining Auxiliary ObjectsabstractMany tracking methods face a fundamental dilemma in practice: tracking has to be computationally efficient but verifying if or not the tracker is following the true target tends to be demanding, especially when the background is cluttered and/or when occlusion occurs. Due to the lack of a good solution to this problem, many existing methods tend to be either computationally intensive with the use of sophisticated image observation models, or vulnerable to the false alarms. This greatly threatens long-duration robust tracking. This paper presents a novel solution to this dilemma by integrating into the tracking process a set of auxiliary objects that are automatically discovered in the video on the fly by data mining. Auxiliary objects have three properties at least in a short time interval: (1) persistent co-occurrence with the target; (2) consistent motion correlation with the target; and (3) easy to track. The collaborative tracking of these auxiliary objects leads to an efficient computation as well as a strong verification. Our extensive experiments have exhibited exciting performance in very challenging real-world testing cases. Ming Yang 0007, Ying Wu 0001, Shihong Lao |
CVPR (1) | 1 |
| 2006 | Tracking Motion-Blurred Targets in VideoabstractMany emerging applications require tracking targets in video. Most existing visual tracking methods do not work well when the target is motion-blurred (especially due to fast motion), because the imperfectness of the target's appearances invalidates the image matching model (or the measurement model) in tracking. This paper presents a novel method to track motion-blurred targets by taking advantage of the blurs without performing image restoration. Unlike the global blur induced by camera motion, this paper is concerned with the local blurs that are due to target's motion. This is a challenging task because the blurs need to be identified blindly. The proposed method addresses this difficulty by integrating signal processing and statistical learning techniques. The estimated blurs are used to reduce the search range by providing strong motion predictions and to localize the best match accurately by modifying the measurement models. Shengyang Dai, Ming Yang 0007, Ying Wu 0001, Aggelos K. Katsaggelos |
ICIP | 2 |
| 2006 | Face detection for automatic exposure control in handheld cameraabstractFace detection is a widely studied topic in computer vision, and advances in algorithms, low cost processing, and CMOS imagers make it practical for embedded consumer applications. As with graphics, the best cost-performance ratio is achieved with dedicated hardware. The challenges of face detection in embedded environments include bandwidth constraints set by low cost memory and a need to find parallelism. Consumer applications need reliability, calling for a hard real-time approach to guarantee that deadlines are met. We present a face detection system for automatic exposure control in a handheld digital camera or camera phone. Contributions include a complexity control scheme to meet hard real-time deadlines, a hardware pipeline design for Haar-like feature calculation, and a system design exploiting several levels of parallelism. The proposed architecture is verified by synthesis to Altera’s low cost Cyclone II FPGA. Simulation results show the algorithm can achieve about 80% detection rate for group portrait pictures. Ming Yang 0007, Ying Wu 0001, Jim E. Crenshaw, Bruce Augustine, Russell Mareachen |
ICVS | 1 |
| 2005 | Multiple Collaborative Kernel TrackingabstractThis paper presents a novel multiple collaborative kernel approach to visual tracking. This approach treats kernel-based tracking in a more general setting, i.e., a relaxation and constraints formulation, in which a complex motion is represented by a set of inter-correlated simpler motions. With this formulation, we present a rigorous analysis on a critical issue of kernel observability and obtain a criterion, based on which we propose a new method using collaborative kernels that has the theoretical guarantee of enhanced observability. This new method has been shown to be computationally efficient in both theory and practice, which can be readily applied to complex motions such as articulated motions. Zhimin Fan 0002, Ying Wu 0001, Ming Yang 0007 |
CVPR (2) | 3 |
| 2005 | Tracking Non-Stationary Appearances and Dynamic Feature SelectionabstractSince the appearance changes of the target jeopardize visual measurements and often lead to tracking failure in practice, trackers need to be adaptive to non-stationary appearances or to dynamically select features to track. However, this idea is threatened by the risk of adaptation drift that roots in its ill-posed nature, unless good constraints are imposed. Different from most existing adaptation schemes, we enforce three novel constraints for the optimal adaptation: (1) negative data, (2) bottom-up pair-wise data constraints, and (3) adaptation dynamics. Substantializing the general adaptation problem as a subspace adaptation problem, this paper gives a closed-form solution as well as a practical iterative algorithm. Extensive experiments have shown that the proposed approach can largely alleviate adaptation drift and achieve better tracking results. Ming Yang 0007, Ying Wu 0001 |
CVPR (2) | 1 |
| 2004 | Fast macroblock mode selection based on motion content classification in H.264/AVCabstractIn H.264/AVC, new coding techniques introduce many optional macroblock encoding modes. Exhaustive search over all possible modes can achieve optimal coding efficiency but is highly computationally expensive. In this paper, a fast macroblock mode selection (FMMS) algorithm based on the motion content classification is proposed. Each macroblock is categorized into complex motion or simple motion contents by a fuzzy classifier, then different mode search orders with distinct early termination schemes are employed according to the classification. The proposed method can be readily incorporated with fast motion estimation algorithms and simulation results show that it can further save 40%-70% of the block distortion calculations on the basis of conventional fast motion estimation algorithms, while maintaining similar rate distortion performance. Ming Yang 0007 |
ICIP | 1 |