VLDB 2026 Research / reviewers in the wild / expert
Guo-Jun Qi
dblp:41/943 · also Guojun Qi
· DBLP profile ↗
183ranked-venue papers
33as first author
64since 2021 · last 2026
0000-0003-3508-1851ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 117 · 15 first-author · 43 since 2021Artificial intelligence and machine learning · 92 · 19 first-author · 40 since 2021Databases, data management, data science and information retrieval · 24 · 11 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-authorComputer networks · 8 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | One-Step Diffusion and Flow Distillation Through Implicit Generator MatchingabstractDespite strong performances on many generative tasks, diffusion and flow matching models require a large number of sampling steps to generate high-quality images. This has motivated the community to develop effective methods to distill pre-trained models into more efficient models. In this paper, we present Implicit Generator Matching (IGM), a systematic approach to distill both pre-trained diffusion/flow matching models into one-step generator models, while maintaining almost the same sample generation ability as the original model, as well as being data-free with no need for training images. The key challenge is that the traditional diffusion/flow-matching loss is intractable to distill a teacher diffusion/flow model with an explicitly defined field into a student generator, whose field is defined implicitly. The main breakthrough, our Implicit Gradient Theorem, provides an exact and efficient gradient to directly optimize the student by aligning this implicit field with the teacher's. IGM shows strong empirical performance for one-step generators, setting new standards. On CIFAR10, our diffusion-based SIM achieves an FID score of 2.06, while flow-based FGM sets a flow-model record with a 3.08 FID. Scaling to text-to-image models, SIM distillation of PixArt-$\alpha$α yields a leading 6.42 aesthetic score, surpassing SDXL-TURBO (5.33), and FGM distillation of SD3 achieves a competitive 0.65 GenEval score against multi-step accelerators like Hyper-SD3 (0.63). Zemin Huang, Weijian Luo, Zhengyang Geng, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Self-Guidance: Boosting Flow and Diffusion Generation on Their OwnabstractProper guidance strategies are essential to achieve high-quality generation results without retraining diffusion and flow-based text-to-image models. Existing guidance either requires specific training or strong inductive biases of diffusion model networks, which potentially limits their ability and application scope. Motivated by the observation that artifact outliers can be detected by a significant decline in the density from a noisier to a cleaner noise level, we propose Self-Guidance (SG), which can significantly improve the quality of the generated image by suppressing the generation of low-quality samples. The biggest difference from existing guidance is that SG only relies on the sampling score function of the original diffusion or flow model at different noise levels, with no need for any tricky and expensive guidance-specific training. This makes SG highly flexible to be used in a plug-and-play manner by any diffusion or flow models. We also introduce an efficient variant of SG, named SG-prev, which reuses the output from the immediately previous diffusion step to avoid additional forward passes of the diffusion network. We conduct extensive experiments on text-to-image and text-to-video generation with different architectures, including UNet and transformer models. With open-sourced diffusion models such as Stable Diffusion 3.5 and FLUX, SG exceeds existing algorithms on multiple metrics, including both FID and Human Preference Score. SG-prev also achieves strong results over both the baseline and the SG, with 50 percent more efficiency. Moreover, we find that SG and SG-prev both have a surprisingly positive effect on the generation of physiologically correct human body structures such as hands, faces, and arms, showing their ability to eliminate human body artifacts with minimal efforts. Weijian Luo, Zhiyang Chen 0002, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | SCASeg: Strip Cross-Attention for Efficient Semantic SegmentationabstractThe Vision Transformer (ViT) has achieved notable success in computer vision, with its variants widely validated across various downstream tasks, including semantic segmentation. However, as general-purpose visual encoders, ViT backbones often do not fully address the specific requirements of task decoders, highlighting opportunities for designing decoders optimized for efficient semantic segmentation. This paper proposes Strip Cross-Attention (SCASeg), an innovative decoder head specifically designed for semantic segmentation. Instead of relying on the conventional skip connections, we utilize lateral connections between encoder and decoder stages, leveraging encoder features as Queries in cross-attention modules. Additionally, we introduce a Cross-Layer Block (CLB) that integrates hierarchical feature maps from various encoder and decoder stages to form a unified representation for Keys and Values. The CLB also incorporates the local perceptual strengths of convolution, enabling SCASeg to capture both global and local context dependencies across multiple layers, thus enhancing feature interaction at different scales and improving overall efficiency. To further optimize computational efficiency, SCASeg compresses the channels of queries and keys into one dimension, creating strip-like patterns that reduce memory usage and increase inference speed compared to traditional vanilla cross-attention. Experiments show that SCASeg's adaptable decoder delivers competitive performance across various setups, outperforming leading segmentation architectures on benchmark datasets, including ADE20K, Cityscapes, COCO-Stuff 164k, and Pascal VOC2012, even under diverse computational constraints. Guoan Xu, Jiaming Chen 0001, Wenfeng Huang, Wenjing Jia, Guangwei Gao, Guo-Jun Qi |
IEEE Trans. Image Process. | 6 |
| 2025 | Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image GenerationabstractDiffusion and flow matching models have achieved remarkable success in text-to-image generation. However, these models typically rely on the predetermined denoising schedules for all prompts. The multi-step reverse diffusion process can be regarded as a kind of chain-of-thought for generating high-quality images step by step. Therefore, diffusion models should reason for each instance to determine the optimal noise schedule adaptively, achieving high generation quality with sampling efficiency. In this paper, we introduce the Time Prediction Diffusion Model (TPDM) for this. TPDM employs a plug-and-play Time Prediction Module (TPM) that predicts the next noise level based on current latent features at each denoising step. We train the TPM using reinforcement learning to maximize a reward that encourages high final image quality while penalizing excessive denoising steps. With such an adaptive scheduler, TPDM not only generates high-quality images that are aligned closely with human preferences but also adjusts diffusion time and the number of denoising steps on the fly, enhancing both performance and efficiency. With Stable Diffusion 3 Medium architecture, TPDM achieves an aesthetic score of 5.44 and a human preference score (HPS) of 29.59, while using around 50% fewer denoising steps to achieve better performance. Zilyu Ye, Zhiyang Chen 0002, Zemin Huang, Weijian Luo, Guo-Jun Qi |
CVPR | 6 |
| 2025 | DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion TransformerabstractSpeech-driven 3D facial animation is important for many multimedia applications. Recent work has shown promise in using either Diffusion models or Transformer architectures for this task. However, their mere aggregation does not lead to improved performance. We suspect this is due to a shortage of paired audio-4D data, which is crucial for the Transformer to effectively perform as a denoiser within the Diffusion framework. To tackle this issue, we present DiffSpeaker, a Transformer-based network equipped with novel biased conditional attention modules. These modules serve as substitutes for the traditional self/cross-attention in standard Transformers, incorporating thoughtfully designed biases that steer the attention mechanisms to concentrate on both the relevant task-specific and diffusion-related conditions. We also explore the trade-off between accurate lip synchronization and non-verbal facial expressions within the Diffusion paradigm. Experiments show our model achieves state-of-the-art performance on existing benchmarks, and fast inference speed owing to its ability to generate facial motions in parallel. Our code is avalable at https://github.com/theEricMa/DiffSpeaker. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Chen Qian 0006, Shukai Chen, Guo-Jun Qi, Zhaoxiang Zhang 0001, Zhen Lei 0001 |
IJCB | 6 |
| 2025 | Learning Dual-Domain Multi-Scale Representations for Single Image DerainingabstractExisting image deraining methods typically rely on single-input, single-output, and single-scale architectures, which overlook the joint multi-scale information between external and internal features. Furthermore, single-domain representations are often too restrictive, limiting their ability to handle the complexities of real-world rain scenarios. To address these challenges, we propose a novel Dual-Domain Multi-Scale Representation Network (DMSR). The key idea is to exploit joint multi-scale representations from both external and internal domains in parallel while leveraging the strengths of both spatial and frequency domains to capture more comprehensive properties. Specifically, our method consists of two main components: the Multi-Scale Progressive Spatial Refinement Module (MPSRM) and the Frequency Domain Scale Mixer (FDSM). The MPSRM enables the interaction and coupling of multi-scale expert information within the internal domain using a hierarchical modulation and fusion strategy. The FDSM extracts multi-scale local information in the spatial domain, while also modeling global dependencies in the frequency domain. Extensive experiments show that our model achieves state-of-the-art performance across six benchmark datasets. Shun Zou, Mingya Zhang, Shipeng Luo, Guangwei Gao, Guo-Jun Qi |
ICME | 6 |
| 2025 | DynaMind: Reasoning over Abstract Video Dynamics for Embodied Decision-MakingabstractIntegrating natural language instructions and visual perception with decision-making is a critical challenge for embodied agents. Existing methods often struggle to balance the conciseness of language commands with the richness of video content. To bridge the gap between modalities, we propose extracting key spatiotemporal patterns from video that capture visual saliency and temporal evolution, referred to as dynamic representation. Building on this, we introduce DynaMind, a framework that enhances decision-making through dynamic reasoning. Specifically, we design an adaptive FrameScorer to evaluate video frames based on semantic consistency and visual saliency, assigning each frame an importance score. These scores are used to filter redundant video content and synthesize compact dynamic representations. Leveraging these representations, we predict critical future dynamics and apply a dynamic-guided policy to generate coherent and context-aware actions. Extensive results demonstrate that DynaMind significantly outperforms the baselines across several simulation benchmarks and real-world scenarios. Ziru Wang, Mengmeng Wang 0005, Jade Dai, Teli Ma, Guo-Jun Qi, Yong Liu 0007, Guang Dai, Jingdong Wang 0001 |
ICML | 5 |
| 2025 | Cross Paradigm Representation and Alignment Transformer for Image DerainingabstractTransformer-based networks have achieved strong performance in low-level vision tasks like image deraining by utilizing spatial or channel-wise self-attention. However, irregular rain patterns and complex geometric overlaps challenge single-paradigm architectures, necessitating a unified framework to integrate complementary global-local and spatial-channel representations. To address this, we propose a novel Cross Paradigm Representation and Alignment Transformer (CPRAformer). Its core idea is the hierarchical representation and alignment, leveraging the strengths of both paradigms (spatial-channel and global-local) to aid image reconstruction. It bridges the gap within and between paradigms, aligning and coordinating them to enable deep interaction and fusion of features. Specifically, we use two types of self-attention in the Transformer blocks: sparse prompt channel self-attention (SPC-SA) and spatial pixel refinement self-attention (SPR-SA). SPC-SA enhances global channel dependencies through dynamic sparsity, while SPR-SA focuses on spatial rain distribution and fine-grained texture recovery. To address the feature misalignment and knowledge differences between them, we introduce the Adaptive Alignment Frequency Module (AAFM), which aligns and interacts with features in a two-stage progressive manner, enabling adaptive guidance and complementarity. This reduces the information gap within and between paradigms. Through this unified cross-paradigm dynamic interaction framework, we achieve the extraction of the most valuable interactive fusion information from the two paradigms. Extensive experiments demonstrate that our model achieves state-of-the-art performance on eight benchmark datasets and further validates CPRAformer's robustness in other image restoration tasks and downstream applications. Shun Zou, Juncheng Li 0003, Guangwei Gao, Guo-Jun Qi |
ACM Multimedia | 5 |
| 2025 | Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language ModelsabstractWe introduce the Diffusion Chain of Lateral Thought (DCoLT), a reasoning framework for diffusion language models. DCoLT treats each intermediate step in the reverse diffusion process as a latent "thinking" action and optimizes the entire reasoning trajectory to maximize the reward on the correctness of the final answer with outcome-based Reinforcement Learning (RL). Unlike traditional Chain-of-Thought (CoT) methods that follow a causal, linear thinking process, DCoLT allows bidirectional, non-linear reasoning with no strict rule on grammatical correctness amid its intermediate steps of thought. We implement DCoLT on two representative Diffusion Language Models (DLMs). First, we choose SEDD as a representative continuous-time discrete diffusion model, where its concrete score derives a probabilistic policy to maximize the RL reward over the entire sequence of intermediate diffusion steps. We further consider the discrete-time masked diffusion language model -- LLaDA, and find that the order to predict and unmask tokens plays an essential role to optimize its RL action resulting from the ranking-based Unmasking Policy Module (UPM) defined by the Plackett-Luce model. Experiments on both math and code generation tasks show that using only public data and 16 H800 GPUs, DCoLT-reinforced DLMs outperform other DLMs trained by SFT or RL or even both. Notably, DCoLT-reinforced LLaDA boosts its reasoning accuracy by +9.8%, +5.7%, +11.4%, +19.5% on GSM8K, MATH, MBPP, and HumanEval. Zemin Huang, Zhiyang Chen 0002, Guo-Jun Qi |
NeurIPS | 5 |
| 2025 | Enhancing semantic audio-visual representation learning with supervised multi-scale attention
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Guo-Jun Qi, Haiyuan Wu, Hirotaka Hachiya |
Pattern Anal. Appl. | 4 |
| 2025 | S2AFormer: Strip Self-Attention for Efficient Vision TransformerabstractThe Vision Transformer (ViT) has achieved remarkable success in computer vision due to its powerful token mixer, which effectively captures global dependencies among all tokens. However, the quadratic complexity of standard self-attention with respect to the number of tokens severely hampers its computational efficiency in practical deployment. Although recent hybrid approaches have sought to combine the strengths of convolutions and self-attention to improve the performance-efficiency trade-off, the costly pairwise token interactions and heavy matrix operations in conventional self-attention remain a critical bottleneck. To overcome this limitation, we introduce S2AFormer, an efficient Vision Transformer architecture built around a novel Strip Self-Attention (SSA) mechanism. Our design incorporates lightweight yet effective Hybrid Perception Blocks (HPBs) that seamlessly fuse the local inductive biases of CNNs with the global modeling capability of Transformer-style attention. The core innovation of SSA lies in simultaneously reducing the spatial resolution of the key ( $K$ ) and value ( $V$ ) tensors while compressing the channel dimension of the query ( $Q$ ) and key ( $K$ ) tensors. This joint spatial-and-channel compression dramatically lowers computational cost without sacrificing representational power, achieving an excellent balance between accuracy and efficiency. We extensively evaluate S2AFormer on a wide range of vision tasks, including image classification (ImageNet-1K), semantic segmentation (ADE20K), and object detection/instance segmentation (COCO). Experimental results consistently show that S2AFormer delivers substantial accuracy improvements together with superior inference speed and throughput across both GPU and non-GPU platforms, establishing it as a highly competitive solution in the landscape of efficient Vision Transformers. Guoan Xu, Wenfeng Huang, Wenjing Jia, Jiamao Li, Guangwei Gao, Guo-Jun Qi |
IEEE Trans. Image Process. | 6 |
| 2025 | Point Evolution Hierarchy Network for Weak Single-Point Human ParsingabstractSparsely single-point human parsing aims at segmenting the human body into fine-grained categories via weak point-level labels (e.g., point-level, scribble-level, or image-level, etc). The point-level label, especially single-point supervision, can simultaneously preserve spatial positions as well as take light annotation time, which is particularly advantageous in alleviating the human labeling burden. However, how to obtain satisfactory parsing performance under limited sparse point annotations is challenging, which requires further investigation. In this paper, we propose a novel end-to-end Point Evolution Hierarchy human parsing Network (PEHNet) for fine-grained human parsing task that just leverages single-point supervision. Motivated by the concept of a divide-and-conquer strategy, we partition all pixels into three distinct groups, i.e., single-point labels, pseudo-region labels, and unlabeled pixels, then optimize each group with suitable mechanisms. To expand the coverage of single-point labels, we introduce a point dissemination module that generates high-quality pseudo-region labels. Furthermore, the point-level spatial position information inherently preserves the structural characteristics of the human body. Inspired by this hierarchical property, we devise a point-level human hierarchy-wise constraint that guides the prediction probabilities to align with the inherent hierarchy of the human body. Experimental results demonstrate that the proposed PEHNet outperforms state-of-the-art parsing methods on two popular human parsing benchmark datasets (LIP and ATR) and one semantic segmentation dataset (Pascal VOC 2012). Sanyi Zhang, Xiaochun Cao, Long Ye, Zhanjie Song, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 5 |
| 2025 | Jointly Understand Your Command and Intention: Reciprocal Co-Evolution Between Scene-Aware 3D Human Motion Synthesis and AnalysisabstractAs two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, we integrate these two paired processes into a Co-Evolving Synthesis-Analysis (CESA) pipeline and mutually benefit their learning. Specifically, scene aware text-to-human synthesis generates diverse indoor motion samples from the same textual description to enrich human scene interaction intra-class diversity, thus significantly benefiting training a robust human motion analysis system. Reciprocally, human motion analysis would enforce semantic scrutiny on each synthesized motion sample to ensure its semantic consistency with the given textual description, thus improving realistic motion synthesis. Considering that real-world indoor human motions are goal-oriented and path-guided, we propose a cascaded generation strategy that factorizes text-driven scene-specific human motion generation into three stages: goal inferring, path planning, and pose synthesizing. Coupling CESA with this powerful cascaded motion synthesis model, we jointly improve realistic human motion synthesis and robust human motion analysis in 3D scenes. Xuehao Gao, Yang Yang 0066, Shaoyi Du, Guo-Jun Qi, Junwei Han 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Efficient Image Super-Resolution With Feature Interaction Weighted Hybrid NetworkabstractLightweight image super-resolution aims to reconstruct high-resolution images from low-resolution images using low computational costs. However, existing methods result in the loss of middle-layer features due to activation functions. To minimize the impact of intermediate feature loss on reconstruction quality, we propose a Feature Interaction Weighted Hybrid Network (FIWHN), which comprises a series of Wide-residual Distillation Interaction Block (WDIB) as the backbone. Every third WDIB forms a Feature Shuffle Weighted Group (FSWG) by applying mutual information shuffle and fusion. Moreover, to mitigate the negative effects of intermediate feature loss, we introduce Wide Residual Weighting units within WDIB. These units effectively fuse features of varying levels of detail through a Wide-residual Distillation Connection (WRDC) and a Self-Calibrating Fusion (SCF). To compensate for global feature deficiencies, we incorporate a Transformer and explore a novel architecture to combine CNN and Transformer. We show that our FIWHN achieves a favorable balance between performance and efficiency through extensive experiments on low-level and high-level tasks. Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jian Yang 0003, Guo-Jun Qi, Chia-Wen Lin |
IEEE Trans. Multim. | 6 |
| 2025 | Knowledge-Guided Semantic Transfer Network for Few-Shot Image RecognitionabstractDeep learning-based models have been shown to outperform human beings in many computer vision tasks with massive available labeled training data in learning. However, humans have an amazing ability to easily recognize images of novel categories by browsing only a few examples of these categories. In this case, few-shot learning comes into being to make machines learn from extremely limited labeled examples. One possible reason why human beings can well learn novel concepts quickly and efficiently is that they have sufficient visual and semantic prior knowledge. Toward this end, this work proposes a novel knowledge-guided semantic transfer network (KSTNet) for few-shot image recognition from a supplementary perspective by introducing auxiliary prior knowledge. The proposed network jointly incorporates vision inferring, knowledge transferring, and classifier learning into one unified framework for optimal compatibility. A category-guided visual learning module is developed in which a visual classifier is learned based on the feature extractor along with the cosine similarity and contrastive loss optimization. To fully explore prior knowledge of category correlations, a knowledge transfer network is then developed to propagate knowledge information among all categories to learn the semantic-visual mapping, thus inferring a knowledge-based classifier for novel categories from base categories. Finally, we design an adaptive fusion scheme to infer the desired classifiers by effectively integrating the above knowledge and visual information. Extensive experiments are conducted on two widely used Mini-ImageNet and Tiered-ImageNet benchmarks to validate the effectiveness of KSTNet. Compared with the state of the art, the results show that the proposed method achieves favorable performance with minimal bells and whistles, especially in the case of one-shot learning. Zechao Li, Hao Tang 0007, Zhimao Peng, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | BARET: Balanced Attention Based Real Image Editing Driven by Target-Text InversionabstractImage editing approaches with diffusion models have been rapidly developed, yet their applicability are subject to requirements such as specific editing types (e.g., foreground or background object editing, style transfer), multiple conditions (e.g., mask, sketch, caption), and time consuming fine-tuning of diffusion models. For alleviating these limitations and realizing efficient real image editing, we propose a novel editing technique that only requires an input image and target text for various editing types including non-rigid edits without fine-tuning diffusion model. Our method contains three novelties: (I) Target-text Inversion Schedule (TTIS) is designed to fine-tune the input target text embedding to achieve fast image reconstruction without image caption and acceleration of convergence. (II) Progressive Transition Scheme applies progressive linear interpolation between target text embedding and its fine-tuned version to generate transition embedding for maintaining non-rigid editing capability. (III) Balanced Attention Module (BAM) balances the tradeoff between textual description and image semantics. By the means of combining self-attention map from reconstruction process and cross-attention map from transition process, the guidance of target text embeddings in diffusion process is optimized. In order to demonstrate editing capability, effectiveness and efficiency of the proposed BARET, we have conducted extensive qualitative and quantitative experiments. Moreover, results derived from user study and ablation study further prove the superiority over other methods. Yuming Qiao, Fanyi Wang, Jingwen Su, Yunjie Yu, Guo-Jun Qi |
AAAI | 7 |
| 2024 | OmniMotionGPT: Animal Motion Generation with Limited DataabstractOur paper aims to generate diverse and realistic ani-mal motion sequences from textual descriptions, without a large-scale animal text-motion dataset. While the task of text-driven human motion synthesis is already extensively studied and benchmarked, it remains challenging to transfer this success to other skeleton structures with limited data. In this work, we design a model architecture that imitates Generative Pretraining Transformer (GPT), utilizing prior knowledge learned from human data to the animal domain. We jointly train motion autoencoders for both animal and human motions and at the same time optimize through the similarity scores among human motion encoding, animal motion encoding, and text CLIP embedding. Presenting the first solution to this problem, we are able to generate animal motions with high diversity and fidelity, quantitatively and qualitatively outperforming the results of training human motion generation baselines on animal data. Additionally, we introduce AnimalML3D, the first text-animal motion dataset with 1240 animation sequences spanning 36 different ani-mal identities. We hope this dataset would mediate the data scarcity problem in text-driven animal motion generation, providing a new playground for the research community. Zhangsihao Yang, Mingyuan Zhou, Mengyi Shan, Bingbing Wen, Ziwei Xuan, Mitch Hill, Guo-Jun Qi, Yalin Wang 0001 |
CVPR | 8 |
| 2024 | UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided TexturesabstractRecent advances in 3D avatar generation have gained significant attention. These breakthroughs aim to produce more realistic animatable avatars, narrowing the gap between virtual and real-world experiences. Most of existing works employ Score Distillation Sampling (SDS) loss, combined with a differentiable renderer and text condition, to guide a diffusion model in generating 3D avatars. How-ever, SDS often generates over-smoothed results with few facial details, thereby lacking the diversity compared with ancestral sampling. On the other hand, other works gen-erate 3D avatar from a single image, where the challenges of unwanted lighting effects, perspective views, and inferior image quality make them difficult to reliably reconstruct the 3D face meshes with the aligned complete textures. In this paper, we propose a novel 3D avatar generation approach termed UltrAvatar with enhanced fidelity of geometry, and superior quality of physically based rendering (PBR)textures without unwanted lighting. To this end, the proposed approach presents a diffuse color extraction model and an authenticity guided texture diffusion model. The former removes the unwanted lighting effects to reveal true diffuse colors, so that the generated avatars can be rendered under various lighting conditions. The latter follows two gradient-based guidances for generating PBR textures to render diverse face-identity features and details better aligning with 3D mesh geometry. We demonstrate the effectiveness and robustness of the proposed method, outperforming the state-of-the-art methods by a large margin in the experiments. Mingyuan Zhou, Rakib Hyder, Ziwei Xuan, Guo-Jun Qi |
CVPR | 4 |
| 2024 | Towards Open Domain Text-Driven Synthesis of Multi-person Motions
Mengyi Shan, Lu Dong 0004, Yutao Han, Yuan Yao 0001, Ifeoma Nwogu, Guo-Jun Qi, Mitch Hill |
ECCV (65) | 7 |
| 2024 | Lightweight High-Resolution Subject Matting in the Real WorldabstractExisting saliency object detection (SOD) methods struggle to satisfy fast inference and accurate results simultaneously in high resolution scenes. They are limited by the quality of public datasets and efficient network modules for high-resolution images. To alleviate these issues, we propose to construct a saliency object matting dataset HRSOM and a lightweight network PSUNet. Considering efficient inference of mobile depolyment framework, we design a symmetric pixel shuffle module and a lightweight module TRSU. Compared to 13 SOD methods, the proposed PSUNet has the best objective performance on the high-resolution benchmark dataset. Evaluation results of objective assessment are superior compared to U2Net that has 10 times of parameter amount of our network. On Snapdragon 8 Gen 2 Mobile Platform, inference a single 640 × 640 image only takes 113ms. And on the subjective assessment, evaluation results are better than the industry benchmark IOS16 (Lift subject from background). Fanyi Wang, Jingwen Su, Guo-Jun Qi |
ICASSP | 5 |
| 2024 | Zero-shot High-fidelity and Pose-controllable Character Animation
Bingwen Zhu, Fanyi Wang, Jingwen Su, Jinxiu Liu, Zuxuan Wu, Guo-Jun Qi, Yu-Gang Jiang 0001 |
IJCAI | 9 |
| 2024 | Equilibrated Diffusion: Frequency-aware Textual Embedding for Equilibrated Image CustomizationabstractImage customization involves learning the subject from provided concept images and generating it within textual contexts, typically yielding alterations of attributes such as style or background. Prevailing methods primarily rely on fine-tuning technique, wherein a unified latent embedding is employed to characterize various concept attributes. However, the attribute entanglement renders customized result challenging to mitigate the influence of subject-irrelevant attributes (e.g., style and background). To overcome these issues, we propose Equilibrated Diffusion, an innovative method that achieves equilibrated image customization by decoupling entangled concept attributes from a frequency-aware perspective, thus harmonizing textual and visual consistency. Unlike conventional approaches that employ a shared latent embedding and tuning process to learn concept, our Equilibrated Diffusion draws inspiration from the correlation between high- and low-frequency components with image style and content, decomposing concept accordingly in the frequency domain. Through independently optimizing concept embeddings in the frequency domain, the denoising model not only enriches its comprehension of style attribute irrelevant to subject identity but also inherently augments its aptitude for accommodating novel stylized descriptions. Furthermore, by combining different frequency embeddings, our model retains the spatially original customization capability. We further design a diffusion process guided by subject masks to alleviate the influence of background attribute, thereby strengthening text alignment. To ensure subject-related information consistency, Residual Reference Attention (RRA) is incorporated into the denoising model of spatial attention computation, effectively preserving structural details. Experimental results demonstrate that Equilibrated Diffusion surpasses other competitors with better subject consistency while closely adhering to text descriptions, thus validating the superiority of our approach. The code is available at https://github.com/maple-research-lab/EqDiff. Xueji Fang, Guo-Jun Qi |
ACM Multimedia | 3 |
| 2024 | One-Step Diffusion Distillation through Score Implicit MatchingabstractDespite their strong performances on many generative tasks, diffusion models require a large number of sampling steps in order to generate realistic samples. This has motivated the community to develop effective methods to distill pre-trained diffusion models into more efficient models, but these methods still typically require few-step inference or perform substantially worse than the underlying model. In this paper, we present Score Implicit Matching (SIM) a new approach to distilling pre-trained diffusion models into single-step generator models, while maintaining almost the same sample generation ability as the original model as well as being data-free with no need of training samples for distillation. The method rests upon the fact that, although the traditional score-based loss is intractable to minimize for generator models, under certain conditions we \emph{can} efficiently compute the \emph{gradients} for a wide class of score-based divergences between a diffusion model and a generator. SIM shows strong empirical performances for one-step generators: on the CIFAR10 dataset, it achieves an FID of 2.06 for unconditional generation and 1.96 for class-conditional generation. Moreover, by applying SIM to a leading transformer-based diffusion model, we distill a single-step generator for text-to-image (T2I) generation that attains an aesthetic score of 6.42 with no performance decline over the original multi-step counterpart, clearly outperforming the other one-step generators including SDXL-TURBO of 5.33, SDXL-LIGHTNING of 5.34 and HYPER-SDXL of 5.85. We will release this industry-ready one-step transformer-based T2I generator along with this paper. Weijian Luo, Zemin Huang, Zhengyang Geng, J. Zico Kolter, Guo-Jun Qi |
NeurIPS | 5 |
| 2024 | Multi-Constraint Transferable Generative Adversarial Networks for Cross-Modal Brain Image Synthesis
Yawen Huang, Hao Zheng 0008, Yuexiang Li, Feng Zheng 0001, Xiantong Zhen, Guo-Jun Qi, Ling Shao 0001, Yefeng Zheng 0001 |
Int. J. Comput. Vis. | 6 |
| 2024 | A Bayesian Federated Learning Framework With Online Laplace ApproximationabstractFederated learning (FL) allows multiple clients to collaboratively learn a globally shared model through cycles of model aggregation and local model training, without the need to share data. Most existing FL methods train local models separately on different clients, and then simply average their parameters to obtain a centralized model on the server side. However, these approaches generally suffer from large aggregation errors and severe local forgetting, which are particularly bad in heterogeneous data settings. To tackle these issues, in this paper, we propose a novel FL framework that uses online Laplace approximation to approximate posteriors on both the client and server side. On the server side, a multivariate Gaussian product mechanism is employed to construct and maximize a global posterior, largely reducing the aggregation errors induced by large discrepancies between local models. On the client side, a prior loss that uses the global posterior probabilistic parameters delivered from the server is designed to guide the local training. Binding such learning constraints from other clients enables our method to mitigate local forgetting. Finally, we achieve state-of-the-art results on several benchmarks, clearly demonstrating the advantages of the proposed method. Liangxi Liu, Xi Jiang 0009, Feng Zheng 0001, Hong Chen 0004, Guo-Jun Qi, Heng Huang 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion PredictionabstractInferring 3D human motion is fundamental in many applications, including understanding human activity and analyzing one's intention. While many fruitful efforts have been made to human motion prediction, most approaches focus on pose-driven prediction and inferring human motion in isolation from the contextual environment, thus leaving the body location movement in the scene behind. However, real-world human movements are goal-directed and highly influenced by the spatial layout of their surrounding scenes. In this paper, instead of planning future human motion in a "dark" room, we propose a Multi-Condition Latent Diffusion network (MCLD) that reformulates the human motion prediction task as a multi-condition joint inference problem based on the given historical 3D body motion and the current 3D scene contexts. Specifically, instead of directly modeling joint distribution over the raw motion sequences, MCLD performs a conditional diffusion process within the latent embedding space, characterizing the cross-modal mapping from the past body movement and current scene context condition embeddings to the future human motion embedding. Extensive experiments on large-scale human motion prediction datasets demonstrate that our MCLD achieves significant improvements over the state-of-the-art methods on both realistic and diverse predictions. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du, Guo-Jun Qi |
IEEE Trans. Image Process. | 5 |
| 2024 | Cross-Receptive Focused Inference Network for Lightweight Image Super-ResolutionabstractRecently, Transformer-based methods have shown impressive performance in single image super-resolution (SISR) tasks due to the ability of global feature extraction. However, the capabilities of Transformers that need to incorporate contextual information to extract features dynamically are neglected. To address this issue, we propose a lightweight Cross-receptive Focused Inference Network (CFIN) that consists of a cascade of CT Blocks mixed with CNN and Transformer. Specifically, in the CT block, we first propose a CNN-based Cross-Scale Information Aggregation Module (CIAM) to enable the model to better focus on potentially helpful information to improve the efficiency of the Transformer phase. Then, we design a novel Cross-receptive Field Guided Transformer (CFGT) to enable the selection of contextual information required for reconstruction by using a modulated convolutional kernel that understands the current semantic information and exploits the information interaction within different self-attention. Extensive experiments have shown that our proposed CFIN can effectively reconstruct images using contextual information, and it can strike a good balance between computational cost and model performance as an efficient model. Juncheng Li 0003, Guangwei Gao, Weihong Deng, Jiantao Zhou 0001, Jian Yang 0003, Guo-Jun Qi |
IEEE Trans. Multim. | 7 |
| 2024 | Self-supervised Multi-view Learning via Auto-encoding 3D Transformationsabstract3D object representation learning is a fundamental challenge in computer vision to infer about the 3D world. Recent advances in deep learning have shown their efficiency in 3D object recognition, among which view-based methods have performed best so far. However, feature learning of multiple views in existing methods is mostly performed in a supervised fashion, which often requires a large amount of data labels with high costs. In contrast, self-supervised learning aims to learn multi-view feature representations without involving labeled data. To this end, we propose a novel self-supervised framework to learn Multi-View Transformation Equivariant Representations (MV-TER), exploring the equivariant transformations of a 3D object and its projected multiple views that we derive. Specifically, we perform a 3D transformation on a 3D object and obtain multiple views before and after the transformation via projection. Then, we train a representation encoding module to capture the intrinsic 3D object representation by decoding 3D transformation parameters from the fused feature representations of multiple views before and after the transformation. Experimental results demonstrate that the proposed MV-TER significantly outperforms the state-of-the-art view-based approaches in 3D object classification and retrieval tasks and show the generalization to real-world datasets. The code is available at https://github.com/gyshgx868/mvter . Xiang Gao 0014, Wei Hu 0003, Guo-Jun Qi |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | High-Fidelity Clothed Avatar Reconstruction from a Single ImageabstractThis paper presents a framework for efficient 3D clothed avatar reconstruction. By combining the advantages of the high accuracy of optimization-based methods and the efficiency of learning-based methods, we propose a coarse-to-fine way to realize a high-fidelity clothed avatar reconstruction (CAR) from a single image. At the first stage, we use an implicit model to learn the general shape in the canonical space of a person in a learning-based way, and at the second stage, we refine the surface detail by estimating the non-rigid deformation in the posed space in an optimization way. A hyper-network is utilized to generate a good initialization so that the convergence of the optimization process is greatly accelerated. Extensive experiments on various datasets show that the proposed CAR successfully produces high-fidelity avatars for arbitrarily clothed humans in real scenes. The codes will be released in https://github.com/TingtingLiao/CAR. Tingting Liao, Yuliang Xiu, Hongwei Yi, Xudong Liu 0006, Guo-Jun Qi, Yong Zhang 0034, Xuan Wang 0009, Xiangyu Zhu 0001, Zhen Lei 0001 |
CVPR | 6 |
| 2023 | OTAvatar: One-Shot Talking Face Avatar with Controllable Tri-Plane RenderingabstractControllability, generalizability and efficiency are the major objectives of constructing face avatars represented by neural implicit field. However, existing methods have not managed to accommodate the three requirements simultaneously. They either focus on static portraits, restricting the representation ability to a specific subject, or suffer from substantial computational cost, limiting their flexibility. In this paper, we propose One-shot Talking face Avatar (OTAvatar), which constructs face avatars by a generalized controllable tri-plane rendering solution so that each personalized avatar can be constructed from only one portrait as the reference. Specifically, OTAvatar first inverts a portrait image to a motion-free identity code. Second, the identity code and a motion code are utilized to modulate an efficient CNN to generate a tri-plane formulated volume, which encodes the subject in the desired motion. Finally, volume rendering is employed to generate an image in any view. The core of our solution is a novel decoupling-by-inverting strategy that disentangles identity and motion in the latent code via optimization-based inversion. Benefiting from the efficient tri-plane representation, we achieve controllable rendering of generalized face avatar at 35 FPS on AIOO. Experiments show promising performance of crossidentity reenactment on subjects out of the training set and better 3D consistency. The code is available at https://github.com/theEricMaIOTAvatar. Zhiyuan Ma 0002, Xiangyu Zhu 0001, Guo-Jun Qi, Zhen Lei 0001, Lei Zhang 0006 |
CVPR | 3 |
| 2023 | POTTER: Pooling Attention Transformer for Efficient Human Mesh RecoveryabstractTransformer architectures have achieved SOTA performance on the human mesh recovery (HMR) from monocular images. However, the performance gain has come at the cost of substantial memory and computational overhead. A lightweight and efficient model to reconstruct accurate human mesh is needed for real-world applications. In this paper, we propose a pure transformer architecture named POoling aTtention TransformER (POTTER) for the HMR task from single images. Observing that the conventional attention module is memory and computationally expensive, we propose an efficient pooling attention module, which significantly reduces the memory and computational cost without sacrificing performance. Furthermore, we design a new transformer architecture by integrating a High-Resolution (HR) stream for the HMR task. The high-resolution local and global features from the HR stream can be utilized for recovering more accurate human mesh. Our POTTER outperforms the SOTA method METRO by only requiring 7% of total parameters and 14% of the Multiply-Accumulate Operations on the Human3.6M (PA-MPJPE metric) and 3DPW (all three metrics) datasets. The project webpage is https://zczcwh.github.io/potter_page/. Xianpeng Liu, Guo-Jun Qi, Chen Chen 0001 |
CVPR | 3 |
| 2023 | FeatER: An Efficient Network for Human Reconstruction via Feature Map-Based TransformERabstractRecently, vision transformers have shown great success in a set of human reconstruction tasks such as 2D/3D human pose estimation (2D/3D HPE) and human mesh reconstruction (HMR) tasks. In these tasks, feature map representations of the human structural information are often extracted first from the image by a CNN (such as HRNet), and then further processed by transformer to predict the heatmaps for HPE or HMR. However, existing transformer architectures are not able to process these feature map inputs directly, forcing an unnatural flattening of the location-sensitive human structural information. Furthermore, much of the performance benefit in recent HPE and HMR methods has come at the cost of ever-increasing computation and memory needs. Therefore, to simultaneously address these problems, we propose FeatER, a novel transformer design that preserves the inherent structure of feature map representations when modeling attention while reducing memory and computational costs. Taking advantage of FeatER, we build an efficient network for a set of human reconstruction tasks including 2D HPE, 3D HPE, and HMR. A feature map reconstruction module is applied to improve the performance of the estimated human pose and mesh. Extensive experiments demonstrate the effectiveness of FeatER on various human pose and mesh datasets. For instance, FeatER outperforms the SOTA method Mesh- Graphormer by requiring 5% of Params and 16% of MACs on Human3.6M and 3DPW datasets. The project webpage is https://zczcwh.github.io/feater_page/. Matías Mendieta, Taojiannan Yang, Guo-Jun Qi, Chen Chen 0001 |
CVPR | 4 |
| 2023 | Monocular 3D Object Detection with Bounding Box Denoising in 3D by PerceiverabstractThe main challenge of monocular 3D object detection is the accurate localization of 3D center. Motivated by a new and strong observation that this challenge can be remedied by a 3D-space local-grid search scheme in an ideal case, we propose a stage-wise approach, which combines the information flow from 2D-to-3D (3D bounding box proposal generation with a single 2D image) and 3D-to-2D (proposal verification by denoising with 3D-to-2D contexts) in a top-down manner. Specifically, we first obtain initial proposals from off-the-shelf backbone monocular 3D detectors. Then, we generate a 3D anchor space by local-grid sampling from the initial proposals. Finally, we perform 3D bounding box denoising at the 3D-to-2D proposal verification stage. To effectively learn discriminative features for denoising highly overlapped proposals, this paper presents a method of using the Perceiver I/O model [20] to fuse the 3D-to-2D geometric information and the 2D appearance information. With the encoded latent representation of a proposal, the verification head is implemented with a self-attention module. Our method, named as MonoXiver, is generic and can be easily adapted to any backbone monocular 3D detectors. Experimental results on the well-established KITTI dataset and the challenging large-scale Waymo dataset show that MonoXiver consistently achieves improvement with limited computation overhead. Xianpeng Liu, Kelvin Cheng 0003, Nan Xue 0001, Guo-Jun Qi, Tianfu Wu 0001 |
ICCV | 5 |
| 2023 | Self-similarity Driven Scale-invariant Learning for Weakly Supervised Person SearchabstractWeakly supervised person search aims to jointly detect and match persons with only bounding box annotations. Existing approaches typically focus on improving the features by exploring the relations of persons. However, scale variation problem is a more severe obstacle and under-studied that a person often owns images with different scales (resolutions). For one thing, small-scale images contain less information of a person, thus affecting the accuracy of the generated pseudo labels. For another, different similarities between cross-scale images of a person increase the difficulty of matching. In this paper, we address it by proposing a novel one-step framework, named Self-similarity driven Scale-invariant Learning (SSL). Scale invariance can be explored based on the self-similarity prior that it shows the same statistical properties of an image at different scales. To this end, we introduce a Multi-scale Exemplar Branch to guide the network in concentrating on the foreground and learning scale-invariant features by hard exemplars mining. To enhance the discriminative power of the learned features, we further introduce a dynamic pseudo label prediction that progressively seeks true labels for training. Experimental results on two standard benchmarks, i.e., PRW and CUHK-SYSU datasets, demonstrate that the proposed method can solve scale variation problem effectively and perform favorably against state-of-the-art methods. Code is available at https://github.com/Wangbenzhi/SSL.git. Benzhi Wang, Yang Yang 0062, Jinlin Wu, Guo-Jun Qi, Zhen Lei 0001 |
ICCV | 4 |
| 2023 | HiEve: A Large-Scale Benchmark for Human-Centric Video Analysis in Complex Events
Weiyao Lin, Huabin Liu 0001, Shizhan Liu, Yuxi Li 0009, Hongkai Xiong, Guo-Jun Qi, Nicu Sebe |
Int. J. Comput. Vis. | 6 |
| 2023 | CaCo: Both Positive and Negative Samples are Directly Learnable via Cooperative-Adversarial Contrastive LearningabstractAs a representative self-supervised method, contrastive learning has achieved great successes in unsupervised training of representations. It trains an encoder by distinguishing positive samples from negative ones given query anchors. These positive and negative samples play critical roles in defining the objective to learn the discriminative encoder, avoiding it from learning trivial features. While existing methods heuristically choose these samples, we present a principled method where both positive and negative samples are directly learnable end-to-end with the encoder. We show that the positive and negative samples can be cooperatively and adversarially learned by minimizing and maximizing the contrastive loss, respectively. This yields cooperative positives and adversarial negatives with respect to the encoder, which are updated to continuously track the learned representation of the query anchors over mini-batches. The proposed method achieves 71.3% and 75.3% in top-1 accuracy respectively over 200 and 800 epochs of pre-training ResNet-50 backbone on ImageNet1K without tricks such as multi-crop or stronger augmentations. With Multi-Crop, it can be further boosted into 75.7%. The source code and pre-trained model are released in https://github.com/maple-research-lab/caco. Xiao Wang 0013, Yuhang Huang 0006, Dan Zeng 0001, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Contrastive Learning With Stronger AugmentationsabstractRepresentation learning has significantly been developed with the advance of contrastive learning methods. Most of those methods are benefited from various data augmentations that are carefully designated to maintain their identities so that the images transformed from the same instance can still be retrieved. However, those carefully designed transformations limited us to further explore the novel patterns exposed by other transformations. Meanwhile, as shown in our experiments, direct contrastive learning for stronger augmented images can not learn representations effectively. Thus, we propose a general framework called Contrastive Learning with Stronger Augmentations (CLSA) to complement current contrastive learning approaches. Here, the distribution divergence between the weakly and strongly augmented images over the representation bank is adopted to supervise the retrieval of strongly augmented queries from a pool of instances. Experiments on the ImageNet dataset and downstream datasets showed the information from the strongly augmented images can significantly boost the performance. For example, CLSA achieves top-1 accuracy of 76.2% on ImageNet with a standard ResNet-50 architecture with a single-layer classifier fine-tuned, which is almost the same level as 76.5% of supervised results. Xiao Wang 0013, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Continuous cross-modal hashing
Hao Zheng 0008, Jinbao Wang 0001, Xiantong Zhen, Jingkuan Song, Feng Zheng 0001, Ke Lu 0002, Guo-Jun Qi |
Pattern Recognit. | 7 |
| 2023 | CTCNet: A CNN-Transformer Cooperation Network for Face Image Super-ResolutionabstractRecently, deep convolution neural networks (CNNs) steered face super-resolution methods have achieved great progress in restoring degraded facial details by joint training with facial priors. However, these methods have some obvious limitations. On the one hand, multi-task joint learning requires additional marking on the dataset, and the introduced prior network will significantly increase the computational cost of the model. On the other hand, the limited receptive field of CNN will reduce the fidelity and naturalness of the reconstructed facial images, resulting in suboptimal reconstructed images. In this work, we propose an efficient CNN-Transformer Cooperation Network (CTCNet) for face super-resolution tasks, which uses the multi-scale connected encoder-decoder architecture as the backbone. Specifically, we first devise a novel Local-Global Feature Cooperation Module (LGCM), which is composed of a Facial Structure Attention Unit (FSAU) and a Transformer block, to promote the consistency of local facial detail and global facial structure restoration simultaneously. Then, we design an efficient Feature Refinement Module (FRM) to enhance the encoded features. Finally, to further improve the restoration of fine facial details, we present a Multi-scale Feature Fusion Unit (MFFU) to adaptively fuse the features from different stages in the encoder procedure. Extensive evaluations on various datasets have assessed that the proposed CTCNet can outperform other state-of-the-art methods significantly. Source code will be available at https://github.com/IVIPLab/CTCNet. Guangwei Gao, Zixiang Xu, Juncheng Li 0003, Jian Yang 0003, Tieyong Zeng, Guo-Jun Qi |
IEEE Trans. Image Process. | 6 |
| 2023 | Adversarial Dense Contrastive Learning for Semi-Supervised Semantic SegmentationabstractSemi-supervised dense prediction tasks, such as semantic segmentation, can be greatly improved through the use of contrastive learning. However, this approach presents two key challenges: selecting informative negative samples from a highly redundant pool and implementing effective data augmentation. To address these challenges, we present an adversarial contrastive learning method specifically for semi-supervised semantic segmentation. Direct learning of adversarial negatives is adopted to retain discriminative information from the past, leading to higher learning efficiency. Our approach also leverages an advanced data augmentation strategy called AdverseMix, which combines information from under-performing classes to generate more diverse and challenging samples. Additionally, we use auxiliary labels and classifiers to prevent over-adversarial negatives from affecting the learning process. Our experiments on the Pascal VOC and Cityscapes datasets demonstrate that our method outperforms the state-of-the-art by a significant margin, even when using a small fraction of labeled data. Ying Wang 0051, Ziwei Xuan, Chiuman Ho, Guo-Jun Qi |
IEEE Trans. Image Process. | 4 |
| 2023 | Exploring the Robustness of Human Parsers Toward Common CorruptionsabstractHuman parsing aims to segment each pixel of the human image with fine-grained semantic categories. However, current human parsers trained with clean data are easily confused by numerous image corruptions such as blur and noise. To improve the robustness of human parsers, in this paper, we construct three corruption robustness benchmarks, termed LIP-C, ATR-C, and Pascal-Person-Part-C, to assist us in evaluating the risk tolerance of human parsing models. Inspired by the data augmentation strategy, we propose a novel heterogeneous augmentation-enhanced mechanism to bolster robustness under commonly corrupted conditions. Specifically, two types of data augmentations from different views, i.e., image-aware augmentation and model-aware image-to-image transformation, are integrated in a sequential manner for adapting to unforeseen image corruptions. The image-aware augmentation can enrich the high diversity of training images with the help of common image operations. The model-aware augmentation strategy that improves the diversity of input data by considering the model's randomness. The proposed method is model-agnostic, and it can plug and play into arbitrary state-of-the-art human parsing frameworks. The experimental results show that the proposed method demonstrates good universality which can improve the robustness of the human parsing models and even the semantic segmentation models when facing various image common corruptions. Meanwhile, it can still obtain approximate performance on clean data. Sanyi Zhang, Xiaochun Cao, Rui Wang 0032, Guo-Jun Qi, Jie Zhou 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Self-Supervised Graph Representation Learning via Topology TransformationsabstractWe present the Topology Transformation Equivariant Representation learning, a general paradigm of self-supervised learning for node representations of graph data to enable the wide applicability of Graph Convolutional Neural Networks (GCNNs). We formalize the proposed model from an information-theoretic perspective, by maximizing the mutual information between topology transformations and node representations before and after the transformations. We derive that maximizing such mutual information can be relaxed to minimizing the cross entropy between the applied topology transformation and its estimation from node representations. In particular, we seek to sample a subset of node pairs from the original graph and flip the edge connectivity between each pair to transform the graph topology. Then, we self-train a representation encoder to learn node representations by reconstructing the topology transformations from the feature representations of the original and transformed graphs. In experiments, we apply the proposed model to the downstream node classification, graph classification and link prediction tasks, and results show that the proposed method outperforms the state-of-the-art unsupervised approaches. Xiang Gao 0014, Wei Hu 0003, Guo-Jun Qi |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2023 | Cross-Domain Contrastive Learning for Unsupervised Domain AdaptationabstractUnsupervised domain adaptation (UDA) aims to transfer knowledge learned from a fully-labeled source domain to a different unlabeled target domain. Most existing UDA methods learn domain-invariant feature representations by minimizing feature distances across domains. In this work, we build upon contrastive self-supervised learning to align features so as to reduce the domain discrepancy between training and testing sets. Exploring the same set of categories shared by both domains, we introduce a simple yet effective framework CDCL, for domain alignment. In particular, given an anchor image from one domain, we minimize its distances to cross-domain samples from the same class relative to those from different categories. Since target labels are unavailable, we use a clustering-based approach with carefully initialized centers to produce pseudo labels. In addition, we demonstrate that CDCL is a general framework and can be adapted to the data-free setting, where the source data are unavailable during training, with minimal modification. We conduct experiments on two widely used domain adaptation benchmarks, i.e., Office-31 and VisDA-2017, for image classification tasks, and demonstrate that CDCL achieves state-of-the-art performance on both datasets. Rui Wang 0095, Zuxuan Wu, Zejia Weng, Jingjing Chen 0001, Guo-Jun Qi, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Exploring Resolution and Degradation Clues as Self-supervised Signal for Low Quality Object Detection
Ziteng Cui, Yingying Zhu 0004, Lin Gu 0003, Guo-Jun Qi, Renrui Zhang, Zenghui Zhang, Tatsuya Harada |
ECCV (9) | 4 |
| 2022 | Small Data Challenges in Big Data Era: A Survey of Recent Progress on Unsupervised and Semi-Supervised MethodsabstractRepresentation learning with small labeled data have emerged in many problems, since the success of deep neural networks often relies on the availability of a huge amount of labeled data that is expensive to collect. To address it, many efforts have been made on training sophisticated models with few labeled data in an unsupervised and semi-supervised fashion. In this paper, we will review the recent progresses on these two major categories of methods. A wide spectrum of models will be categorized in a big picture, where we will show how they interplay with each other to motivate explorations of new ideas. We will review the principles of learning the transformation equivariant, disentangled, self-supervised and semi-supervised representations, all of which underpin the foundation of recent progresses. Many implementations of unsupervised and semi-supervised generative models have been developed on the basis of these criteria, greatly expanding the territory of existing autoencoders, generative adversarial nets (GANs) and other deep networks by exploring the distribution of unlabeled data for more powerful representations. We will discuss emerging topics by revealing the intrinsic connections between unsupervised and semi-supervised learning, and propose in future directions to bridge the algorithmic and theoretical gap between transformation equivariance for unsupervised learning and supervised invariance for supervised learning, and unify unsupervised pretraining and supervised finetuning. We will also provide a broader outlook of future directions to unify transformation and instance equivariances for representation learning, connect unsupervised and semi-supervised augmentations, and explore the role of the self-supervised regularization for many learning problems. Guo-Jun Qi, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Learning Generalized Transformation Equivariant Representations Via AutoEncoding TransformationsabstractTransformation equivariant representations (TERs) aim to capture the intrinsic visual structures that equivary to various transformations by expanding the notion of translation equivariance underlying the success of convolutional neural networks (CNNs). For this purpose, we present both deterministic AutoEncoding Transformations (AET) and probabilistic AutoEncoding Variational Transformations (AVT) models to learn visual representations from generic groups of transformations. While the AET is trained by directly decoding the transformations from the learned representations, the AVT is trained by maximizing the joint mutual information between the learned representation and transformations. This results in generalized TERs (GTERs) equivariant against transformations in a more general fashion by capturing complex patterns of visual structures beyond the conventional linear equivariance under a transformation group. The presented approach can be extended to (semi-)supervised models by jointly maximizing the mutual information of the learned representation with both labels and transformations. Experiments demonstrate the proposed models outperform the state-of-the-art models in both unsupervised and (semi-)supervised tasks. Moreover, we show that the unsupervised representation can even surpass the fully supervised representation pretrained on ImageNet when they are fine-tuned for the object detection task. Guo-Jun Qi, Liheng Zhang, Feng Lin 0009, Xiao Wang 0013 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Spatiotemporal Co-Attention Recurrent Neural Networks for Human-Skeleton Motion PredictionabstractHuman motion prediction aims to generate future motions based on the observed human motions. Witnessing the success of Recurrent Neural Networks (RNN) in modeling sequential data, recent works utilize RNNs to model human-skeleton motions on the observed motion sequence and predict future human motions. However, these methods disregard the existence of the spatial coherence among joints and the temporal evolution among skeletons, which reflects the crucial characteristics of human motions in spatiotemporal space. To this end, we propose a novel Skeleton-Joint Co-Attention Recurrent Neural Networks (SC-RNN) to capture the spatial coherence among joints, and the temporal evolution among skeletons simultaneously on a skeleton-joint co-attention feature map in spatiotemporal space. First, a skeleton-joint feature map is constructed as the representation of the observed motion sequence. Second, we design a new Skeleton-Joint Co-Attention (SCA) mechanism to dynamically learn a skeleton-joint co-attention feature map of this skeleton-joint feature map, which can refine the useful observed motion information to predict one future motion. Third, a variant of GRU embedded with SCA collaboratively models the human-skeleton motion and human-joint motion in spatiotemporal space by regarding the skeleton-joint co-attention feature map as the motion context. Experimental results of human motion prediction demonstrate that the proposed method outperforms the competing methods. Xiangbo Shu, Liyan Zhang 0001, Guo-Jun Qi, Wei Liu 0005, Jinhui Tang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | $K$K-Shot Contrastive Learning of Visual Features With Multiple Instance AugmentationsabstractIn this paper, we propose the K-Shot Contrastive Learning (KSCL) of visual features by applying multiple augmentations to investigate the sample variations within individual instances. It aims to combine the advantages of inter-instance discrimination by learning discriminative features to distinguish between different instances, as well as intra-instance variations by matching queries against the variants of augmented samples over instances. Particularly, for each instance, it constructs an instance subspace to model the configuration of how the significant factors of variations in K-shot augmentations can be combined to form the variants of augmentations. Given a query, the most relevant variant of instances is then retrieved by projecting the query onto their subspaces to predict the positive instance class. This generalizes the existing contrastive learning that can be viewed as a special one-shot case. An eigenvalue decomposition is performed to configure instance subspaces, and the embedding network can be trained end-to-end through the differentiable subspace configuration. Experiment results demonstrate the proposed K-shot contrastive learning achieves superior performances to the state-of-the-art unsupervised methods. Haohang Xu, Hongkai Xiong, Guo-Jun Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Hierarchical Deep CNN Feature Set-Based Representation Learning for Robust Cross-Resolution Face RecognitionabstractCross-resolution face recognition (CRFR), which is important in intelligent surveillance and biometric forensics, refers to the problem of matching a low-resolution (LR) probe face image against high-resolution (HR) gallery face images. Existing shallow learning-based and deep learning-based methods focus on mapping the HR-LR face pairs into a joint feature space where the resolution discrepancy is mitigated. However, little works consider how to extract and utilize the intermediate discriminative features from the noisy LR query faces to further mitigate the resolution discrepancy due to the resolution limitations. In this study, we desire to fully exploit the multi-level deep convolutional neural network (CNN) feature set for robust CRFR. In particular, our contributions are threefold. (i) To learn more robust and discriminative features, we desire to adaptively fuse the contextual features from different layers. (ii) To fully exploit these contextual features, we design a feature set-based representation learning (FSRL) scheme to collaboratively represent the hierarchical features for more accurate recognition. Moreover, FSRL utilizes the primitive form of feature maps to keep the latent structural information, especially in noisy cases. (iii) To further promote the recognition performance, we desire to fuse the hierarchical recognition outputs from different stages. Meanwhile, the discriminability from different scales can also be fully integrated. By exploiting these advantages, the efficiency of the proposed method can be delivered. Experimental results on several face datasets have verified the superiority of the presented algorithm to the other competitive CRFR approaches. Guangwei Gao, Yi Yu 0001, Jian Yang 0003, Guo-Jun Qi, Meng Yang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Spatial-Temporal Pyramid Graph Reasoning for Action RecognitionabstractSpatial-temporal relation reasoning is a significant yet challenging problem for video action recognition. Previous works typically apply local operations like 2D or 3D CNNs to conduct space-time interactions in video sequences, or simply capture space-time long-range relations of a single fixed scale. However, this is inadequate for obtaining a comprehensive action representation. Besides, most models treat all input frames equally for the final classification, without selecting key frames and motion-sensitive regions. This introduces irrelevant video content and hurts the performance of models. In this paper, we propose a generic Spatial-Temporal Pyramid Graph Network (STPG-Net) to adaptively capture long-range spatial-temporal relations in video sequences at multiple scales. Specifically, we design a temporal attention (TA) module and a spatial-temporal attention (STA) module to learn the contribution of each frame and each space-time region to an action at a feature level, respectively. We then apply the selected key information to build spatial-temporal pyramid graphs for long-range relation reasoning and more comprehensive action representation learning. STPG-Net can be flexibly integrated into 2D and 3D backbone networks in a plug-and-play manner. Extensive experiments show that it brings consistent improvements over many challenging baselines on several standard action recognition benchmarks (i.e., Something-Something V1 & V2, and FineGym), demonstrating the effectiveness of our approach. Tiantian Geng, Feng Zheng 0001, Xiaorong Hou, Ke Lu 0002, Guo-Jun Qi, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | AIParsing: Anchor-Free Instance-Level Human ParsingabstractMost state-of-the-art instance-level human parsing models adopt two-stage anchor-based detectors and, therefore, cannot avoid the heuristic anchor box design and the lack of analysis on a pixel level. To address these two issues, we have designed an instance-level human parsing network which is anchor-free and solvable on a pixel level. It consists of two simple sub-networks: an anchor-free detection head for bounding box predictions and an edge-guided parsing head for human segmentation. The anchor-free detector head inherits the pixel-like merits and effectively avoids the sensitivity of hyper-parameters as proved in object detection applications. By introducing the part-aware boundary clue, the edge-guided parsing head is capable to distinguish adjacent human parts from among each other up to 58 parts in a single human instance, even overlapping instances. Meanwhile, a refinement head integrating box-level score and part-level parsing quality is exploited to improve the quality of the parsing results. Experiments on two multiple human parsing datasets (i.e., CIHP and LV-MHP-v2.0) and one video instance-level human parsing dataset (i.e., VIP) show that our method achieves the best global-level and instance-level performance over state-of-the-art one-stage top-down alternatives. Sanyi Zhang, Xiaochun Cao, Guo-Jun Qi, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | A Unified Multi-Scenario Attacking Network for Visual Object TrackingabstractExisting methods of adversarial attacks successfully generate adversarial examples to confuse Deep Neural Networks (DNNs) of image classification and object detection, resulting in wrong predictions. However, these methods are difficult to attack models of video object tracking, because the tracking algorithms could handle sequential information across video frames and the categories of targets tracked are normally unknown in advance. In this paper, we propose a Unified and Effective Network, named UEN, to attack visual object tracking models. There are several appealing characteristics of UEN: (1) UEN could produce various invisible adversarial perturbations according to different attack settings by using only one simple end-to-end network with three ingenious loss function; (2) UEN could generate general visible adversarial patch patterns to attack the advanced trackers in the real-world; (3) Extensive experiments show that UEN is able to attack many state-of-the-art trackers effectively (e.g. SiamRPN-based networks and DiMP) on popular tracking datasets including OTB100, UAV123, and GOT10K, making online real-time attacks possible. The attack results outperform the introduced baseline in terms of attacking ability and attacking efficiency. Xuesong Chen 0001, Canmiao Fu, Feng Zheng 0001, Yong Zhao 0010, Hongsheng Li 0001, Ping Luo 0002, Guo-Jun Qi |
AAAI | 7 |
| 2021 | Auto-Encoding Transformations in Reparameterized Lie Groups for Unsupervised LearningabstractUnsupervised training of deep representations has demonstrated remarkable potentials in mitigating the prohibitive expenses on annotating labeled data recently. Among them is predicting transformations as a pretext task to self-train representations, which has shown great potentials for unsupervised learning. However, existing approaches in this category learn representations by either treating a discrete set of transformations as separate classes, or using the Euclidean distance as the metric to minimize the errors between transformations. None of them has been dedicated to revealing the vital role of the geometry of transformation groups in learning representations. Indeed, an image must continuously transform along the curved manifold of a transformation group rather than through a straight line in the forbidden ambient Euclidean space. This suggests the use of geodesic distance to minimize the errors between the estimated and groundtruth transformations. Particularly, we focus on homographies, a general group of planar transformations containing the Euclidean, similarity and affine transformations as its special cases. To avoid an explicit computing of intractable Riemannian logarithm, we project homographies onto an alternative group of rotation transformations SR(3) with a tractable form of geodesic distance. Experiments demonstrate the proposed approach to Auto-Encoding Transformations exhibits superior performances on a variety of recognition problems. Feng Lin 0009, Haohang Xu, Houqiang Li, Hongkai Xiong, Guo-Jun Qi |
AAAI | 5 |
| 2021 | AdCo: Adversarial Contrast for Efficient Learning of Unsupervised Representations From Self-Trained Negative AdversariesabstractContrastive learning relies on constructing a collection of negative examples that are sufficiently hard to discriminate against positive queries when their representations are self-trained. Existing contrastive learning methods either maintain a queue of negative samples over minibatches while only a small portion of them are updated in an iteration, or only use the other examples from the current minibatch as negatives. They could not closely track the change of the learned representation over iterations by updating the entire queue as a whole, or discard the useful information from the past minibatches. Alternatively, we present to directly learn a set of negative adversaries playing against the self-trained representation. Two players, the representation network and negative adversaries, are alternately updated to obtain the most challenging negative examples against which the representation of positive queries will be trained to discriminate. We further show that the negative adversaries are updated towards a weighted combination of positive queries by maximizing the adversarial contrastive loss, thereby allowing them to closely track the change of representations over time. Experiment results demonstrate the proposed Adversarial Contrastive (AdCo) model not only achieves superior performances (a top-1 accuracy of 73.2% over 200 epochs and 75.7% over 800 epochs with linear evaluation on ImageNet), but also can be pre-trained more efficiently with much shorter GPU time and fewer epochs. The source code is available at https://github.com/maple-research-lab/AdCo. Qianjiang Hu, Xiao Wang 0004, Wei Hu 0003, Guo-Jun Qi |
CVPR | 4 |
| 2021 | Causal Attention for Vision-Language TasksabstractWe present a novel attention mechanism: Causal Attention (CATT), to remove the ever-elusive confounding effect in existing attention-based vision-language models. This effect causes harmful bias that misleads the attention module to focus on the spurious correlations in training data, damaging the model generalization. As the confounder is unobserved in general, we use the front-door adjustment to realize the causal intervention, which does not require any knowledge on the confounder. Specifically, CATT is implemented as a combination of 1) In-Sample Attention (IS-ATT) and 2) Cross-Sample Attention (CS-ATT), where the latter forcibly brings other samples into every IS-ATT, mimicking the causal intervention. CATT abides by the QK-V convention and hence can replace any attention module such as top-down attention and self-attention in Transformers. CATT improves various popular attention-based vision-language models by considerable margins. In particular, we show that CATT has great potential in large-scale pre-training, e.g., it can promote the lighter LXMERT [57], which uses fewer data and less computational power, comparable to the heavier UNITER [14]. Code is published in https://github.com/yangxuntu/lxmertcatt. Xu Yang 0021, Hanwang Zhang, Guo-Jun Qi, Jianfei Cai 0001 |
CVPR | 3 |
| 2021 | Knowledge-Enriched Distributional Model Inversion AttacksabstractModel inversion (MI) attacks are aimed at reconstructing training data from model parameters. Such attacks have triggered increasing concerns about privacy, especially given a growing number of online model repositories. However, existing MI attacks against deep neural networks (DNNs) have large room for performance improvement. We present a novel inversion-specific GAN that can better distill knowledge useful for performing attacks on private models from public data. In particular, we train the discriminator to differentiate not only the real and fake samples but the soft-labels provided by the target model. Moreover, unlike previous work that directly searches for a single data point to represent a target class, we propose to model a private data distribution for each target class. Our experiments show that the combination of these techniques can significantly boost the success rate of the state-of-the-art MI attacks by 150%, and generalize better to a variety of datasets and models. Our code is available at https://github.com/SCccc21/Knowledge-Enriched-DMI. Si Chen 0008, Mostafa Kahla, Ruoxi Jia 0001, Guo-Jun Qi |
ICCV | 4 |
| 2021 | Multitask AET with Orthogonal Tangent Regularity for Dark Object DetectionabstractDark environment becomes a challenge for computer vision algorithms owing to insufficient photons and undesirable noise. To enhance object detection in a dark environment, we propose a novel multitask auto encoding transformation (MAET) model which is able to explore the intrinsic pattern behind illumination translation. In a self-supervision manner, the MAET learns the intrinsic visual structure by encoding and decoding the realistic illumination-degrading transformation considering the physical noise model and image signal processing (ISP). Based on this representation, we achieve the object detection task by decoding the bounding box coordinates and classes. To avoid the over-entanglement of two tasks, our MAET disentangles the object and degrading features by imposing an orthogonal tangent regularity. This forms a parametric manifold along which multitask predictions can be geometrically formulated by maximizing the orthogonality between the tangents along the outputs of respective tasks. Our framework can be implemented based on the mainstream object detection architecture and directly trained end-to-end using normal target detection datasets, such as VOC and COCO. We have achieved the state-of-the-art performance using synthetic and real-world datasets. Codes will be released at https://github.com/cuiziteng/MAET. Ziteng Cui, Guo-Jun Qi, Lin Gu 0003, Shaodi You, Zenghui Zhang, Tatsuya Harada |
ICCV | 2 |
| 2021 | Prior-knowledge and attention based meta-learning for few-shot learning
Yunxiao Qin, Xiangyu Zhu 0001, Jingping Shi, Guo-Jun Qi, Zhen Lei 0001 |
Knowl. Based Syst. | 7 |
| 2021 | Hierarchical Long Short-Term Concurrent Memory for Human Interaction RecognitionabstractIn this work, we aim to address the problem of human interaction recognition in videos by exploring the long-term inter-related dynamics among multiple persons. Recently, Long Short-Term Memory (LSTM) has become a popular choice to model individual dynamic for single-person action recognition due to its ability to capture the temporal motion information in a range. However, most existing LSTM-based methods focus only on capturing the dynamics of human interaction by simply combining all dynamics of individuals or modeling them as a whole. Such methods neglect the inter-related dynamics of how human interactions change over time. To this end, we propose a novel Hierarchical Long Short-Term Concurrent Memory (H-LSTCM) to model the long-term inter-related dynamics among a group of persons for recognizing human interactions. Specifically, we first feed each person's static features into a Single-Person LSTM to model the single-person dynamic. Subsequently, at one time step, the outputs of all Single-Person LSTM units are fed into a novel Concurrent LSTM (Co-LSTM) unit, which mainly consists of multiple sub-memory units, a new cell gate, and a new co-memory cell. In the Co-LSTM unit, each sub-memory unit stores individual motion information, while this Co-LSTM unit selectively integrates and stores inter-related motion information between multiple interacting persons from multiple sub-memory units via the cell gate and co-memory cell, respectively. Extensive experiments on several public datasets validate the effectiveness of the proposed H-LSTCM by comparing against baseline and state-of-the-art methods. Xiangbo Shu, Jinhui Tang 0001, Guo-Jun Qi, Wei Liu 0005, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Partially-Connected Neural Architecture Search for Reduced Computational RedundancyabstractDifferentiable architecture search (DARTS) enables effective neural architecture search (NAS) using gradient descent, but suffers from high memory and computational costs. In this paper, we propose a novel approach, namely Partially-Connected DARTS (PC-DARTS), to achieve efficient and stable neural architecture search by reducing the channel and spatial redundancies of the super-network. In the channel level, partial channel connection is presented to randomly sample a small subset of channels for operation selection to accelerate the search process and suppress the over-fitting of the super-network. Side operation is introduced for bypassing (non-sampled) channels to guarantee the performance of searched architectures under extremely low sampling rates. In the spatial level, input features are down-sampled to eliminate spatial redundancy and enhance the efficiency of the mixed computation for operation selection. Furthermore, edge normalization is developed to maintain the consistency of edge selection based on channel sampling with the architectural parameters for edges. Theoretical analysis shows that partial channel connection and parameterized side operation are equivalent to regularizing the super-network on the weights and architectural parameters during bilevel optimization. Experimental results demonstrate that the proposed approach achieves higher search speed and training stability than DARTS. PC-DARTS obtains a top-1 error rate of 2.55 percent on CIFAR-10 with 0.07 GPU-days for architecture search, and a state-of-the-art top-1 error rate of 24.1 percent on ImageNet (under the mobile setting) within 2.8 GPU-days. Yuhui Xu 0002, Lingxi Xie, Wenrui Dai, Xiaopeng Zhang 0008, Xin Chen 0033, Guo-Jun Qi, Hongkai Xiong, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Human Parsing With Pyramidical Gather-Excite ContextabstractHuman parsing, especially in the wild, has attracted a lot of attention due to its great potential in many real-world applications. The Pyramid Spatial Parsing (PSP) module has shown superior performances in scene and human parsing tasks. However, the basic AvgPool operation in PSP equally aggregates spatial clues of a local region, and thus mixes up influences of different human parts presented in this region. It results in failures in capturing useful contexts relevant to parsing different parts. To address this problem, a suitable mechanism to collect spatial clues aligning with different human parts is proposed in this paper. We employ a Gather-Excite (GE) operation, a replacement of the AvgPool-Upsample operation in a pyramidical structure, to accurately reflect relevant human parts of various scales. The GE operation contains two steps: the gather operation that adaptively aggregates spatial clues to relevant human parts, and the excite operation that generates new feature maps with the gathered contextual information. This results in a novel Pyramidical Gather-Excite Context (PGEC) module to solve the multi-scale problem and parse person at various scales. The PGEC module is composed of multiple GE operations with different spatial extents and aggregates local and global spatial clues for better modeling multi-scale contextual information in parallel. Moreover, we integrate the PGEC module with fine-grained details, edge preserving module and deep supervision to formulate a novel PGEC Network (PGECNet) for human parsing. The proposed PGECNet has achieved state-of-the-art performance on four single-person human parsing datasets (i.e., LIP, PPSS, ATR and Fashion Clothing) and two multi-person human parsing datasets (i.e., PASCAL-Person-Part and CIHP). The experimental results show that the proposed PGEC is superior to the PSP and ASPP modules especially in single-human parsing task. The source code is publicly available at https://github.com/31sy/PGECNet. Sanyi Zhang, Guo-Jun Qi, Xiaochun Cao, Zhanjie Song, Jie Zhou 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | An End-to-End Foreground-Aware Network for Person Re-IdentificationabstractPerson re-identification is a crucial task of identifying pedestrians of interest across multiple surveillance camera views. For person re-identification, a pedestrian is usually represented with features extracted from a rectangular image region that inevitably contains the scene background, which incurs ambiguity to distinguish different pedestrians and degrades the accuracy. Thus, we propose an end-to-end foreground-aware network to discriminate against the foreground from the background by learning a soft mask for person re-identification. In our method, in addition to the pedestrian ID as supervision for the foreground, we introduce the camera ID of each pedestrian image for background modeling. The foreground branch and the background branch are optimized collaboratively. By presenting a target attention loss, the pedestrian features extracted from the foreground branch become more insensitive to backgrounds, which greatly reduces the negative impact of changing backgrounds on pedestrian matching across different camera views. Notably, in contrast to existing methods, our approach does not require an additional dataset to train a human landmark detector or a segmentation model for locating the background regions. The experimental results conducted on three challenging datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17, demonstrate the effectiveness of our approach. Wengang Zhou 0001, Jianzhuang Liu, Guo-Jun Qi, Qi Tian 0001, Houqiang Li |
IEEE Trans. Image Process. | 4 |
| 2021 | EnAET: A Self-Trained Framework for Semi-Supervised and Supervised Learning With Ensemble TransformationsabstractDeep neural networks have been successfully applied to many real-world applications. However, such successes rely heavily on large amounts of labeled data that is expensive to obtain. Recently, many methods for semi-supervised learning have been proposed and achieved excellent performance. In this study, we propose a new EnAET framework to further improve existing semi-supervised methods with self-supervised information. To our best knowledge, all current semi-supervised methods improve performance with prediction consistency and confidence ideas. We are the first to explore the role of self-supervised representations in semi-supervised learning under a rich family of transformations. Consequently, our framework can integrate the self-supervised information as a regularization term to further improve all current semi-supervised methods. In the experiments, we use MixMatch, which is the current state-of-the-art method on semi-supervised learning, as a baseline to test the proposed EnAET framework. Across different datasets, we adopt the same hyper-parameters, which greatly improves the generalization ability of the EnAET framework. Experiment results on different datasets demonstrate that the proposed EnAET framework greatly improves the performance of current semi-supervised algorithms. Moreover, this framework can also improve supervised learning by a large margin, including the extremely challenging scenarios with only 10 images per class. The code and experiment records are available in https://github.com/maple-research-lab/EnAET. Xiao Wang 0013, Daisuke Kihara, Jiebo Luo 0001, Guo-Jun Qi |
IEEE Trans. Image Process. | 4 |
| 2021 | Progressive Unsupervised Person Re-Identification by Tracklet Association With Spatio-Temporal RegularizationabstractExisting methods for person re-identification (Re-ID) are mostly based on supervised learning which requires numerous manually labeled samples across all camera views for training. Such a paradigm suffers the scalability issue since in real-world Re-ID application, it is difficult to exhaustively label abundant identities over multiple disjoint camera views. To this end, we propose a progressive deep learning method for unsupervised person Re-ID in the wild by Tracklet Association with Spatio-Temporal Regularization (TASTR). In our approach, we first collect tracklet data within each camera by automatic person detection and tracking. Then, an initial Re-ID model is trained based on within-camera triplet construction for person representation learning. After that, based on the person visual feature and spatio-temporal constraint, we associate cross-camera tracklets to generate cross-camera triplets and update the Re-ID model. Lastly, with the refined Re-ID model, better visual feature of person can be extracted, which further promote the association of cross-camera tracklets. The last two steps are iterated multiple times to progressively upgrade the Re-ID model. To facilitate the study, we have collected a new 4K UHD video dataset named Campus4K with full frames and full spatio-temporal information. Experimental results show that with the spatio-temporal constraint in the training phase, the proposed approach outperforms the state-of-the-art unsupervised methods by notable margins on DukeMTMC-reID, and achieves competitive performance to fully supervised methods on both DukeMTMC-reID and Campus4K datasets. Qiaokang Xie, Wengang Zhou 0001, Guo-Jun Qi, Qi Tian 0001, Houqiang Li |
IEEE Trans. Multim. | 3 |
| 2020 | POST: POlicy-Based Switch TrackingabstractIn visual object tracking, by reasonably fusing multiple experts, ensemble framework typically achieves superior performance compared to the individual experts. However, the necessity of parallelly running all the experts in most existing ensemble frameworks heavily limits their efficiency. In this paper, we propose POST, a POlicy-based Switch Tracker for robust and efficient visual tracking. The proposed POST tracker consists of multiple weak but complementary experts (trackers) and adaptively assigns one suitable expert for tracking in each frame. By formulating this expert switch in consecutive frames as a decision-making problem, we learn an agent via reinforcement learning to directly decide which expert to handle the current frame without running others. In this way, the proposed POST tracker maintains the performance merit of multiple diverse models while favorably ensuring the tracking efficiency. Extensive ablation studies and experimental comparisons against state-of-the-art trackers on 5 prevalent benchmarks verify the effectiveness of the proposed method. Ning Wang 0020, Wengang Zhou 0001, Guo-Jun Qi, Houqiang Li |
AAAI | 3 |
| 2020 | GraphTER: Unsupervised Learning of Graph Transformation Equivariant Representations via Auto-Encoding Node-Wise TransformationsabstractRecent advances in Graph Convolutional Neural Networks (GCNNs) have shown their efficiency for nonEuclidean data on graphs, which often require a large amount of labeled data with high cost. It it thus critical to learn graph feature representations in an unsupervised manner in practice. To this end, we propose a novel unsupervised learning of Graph Transformation Equivariant Representations (GraphTER), aiming to capture intrinsic patterns of graph structure under both global and local transformations. Specifically, we allow to sample different groups of nodes from a graph and then transform them node-wise isotropically or anisotropically. Then, we self-train a representation encoder to capture the graph structures by reconstructing these node-wise transformations from the feature representations of the original and transformed graphs. In experiments, we apply the learned GraphTER to graphs of 3D point cloud data, and results on point cloud segmentation/classification show that GraphTER significantly outperforms state-of-the-art unsupervised approaches and pushes greatly closer towards the upper bound set by the fully supervised counterparts. The code is available at: https://github.com/gyshgx868/graph-ter. Xiang Gao 0014, Wei Hu 0003, Guo-Jun Qi |
CVPR | 3 |
| 2020 | Transformation GAN for Unsupervised Image Synthesis and Representation LearningabstractGenerative Adversarial Networks (GAN) have shown promising performance in image synthesis and unsupervised learning (USL). In most cases, however, the representations extracted from unsupervised GAN are usually unsatisfactory in other computer vision tasks. By using conditional GAN (CGAN), this problem could be solved to some extent, but the main drawback of such models is the necessity for labeled data. To improve both image synthesis quality and representation learning performance under the unsupervised setting, in this paper, we propose a simple yet effective Transformation Generative Adversarial Networks (TrGAN). In our approach, instead of capturing the joint distribution of image-label pairs p(x,y) as in conditional GAN, we try to estimate the joint distribution of transformed image t(x) and transformation t. Specifically, given a randomly sampled transformation t, we train the discriminator to give an estimate of input transformation, while following the adversarial training scheme of the original GAN. In addition, intermediate feature matching as well as feature-transform matching methods are introduced to strengthen the regularization on the generated features. To evaluate the quality of both generated samples and extracted representations, extensive experiments are conducted on four public datasets. The experimental results on the quality of both the synthesized images and the extracted representations demonstrate the effectiveness of our method. Wengang Zhou 0001, Guo-Jun Qi, Zhongqian Fu, Qi Tian 0001, Houqiang Li |
CVPR | 3 |
| 2020 | Rotation Equivariant Graph Convolutional Network for Spherical Image ClassificationabstractConvolutional neural networks (CNNs) designed for low-dimensional regular grids will unfortunately lead to non-optimal solutions for analyzing spherical images, due to their different geometrical properties from planar images. In this paper, we generalize the grid-based CNNs to a non-Euclidean space by taking into account the geometry of spherical surfaces and propose a Spherical Graph Convolutional Network (SGCN) to encode rotation equivariant representations. Specifically, we propose a spherical graph construction criterion showing that a graph needs to be regular by evenly covering the spherical surfaces in order to design a rotation equivariant graph convolutional layer. For the practical case where the perfectly regular graph does not exist, we design two quantitative measures to evaluate the degree of irregularity for a spherical graph. The Geodesic ICOsahedral Pixelation (GICOPix) is adopted to construct spherical graphs with the minimum degree of irregularity compared to the current popular pixelation schemes. In addition, we design a hierarchical pooling layer to keep the rotation-equivariance, followed by a transition layer to enforce the invariance to the rotations for spherical image classification. We evaluate the proposed graph convolutional layers with different pixelations schemes in terms of equivariance errors. We also assess the effectiveness of the proposed SGCN in fulfilling rotation-invariance by the invariance error of the transition layers and recognizing the spherical images and 3D objects. Qin Yang 0002, Wenrui Dai, Junni Zou, Guo-Jun Qi, Hongkai Xiong |
CVPR | 5 |
| 2020 | WCP: Worst-Case Perturbations for Semi-Supervised Deep LearningabstractIn this paper, we present a novel regularization mechanism for training deep networks by minimizing the {\em Worse-Case Perturbation} (WCP). It is based on the idea that a robust model is least likely to be affected by small perturbations, such that its output decisions should be as stable as possible on both labeled and unlabeled examples. We will consider two forms of WCP regularizations -- additive and DropConnect perturbations, which impose additive noises on network weights, and make structural changes by dropping the network connections, respectively. We will show that the worse cases of both perturbations can be derived by solving respective optimization problems with spectral methods. The WCP can be minimized on both labeled and unlabeled data so that networks can be trained in a semi-supervised fashion. This leads to a novel paradigm of semi-supervised classifiers by stabilizing the predicted outputs in presence of the worse-case perturbations imposed on the network weights and structures. Liheng Zhang, Guo-Jun Qi |
CVPR | 2 |
| 2020 | Cross-VAE: Towards Disentangling Expression from Identity For Human FacesabstractFacial expression and identity are two independent yet intertwined components for representing a face. For facial expression recognition, identity can contaminate the training procedure by providing tangled but irrelevant information. In this paper, we propose to learn clearly disentangled and discriminative features that are invariant of identities for expression recognition. However, such disentanglement normally requires annotations of both expression and identity on one large dataset, which is often unavailable. Our solution is to extend conditional VAE to a crossed version named Cross-VAE, which is able to use partially labeled data to disentangle expression from identity. We emphasis the following novel characteristics of our Cross-VAE: (1) It is based on an independent assumption that the two latent representations' distributions are orthogonal. This ensures both encoded representations to be disentangled and expressive. (2) It utilizes a symmetric training procedure where the output of each encoder is fed as the condition of the other. Thus two partially labeled sets can be jointly used. Extensive experiments show that our proposed method is capable of encoding expressive and disentangled features for facial expression. Compared with the baseline methods, our model shows an improvement of 3.56% on average in terms of accuracy. Haozhe Wu, Jia Jia 0001, Lingxi Xie, Guo-Jun Qi, Yuanchun Shi, Qi Tian 0001 |
ICASSP | 4 |
| 2020 | PC-DARTS: Partial Channel Connections for Memory-Efficient Architecture Search
Yuhui Xu 0002, Lingxi Xie, Xiaopeng Zhang 0008, Xin Chen 0033, Guo-Jun Qi, Qi Tian 0001, Hongkai Xiong |
ICLR | 5 |
| 2020 | Loss-Sensitive Generative Adversarial Networks on Lipschitz Densities
Guo-Jun Qi |
Int. J. Comput. Vis. | 1 |
| 2020 | Learning Compact Features for Human Activity Recognition Via Probabilistic First-Take-AllabstractWith the popularity of mobile sensor technology, smart wearable devices open a unprecedented opportunity to solve the challenging human activity recognition (HAR) problem by learning expressive representations from the multi-dimensional daily sensor signals. This inspires us to develop a new algorithm applicable to both camera-based and wearable sensor-based HAR systems. Although competitive classification accuracy has been reported, existing methods often face the challenge of distinguishing visually similar activities composed of activity patterns in different temporal orders. In this paper, we propose a novel probabilistic algorithm to compactly encode temporal orders of activity patterns for HAR. Specifically, the algorithm learns an optimal set of latent patterns such that their temporal structures really matter in recognizing different human activities. Then, a novel probabilistic First-Take-All (pFTA) approach is introduced to generate compact features from the orders of these latent patterns to encode the entire sequence, and the temporal structural similarity between different sequences can be efficiently measured by the Hamming distance between compact features. Experiments on three public HAR datasets show the proposed pFTA approach can achieve competitive performance in terms of accuracy as well as efficiency. Jun Ye 0003, Guo-Jun Qi, Naifan Zhuang, Hao Hu 0010, Kien A. Hua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | E-Optimal Sensor Selection for Compressive Sensing-Based PurposesabstractCollaborative estimation of a sparse vector x by M potential measurements is considered. Each measurement is the projection of x obtained by a regressor, i.e., ym1/4 amTx. The problem of selecting K sensor measurements from a set of M potential sensors is studied where K ≫ M and K is less than the dimension of x. In other words, we aim to reduce the problem to an under-determined system of equations in which a sparse solution is desired. This paper suggests selecting sensors in a way that the reduced matrix construct a well conditioned measurement matrix. Our criterion is based on E-optimality, which is highly related to the restricted isometry property that provides some guarantees for sparse solution obtained by ℓ1minimization. The proposed basic E-optimal selection is vulnerable to outlier and noisy data. The robust version of the algorithm is presented for distributed selection for big data sets. Moreover, an online implementation is proposed that involves partially observed measurements in a sequential manner. Our simulation results show the proposed method outperforms the other criteria for collaborative spectrum sensing in cognitive radio networks (CRNs). Our suggested selection method is evaluated in machine learning applications. It is used to pick up the most informative features/data. Specifically, the proposed method is exploited for face recognition with partial training data. Mohsen Joneidi, Alireza Zaeemzadeh, Behzad Shahrasbi, Guo-Jun Qi, Nazanin Rahnavard |
IEEE Trans. Big Data | 4 |
| 2020 | Guest Editorial Introduction to the Special Section on Representation Learning for Visual Content UnderstandingabstractRepresentation learning methods allow a system to automatically learn robust and discriminative features from raw data for given goals, which play an important role in various visual content understanding applications, such as visual object segmentation, detection, tracking, recognition, and search. The performance of visual content understanding tasks is heavily dependent on the choice of data representation (or features) on which they are applied. Conventional feature representation methods usually employ transformations of data that make it easier to extract useful information, such as scale-invariant feature transform (SIFT), local binary patterns (LBP), and histogram of oriented gradients (HOG). In recent years, deep learning techniques have been widely applied to learn data-driven representations with supervised annotations and achieved great success in different visual content understanding tasks. Representative methods include the ResNet method for image classification, the DeepFace method for face recognition, and the feature pyramid networks (FPNs) method for object detection. Despite recent progresses on deep representation learning with a great amount of annotated data, how to effectively learn visual representation with limited data annotations still requires many efforts. This special section focuses on data-effective representation learning methods for visual content understanding. Jiwen Lu, Yuxin Peng 0001, Guo-Jun Qi, Jun Yu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Learning to Adaptively Scale Recurrent Neural NetworksabstractRecent advancements in recurrent neural network (RNN) research have demonstrated the superiority of utilizing multiscale structures in learning temporal representations of time series. Currently, most of multiscale RNNs use fixed scales, which do not comply with the nature of dynamical temporal patterns among sequences. In this paper, we propose Adaptively Scaled Recurrent Neural Networks (ASRNN), a simple but efficient way to handle this problem. Instead of using predefined scales, ASRNNs are able to learn and adjust scales based on different temporal contexts, making them more flexible in modeling multiscale patterns. Compared with other multiscale RNNs, ASRNNs are bestowed upon dynamical scaling capabilities with much simpler structures, and are easy to be integrated with various RNN cells. The experiments on multiple sequence modeling tasks indicate ASRNNs can efficiently adapt scales based on different sequence contexts and yield better performances than baselines without dynamical scaling abilities. Hao Hu 0010, Guo-Jun Qi |
AAAI | 3 |
| 2019 | Task Agnostic Meta-Learning for Few-Shot LearningabstractMeta-learning approaches have been proposed to tackle the few-shot learning problem. Typically, a meta-learner is trained on a variety of tasks in the hopes of being generalizable to new tasks. However, the generalizability on new tasks of a meta-learner could be fragile when it is over-trained on existing tasks during meta-training phase. In other words, the initial model of a meta-learner could be too biased towards existing tasks to adapt to new tasks, especially when only very few examples are available to update the model. To avoid a biased meta-learner and improve its generalizability, we propose a novel paradigm of Task-Agnostic Meta-Learning (TAML) algorithms. Specifically, we present an entropy-based approach that meta-learns an unbiased initial model with the largest uncertainty over the output labels by preventing it from over-performing in classification tasks. Alternatively, a more general inequality-minimization TAML is presented for more ubiquitous scenarios by directly minimizing the inequality of initial losses beyond the classification tasks wherever a suitable loss can be defined. Experiments on benchmarked datasets demonstrate that the proposed approaches outperform compared meta-learning algorithms in both few-shot classification and reinforcement learning tasks. Muhammad Abdullah Jamal, Guo-Jun Qi |
CVPR | 2 |
| 2019 | AET vs. AED: Unsupervised Representation Learning by Auto-Encoding Transformations Rather Than DataabstractThe success of deep neural networks often relies on a large amount of labeled examples, which can be difficult to obtain in many real scenarios. To address this challenge, unsupervised methods are strongly preferred for training neural networks without using any labeled data. In this paper, we present a novel paradigm of unsupervised representation learning by Auto-Encoding Transformation (AET) in contrast to the conventional Auto-Encoding Data (AED) approach. Given a randomly sampled transformation, AET seeks to predict it merely from the encoded features as accurately as possible at the output end. The idea is the following: as long as the unsupervised features successfully encode the essential information about the visual structures of original and transformed images, the transformation can be well predicted. We will show that this AET paradigm allows us to instantiate a large variety of transformations, from parameterized, to non-parameterized and GAN-induced ones. Our experiments show that AET greatly improves over existing unsupervised approaches, setting new state-of-the-art performances being greatly closer to the upper bounds by their fully supervised counterparts on CIFAR-10, ImageNet and Places datasets. Liheng Zhang, Guo-Jun Qi, Jiebo Luo 0001 |
CVPR | 2 |
| 2019 | Few-Shot Image Recognition With Knowledge TransferabstractHuman can well recognize images of novel categories just after browsing few examples of these categories. One possible reason is that they have some external discriminative visual information about these categories from their prior knowledge. Inspired from this, we propose a novel Knowledge Transfer Network architecture (KTN) for few-shot image recognition. The proposed KTN model jointly incorporates visual feature learning, knowledge inferring and classifier learning into one unified framework for their optimal compatibility. First, the visual classifiers for novel categories are learned based on the convolutional neural network with the cosine similarity optimization. To fully explore the prior knowledge, a semantic-visual mapping network is then developed to conduct knowledge inference, which enables to infer the classifiers for novel categories from base categories. Finally, we design an adaptive fusion scheme to infer the desired classifiers by effectively integrating the above knowledge and visual information. Extensive experiments are conducted on two widely-used Mini-ImageNet and ImageNet Few-Shot benchmarks to evaluate the effectiveness of the proposed method. The results compared with the state-of-the-art approaches show the encouraging performance of the proposed method, especially on 1-shot and 2-shot tasks. Zhimao Peng, Zechao Li, Junge Zhang, Guo-Jun Qi, Jinhui Tang 0001 |
ICCV | 5 |
| 2019 | AVT: Unsupervised Learning of Transformation Equivariant Representations by Autoencoding Variational TransformationsabstractThe learning of Transformation-Equivariant Representations (TERs), which is introduced by Hinton et al. \cite{hinton2011transforming}, has been considered as a principle to reveal visual structures under various transformations. It contains the celebrated Convolutional Neural Networks (CNNs) as a special case that only equivary to the translations. In contrast, we seek to train TERs for a generic class of transformations and train them in an {\em unsupervised} fashion. To this end, we present a novel principled method by Autoencoding Variational Transformations (AVT), compared with the conventional approach to autoencoding data. Formally, given transformed images, the AVT seeks to train the networks by maximizing the mutual information between the transformations and representations. This ensures the resultant TERs of individual images contain the {\em intrinsic} information about their visual structures that would equivary {\em extricably} under various transformations in a generalized {\em nonlinear} case. Technically, we show that the resultant optimization problem can be efficiently solved by maximizing a variational lower-bound of the mutual information. This variational approach introduces a transformation decoder to approximate the intractable posterior of transformations, resulting in an autoencoding architecture with a pair of the representation encoder and the transformation decoder. Experiments demonstrate the proposed AVT model sets a new record for the performances on unsupervised tasks, greatly closing the performance gap to the supervised models. Guo-Jun Qi, Liheng Zhang, Chang Wen Chen, Qi Tian 0001 |
ICCV | 1 |
| 2019 | Boundary Perception Guidance: A Scribble-Supervised Semantic Segmentation ApproachabstractSemantic segmentation suffers from the fact that densely annotated masks are expensive to obtain. To tackle this problem, we aim at learning to segment by only leveraging scribbles that are much easier to collect for supervision. To fully explore the limited pixel-level annotations from scribbles, we present a novel Boundary Perception Guidance (BPG) approach, which consists of two basic components, i.e., prediction refinement and boundary regression. Specifically, the prediction refinement progressively makes a better segmentation by adopting an iterative upsampling and a semantic feature enhancement strategy. In the boundary regression, we employ class-agnostic edge maps for supervision to effectively guide the segmentation network in localizing the boundaries between different semantic regions, leading to producing finer-grained representation of feature maps for semantic segmentation. The experiment results on the PASCAL VOC 2012 demonstrate the proposed BPG achieves mIoU of 73.2% without fully connected Conditional Random Field (CRF) and 76.0% with CRF, setting up the new state-of-the-art in literature. Bin Wang 0065, Guo-Jun Qi, Sheng Tang, Tianzhu Zhang 0001, Yunchao Wei, Yongdong Zhang 0001 |
IJCAI | 2 |
| 2019 | Improving Diversity of Image Captioning Through Variational Autoencoders and Adversarial LearningabstractLearning translation from images to human-readable natural language has become a great challenge in computer vision research in recent years. Existing works explore the semantic correlation between the visual and language domains via encoder-to-decoder learning frameworks based on classifying visual features in the language domain. This approach, however, is criticized for its lacking of naturalness and diversity. In this paper, we demonstrate a novel way to learn a semantic connection between visual information and natural language directly based on a Variational Autoencoder (VAE) that is trained in an adversarial routine. Instead of using the classification based discriminator, our method directly learns to estimate the diversity between a hidden vector embedded from a text encoder and an informative feature that is sampled from a learned distribution of the autoencoders. We show that the sentences learned from this matching contains accurate semantic meaning with high diversity in the image captioning task. Our experiments on the popular MSCOCO dataset indicates that our method learns to generate high-quality natural language with competitive scores on both correctness and diversity. Guo-Jun Qi, Kien A. Hua |
WACV | 2 |
| 2019 | Cascade Attention Machine for Occluded Landmark Detection in 2D X-Ray AngiographyabstractIn cardiac interventions, localization of guiding catheter tip in 2D fluoroscopic images is important to specify vessel branches and calibrate vessels with stenosis. While detection of guiding catheter tip is not trivial in contrast-free images due to low dose radiation as well as occlusion by other devices, it is even more challenging in contrast-filled images. As contrast-filled vessels become visible in X-ray imaging, the landmark of guiding catheter tip can often be completely occluded by the contrast medium. It is difficult even for human eyes to precisely localize the catheter tip from a single angiography image. Physicians have to rely on information before the inject of contrast medium to localize the guiding catheter tip occluded by contrast medium. Automatic landmark detection when occlusion happens is important and can significantly simplify the intervention workflow. To address this problem, we propose a novel Cascade Attention Machine (CAM) model. It borrows the idea of how human experts localize the catheter tip by first per-forming landmark detection when occlusion does not happen, then leveraging this information as prior knowledge to assist the occluded detection. Attention maps are computed from non-occluded detection to further refine the heatmaps for occluded detection to guide the inference focusing on related regions. Experiments on X-ray angiography demonstrate the promising performance compared with the state-of-the-art baselines. It shows that the CAM can capture the relation between situations with and without occlusion to achieve precise detection of occluded landmark. Liheng Zhang, Vivek K. Singh 0002, Guo-Jun Qi, Terrence Chen |
WACV | 3 |
| 2019 | Large-Scale Bisample Learning on ID Versus Spot Face Recognition
Xiangyu Zhu 0001, Zhen Lei 0001, Hailin Shi, Fan Yang 0062, Dong Yi, Guo-Jun Qi, Stan Z. Li |
Int. J. Comput. Vis. | 7 |
| 2019 | Image generation via latent space learning using improved combination
Jie Zhou 0001, Guo-Jun Qi |
Neurocomputing | 4 |
| 2019 | Sharp Attention Network via Adaptive Sampling for Person Re-IdentificationabstractIn this paper, we present novel sharp attention networks by adaptively sampling feature maps from convolutional neural networks for person re-identification (re-ID) problems. Due to the introduction of sampling-based attention models, the proposed approach can adaptively generate sharper attention-aware feature masks. This greatly differs from the gating-based attention mechanism that relies on soft gating functions to select the relevant features for person re-ID. In contrast, the proposed sampling-based attention mechanism allows us to effectively trim irrelevant features by enforcing the resultant feature masks to focus on the most discriminative features. It can produce sharper attentions that is more assertive in localizing subtle features relevant to re-identifying people across cameras. For this purpose, a differentiable Gumbel-Softmax sampler is employed to approximate the Bernoulli sampling to train the sharp attention networks. Extensive experimental evaluations demonstrate the superiority of this new sharp attention model for person re-ID over other related existing, published state-of-the-art works on three challenging benchmarks, including CUHK03, Market-1501, and DukeMTMC-reID. Chen Shen 0003, Guo-Jun Qi, Rongxin Jiang 0001, Zhongming Jin 0001, Hongwei Yong, Yaowu Chen, Xian-Sheng Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep Ordinal Hashing With Spatial AttentionabstractHashing has attracted increasing research attention in recent years due to its high efficiency of computation and storage in image retrieval. Recent works have demonstrated the superiority of simultaneous feature representations and hash functions learning with deep neural networks. However, most existing deep hashing methods directly learn the hash functions by encoding the global semantic information, while ignoring the local spatial information of images. The loss of local spatial structure makes the performance bottleneck of hash functions, therefore limiting its application for accurate similarity retrieval. In this paper, we propose a novel deep ordinal hashing (DOH) method, which learns ordinal representations to generate ranking-based hash codes by leveraging the ranking structure of feature space from both local and global views. In particular, to effectively build the ranking structure, we propose to learn the rank correlation space by exploiting the local spatial information from fully convolutional network and the global semantic information from the convolutional neural network simultaneously. More specifically, an effective spatial attention model is designed to capture the local spatial information by selectively learning well-specified locations closely related to target objects. In such hashing framework, the local spatial and global semantic nature of images is captured in an end-to-end ranking-to-hashing manner. Experimental results conducted on three widely used datasets demonstrate that the proposed DOH method significantly outperforms the state-of-the-art hashing methods. Lu Jin 0001, Xiangbo Shu, Kai Li 0005, Zechao Li, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Deep Semantic-Preserving Ordinal Hashing for Cross-Modal Similarity SearchabstractCross-modal hashing has attracted increasing research attention due to its efficiency for large-scale multimedia retrieval. With simultaneous feature representation and hash function learning, deep cross-modal hashing (DCMH) methods have shown superior performance. However, most existing methods on DCMH adopt binary quantization functions (e.g., [Formula: see text]) to generate hash codes, which limit the retrieval performance since binary quantization functions are sensitive to the variations of numeric values. Toward this end, we propose a novel end-to-end ranking-based hashing framework, in this paper, termed as deep semantic-preserving ordinal hashing (DSPOH), to learn hash functions with deep neural networks by exploring the ranking structure of feature dimensions. In DSPOH, the ordinal representation, which encodes the relative rank ordering of feature dimensions, is explored to generate hash codes. Such ordinal embedding benefits from the numeric stability of rank correlation measures. To make the hash codes discriminative, the ordinal representation is expected to well predict the class labels so that the ranking-based hash function learning is optimally compatible with the label predicting. Meanwhile, the intermodality similarity is preserved to guarantee that the hash codes of different modalities are consistent. Importantly, DSPOH can be effectively integrated with different types of network architectures, which demonstrates the flexibility and scalability of our proposed hashing framework. Extensive experiments on three widely used multimodal data sets show that DSPOH outperforms state of the art for cross-modal retrieval tasks. Lu Jin 0001, Kai Li 0005, Zechao Li, Fu Xiao 0001, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2019 | Multi-level Similarity Perception Network for Person Re-identificationabstractIn this article, we propose a novel deep Siamese architecture based on a convolutional neural network (CNN) and multi-level similarity perception for the person re-identification (re-ID) problem. According to the distinct characteristics of diverse feature maps, we effectively apply different similarity constraints to both low-level and high-level feature maps during training stage. Due to the introduction of appropriate similarity comparison mechanisms at different levels, the proposed approach can adaptively learn discriminative local and global feature representations, respectively, while the former is more sensitive in localizing part-level prominent patterns relevant to re-identifying people across cameras. Meanwhile, a novel strong activation pooling strategy is utilized on the last convolutional layer for abstract local-feature aggregation to pursue more representative feature representations. Based on this, we propose final feature embedding by simultaneously encoding original global features and discriminative local features. In addition, our framework has two other benefits: First, classification constraints can be easily incorporated into the framework, forming a unified multi-task network with similarity constraints. Second, as similarity-comparable information has been encoded in the network’s learning parameters via back-propagation, pairwise input is not necessary at test time. That means we can extract features of each gallery image and build an index in an off-line manner, which is essential for large-scale real-world applications. Experimental results on multiple challenging benchmarks demonstrate that our method achieves splendid performance compared with the current state-of-the-art approaches. Chen Shen 0003, Zhongming Jin 0001, Wenqing Chu, Rongxin Jiang 0001, Yaowu Chen, Guo-Jun Qi, Xian-Sheng Hua 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2019 | Rethinking the Combined and Individual Orders of Derivative of States for Differential Recurrent Neural Networks: Deep Differential Recurrent Neural NetworksabstractDue to their special gating schemes, Long Short-Term Memory (LSTM) has shown greater potential to process complex sequential information than the traditional Recurrent Neural Network (RNN). The conventional LSTM, however, fails to take into consideration the impact of salient spatio-temporal dynamics present in the sequential input data. This problem was first addressed by the differential Recurrent Neural Network (dRNN), which uses a differential gating scheme known as Derivative of States (DoS). DoS uses higher orders of internal state derivatives to analyze the change in information gain originated from the salient motions between the successive frames. The weighted combination of several orders of DoS is then used to modulate the gates in dRNN. While each individual order of DoS is good at modeling a certain level of salient spatio-temporal sequences, the sum of all the orders of DoS could distort the detected motion patterns. To address this problem, we propose to control the LSTM gates via individual orders of DoS. To fully utilize the different orders of DoS, we further propose to stack multiple levels of LSTM cells in an increasing order of state derivatives. The proposed model progressively builds up the ability of the LSTM gates to detect salient dynamical patterns in deeper stacked layers modeling higher orders of DoS; thus, the proposed LSTM model is termed deep differential Recurrent Neural Network ( d 2 RNN). The effectiveness of the proposed model is demonstrated on three publicly available human activity datasets: NUS-HGA, Violent-Flows, and UCF101. The proposed model outperforms both LSTM and non-LSTM based state-of-the-art algorithms. Naifan Zhuang, Guo-Jun Qi, The Duc Kieu, Kien A. Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2018 | Global Versus Localized Generative Adversarial NetsabstractIn this paper, we present a novel localized Generative Adversarial Net (GAN) to learn on the manifold of real data. Compared with the classic GAN that globally parameterizes a manifold, the Localized GAN (LGAN) uses local coordinate charts to parameterize distinct local geometry of how data points can transform at different locations on the manifold. Specifically, around each point there exists a local generator that can produce data following diverse patterns of transformations on the manifold. The locality nature of LGAN enables local generators to adapt to and directly access the local geometry without need to invert the generator in a global GAN. Furthermore, it can prevent the manifold from being locally collapsed to a dimensionally deficient tangent subspace by imposing an orthonormality prior between tangents. This provides a geometric approach to alleviating mode collapse at least locally on the manifold by imposing independence between data transformations in different tangent directions. We will also demonstrate the LGAN can be applied to train a robust classifier that prefers locally consistent classification decisions on the manifold, and the resultant regularizer is closely related with the Laplace-Beltrami operator. Our experiments show that the proposed LGANs can not only produce diverse image transformations, but also deliver superior classification performances. Guo-Jun Qi, Liheng Zhang, Hao Hu 0010, Marzieh Edraki, Jingdong Wang 0001, Xian-Sheng Hua 0001 |
CVPR | 1 |
| 2018 | Interleaved Structured Sparse Convolutional Neural NetworksabstractIn this paper, we study the problem of designing efficient convolutional neural network architectures with the interest in eliminating the redundancy in convolution kernels. In addition to structured sparse kernels, low-rank kernels and the product of low-rank kernels, the product of structured sparse kernels, which is a framework for interpreting the recently-developed interleaved group convolutions (IGC) and its variants (e.g., Xception), has been attracting increasing interests. Motivated by the observation that the convolutions contained in a group convolution in IGC can be further decomposed in the same manner, we present a modularized building block, IGC-V2: interleaved structured sparse convolutions. It generalizes interleaved group convolutions, which is composed of two structured sparse kernels, to the product of more structured sparse kernels, further eliminating the redundancy. We present the complementary condition and the balance condition to guide the design of structured sparse kernels, obtaining a balance among three aspects: model size, computation complexity and classification accuracy. Experimental results demonstrate the advantage on the balance among these three aspects compared to interleaved group convolutions and Xception, and competitive performance compared to other state-of-the-art architecture design methods. Guotian Xie, Jingdong Wang 0001, Ting Zhang 0002, Jian-Huang Lai, Richang Hong, Guo-Jun Qi |
CVPR | 6 |
| 2018 | Generalized Loss-Sensitive Adversarial Learning with Manifold Margins
Marzieh Edraki, Guo-Jun Qi |
ECCV (5) | 2 |
| 2018 | An Adversarial Approach to Hard Triplet Generation
Yiru Zhao, Zhongming Jin 0001, Guo-Jun Qi, Hongtao Lu 0001, Xian-Sheng Hua 0001 |
ECCV (9) | 3 |
| 2018 | Automated Pulmonary Nodule Detection: High Sensitivity with Few Candidates
Bin Wang 0065, Guo-Jun Qi, Sheng Tang, Liheng Zhang, Lixi Deng, Yongdong Zhang 0001 |
MICCAI (2) | 2 |
| 2018 | Multimedia Content Understanding by Learning from Very Few Examples: Recent Progress on Unsupervised, Semi-Supervised and Supervised Deep Learning ApproachesabstractIn this tutorial, the speaker will present serval parallel efforts on building deep learning models with very few supervision information, with or without unsupervised data available. In particular, we will discuss in details. (1) Generative Adverbial Nets (GANs) and their applications to unsupervised feature extractions, semi-supervised learning with few labeled examples and a large amount of unlabeled data. We will discuss the state-of-the-art results that have been achieved by the semi-supervised GANs. (2) Low-Shot Learning algorithms to train and test models on disjoint sets of tasks. We will discuss the ideas of how to efficiently adapt models to tasks with very few examples. In particular, we will discuss several paradigms of learning-to-learn approaches. (3) We will also discuss how to transfer models across modalities by leveraging abundant labels from one modality to train a model for other modalities with few labels. We will discuss in details the cross-modal label transfer approach. Guo-Jun Qi |
ICMR | 1 |
| 2018 | Previewer for Multi-Scale Object DetectorabstractMost multi-scale detectors face a challenge of small-size false positives due to the inadequacy of low-level features, which have small receptive field sizes and weak semantic capabilities. This paper demonstrates independent predictions from different feature layers on the same region is beneficial for reducing false positives. We propose a novel light-weight previewer block, which previews the objectness probability for the potential regression region of each prior box, using the stronger features with larger receptive fields and more contextual information for better predictions. This previewer block is generic and can be easily implemented in multi-scale detectors, such as SSD, RFBNet and MS-CNN. Extensive experiments are conducted on PASCAL VOC and KITTI pedestrian benchmark to show the superiority of the proposed method. Zhihang Fu, Zhongming Jin 0001, Guo-Jun Qi, Chen Shen 0003, Rongxin Jiang 0001, Yaowu Chen, Xian-Sheng Hua 0001 |
ACM Multimedia | 3 |
| 2018 | CapProNet: Deep Feature Learning via Orthogonal Projections onto Capsule SubspacesabstractIn this paper, we formalize the idea behind capsule nets of using a capsule vector rather than a neuron activation to predict the label of samples. To this end, we propose to learn a group of capsule subspaces onto which an input feature vector is projected. Then the lengths of resultant capsules are used to score the probability of belonging to different classes. We train such a Capsule Projection Network (CapProNet) by learning an orthogonal projection matrix for each capsule subspace, and show that each capsule subspace is updated until it contains input feature vectors corresponding to the associated class. With low dimensionality of capsule subspace as well as an iterative method to estimate the matrix inverse, only a small negligible computing overhead is incurred to train the network. Experiment results on image datasets show the presented network can greatly improve the performance of state-of-the-art Resnet backbones by $10-20\%$ with almost the same computing cost. Liheng Zhang, Marzieh Edraki, Guo-Jun Qi |
NeurIPS | 3 |
| 2018 | Weakly Supervised Facial Attribute Manipulation via Deep Adversarial NetworkabstractAutomatically manipulating facial attributes is challenging because it needs to modify the facial appearances, while keeping not only the person's identity but also the realism of the resultant images. Unlike the prior works on the facial attribute parsing, we aim at an inverse and more challenging problem called attribute manipulation by modifying a facial image in line with a reference facial attribute. Given a source input image and reference images with a target attribute, our goal is to generate a new image (i.e., target image) that not only possesses the new attribute but also keeps the same or similar content with the source image. In order to generate new facial attributes, we train a deep neural network with a combination of a perceptual content loss and two adversarial losses, which ensure the global consistency of the visual content while implementing the desired attributes often impacting on local pixels. The model automatically adjusts the visual attributes on facial appearances and keeps the edited images as realistic as possible. The evaluation shows that the proposed model can provide a unified solution to both local and global facial attribute manipulation such as expression change and hair style transfer. Moreover, we further demonstrate that the learned attribute discriminator can be used for attribute localization. Yilin Wang 0002, Suhang Wang, Guo-Jun Qi, Jiliang Tang, Baoxin Li |
WACV | 3 |
| 2018 | Image Classification With Tailored Fine-Grained DictionariesabstractIn this paper, we propose a novel fine-grained dictionary learning method for image classification. To learn a high-quality discriminative dictionary, three types of multispecific subdictionaries, i.e., class-specific dictionaries (CSDs), universal dictionary (UD), and family-specific dictionaries (FSDs), are simultaneously uncovered. Here, CSDs and UD, respectively, model the patterns for each class and the patterns irrespective of any class. FSDs can help reveal the shared patterns between multiple image classes, by filling the gap between the patterns in CSDs and UD. The dependence among image classes is revealed by the shared FSDs, and a common FSD can be assigned to several classes to represent their residual. Finally, the most discriminative FSD for each class is identified by minimizing the sparse reconstruction error. Extensive experiments are conducted on different widely used data sets for image classification. The results demonstrate the superior performance of the proposed method over some state-of-the-art methods. Xiangbo Shu, Jinhui Tang 0001, Guo-Jun Qi, Zechao Li, Yu-Gang Jiang 0001, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Semantic Neighbor Graph Hashing for Multimodal RetrievalabstractHashing methods have been widely used for approximate nearest neighbor search in recent years due to its computational and storage effectiveness. Most existing multimodal hashing methods try to preserve the similarity relationship based on either metric distances or semantic labels in a procrustean way, while ignoring the intra-class and inter-class variations inherent in the metric space. In this paper, we propose a novel multimodal hashing method, termed as semantic neighbor graph hashing (SNGH), which aims to preserve the fine-grained similarity metric based on the semantic graph that is constructed by jointly pursuing the semantic supervision and the local neighborhood structure. Specifically, the semantic graph is constructed to capture the local similarity structure for the image modality and the text modality, respectively. Furthermore, we define a function based on the local similarity in particular to adaptively calculate multi-level similarities by encoding the intra-class and inter-class variations. After obtaining the unified hash codes, the logistic regression with kernel trick is employed to learn view-specific hash functions independently for each modality. Extensive experiments are conducted on four widely used multimodal data sets. The experimental results demonstrate the superiority of the proposed SNGH method compared with the state-of-the-art multimodal hashing methods. Lu Jin 0001, Kai Li 0005, Hao Hu 0010, Guo-Jun Qi, Jinhui Tang 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Learning Label Preserving Binary Codes for Multimedia Retrieval: A General ApproachabstractLearning-based hashing has been researched extensively in the past few years due to its great potential in fast and accurate similarity search among huge volumes of multimedia data. In this article, we present a novel multimedia hashing framework, called Label Preserving Multimedia Hashing (LPMH) for multimedia similarity search. In LPMH, a general optimization method is used to learn the joint binary codes of multiple media types by explicitly preserving semantic label information. Compared with existing hashing methods which are typically developed under and thus restricted to some specific objective functions, the proposed optimization strategy is not tied to any specific loss function and can easily incorporate bit balance constraints to produce well-balanced binary codes. Specifically, our formulation leads to a set of Binary Integer Programming (BIP) problems that have exact solutions both with and without bit balance constraints. These problems can be solved extremely fast and the solution can easily scale up to large-scale datasets. In the hash function learning stage, the boosted decision trees algorithm is utilized to learn multiple media-specific hash functions that can map heterogeneous data sources into a homogeneous Hamming space for cross-media retrieval. We have comprehensively evaluated the proposed method using a range of large-scale datasets in both single-media and cross-media retrieval tasks. The experimental results demonstrate that LPMH is competitive with state-of-the-art methods in both speed and accuracy. Kai Li 0005, Guo-Jun Qi, Kien A. Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | CLARE: A Joint Approach to Label Classification and Tag RecommendationabstractData classification and tag recommendation are both important and challenging tasks in social media. These two tasks are often considered independently and most efforts have been made to tackle them separately. However, labels in data classification and tags in tag recommendation are inherently related. For example, a Youtube video annotated with NCAA, stadium, pac12 is likely to be labeled as football, while a video/image with the class label of coast is likely to be tagged with beach, sea, water and sand. The existence of relations between labels and tags motivates us to jointly perform classification and tag recommendation for social media data in this paper. In particular, we provide a principled way to capture the relations between labels and tags, and propose a novel framework CLARE, which fuses data CLAssification and tag REcommendation into a coherent model. With experiments on three social media datasets, we demonstrate that the proposed framework CLARE achieves superior performance on both tasks compared to the state-of-the-art methods. Yilin Wang 0002, Suhang Wang, Jiliang Tang, Guo-Jun Qi, Huan Liu 0001, Baoxin Li |
AAAI | 4 |
| 2017 | Interleaved Group ConvolutionsabstractIn this paper, we present a simple and modularized neural network architecture, named interleaved group convolutional neural networks (IGCNets). The main point lies in a novel building block, a pair of two successive interleaved group convolutions: primary group convolution and secondary group convolution. The two group convolutions are complementary: (i) the convolution on each partition in primary group convolution is a spatial convolution, while on each partition in secondary group convolution, the convolution is a point-wise convolution; (ii) the channels in the same secondary partition come from different primary partitions. We discuss one representative advantage: Wider than a regular convolution with the number of parameters and the computation complexity preserved. We also show that regular convolutions, group convolution with summation fusion, and the Xception block are special cases of interleaved group convolutions. Empirical results over standard benchmarks, CIFAR-10, CIFAR-100, SVHN and ImageNet demonstrate that our networks are more efficient in using parameters and computation complexity with similar or higher accuracy. Ting Zhang 0002, Guo-Jun Qi, Bin Xiao 0001, Jingdong Wang 0001 |
ICCV | 2 |
| 2017 | Context-patch based face hallucination via thresholding locality-constrained representation and reproducing learningabstractFace hallucination, which refers to predicting a HighResolution (HR) face image from an observed Low-Resolution (LR) one, is a challenging problem. Most state-of-the-arts employ local face structure prior to estimate the optimal representations for each patch by the training patches of the same position, and achieve good reconstruction performance. However, they do not take into account the contextual information of image patch, which is very useful for the expression of human face. Different from position-patch based methods, in this paper we leverage the contextual information and develop a robust and efficient context-patch face hallucination algorithm, called Thresholding Locality-constrained Representation with Reproducing learning (TLcR-RL). In TLcR-RL, we use a thresholding strategy to enhance the stability of patch representation and the reconstruction accuracy. Additionally, we develop a reproducing learning to iteratively enhance the estimated result by adding the estimated HR face to the training set. Experiments demonstrate that the performance of our proposed framework has a substantial increase when compared to state-of-the-arts, including recently proposed deep learning based method. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Guo-Jun Qi, Akiko Aizawa |
ICME | 5 |
| 2017 | State-Frequency Memory Recurrent Neural NetworksabstractModeling temporal sequences plays a fundamental role in various modern applications and has drawn more and more attentions in the machine learning community. Among those efforts on improving the capability to represent temporal data, the Long Short-Term Memory (LSTM) has achieved great success in many areas. Although the LSTM can capture long-range dependency in the time domain, it does not explicitly model the pattern occurrences in the frequency domain that plays an important role in tracking and predicting data points over various time cycles. We propose the State-Frequency Memory (SFM), a novel recurrent architecture that allows to separate dynamic patterns across different frequency components and their impacts on modeling the temporal contexts of input sequences. By jointly decomposing memorized dynamics into state-frequency components, the SFM is able to offer a fine-grained analysis of temporal sequences by capturing the dependency of uncovered patterns in both time and frequency domains. Evaluations on several temporal modeling tasks demonstrate the SFM can yield competitive performances, in particular as compared with the state-of-the-art LSTM models. Hao Hu 0010, Guo-Jun Qi |
ICML | 2 |
| 2017 | Mixture Factorized Ornstein-Uhlenbeck Processes for Time-Series ForecastingabstractForecasting the future observations of time-series data can be performed by modeling the trend and fluctuations from the observed data. Many classical time-series analysis models like Autoregressive model (AR) and its variants have been developed to achieve such forecasting ability. While they are often based on the white noise assumption to model the data fluctuations, a more general Brownian motion has been adopted that results in Ornstein-Uhlenbeck (OU) process. The OU process has gained huge successes in predicting the future observations over many genres of time series, however, it is still limited in modeling simple diffusion dynamics driven by a single persistent factor that never evolves over time. However, in many real problems, a mixture of hidden factors are usually present, and when and how frequently they appear or disappear are unknown ahead of time. This imposes a challenge that inspires us to develop a Mixture Factorized OU process (MFOUP) to model evolving factors. The new model is able to capture the changing states of multiple mixed hidden factors, from which we can infer their roles in driving the movements of time series. We conduct experiments on three forecasting problems, covering sensor and market data streams. The results show its competitive performance on predicting future observations and capturing evolution patterns of hidden factors as compared with the other algorithms. Guo-Jun Qi, Jiliang Tang, Jingdong Wang 0001, Jiebo Luo 0001 |
KDD | 1 |
| 2017 | Stock Price Prediction via Discovering Multi-Frequency Trading PatternsabstractStock prices are formed based on short and/or long-term commercial and trading activities that reflect different frequencies of trading patterns. However, these patterns are often elusive as they are affected by many uncertain political-economic factors in the real world, such as corporate performances, government policies, and even breaking news circulated across markets. Moreover, time series of stock prices are non-stationary and non-linear, making the prediction of future price trends much challenging. To address them, we propose a novel State Frequency Memory (SFM) recurrent network to capture the multi-frequency trading patterns from past market data to make long and short term predictions over time. Inspired by Discrete Fourier Transform (DFT), the SFM decomposes the hidden states of memory cells into multiple frequency components, each of which models a particular frequency of latent trading pattern underlying the fluctuation of stock price. Then the future stock prices are predicted as a nonlinear mapping of the combination of these components in an Inverse Fourier Transform (IFT) fashion. Modeling multi-frequency trading patterns can enable more accurate predictions for various time ranges: while a short-term prediction usually depends on high frequency trading patterns, a long-term prediction should focus more on the low frequency trading patterns targeting at long-term return. Unfortunately, no existing model explicitly distinguishes between various frequencies of trading patterns to make dynamic predictions in literature. The experiments on the real market data also demonstrate more competitive performance by the SFM as compared with the state-of-the-art methods. Liheng Zhang, Charu C. Aggarwal, Guo-Jun Qi |
KDD | 3 |
| 2017 | Linear Subspace Ranking Hashing for Cross-Modal RetrievalabstractHashing has attracted a great deal of research in recent years due to its effectiveness for the retrieval and indexing of large-scale high-dimensional multimedia data. In this paper, we propose a novel ranking-based hashing framework that maps data from different modalities into a common Hamming space where the cross-modal similarity can be measured using Hamming distance. Unlike existing cross-modal hashing algorithms where the learned hash functions are binary space partitioning functions, such as the sign and threshold function, the proposed hashing scheme takes advantage of a new class of hash functions closely related to rank correlation measures which are known to be scale-invariant, numerically stable, and highly nonlinear. Specifically, we jointly learn two groups of linear subspaces, one for each modality, so that features' ranking orders in different linear subspaces maximally preserve the cross-modal similarities. We show that the ranking-based hash function has a natural probabilistic approximation which transforms the original highly discontinuous optimization problem into one that can be efficiently solved using simple gradient descent algorithms. The proposed hashing framework is also flexible in the sense that the optimization procedures are not tied up to any specific form of loss function, which is typical for existing cross-modal hashing methods, but rather we can flexibly accommodate different loss functions with minimal changes to the learning steps. We demonstrate through extensive experiments on four widely-used real-world multimodal datasets that the proposed cross-modal hashing method can achieve competitive performance against several state-of-the-arts with only moderate training and testing time. Kai Li 0005, Guo-Jun Qi, Jun Ye 0003, Kien A. Hua |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Joint Intermodal and Intramodal Label Transfers for Extremely Rare or Unseen ClassesabstractIn this paper, we present a label transfer model from texts to images for image classification tasks. The problem of image classification is often much more challenging than text classification. On one hand, labeled text data is more widely available than the labeled images for classification tasks. On the other hand, text data tends to have natural semantic interpretability, and they are often more directly related to class labels. On the contrary, the image features are not directly related to concepts inherent in class labels. One of our goals in this paper is to develop a model for revealing the functional relationships between text and image features as to directly transfer intermodal and intramodal labels to annotate the images. This is implemented by learning a transfer function as a bridge to propagate the labels between two multimodal spaces. However, the intermodal label transfers could be undermined by blindly transferring the labels of noisy texts to annotate images. To mitigate this problem, we present an intramodal label transfer process, which complements the intermodal label transfer by transferring the image labels instead when relevant text is absent from the source corpus. In addition, we generalize the inter-modal label transfer to zero-shot learning scenario where there are only text examples available to label unseen classes of images without any positive image examples. We evaluate our algorithm on an image classification task and show the effectiveness with respect to the other compared algorithms. Guo-Jun Qi, Wei Liu 0005, Charu C. Aggarwal, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2017 | Tri-Clustered Tensor Completion for Social-Aware Image Tag RefinementabstractSocial image tag refinement, which aims to improve tag quality by automatically completing the missing tags and rectifying the noise-corrupted ones, is an essential component for social image search. Conventional approaches mainly focus on exploring the visual and tag information, without considering the user information, which often reveals important hints on the (in)correct tags of social images. Towards this end, we propose a novel tri-clustered tensor completion framework to collaboratively explore these three kinds of information to improve the performance of social image tag refinement. Specifically, the inter-relations among users, images and tags are modeled by a tensor, and the intra-relations between users, images and tags are explored by three regularizations respectively. To address the challenges of the super-sparse and large-scale tensor factorization that demands expensive computing and memory cost, we propose a novel tri-clustering method to divide the tensor into a certain number of sub-tensors by simultaneously clustering users, images and tags into a bunch of tri-clusters. And then we investigate two strategies to complete these sub-tensors by considering (in)dependence between the sub-tensors. Experimental results on a real-world social image database demonstrate the superiority of the proposed method compared with the state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Guo-Jun Qi, Zechao Li, Meng Wang 0001, Shuicheng Yan, Ramesh Jain 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2017 | Guest Editorial: Big Media Data: Understanding, Search, and MiningabstractThe paper in this special section focus on big data in the media industry. Big media data is a new research area, and has been attracting a lot of research interests in both industry and academia. This editorial, as the fourth part of this special issue, introduces four papers on fine-grained image classification, object retrieval, and applications of social media. Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal |
IEEE Trans. Big Data | 2 |
| 2017 | LEGO-MM: LEarning Structured Model by Probabilistic loGic Ontology Tree for MultiMediaabstractRecent advances in multimedia ontology have resulted in a number of concept models, e.g., large-scale concept for multimedia and Mediamill 101, which are accessible and public to other researchers. However, most current research effort still focuses on building new concepts from scratch, very few work explores the appropriate method to construct new concepts upon the existing models already in the warehouse. To address this issue, we propose a new framework in this paper, termed LEarning Structured Model by Probabilistic loGic Ontology Tree for MultiM edia (LEGO1-MM), which can seamlessly integrate both the new target training examples and the existing primitive concept models to infer the more complex concept models. LEGO-MM treats the primitive concept models as the lego toy to potentially construct an unlimited vocabulary of new concepts. Specifically, we first formulate the logic operations to be the lego connectors to combine the existing concept models hierarchically in probabilistic logic ontology trees. Then, we incorporate new target training information simultaneously to efficiently disambiguate the underlying logic tree and correct the error propagation. Extensive experiments are conducted on a large vehicle domain data set from ImageNet. The results demonstrate that LEGO-MM has significantly superior performance over the existing state-of-the-art methods, which build new concept models from scratch. Jinhui Tang 0001, Shiyu Chang, Guo-Jun Qi, Qi Tian 0001, Yong Rui, Thomas S. Huang |
IEEE Trans. Image Process. | 3 |
| 2017 | A Temporal Order Modeling Approach to Human Action Recognition from Multimodal Sensor DataabstractFrom wearable devices to depth cameras, researchers have exploited various multimodal data to recognize human actions for applications, such as video gaming, education, and healthcare. Although there many successful techniques have been presented in the literature, most current approaches have focused on statistical or local spatiotemporal features and do not explicitly explore the temporal dynamics of the sensor data. However, human action data contain rich temporal structure information that can characterize the unique underlying patterns of different action categories. From this perspective, we propose a novel temporal order modeling approach to human action recognition. Specifically, we explore subspace projections to extract the latent temporal patterns from different human action sequences. The temporal order between these patterns are compared, and the index of the pattern that appears first is used to encode the entire sequence. This process is repeated multiple times and produces a compact feature vector representing the temporal dynamics of the sequence. Human action recognition can then be efficiently solved by the nearest neighbor search based on the Hamming distance between these compact feature vectors. We further introduce a sequential optimization algorithm to learn the optimized projections that preserve the pairwise label similarity of the action sequences. Experimental results on two public human action datasets demonstrate the superior performance of the proposed technique in both accuracy and efficiency. Jun Ye 0003, Hao Hu 0010, Guo-Jun Qi, Kien A. Hua |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2016 | Hierarchically Gated Deep Networks for Semantic SegmentationabstractSemantic segmentation aims to parse the scene structure of images by annotating the labels to each pixel so that images can be segmented into different regions. While image structures usually have various scales, it is difficult to use a single scale to model the spatial contexts for all individual pixels. Multi-scale Convolutional Neural Networks (CNNs) and their variants have made striking success for modeling the global scene structure for an image. However, they are limited in labeling fine-grained local structures like pixels and patches, since spatial contexts might be blindly mixed up without appropriately customizing their scales. To address this challenge, we develop a novel paradigm of multiscale deep network to model spatial contexts surrounding different pixels at various scales. It builds multiple layers of memory cells, learning feature representations for individual pixels at their customized scales by hierarchically absorbing relevant spatial contexts via memory gates between layers. Such Hierarchically Gated Deep Networks (HGDNs) can customize a suitable scale for each pixel, thereby delivering better performance on labeling scene structures of various scales. We conduct the experiments on two datasets, and show competitive results compared with the other multi-scale deep networks on the semantic segmentation task. Guo-Jun Qi |
CVPR | 1 |
| 2016 | Supervised Quantization for Similarity SearchabstractIn this paper, we address the problem of searching for semantically similar images from a large database. We present a compact coding approach, supervised quantization. Our approach simultaneously learns feature selection that linearly transforms the database points into a low-dimensional discriminative subspace, and quantizes the data points in the transformed space. The optimization criterion is that the quantized points not only approximate the transformed points accurately, but also are semantically separable: the points belonging to a class lie in a cluster that is not overlapped with other clusters corresponding to other classes, which is formulated as a classification problem. The experiments on several standard datasets show the superiority of our approach over the state-of-the art supervised hashing and unsupervised quantization algorithms. Ting Zhang 0002, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001 |
CVPR | 3 |
| 2016 | Cross-modal hashing through ranking subspace learningabstractHashing has been widely used for approximate nearest neighbor search of high-dimensional multimedia data. In this paper, we propose a novel hash learning framework that maps high-dimensional multimodal data into a common Hamming space where the cross-modal similarity can be measured using Hamming distance. Unlike existing cross-modal hashing methods that learn hash functions in the form of numeric quantization of linear projections, the proposed hash learning algorithm encodes features' ranking properties and takes advantage of rank correlations which are known to be scale-invariant, numerically stable and highly nonlinear. Specifically, we learn two groups of subspaces jointly, one for each modality, so that the ranking orders in those subspaces maximally preserve the cross-modal similarity. Extensive experiments on realworld datasets demonstrate superiority of the proposed methods compared to state-of-the-arts. Kai Li 0005, Guo-Jun Qi, Jun Ye 0003, Kien A. Hua |
ICME | 2 |
| 2016 | Supervised Ranking Hash for Semantic Similarity SearchabstractThe era of big data has spawned unprecedented interests in developing hashing algorithms for their storage efficiency and effectiveness in fast nearest neighbor search in large-scale databases. Most of the existing hash learning algorithms focus on learning hash functions which generate binary codes by numeric quantization of some projected feature space. In this work, we propose a novel hash learning framework that encodes features' ranking orders instead of quantizing their numeric values in a number of optimal low-dimensional ranking subspaces. We formulate the ranking-based hash learning problem as the optimization of a continuous probabilistic error function using softmax approximation and present an efficient learning algorithm to solve the problem. We extensively evaluate the proposed algorithm in several datasets and demonstrate superior performance against several state-of-the-arts. Kai Li 0005, Guo-Jun Qi, Jun Ye 0003, Tuoerhongjiang Yusufu, Kien A. Hua |
ISM | 2 |
| 2016 | Temporal Order-based First-Take-All Hashing for Fast Attention-Deficit-Hyperactive-Disorder DetectionabstractAttention Deficit Hyperactive Disorder (ADHD) is one of the most common childhood disorders and can continue through adolescence and adulthood. Although the root cause of the problem still remains unknown, recent advancements in brain imaging technology reveal there exists differences between neural activities of Typically Developing Children (TDC) and ADHD subjects. Inspired by this, we propose a novel First-Take-All (FTA) hashing framework to investigate the problem of fast ADHD subjects detection through the fMRI time-series of neuron activities. By hashing time courses from regions of interests (ROIs) in the brain into fixed-size hash codes, FTA can compactly encode the temporal order differences between the neural activity patterns that are key to distinguish TDC and ADHD subjects. Such patterns can be directly learned via minimizing the training loss incurred by the generated FTA codes. By conducting similarity search on the resultant FTA codes, data-driven ADHD detection can be achieved in an efficient fashion. The experiments' results on real-world ADHD detection benchmarks demonstrate the FTA can outperform the state-of-the-art baselines using only neural activity time series without any phenotypic information. Hao Hu 0010, Joey Velez-Ginorio, Guo-Jun Qi |
KDD | 3 |
| 2016 | Large-scale supervised similarity learning in networks
Shiyu Chang, Guo-Jun Qi, Yingzhen Yang, Charu C. Aggarwal, Meng Wang 0001, Thomas S. Huang |
Knowl. Inf. Syst. | 2 |
| 2016 | Guest Editorial: Big Media Data: Understanding, Search, and MiningabstractThe two papers in this special section focus on big media data. This is a new research area, and has been attracting a lot of research interests in both industry and academia. This editorial, as the third part of this special issue, introduces two papers on video copy detection data and personalized travel sequence recommendation from social media. Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal |
IEEE Trans. Big Data | 2 |
| 2016 | Generalized Deep Transfer Networks for Knowledge Propagation in Heterogeneous DomainsabstractIn recent years, deep neural networks have been successfully applied to model visual concepts and have achieved competitive performance on many tasks. Despite their impressive performance, traditional deep networks are subjected to the decayed performance under the condition of lacking sufficient training data. This problem becomes extremely severe for deep networks trained on a very small dataset, making them overfitting by capturing nonessential or noisy information in the training set. Toward this end, we propose a novel generalized deep transfer networks (DTNs), capable of transferring label information across heterogeneous domains, textual domain to visual domain. The proposed framework has the ability to adequately mitigate the problem of insufficient training images by bringing in rich labels from the textual domain. Specifically, to share the labels between two domains, we build parameter- and representation-shared layers. They are able to generate domain-specific and shared interdomain features, making this architecture flexible and powerful in capturing complex information from different domains jointly. To evaluate the proposed method, we release a new dataset extended from NUS-WIDE at http://imag.njust.edu.cn/NUS-WIDE-128.html. Experimental results on this dataset show the superior performance of the proposed DTNs compared to existing state-of-the-art methods. Jinhui Tang 0001, Xiangbo Shu, Zechao Li, Guo-Jun Qi, Jingdong Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2015 | Sparse composite quantizationabstractThe quantization techniques have shown competitive performance in approximate nearest neighbor search. The state-of-the-art algorithm, composite quantization, takes advantage of the compositionabity, i.e., the vector approximation accuracy, as opposed to product quantization and Cartesian k-means. However, we have observed that the runtime cost of computing the distance table in composite quantization, which is used as a lookup table for fast distance computation, becomes nonnegligible in real applications, e.g., reordering the candidates retrieved from the inverted index when handling very large scale databases. To address this problem, we develop a novel approach, called sparse composite quantization, which constructs sparse dictionaries. The benefit is that the distance evaluation between the query and the dictionary element (a sparse vector) is accelerated using the efficient sparse vector operation, and thus the cost of distance table computation is reduced a lot. Experiment results on large scale ANN retrieval tasks (1M SIFTs and 1B SIFTs) and applications to object retrieval show that the proposed approach yields competitive performance: superior search accuracy to product quantization and Cartesian k-means with almost the same computing cost, and much faster ANN search than composite quantization with the same level of accuracy. Ting Zhang 0002, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001 |
CVPR | 2 |
| 2015 | Differential Recurrent Neural Networks for Action RecognitionabstractThe long short-term memory (LSTM) neural network is capable of processing complex sequential information since it utilizes special gating schemes for learning representations from long input sequences. It has the potential to model any time-series or sequential data, where the current hidden state has to be considered in the context of the past hidden states. This property makes LSTM an ideal choice to learn the complex dynamics of various actions. Unfortunately, the conventional LSTMs do not consider the impact of spatio-temporal dynamics corresponding to the given salient motion patterns, when they gate the information that ought to be memorized through time. To address this problem, we propose a differential gating scheme for the LSTM neural network, which emphasizes on the change in information gain caused by the salient motions between the successive frames. This change in information gain is quantified by Derivative of States (DoS), and thus the proposed LSTM model is termed as differential Recurrent Neural Network (dRNN). We demonstrate the effectiveness of the proposed model by automatically recognizing actions from the real-world 2D and 3D human action datasets. Our study is one of the first works towards demonstrating the potential of learning complex time-series representations via high-order derivatives of states. Vivek Veeriah, Naifan Zhuang, Guo-Jun Qi |
ICCV | 3 |
| 2015 | Heterogeneous Network Embedding via Deep ArchitecturesabstractData embedding is used in many machine learning applications to create low-dimensional feature representations, which preserves the structure of data points in their original space. In this paper, we examine the scenario of a heterogeneous network with nodes and content of various types. Such networks are notoriously difficult to mine because of the bewildering combination of heterogeneous contents and structures. The creation of a multidimensional embedding of such data opens the door to the use of a wide variety of off-the-shelf mining techniques for multidimensional data. Despite the importance of this problem, limited efforts have been made on embedding a network of scalable, dynamic and heterogeneous data. In such cases, both the content and linkage structure provide important cues for creating a unified feature representation of the underlying network. In this paper, we design a deep embedding algorithm for networked data. A highly nonlinear multi-layered embedding function is used to capture the complex interactions between the heterogeneous data in a network. Our goal is to create a multi-resolution deep embedding function, that reflects both the local and global network structures, and makes the resulting embedding useful for a variety of data mining tasks. In particular, we demonstrate that the rich content and linkage information in a heterogeneous network can be captured by such an approach, so that similarities among cross-modal data can be measured directly in a common embedding space. Once this goal has been achieved, a wide variety of data mining problems can be solved by applying off-the-shelf algorithms designed for handling vector representations. Our experiments on real-world network datasets show the effectiveness and scalability of the proposed algorithm as compared to the state-of-the-art embedding methods. Shiyu Chang, Wei Han 0002, Jiliang Tang, Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
KDD | 4 |
| 2015 | State-Driven Dynamic Sensor Selection and Prediction with State-Stacked SparsenessabstractAn important problem in large-scale sensor mining is that of selecting relevant sensors for prediction purposes. Selecting small subsets of sensors, also referred to as active sensors, often leads to lower operational costs, and it reduces the noise and information overload for prediction. Existing sensor selection and prediction models either select a set of sensors a priori, or they use adaptive algorithms to determine the most relevant sensors for prediction. Sensor data sets often show dynamically varying patterns, because of which it is suboptimal to select a fixed subset of active sensors. To address this problem, we develop a novel dynamic prediction model that uses the notion of hidden system states to dynamically select a varying subset of sensors. These hidden system states are automatically learned by our model in a data-driven manner. The proposed algorithm can rapidly switch between different sets of active sensors when the model detects the (periodic or intermittent) change in the system state. We derive the dynamic sensor selection strategy by minimizing the error rates in tracking and predicting sensor readings over time. We introduce the notion of state-stacked sparseness to select a subset of the most critical sensors as a function of evolving system state. We present experimental results on two real sensor datasets, corresponding to oil drilling rig sensors and intensive care unit (ICU) sensors, and demonstrate the superiority of our approach with respect to other models. Guo-Jun Qi, Charu C. Aggarwal, Deepak S. Turaga, Daby M. Sow, Phil Anno |
KDD | 1 |
| 2015 | Deep Learning Architecture with Dynamically Programmed Layers for Brain Connectome PredictionabstractThis paper explores the idea of using deep neural network architecture with dynamically programmed layers for brain connectome prediction problem. Understanding the brain connectome structure is a very interesting and a challenging problem. It is critical in the research for epilepsy and other neuropathological diseases. We introduce a new deep learning architecture that exploits the spatial and temporal nature of the neuronal activation data. The architecture consists of a combination of Convolutional layer and a Recurrent layer for predicting the connectome of neurons based on their time-series of activation data. The key contribution of this paper is a dynamically programmed layer that is critical in determining the alignment between the neuronal activations of pair-wise combinations of neurons. Vivek Veeriah, Rohit Durvasula, Guo-Jun Qi |
KDD | 3 |
| 2015 | Temporal Order-Preserving Dynamic Quantization for Human Action Recognition from Multimodal Sensor StreamsabstractRecent commodity depth cameras have been widely used in the applications of video games, business, surveillance and have dramatically changed the way of human-computer interaction. They provide rich multimodal information that can be used to interpret the human-centric environment. However, it is still of great challenge to model the temporal dynamics of the human actions and great potential can be exploited to further enhance the retrieval accuracy by adequately modeling the patterns of these actions. To address this challenge, we propose a temporal order-preserving dynamic quantization method to extract the most discriminative patterns of the action sequence. We further present a multimodal feature fusion method that can be derived in this dynamic quantization framework to exploit different discriminative capability of features from multiple modalities. Experiments based on three public human action datasets show that the proposed technique has achieved state-of-the-art performance. Jun Ye 0003, Kai Li 0005, Guo-Jun Qi, Kien A. Hua |
ICMR | 3 |
| 2015 | Weakly-Shared Deep Transfer Networks for Heterogeneous-Domain Knowledge PropagationabstractIn recent years, deep networks have been successfully applied to model image concepts and achieved competitive performance on many data sets. In spite of impressive performance, the conventional deep networks can be subjected to the decayed performance if we have insufficient training examples. This problem becomes extremely severe for deep networks with powerful representation structure, making them prone to over fitting by capturing nonessential or noisy information in a small data set. In this paper, to address this challenge, we will develop a novel deep network structure, capable of transferring labeling information across heterogeneous domains, especially from text domain to image domain. This weakly-shared Deep Transfer Networks (DTNs) can adequately mitigate the problem of insufficient image training data by bringing in rich labels from the text domain. Xiangbo Shu, Guo-Jun Qi, Jinhui Tang 0001, Jingdong Wang 0001 |
ACM Multimedia | 2 |
| 2015 | Guest Editorial: Big Media Data: Understanding, Search, and MiningabstractThe papers in this special section focus on the topic of Big Data. There has been an explosive increase in media data, such as images, videos and social media in the Internet, mobile devices, and desktops. Accordingly, the area of big media data has been attracting a lot of research interest from both academia and industry. Big media data offers a promising possibility for in-depth media understanding, as well as exploring the very big scale media data to bridge the well-known semantic gap between high-level semantics and low-level features. It provides richer information, ranging from social relations to context information associated to rich media data of diverse modalities. It also provides us the opportunity to mine reliable and helpful knowledge from big media data for a wide variety of applications. Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal |
IEEE Trans. Big Data | 2 |
| 2015 | Guest Editorial: Big Media Data: Understanding, Search, and Mining (Part 2)abstractThe articles in this special section focus on Big Data. In the first part of this special issue, we have introduced three papers on large scale similar image search, image search quality improvement, and semi-supervised multi-label image annotation. This second part of this special issue includes two examples on large scale visual retrieval. The contributions cover two themes: visual semantic relationship learning and large scale indexing. Jingdong Wang 0001, Guo-Jun Qi, Nicu Sebe, Charu C. Aggarwal |
IEEE Trans. Big Data | 2 |
| 2015 | Breaking the Barrier to Transferring Link Information across NetworksabstractLink prediction is one of the most fundamental problems in graph modeling and mining. It has been studied in a wide range of scenarios, from uncovering missing links between different entities in databases, to recommending relations between people in social networks. In this problem, we wish to predict unseen links in a growing target network by exploiting existing structures in source networks. Most of the existing methods often assume that abundant links are available in the target network to build a model for link prediction. However, in many scenarios, the target network may be too sparse to enable robust inference process, which makes link prediction challenging with the paucity of link data. On the other hand, in many cases, other (more densely linked) auxiliary networks can be available that contains similar link structure relevant to that in the target network. The linkage information in the existing networks can be used in conjunction with the node attribute information in both networks in order to make more accurate link recommendations. Thus, this paper proposes the use of learning methods to perform link inference by transferring the link information from the source network to the target network. We also note that the source network may contain the link information irrelevant to the target network. This leads to cross-network bias between the networks, which makes the link model built upon the source network misaligned with the link structure of the target network. Therefore, we re-sample the source network to rectify such cross-network bias by maximizing the cross-network relevance measured by the node attributes, as well as preserving as rich link information as possible to avoid the loss of source link structure caused by the re-sampling algorithm. The link model based on the re-sampled source network can make more accurate link predictions on the target network with aligned link structures across the networks. We present experimental results illustrating the effectiveness of the approach. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2015 | Guest Editorial: Deep Learning for Multimedia ComputingabstractThe twenty papers in this special section aim at providing a forum to present recent advancements in deep learning research that directly concerns the multimedia community. Specifically, deep learning has successfully designed algorithms that can build deep nonlinear representations to mimic how the brain perceives and understands multimodal information, ranging from low-level signals like images and audios, to high-level semantic data like natural language. For multimedia research, it is especially important to develop deep networks to capture the dependencies between different genres of data, building joint deep representation for diverse modalities. Guo-Jun Qi, Hugo Larochelle, Benoit Huet, Jiebo Luo 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Factorized Similarity Learning in NetworksabstractThe problem of similarity learning is relevant to many data mining applications, such as recommender systems, classification, and retrieval. This problem is particularly challenging in the context of networks, which contain different aspects such as the topological structure, content, and user supervision. These different aspects need to be combined effectively, in order to create a holistic similarity function. In particular, while most similarity learning methods in networks such as Sim Rank utilize the topological structure, the user supervision and content are rarely considered. In this paper, a Factorized Similarity Learning (FSL) is proposed to integrate the link, node content, and user supervision into an uniform framework. This is learned by using matrix factorization, and the final similarities are approximated by the span of low rank matrices. The proposed framework is further extended to a noise-tolerant version by adopting a hinge-loss alternatively. To facilitate efficient computation on large scale data, a parallel extension is developed. Experiments are conducted on the DBLP and CoRA datasets. The results show that FSL is robust, efficient, and outperforms the state-of-the-art. Shiyu Chang, Guo-Jun Qi, Charu C. Aggarwal, Meng Wang 0001, Thomas S. Huang |
ICDM | 2 |
| 2014 | Social media mining and knowledge discovery
Jinhui Tang 0001, Dacheng Tao, Guo-Jun Qi, Benoit Huet |
Multim. Syst. | 3 |
| 2014 | Multi-Label Image Categorization With Sparse Factor RepresentationabstractThe goal of multilabel classification is to reveal the underlying label correlations to boost the accuracy of classification tasks. Most of the existing multilabel classifiers attempt to exhaustively explore dependency between correlated labels. It increases the risk of involving unnecessary label dependencies, which are detrimental to classification performance. Actually, not all the label correlations are indispensable to multilabel model. Negligible or fragile label correlations cannot be generalized well to the testing data, especially if there exists label correlation discrepancy between training and testing sets. To minimize such negative effect in the multilabel model, we propose to learn a sparse structure of label dependency. The underlying philosophy is that as long as the multilabel dependency cannot be well explained, the principle of parsimony should be applied to the modeling process of the label correlations. The obtained sparse label dependency structure discards the outlying correlations between labels, which makes the learned model more generalizable to future samples. Experiments on real world data sets show the competitive results compared with existing algorithms. Fuming Sun, Jinhui Tang 0001, Guo-Jun Qi, Thomas S. Huang |
IEEE Trans. Image Process. | 4 |
| 2013 | Link prediction across networks by biased cross-network samplingabstractThe problem of link inference has been widely studied in a variety of social networking scenarios. In this problem, we wish to predict future links in a growing network with the use of the existing network structure. However, most of the existing methods work well only if a significant number of links are already available in the network for the inference process. In many scenarios, the existing network may be too sparse, and may have too few links to enable meaningful learning mechanisms. This paucity of linkage information can be challenging for the link inference problem. However, in many cases, other (more densely linked) networks may be available which show similar linkage structure in terms of underlying attribute information in the nodes. The linkage information in the existing networks can be used in conjunction with the node attribute information in both networks in order to make meaningful link recommendations. Thus, this paper introduces the use of transfer learning methods for performing cross-network link inference. We present experimental results illustrating the effectiveness of the approach. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
ICDE | 1 |
| 2013 | Multimedia LEGO: Learning Structured Model by Probabilistic Logic Ontology TreeabstractRecent advances in Multimedia research have generated a large collection of concept models, e.g., LSCOM and Media mill 101, which become accessible to other researchers. While most current research effort still focuses on building new concepts from scratch, little effort has been made on constructing new concepts upon the existing models already in the warehouse. To address this issue, we develop a new framework in this paper, termed LEGO, to seamlessly integrate both the new target training examples and the existing primitive concept models. LEGO treats the primitive concept models as a lego toy to potentially construct an unlimited vocabulary of new concepts. Specifically, LEGO first formulates the logic operations to be the lego connectors to combine existing concept models hierarchically in probabilistic logic ontology trees. LEGO then simultaneously incorporates new target training information to efficiently disambiguate the underlying logic tree and correct the error propagation. We present extensive experimental results on a large vehicle domain data set from Image Net, and demonstrate significantly superior performance over existing state-of-the-art approaches which build new concept models from scratch. Shiyu Chang, Guo-Jun Qi, Jinhui Tang 0001, Qi Tian 0001, Yong Rui, Thomas S. Huang |
ICDM | 2 |
| 2013 | Online community detection in social sensingabstractThe proliferation of location and GPS data streams which are collected in a wide variety of participatory sensing applications has created numerous possibilities for analysis of the underlying patterns of activity. Typically, the spatio-temporal patterns arising from such activity can be analyzed in order to determine the latent community structure in the underlying data. In this paper, we will examine the problem of online community detection from the location data collected from such social sensing applications in real time. Such data brings numerous challenges associated with it, in that they can be of a relatively large scale, and can be extremely noisy from the perspective of both data representation and analysis. Furthermore, the community structure in the underlying data cannot be directly inferred from the shape of the underlying trajectories, since a considerable amount of variation may exist in terms of trajectories of individuals belonging to the same community. In this paper, we will design online algorithms for community detection in social sensing applications. Our algorithm uses a robust and efficiently updateable model with the use of Gibbs sampling, and we will show its effectiveness and efficiency for social sensing applications. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
WSDM | 1 |
| 2013 | Mining collective intelligence in diverse groupsabstractCollective intelligence, which aggregates the shared information from large crowds, is often negatively impacted by unreliable information sources with the low quality data. This becomes a barrier to the effective use of collective intelligence in a variety of applications. In order to address this issue, we propose a probabilistic model to jointly assess the reliability of sources and find the true data. We observe that different sources are often not independent of each other. Instead, sources are prone to be mutually influenced, which makes them dependent when sharing information with each other. High dependency between sources makes collective intelligence vulnerable to the overuse of redundant (and possibly incorrect) information from the dependent sources. Thus, we reveal the latent group structure among dependent sources, and aggregate the information at the group level rather than from individual sources directly. This can prevent the collective intelligence from being inappropriately dominated by dependent sources. We will also explicitly reveal the reliability of groups, and minimize the negative impacts of unreliable groups. Experimental results on real-world data sets show the effectiveness of the proposed approach with respect to existing algorithms. Guo-Jun Qi, Charu C. Aggarwal, Jiawei Han 0001, Thomas S. Huang |
WWW | 1 |
| 2013 | Cross-Space Affinity Learning with Its Application to Movie RecommendationabstractIn this paper, we propose a novel cross-space affinity learning algorithm over different spaces with heterogeneous structures. Unlike most of affinity learning algorithms on the homogeneous space, we construct a cross-space tensor model to learn the affinity measures on heterogeneous spaces subject to a set of order constraints from the training pool. We further enhance the model with a factorization form which greatly reduces the number of parameters of the model with a controlled complexity. Moreover, from the practical perspective, we show the proposed factorized cross-space tensor model can be efficiently optimized by a series of simple quadratic optimization problems in an iterative manner. The proposed cross-space affinity learning algorithm can be applied to many real-world problems, which involve multiple heterogeneous data objects defined over different spaces. In this paper, we apply it into the recommendation system to measure the affinity between users and the product items, where a higher affinity means a higher rating of the user on the product. For an empirical evaluation, a widely used benchmark movie recommendation data set—MovieLens—is used to compare the proposed algorithm with other state-of-the-art recommendation algorithms and we show that very competitive results can be obtained. Jinhui Tang 0001, Guo-Jun Qi, Liyan Zhang 0001, Changsheng Xu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2012 | Community Detection with Edge Content in Social Media NetworksabstractThe problem of community detection in social media has been widely studied in the social networking community in the context of the structure of the underlying graphs. Most community detection algorithms use the links between the nodes in order to determine the dense regions in the graph. These dense regions are the communities of social media in the graph. Such methods are typically based purely on the linkage structure of the underlying social media network. However, in many recent applications, edge content is available in order to provide better supervision to the community detection process. Many natural representations of edges in social interactions such as shared images and videos, user tags and comments are naturally associated with content on the edges. While some work has been done on utilizing node content for community detection, the presence of edge content presents unprecedented opportunities and flexibility for the community detection process. We will show that such edge content can be leveraged in order to greatly improve the effectiveness of the community detection process in social media networks. We present experimental results illustrating the effectiveness of our approach. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
ICDE | 1 |
| 2012 | Transfer Learning of Distance Metrics by Cross-Domain Metric Sampling across Heterogeneous SpacesabstractThe problem of transfer learning has recently been of great interest in a variety of machine learning applications. In this paper, we examine a new angle to the transfer learning problem, where we examine the problem of distance function learning. Specifically, we focus on the problem of how our knowledge of distance functions in one domain can be transferred to a new domain. A good semantic understanding of the feature space is critical in providing the domain specific understanding for setting up good distance functions. Unfortunately, not all domains have feature representations which are equally interpretable. For example, in some domains such as text, the semantics of the feature representation are clear, as a result of which it is easy for a domain expert to set up distance functions for specific kinds of semantics. In the case of image data, the features are semantically harder to interpret, and it is harder to set up distance functions, especially for particular semantic criteria. In this paper, we focus on the problem of transfer learning as a way to close the semantic gap between different domains, and show how to use correspondence information between two domains in order to set up distance functions for the semantically more challenging domain. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
SDM | 1 |
| 2012 | On clustering heterogeneous social media objects with outlier linksabstractThe clustering of social media objects provides intrinsic understanding of the similarity relationships between documents, images, and their contextual sources. Both content and link structure provide important cues for an effective clustering algorithm of the underlying objects. While link information provides useful hints for improving the clustering process, it also contains a significant amount of noisy information. Therefore, a robust clustering algorithm is required to reduce the impact of noisy links. In order to address the aforementioned problems, we propose heterogeneous random fields to model the structure and content of social media networks. We design a probability measure on the social media networks which output a configuration of clusters that are consistent with both content and link structure. Furthermore, noisy links can also be detected, and their impact on the clustering algorithm can be significantly reduced. We conduct experiments on a real social media network and show the advantage of the method over other state-of-the-art algorithms. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
WSDM | 1 |
| 2012 | Recommending Flickr groups with social topic model
Jingdong Wang 0001, Zhe Zhao 0001, Jiazhen Zhou, Hao Wang 0005, Bin Cui 0001, Guo-Jun Qi |
Inf. Retr. | 6 |
| 2012 | Exploring Context and Content Links in Social Media: A Latent Space MethodabstractSocial media networks contain both content and context-specific information. Most existing methods work with either of the two for the purpose of multimedia mining and retrieval. In reality, both content and context information are rich sources of information for mining, and the full power of mining and processing algorithms can be realized only with the use of a combination of the two. This paper proposes a new algorithm which mines both context and content links in social media networks to discover the underlying latent semantic space. This mapping of the multimedia objects into latent feature vectors enables the use of any off-the-shelf multimedia retrieval algorithms. Compared to the state-of-the-art latent methods in multimedia analysis, this algorithm effectively solves the problem of sparse context links by mining the geometric structure underlying the content links between multimedia objects. Specifically for multimedia annotation, we show that an effective algorithm can be developed to directly construct annotation models by simultaneously leveraging both context and content information based on latent structure between correlated semantic concepts. We conduct experiments on the Flickr data set, which contains user tags linked with images. We illustrate the advantages of our approach over the state-of-the-art multimedia retrieval techniques. Guo-Jun Qi, Charu C. Aggarwal, Qi Tian 0001, Heng Ji 0001, Thomas S. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | Web-Scale Multimedia Information NetworksabstractThe abundance of multimedia data on the Web presents both challenges (how to annotate, search, and mine) and opportunities (crawling the Web to create large structured multimedia data bases which can be used to do inference effectively). Because of the huge data volume, considering all semantic concepts as on the same (flat) level is not viable. In this paper, we introduce a unified STRUCTURED representation called multimedia information networks (MINets), which incorporates ontology and cross-media links, covering both content and context knowledge. Ontology and cross-media structures are constructed and expanded by automatically constructing MINets from web-scale data by state-of-the-art information extraction and knowledge-based population techniques. The resultant MINet will contain a wide range of linkages, including logical, statistical, and semantic relations among informative concept nodes, which connects proliferative ontology as well as cross-media web-scale resources together. The raw data collected in construction phase often contain much noisy, incomplete, or even conflicting information which could be detrimental to information extraction and utilization. Then, the redundant link structure can be utilized to distill MINets and improve quality of information (QoI). Moreover, advanced inference theory and system can be built upon the linked MINets, and then high-level ontological knowledge can be inferred and integrated in a logically harmonious network structure in MINets which is consistent with human cognition. Even more, as information channels, the ontology and cross-media links in MINets connect informative knowledge resources together, which makes it possible to increase the portability of information between different resources to increase information utilization levels. Guo-Jun Qi, Min-Hsuan Tsai, Shen-Fu Tsai, Liangliang Cao, Thomas S. Huang |
Proc. IEEE | 1 |
| 2011 | Towards cross-category knowledge propagation for learning visual conceptsabstractIn recent years, knowledge transfer algorithms have become one of most the active research areas in learning visual concepts. Most of the existing learning algorithms focuses on leveraging the knowledge transfer process which is specific to a given category. However, in many cases, such a process may not be very effective when a particular target category has very few samples. In such cases, it is interesting to examine, whether it is feasible to use cross-category knowledge for improving the learning process by exploring the knowledge in correlated categories. Such a task can be quite challenging due to variations in semantic similarities and differences between categories, which could either help or hinder the cross-category learning process. In order to address this challenge, we develop a cross-category label propagation algorithm, which can directly propagate the inter-category knowledge at instance level between the source and the target categories. Furthermore, this algorithm can automatically detect conditions under which the transfer process can be detrimental to the learning process. This provides us a way to know when the transfer of cross-category knowledge is both useful and desirable. We present experimental results on real image and video data sets in order to demonstrate the effectiveness of our approach. Guo-Jun Qi, Charu C. Aggarwal, Yong Rui, Qi Tian 0001, Shiyu Chang, Thomas S. Huang |
CVPR | 1 |
| 2011 | Locality-sensitive support vector machine by exploring local correlation and global regularizationabstractLocal classifiers have obtained great success in classification task due to its powerful discriminating ability on local regions. However, most of them still have restricted generalization in twofold: (1) each local classifier is sensitive to noise in local regions which leads to overfitting phenomenon in local classifiers; (2) the local classifiers also ignore the local correlation determined by the sample distribution in each local region. To overcome the above two problems, we present a novel locality-sensitive support vector machine (LSSVM) in this paper for image retrieval problem. This classifier applies locality-sensitive hashing (LSH) to divide the whole feature space into a number of local regions, on each of them a local model can be better constructed due to smaller within-class variation on it. To avoid these local models from overfitting into locality-sensitive structures, it imposes a global regularizer across local regions so that local classifiers are smoothly glued together to form a regularized overall classifier. local correlation is modeled to capture the sample distribution that determines the locality structure of each local region, which can increase the discriminating ability of the algorithm. To evaluate the performance, we apply the proposed algorithm into image retrieval task and competitive results are obtained on the real-world web image data set. Guo-Jun Qi, Qi Tian 0001, Thomas S. Huang |
CVPR | 1 |
| 2011 | PhotoNet: A similarity-aware image delivery service for situation awareness
Md. Yusuf Sarwar Uddin, Guo-Jun Qi, Tom Huang, Tarek F. Abdelzaher, Guohong Cao |
IPSN | 3 |
| 2011 | PhotoNet: A Similarity-Aware Picture Delivery Service for Situation AwarenessabstractWe propose PhotoNet, a picture delivery service for camera sensor networks. PhotoNet is motivated by the needs of disaster-response applications, where a group of survivors and first responders may survey damage and send images to a rescue center in the absence of a functional communication infrastructure. The protocol runs on mobile devices, handling opportunistic forwarding (when they come in contact) and in-network storage. It assigns priorities to images for forwarding and replacement depending on the degree of similarity (or dissimilarity) among them, such that scarce resources are assigned to delivery of most ``deserving'' content first. Prioritization aims at reducing semantic redundancy such as that between pictures of the same scene at the same location taken from slightly different angles. This is in contrast to redundancy among identical objects and among time series data. PhotoNet delivers more diverse pictures in terms of event coverage suppressing logically redundant content belonging to the same event. We show that, in resource constrained networks, reducing semantic redundancy can significantly improve the utility of the service. Md. Yusuf Sarwar Uddin, Fatemeh Saremi, Guo-Jun Qi, Tarek F. Abdelzaher, Thomas S. Huang |
RTSS | 4 |
| 2011 | Towards semantic knowledge propagation from text corpus to web imagesabstractIn this paper, we study the problem of transfer learning from text to images in the context of network data in which link based bridges are available to transfer the knowledge between the different domains. The problem of classification of image data is often much more challenging than text data because of the following two reasons: (a) Labeled text data is very widely available for classification purposes. On the other hand, this is often not the case for image data, in which a lot of images are available from many sources, but many of them are often not labeled. (b) The image features are not directly related to semantic concepts inherent in class labels. On the other hand, since text data tends to have natural semantic interpretability (because of their human origins), they are often more directly related to class labels. Therefore, the relationships between the images and text features also provide additional hints for the classification process in terms of the image feature transformations which provide the most effective results. Guo-Jun Qi, Charu C. Aggarwal, Thomas S. Huang |
WWW | 1 |
| 2011 | Image annotation by kNN-sparse graph-based label propagation over noisily tagged web imagesabstractIn this article, we exploit the problem of annotating a large-scale image corpus by label propagation over noisily tagged web images. To annotate the images more accurately, we propose a novel k NN-sparse graph-based semi-supervised learning approach for harnessing the labeled and unlabeled data simultaneously. The sparse graph constructed by datum-wise one-vs- k NN sparse reconstructions of all samples can remove most of the semantically unrelated links among the data, and thus it is more robust and discriminative than the conventional graphs. Meanwhile, we apply the approximate k nearest neighbors to accelerate the sparse graph construction without loosing its effectiveness. More importantly, we propose an effective training label refinement strategy within this graph-based learning framework to handle the noise in the training labels, by bringing in a dual regularization for both the quantity and sparsity of the noise. We conduct extensive experiments on a real-world image database consisting of 55,615 Flickr images and noisily tagged training labels. The results demonstrate both the effectiveness and efficiency of the proposed approach and its capability to deal with the noise in the training labels. Jinhui Tang 0001, Richang Hong, Shuicheng Yan, Tat-Seng Chua, Guo-Jun Qi, Ramesh Jain 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2010 | Image Classification With Kernelized Spatial-ContextabstractThe goal of image classification is to classify a collection of unlabeled images into a set of semantic classes. Many methods have been proposed to approach this goal by leveraging visual appearances of local patches in images. However, the spatial context between these local patches also provides significant information to improve the classification accuracy. Traditional spatial contextual models, such as two-dimensional hidden Markov model, attempt to construct one common model for each image category to depict the spatial structures of the images in this class. However due to large intra-class variances in an image category, one single model has difficulties in representing various spatial contexts in different images. In contrast, we propose to construct a prototype set of spatial contextual models by leveraging the kernel methods rather than only one model. Such an algorithm combines the advantages of rich representation ability of spatial contextual models as well as the powerful classification ability of kernel method. In particular, we propose a new distance measure between different spatial contextual models by integrating joint appearance-spatial image features. Such a distance measure can be efficiently computed in a recursive formulation that scales well to image size. Extensive experiments demonstrate that the proposed approach significantly outperforms the state-of-the-art approaches. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, HongJiang Zhang |
IEEE Trans. Multim. | 1 |
| 2010 | Image Annotation by Graph-Based Inference With Integrated Multiple/Single Instance RepresentationsabstractIn most of the learning-based image annotation approaches, images are represented using multiple-instance (local) or single-instance (global) features. Their performances, however, are mixed as for certain concepts, the single-instance representations of images are more suitable, while for others, the multiple-instance representations are better. Thus this paper explores a unified learning framework that combines the multiple-instance and single-instance representations for image annotation. More specifically, we propose an integrated graph-based semi-supervised learning framework to utilize these two types of representations simultaneously. We further explore three strategies to convert from multiple-instance representation into a single-instance one. Experiments conducted on the COREL image dataset demonstrate the effectiveness and efficiency of the proposed integrated framework and the conversion strategies. Jinhui Tang 0001, Guo-Jun Qi, Tat-Seng Chua |
IEEE Trans. Multim. | 3 |
| 2009 | An efficient sparse metric learning in high-dimensional space via l1-penalized log-determinant regularizationabstractThis paper proposes an efficient sparse metric learning algorithm in high dimensional space via an l1-penalized log-determinant regularization. Compare to the most existing distance metric learning algorithms, the proposed algorithm exploits the sparsity nature underlying the intrinsic high dimensional feature space. This sparsity prior of learning distance metric serves to regularize the complexity of the distance model especially in the "less example number p and high dimension d" setting. Theoretically, by analogy to the covariance estimation problem, we find the proposed distance learning algorithm has a consistent result at rate O (√m2 log d)/n) to the target distance matrix with at most m nonzeros per row. Moreover, from the implementation perspective, this l1-penalized log-determinant formulation can be efficiently optimized in a block coordinate descent fashion which is much faster than the standard semi-definite programming which has been widely adopted in many other advanced distance learning algorithms. We compare this algorithm with other state-of-the-art ones on various datasets and competitive results are obtained. Guo-Jun Qi, Jinhui Tang 0001, Zhengjun Zha, Tat-Seng Chua, HongJiang Zhang |
ICML | 1 |
| 2009 | Learning semantic distance from community-tagged media collectionabstractThis paper proposes a novel semantic-aware distance met-ric for images by mining multimedia data on the Internet, in particular, web images and their associated tags. As well known, a proper distance metric between images is a key ingredient in many realistic web image retrieval engines, as well many image understanding techniques. In this pa-per, we attempt to mine a novel distance metric from the web images by integrating their visual content as well as the associated user tags. Different from many existing dis-tance metric learning algorithms which utilize the dissimilar or similar information between images pixels or features in signal level, the proposed scheme also takes the associated user-input tags into consideration. The visual content of images is also leveraged to respect an intuitive assumption that the visual similar images ought to have a smaller dis-tance. A semi-definite programming is formulated to encode the above two aspects of criteria to learn the distance met-ric and we show such an optimization problem can be effi-ciently solved with a closed-form solution. We evaluate the proposed algorithm on two datasets. One is the benchmark Corel dataset and the other is a real-world dataset crawled from the image sharing website Flickr. By comparison with other existing distance learning algorithms, competitive re-sults are obtained by the proposed algorithm in experiments. Guo-Jun Qi, Xian-Sheng Hua 0001, HongJiang Zhang |
ACM Multimedia | 1 |
| 2009 | Inferring semantic concepts from community-contributed images and noisy tagsabstract10.1145/1631272.1631305 Jinhui Tang 0001, Shuicheng Yan, Richang Hong, Guo-Jun Qi, Tat-Seng Chua |
ACM Multimedia | 4 |
| 2009 | Semi-supervised kernel density estimation for video annotation
Meng Wang 0001, Xian-Sheng Hua 0001, Tao Mei 0001, Richang Hong, Guo-Jun Qi, Yan Song 0001, Li-Rong Dai 0001 |
Comput. Vis. Image Underst. | 5 |
| 2009 | Two-Dimensional Multilabel Active Learning with an Efficient Online Adaptation Model for Image ClassificationabstractConventional active learning dynamically constructs the training set only along the sample dimension. While this is the right strategy in binary classification, it is suboptimal for multilabel image classification. We argue that for each selected sample, only some effective labels need to be annotated while others can be inferred by exploring the label correlations. The reason is that the contributions of different labels to minimizing the classification error are different due to the inherent label correlations. To this end, we propose to select sample-label pairs, rather than only samples, to minimize a multilabel Bayesian classification error bound. We call it two-dimensional active learning because it considers both the sample dimension and the label dimension. Furthermore, as the number of training samples increases rapidly over time due to active learning, it becomes intractable for the offline learner to retrain a new model on the whole training set. So we develop an efficient online learner to adapt the existing model with the new one by minimizing their model distance under a set of multilabel constraints. The effectiveness and efficiency of the proposed method are evaluated on two benchmark data sets and a realistic image collection from a real-world image sharing Web site-Corbis. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, HongJiang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Video semantic analysis based on structure-sensitive anisotropic manifold ranking
Jinhui Tang 0001, Guo-Jun Qi, Meng Wang 0001, Xian-Sheng Hua 0001 |
Signal Process. | 2 |
| 2009 | Unified Video Annotation via Multigraph LearningabstractLearning-based video annotation is a promising approach to facilitating video retrieval and it can avoid the intensive labor costs of pure manual annotation. But it frequently encounters several difficulties, such as insufficiency of training data and the curse of dimensionality. In this paper, we propose a method named optimized multigraph-based semi-supervised learning (OMG-SSL), which aims to simultaneously tackle these difficulties in a unified scheme. We show that various crucial factors in video annotation, including multiple modalities, multiple distance functions, and temporal consistency, all correspond to different relationships among video units, and hence they can be represented by different graphs. Therefore, these factors can be simultaneously dealt with by learning with multiple graphs, namely, the proposed OMG-SSL approach. Different from the existing graph-based semi-supervised learning methods that only utilize one graph, OMG-SSL integrates multiple graphs into a regularization framework in order to sufficiently explore their complementation. We show that this scheme is equivalent to first fusing multiple graphs and then conducting semi-supervised learning on the fused graph. Through an optimization approach, it is able to assign suitable weights to the graphs. Furthermore, we show that the proposed method can be implemented through a computationally efficient iterative process. Extensive experiments on the TREC video retrieval evaluation (TRECVID) benchmark have demonstrated the effectiveness and efficiency of our proposed approach. Meng Wang 0001, Xian-Sheng Hua 0001, Richang Hong, Jinhui Tang 0001, Guo-Jun Qi, Yan Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2009 | Correlative Linear Neighborhood Propagation for Video AnnotationabstractRecently, graph-based semisupervised learning methods have been widely applied in multimedia research area. However, for the application of video semantic annotation in multilabel setting, these methods neglect an important characteristic of video data: The semantic concepts appear correlatively and interact naturally with each other rather than exist in isolation. In this paper, we adapt this semantic correlation into graph-based semisupervised learning and propose a novel method named correlative linear neighborhood propagation to improve annotation performance. Experiments conducted on the Text REtrieval Conference VIDeo retrieval evaluation data set have demonstrated its effectiveness and efficiency. Jinhui Tang 0001, Xian-Sheng Hua 0001, Meng Wang 0001, Zhiwei Gu, Guo-Jun Qi, Xiuqing Wu |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2008 | Two-Dimensional Active Learning for image classificationabstractIn this paper, we propose a two-dimensional active learning scheme and show its application in image classification. Traditional active learning methods select samples only along the sample dimension. While this is the right strategy in binary classification, it is sub-optimal for multi-label classification. In multi-label classification, we argue that, for each selected sample, only a part of more effective labels are necessary to be annotated while others can be inferred by exploring the correlations among the labels. The reason is that the contributions of different labels to minimizing the classification error are different due to the inherent label correlations. To this end, we propose to select sample-label pairs, rather than only samples, to minimize a multi-label Bayesian classification error bound. This new active learning strategy not only considers the sample dimension but also the label dimension, and we call it Two-Dimensional Active Learning (2DAL). We also show that the traditional active learning formulation is a special case of 2DAL when there is only one label. Extensive experiments conducted on two real-world applications show that the 2DAL significantly outperforms the best existing approaches which did not take label correlation into account. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, HongJiang Zhang |
CVPR | 1 |
| 2008 | A joint appearance-spatial distance for kernel-based image categorizationabstractThe goal of image categorization is to classify a collection of unlabeled images into a set of predefined classes to support semantic-level image retrieval. The distance measures used in most existing approaches either ignored the spatial structures or used them in a separate step. As a result, these distance measures achieved only limited success. To address these difficulties, in this paper, we propose a new distance measure that integrates joint appearance-spatial image features. Such a distance measure is computed as an upper bound of an information-theoretic discrimination, and can be computed efficiently in a recursive formulation that scales well to image size. In addition, the upper bound approximation can be further tightened via adaption learning from a universal reference model. Extensive experiments on two widely-used data sets show that the proposed approach significantly outperforms the state-of-the-art approaches. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, Zhengjun Zha, HongJiang Zhang |
CVPR | 1 |
| 2008 | Joint multi-label multi-instance learning for image classificationabstractIn real world, an image is usually associated with multiple labels which are characterized by different regions in the image. Thus image classification is naturally posed as both a multi-label learning and multi-instance learning problem. Different from existing research which has considered these two problems separately, we propose an integrated multi-label multi-instance learning (MLMIL) approach based on hidden conditional random fields (HCRFs), which simultaneously captures both the connections between semantic labels and regions, and the correlations among the labels in a single formulation. We apply this MLMIL framework to image classification and report superior performance compared to key existing approaches over the MSR Cambridge (MSRC) and Corel data sets. Zhengjun Zha, Xian-Sheng Hua 0001, Tao Mei 0001, Jingdong Wang 0001, Guo-Jun Qi, Zengfu Wang |
CVPR | 5 |
| 2008 | Query-independent learning for video searchabstractMost of existing learning-based methods for query-by-example take the query examples as ldquopositiverdquo and build a model for each query. These methods, referred to as query-dependent, only achieved limited success as they can hardly be applied to real-world applications, in which an arbitrary query is usually given. To address this problem, we propose to learn a query-independent model by exploiting the relevance information which exists in the pair of query-document. The proposed approach takes a query-document pair as a sample and extracts a set of query-independent textual and visual features from each pair. It is general and suitable for a real-world video search system since the learned relevance relation is independent on any query. We conducted extensive experiments over TRECVID 2005-2007 corpus and shown superior performance (+37% in Mean Average Precision) to the query-dependent learning approaches. Yuan Liu 0017, Tao Mei 0001, Guo-Jun Qi, Xiuqing Wu, Xian-Sheng Hua 0001 |
ICME | 3 |
| 2008 | Online multi-label active annotation: towards large-scale content-based video searchabstractExisting video search engines have not taken the advantages of video content analysis and semantic understanding. Video search in academia uses semantic annotation to approach content-based indexing. We argue this is a promising direction to enable real content-based video search. However, due to the complexity of both video data and semantic concepts, existing techniques on automatic video annotation are still not able to handle large-scale video set and large-scale concept set, in terms of both annotation accuracy and computation cost. To address this problem, in this paper, we propose a scalable framework for annotation-based video search, as well as a novel approach to enable large-scale semantic concept annotation, that is, online multi-label active learning. This framework is scalable to both the video sample dimension and concept label dimension. Large-scale unlabeled video samples are assumed to arrive consecutively in batches with an initial pre-labeled training set, based on which a preliminary multi-label classifier is built. For each arrived batch, a multi-label active learning engine is applied, which automatically selects and manually annotates a set of unlabeled sample-label pairs. And then an online learner updates the original classifier by taking the newly labeled sample-label pairs into consideration. This process repeats until all data are arrived. During the process, new labels, even without any pre-labeled training samples, can be incorporated into the process anytime. Experiments on TRECVID dataset demonstrate the effectiveness and efficiency of the proposed framework. Xian-Sheng Hua 0001, Guo-Jun Qi |
ACM Multimedia | 2 |
| 2008 | Integrated graph-based semi-supervised multiple/single instance learning framework for image annotationabstractRecently, many learning methods based on multiple-instance (local) or single-instance (global) representations of images have been proposed for image annotation. Their performances on image annotation, however, are mixed as for certain concepts the single-instance representations of images are more suitable, while for some other concepts the multiple-instance representations are better. Thus in this paper, we explore an unified learning framework that combines the multiple-instance and single-instance representations for image annotation. More specifically, we propose an integrated graph-based semi-supervised learning framework to utilize these two types of representations simultaneously, and explore an effective and computationally efficient strategy to convert the multiple-instance representation into a single-instance one. Experiments conducted on the Coral image dataset show the effectiveness and efficiency of the proposed integrated framework. Jinhui Tang 0001, Guo-Jun Qi, Tat-Seng Chua |
ACM Multimedia | 3 |
| 2008 | Video Annotation Based on Kernel Linear Neighborhood PropagationabstractThe insufficiency of labeled training data for representing the distribution of the entire dataset is a major obstacle in automatic semantic annotation of large-scale video database. Semi-supervised learning algorithms, which attempt to learn from both labeled and unlabeled data, are promising to solve this problem. In this paper, a novel graph-based semi-supervised learning method namedkernellinearneighborhoodpropagation(KLNP) is proposed and applied to video annotation. This approach combines theconsistencyassumption, which is the basic assumption in semi-supervised learning, and thelocallinearembedding(LLE) method in a nonlinear kernel-mapped space. KLNP improves a recently proposed methodlinearneighborhoodpropagation(LNP) by tackling the limitation of its local linear assumption on the distribution of semantics. Experiments conducted on the TRECVID data set demonstrate that this approach outperforms other popular graph-based semi-supervised learning methods for video semantic annotation. Jinhui Tang 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Yan Song 0001, Xiuqing Wu |
IEEE Trans. Multim. | 3 |
| 2008 | Correlative multilabel video annotation with temporal kernelsabstractAutomatic video annotation is an important ingredient for semantic-level video browsing, search and navigation. Much attention has been paid to this topic in recent years. These researches have evolved through two paradigms. In the first paradigm, each concept is individually annotated by a pre-trained binary classifier. However, this method ignores the rich information between the video concepts and only achieves limited success. Evolved from the first paradigm, the methods in the second paradigm add an extra step on the top of the first individual classifiers to fuse the multiple detections of the concepts. However, the performance of these methods can be degraded by the error propagation incurred in the first step to the second fusion one. In this article, another paradigm of the video annotation method is proposed to address these problems. It simultaneously annotates the concepts as well as model correlations between them in one step by the proposed Correlative Multilabel (CML) method, which benefits from the compensation of complementary information between different labels. Furthermore, since the video clips are composed by temporally ordered frame sequences, we extend the proposed method to exploit the rich temporal information in the videos. Specifically, a temporal-kernel is incorporated into the CML method based on the discriminative information between Hidden Markov Models (HMMs) that are learned from the videos. We compare the performance between the proposed approach and the state-of-the-art approaches in the first and second paradigms on the widely used TRECVID data set. As to be shown, superior performance of the proposed method is gained. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, Tao Mei 0001, Meng Wang 0001, HongJiang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2007 | Concurrent Multiple Instance Learning for Image CategorizationabstractWe propose a new multiple instance learning (MIL) algorithm to learn image categories. Unlike existing MIL algorithms, in which the individual instances in a bag are assumed to be independent with each other, we develop concurrent tensors to explicitly model the inter-dependency between the instances to better capture image's inherent semantics. Rank-1 tensor factorization is then applied to obtain the label of each instance. Furthermore, we formulate the classification problem in the reproducing kernel Hilbert space (RKHS) to extend instance label prediction to the whole feature space. Finally, a regularizer is introduced, which avoids overfitting and significantly improves learning machine's generalization capability, similar to that in SVMs. We report superior categorization performances compared with key existing approaches on both the COREL and the Caltech datasets. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Tao Mei 0001, Jinhui Tang 0001, HongJiang Zhang |
CVPR | 1 |
| 2007 | Temporally Consistent Gaussian Random Field for Video Semantic AnalysisabstractAs a major family of semi-supervised learning, graph based semi-supervised learning methods have attracted lots of interests in the machine learning community as well as many application areas recently. However, for the application of video semantic annotation, these methods only consider the relations among samples in the feature space and neglect an intrinsic property of video data: the temporally adjacent video segments (e.g., shots) usually have similar semantic concept. In this paper, we adapt this temporal consistency property of video data into graph based semi-supervised learning and propose a novel method named temporally consistent Gaussian random field (TCGRF) to improve the annotation results. Experiments conducted on the TREC VID data set have demonstrated its effectiveness. Jinhui Tang 0001, Xian-Sheng Hua 0001, Tao Mei 0001, Guo-Jun Qi, Shipeng Li 0001, Xiuqing Wu |
ICIP (4) | 4 |
| 2007 | Transductive Inference with Hierarchical Clustering for Video AnnotationabstractIn this paper, we present a novel framework for video semantic detection based on transductive inference and hierarchical clustering, which directly focuses on predicting the available samples in a current unlabeled pool, instead of trying to build a classifier workable for any unavailable data. In this framework, a number of hierarchical clustering results are constructed from theentire video datasetcontaining both labeled and unlabeled examples. We aim to make the clusters aspureas possible, i.e., samples in a same cluster mostly have a same label. To furtherpurifythese hierarchical clustering results, an EM based cluster-tuning algorithm is iteratively employed. Based on these clustering results, several hypotheses are generated byprobability votingamong labeled samples in the obtained clusters. From these hypotheses, one of them is chosen according to theVapnik combined bound, and it is then applied to predict the labels of unlabeled samples. This selected transductive hypothesis, which is only interested in predicting the available unlabeled samples in test set rather than producing a general classifier like inductive inference learning, exploits the structure and distribution of the unlabeled pool to achieve a minimal test error bound. Thus it can have better generalization ability for video annotation both theoretically and experimentally. This is also shown by our experiment results. Guo-Jun Qi, Xian-Sheng Hua 0001, Yan Song 0001, Jinhui Tang 0001, HongJiang Zhang |
ICME | 1 |
| 2007 | Beyond Accuracy: Typicality Ranking for Video AnnotationabstractIn this paper, we address the issue of typicality ranking for video annotation and propose to use a novel criterion, average typicality precision (ATP), to replace the frequently used one, average precision (AP), for evaluating the performance of video annotation algorithms. General annotation methods just care the number of true-positive samples at the top of the ranked list; they actually do not care the order of these samples. We argue that it is more reasonable to rank "typical" true-positive samples higher than non-typical ones, which can be evaluated by our proposed ATP. However, generally the labels of the training data only differentiate true from false; that is to say, typical or non-typical training samples have the same contribution to the learning process. Therefore, the labels of the unlabeled data learned from these training data can not well measure the typicality. In this paper, we relax the labels of the training data to real-valued typicality scores by a pre-processing stage, which is accomplished by three approaches, including density estimation, user feedback and active learning. Then the typicality scores of the training data are propagated to unlabeled data using manifold-ranking. Experiments conducted on the TRECVID data set demonstrate that this typicality ranking scheme is more consistent with human perception than normal accuracy based ranking schemes. Jinhui Tang 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Zhiwei Gu, Xiuqing Wu |
ICME | 3 |
| 2007 | Anisotropic Manifold Ranking for Video AnnotationabstractGraph-based semi-supervised learning (SSL) has attracted lots of interests in machine learning community as well as many application areas including video annotation recently. However, one of the two basic assumptions, structure assumption, which is an essential point of graph-based SSL, is not embedded into the pairwise similarity measure. Accordingly, we propose a novel graph-based SSL method for video annotation, named anisotropic manifold ranking (AniMR), based on a structure-related similarity measure. This method takes the influence of the density difference between samples into account to improve the pairwise similarity. Furthermore, we will show that AniMR can also be deduced from partial differential equation (PDE) based anisotropic diffusion. It demonstrates that the label propagation in AniMR is anisotropic, which is intrinsically different from the isotropic label propagation process in general graph-based SSL methods. Experiments conducted on the TRECVID data set show this approach outperforms ordinary graph-based SSL methods and is effective for video semantic annotation. Jinhui Tang 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Tao Mei 0001, Xiuqing Wu |
ICME | 3 |
| 2007 | A Novel Multiple Instance Learning Approach for Image Retrieval Based on Adaboost Feature SelectionabstractIn image retrieval, the concepts are usually in region-level but annotated in image-level, which leads to a major difficulty in learning the target concepts. In this paper, we formulate region-based image retrieval as a multiple-instance learning (MIL) problem, and propose an efficient and effective algorithm, named MI-AdaBoost, to solve it. The algorithm firstly maps each bag into a new bag feature space using a certain set of instance prototypes, and then adopts AdaBoost to select the bag features and build classifiers simultaneously. Experiments on both COREL and MUSK datasets show the proposed scheme is much more efficient than some typical existing MIL algorithms while has comparable results. Xun Yuan 0001, Xian-Sheng Hua 0001, Meng Wang 0001, Guo-Jun Qi, Xiuqing Wu |
ICME | 4 |
| 2007 | Correlative multi-label video annotationabstractAutomatically annotating concepts for video is a key to semantic-level video browsing, search and navigation. The research on this topic evolved through two paradigms. The first paradigm used binary classification to detect each individual concept in a concept set. It achieved only limited success, as it did not model the inherent correlation between concepts, e.g., urban and building. The second paradigm added a second step on top of the individual concept detectors to fuse multiple concepts. However, its performance varies because the errors incurred in the first detection step can propagate to the second fusion step and therefore degrade the overall performance. To address the above issues, we propose a third paradigm which simultaneously classifies concepts and models correlations between them in a single step by using a novel Correlative Multi-Label (CML) framework. We compare the performance between our proposed approach and the state-of-the-art approaches in the first and second paradigms on the widely used TRECVID data set. We report superior performance from the proposed approach. Guo-Jun Qi, Xian-Sheng Hua 0001, Yong Rui, Jinhui Tang 0001, Tao Mei 0001, HongJiang Zhang |
ACM Multimedia | 1 |
| 2007 | Typicality ranking via semi-supervised multiple-instance learningabstractMost of the existing methods for natural scene categorization only consider whether a sample is relevant or irrelevant to a particular concept. However, for the samples relevant to a certain concept, their typicalities or relevancy scores to the concept generally are different. Typicality measure should be taken into account to make the categorization results more consistent with human's perception. In this paper, we propose a novel typicality ranking scheme for categorizing natural scenes through a two-stage semi-supervised multiple-instance learning method. The first stage infers the typicalities of the underlying positive instances (i.e., regions in images) in the training dataset and the second one predicts the typicality of each bag (i.e., image) in a semi-supervised manner. Compared to existing typicality ranking approaches, the main advantages of the proposed method lie in twofold. First, it only needs image-level labels instead of region-level ones in the training stage. Second, it is fully automated and no human feedback is required. Experiments conducted on a COREL image dataset demonstrate the effectiveness of the proposed approach. Jinhui Tang 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Xiuqing Wu |
ACM Multimedia | 3 |
| 2007 | Structure-sensitive manifold ranking for video concept detectionabstractPairwise similarity of samples is an essential factor in graph propagation based semi-supervised learning methods. Usually it is estimated based on Euclidean distance. However, the structural assumption, which is a basic assumption in these methods, has not been taken into consideration in the normal pairwise similarity measure. In this paper, we propose a novel graph-based learning approach, named Structure-Sensitive Manifold Ranking (SSMR),based on a structure-sensitive similarity measure. Instead of using distance only, SSMR takes local distribution differences into account to more accurately measure pairwise similarity. Furthermore, we show that SSMR can also be deduced from a partial differential equation based anisotropic diffusion. Experiments conducted on the TRECVID dataset show that this approach significantly outperforms existing graph-based semi-supervised learning methods for video semantic concept detection. Jinhui Tang 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Meng Wang 0001, Tao Mei 0001, Xiuqing Wu |
ACM Multimedia | 3 |
| 2007 | Refining video annotation by exploiting pairwise concurrent relationabstractVideo annotation is a promising and essential step for content-based video search and retrieval. Most of the state-of-the-art video annotation approaches detect multiple semantic concepts in an isolated manner, which neglect the fact that video concepts are usually correlated in semantic nature. In this paper, we propose to refine video annotation by leveraging the pairwise concurrent relation among video concepts. Such concurrent relation is explicitly modeled by a concurrent matrix and then a propagation strategy is adopted to refine the annotations. Through spreading the scores of all related concepts to each other iteratively, the detection results approach stable and optimal. In contrast with existing concept fusion methods, the proposed approach is computationally more efficient and easy to implement, not requiring to construct any contextual model. Furthermore, we show its intuitive connection with the PageRank algorithm. We conduct the experiments on TRECVID 2005 corpus and report superior performance compared to existing key approaches. Zhengjun Zha, Tao Mei 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Zengfu Wang |
ACM Multimedia | 4 |
| 2007 | Kernel-Based Linear Neighborhood Propagation for Semantic Video Annotation
Jinhui Tang 0001, Xian-Sheng Hua 0001, Yan Song 0001, Guo-Jun Qi, Xiuqing Wu |
PAKDD | 4 |
| 2006 | Video Annotation by Active Learning and Semi-Supervised EnsemblingabstractSupervised and semi-supervised learning are frequently applied methods to annotate videos by mapping low-level features into semantic concepts. Due to the large semantic gap, the main constraint of these methods is that the information contained in a limited-size labeled dataset can hardly represent the distributions of the semantic concepts. In this paper, we propose a novel semi-automatic video annotation framework, active learning with semi-supervised ensembling, which tries to tackle the disadvantages of current video annotation solutions. Firstly the initial training set is constructed based on distribution analysis of the entire video dataset and then an active learning scheme is combined into a semi-supervised ensembling framework, which selects the samples to maximize the margin of the ensemble classifier based on both labeled and unlabeled data. Experimental results show that the proposed method performs superior to general semi-supervised learning algorithms and typical active learning algorithms in terms of annotation accuracy and stability Yan Song 0001, Guo-Jun Qi, Xian-Sheng Hua 0001, Li-Rong Dai 0001, Renhua Wang |
ICME | 2 |