EDBT 2026 Demo / reviewers in the wild / expert
Fang Zhao 0006
dblp:72/4898-6
· DBLP profile ↗
45ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-6772-8042ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 15 since 2021Artificial intelligence and machine learning · 29 · 9 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Foundation Model for Skeleton-Based Human Action UnderstandingabstractHuman action understanding serves as a foundational pillar in the field of intelligent motion perception.Skeletons serve as a modality- and device-agnostic representation for human modeling, and skeleton-based action understanding has potential applications in humanoid robot control and interaction. However, existing works often lack the scalability and generalization required to handle diverse action understanding tasks. There is no skeleton foundation model that can be adapted to a wide range of action understanding tasks. This paper presents a Unified Skeleton-based Dense Representation Learning (USDRL) framework, which serves as a foundational model for skeleton-based human action understanding. USDRL consists of a Transformer-based Dense Spatio-Temporal Encoder (DSTE), Multi-Grained Feature Decorrelation (MG-FD), and Multi-Perspective Consistency Training (MPCT). The DSTE module adopts two parallel streams to learn temporal dynamic and spatial structure features. The MG-FD module collaboratively performs feature decorrelation across temporal, spatial, and instance domains to reduce dimensional redundancy and enhance information extraction. The MPCT module employs both multi-view and multi-modal self-supervised consistency training. The former enhances the learning of high-level semantics and mitigates the impact of low-level discrepancies, while the latter effectively facilitates the learning of informative multimodal features. We perform extensive experiments on 25 benchmarks across across 9 skeleton-based action understanding tasks, covering coarse prediction, dense prediction, and transferred prediction. Our approach significantly outperforms the current state-of-the-art methods. We hope that this work would broaden the scope of research in skeleton-based action understanding and encourage more attention to dense prediction tasks. Hongsong Wang 0001, Wanjiang Weng, Junbo Wang 0003, Fang Zhao 0006, Guosen Xie, Xin Geng 0001, Liang Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | MambaPTP: Exploring the Potential of Mamba for Pedestrian Trajectory PredictionabstractPedestrian Trajectory Prediction (PTP) aims to predict the future trajectory of pedestrians based on a historical trajectory. Transformer-based approaches have demonstrated unparalleled performance for PTP tasks, encoding long-term temporal dependencies and heterogeneous spatial interactions of pedestrians. However, Transformer often involves redundant information and noisy interactions from irrelevant regions by considering all available trajectory features. Recently, the structured state space model, Mamba has been proposed, which captures long-range dependency in sequences with a selective mechanism to filter out redundant information. To further tap into the potential of the novel Mamba architecture for the PTP task, in this paper, we presentMambaPTP, which predicts future trajectories based purely on Mamba mechanisms, to mitigate the noisy interactions of irrelevant trajectory features and avoid repetitive trajectory modeling, while maintaining high-performance trajectory prediction. Specifically, we propose a new Bidirectional Gating Mamba (BGM) module with bidirectional state space models, which leverages the sparse gate mechanism to select informative temporal patterns and spatial interactions. Moreover, we design a Bidirectional Trajectory Alignment (BTA) module towards aligning the predicted trajectory to the ground truth, ensuring that the model to learn the effective sparse feature representation of trajectories. We conduct extensive experiments on several mainstream pedestrian trajectory prediction datasets. The results demonstrate that the proposed MambaPTP achieves competitive performance compared to advanced Transformer-based models. We hope this paper can further inspire research in Mamba for the PTP task, leading to a tighter integration of the Mamba and PTP communities. Shuangqing Zhang, Gangming Zhao, Fan Lyu, Songping Wang, Zhang Zhang 0001, Fang Zhao 0006, Caifeng Shan, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Through the Looking Glass: A Dual Perspective on Weakly Supervised Few-Shot SegmentationabstractMeta-learning aims to uniformly sample homologous support-query pairs, characterized by the same categories and similar attributes, and extract useful inductive biases through identical network architectures. However, this identical network design results in over-semantic homogenization. To address this, we propose a novel homologous but heterogeneous network. By treating support-query pairs as dual perspectives, we introduce heterogeneous visual aggregation (HA) modules to enhance complementarity while preserving semantic commonality. To further reduce semantic noise and amplify the uniqueness of heterogeneous semantics, we design a heterogeneous transport (HT) module. Finally, we propose heterogeneous CLIP (HC) textual information to enhance the generalization capability of multimodal models. In the weakly-supervised few-shot semantic segmentation (WFSS) task, with only 1/24 of the parameters of existing state-of-the-art models, TLG achieves a 13.2% improvement on Pascal- $5{^{\text {i}}}$ and a 7.9% improvement on COCO- $20{^{\text {i}}}$ . To the best of our knowledge, TLG is also the first weakly-supervised (image-level) model that outperforms fully supervised (pixel-level) models under the same backbone architectures. The code is available at https://github.com/jarch-ma/TLG. Jiaqi Ma 0006, Guosen Xie, Fang Zhao 0006, Zechao Li |
IEEE Trans. Image Process. | 3 |
| 2025 | Kernel-Aware Graph Prompt Learning for Few-Shot Anomaly DetectionabstractFew-shot anomaly detection (FSAD) aims to detect unseen anomaly regions with the guidance of very few normal support images from the same class. Existing FSAD methods usually find anomalies by directly designing complex text prompts to align them with visual features under the prevailing large vision-language model paradigm. However, these methods, almost always, neglect intrinsic contextual information in visual features, e.g., the interaction relationships between different vision layers, which is an important clue for detecting anomalies comprehensively. To this end, we propose a kernel-aware graph prompt learning framework, termed as KAG-prompt, by reasoning the cross-layer relations among visual features for FSAD. Specifically, a kernel-aware hierarchical graph is built by taking the different layer features focusing on anomalous regions of different sizes as nodes, meanwhile, the relationships between arbitrary pairs of nodes stand for the edges of the graph. By message passing over this graph, KAG-prompt can capture cross-layer contextual information, thus leading to more accurate anomaly prediction. Moreover, to integrate the information of multiple important anomaly signals in the prediction map, we propose a novel image-level scoring method based on multi-level information fusion. Extensive experiments on MVTecAD and VisA datasets show that KAG-prompt achieves state-of-the-art FSAD results for image-level/pixel-level anomaly detection. Fenfang Tao, Guosen Xie, Fang Zhao 0006, Xiangbo Shu |
AAAI | 3 |
| 2025 | Graph Interaction Prompt Network for Few-Shot Medical Image Anomaly DetectionabstractFew-shot medical image anomaly detection aims to detect and locate anomalies with limited data, playing a crucial role in clinical practice. In recent years, the large pre-trained vision-language model CLIP has demonstrated impressive performance across a variety of few-shot and zero-shot downstream tasks. However, CLIP mainly focuses on aligning text and images, emphasizing the semantics of global foreground objects rather than distinguishing local subtle normal or abnormal areas in the images. To address this challenge, we propose the Graph Interaction Prompt Network (GIPN), a framework that leverages graph interaction and text-prompt learning for precise anomaly detection in medical images. Specifically, we introduce a graph interaction prompt module that enables cross-layer interactions among visual features in a latent graph space, guiding the model to focus on challenging anomalous regions and enhancing feature representations. Additionally, we develop a dual-stream fusion strategy, which merges the hierarchical graph interaction features with the original vision-language features, better capturing critical cues for anomaly prediction under the guidance of text prompts. Extensive experiments on three medical anomaly datasets demonstrate that GIPN outperforms the current state-of-the-art few-shot medical anomaly detection approaches. Our code is available at https://github.com/CVL-hub/GIPN. Fenfang Tao, Tian-Zhu Xiang, Fang Zhao 0006, Guosen Xie |
BIBM | 4 |
| 2025 | Region-Aware Compositional Context Prompting for Zero-Shot Anomaly DetectionabstractZero-shot anomaly detection (ZSAD) aims to identify anomalies of unseen classes without requiring samples from those classes. Existing methods typically rely on pre-trained visual language models, such as CLIP, to detect anomalies by designing or learning generic text prompts and computing similarities with image features, which often fail to address the complexity and novelty of anomaly patterns, especially when the target domain exhibits significant differences from the source domain. To address the problems, we propose Region-aware Compositional Context Prompting (ReCo-CoP) for ZSAD, which dynamically generates contextual prompts by integrating both global and local visual information. Specifically, we introduce a Compositional Context Prompting (CCP) module that incorporates global visual features into the context through a set of basis vectors shared among images, and a Regional Context Prompting (RCP) module that optimizes the context based on image patch features, thereby enhancing the model’s ability to perceive local abnormal regions. Additionally, we combine dynamically generated prompts with static generic prompts to prevent the model from losing the essential general knowledge. Extensive experiments on 12 datasets from industrial and medical domains demonstrate the superior zero-shot detection performance of our model. The code is available at https://github.com/WenDongyp/ReCoCoP Guanglei Chu, Guosen Xie, Caifeng Shan, Fang Zhao 0006 |
ECAI | 6 |
| 2025 | AnomalyControl: Highly-Aligned Anomalous Image Generation with Controlled Diffusion ModelabstractIn industrial scenarios, diverse anomalous images are difficult to acquire, significantly limiting the performance of industrial anomaly detection methods. Automatically generating anomalous images for anomaly detection has the potential to solve the above problem. However, existing anomaly generation models are still not satisfactory regarding the authenticity and controllability of anomaly generation. In this paper, we propose a controlled anomaly generation model named AnomalyControl to generate realistic anomalous images aligned highly with both text prompts and anomaly masks. First, we introduce a CLIP-guided anomaly prompt generator that leverages a CLIP text encoder to find anomaly text prompts most aligned with real anomalous images. Secondly, we propose an anomaly appearance and shape decoupling mechanism, which designs an embedding similarity loss to enforce the alignment between the anomaly text prompt and anomalies generated with different shapes at the same location, making the appearance of generated anomalies better maintain semantic consistency when the anomaly shape changes. Then, a training-free local control enhancement strategy is employed to provide stronger control intensity to anomaly regions during inference for finer alignment with anomaly masks. Finally, a hard sample generation module is proposed to create anomalous samples with subtle shapes and imperceptible anomaly appearances, enabling the downstream anomaly detection model to focus on learning low-saliency anomaly features. Extensive experiments demonstrate that anomalous images generated by our model outperform the state-of-the-art anomaly generation methods in terms of authenticity and consistency, and can significantly improve the performance of downstream anomaly detection tasks, especially anomaly localization. Yuanyi Duan, Qinlong Wu, Guosen Xie, Fang Zhao 0006, Caifeng Shan |
ACM Multimedia | 5 |
| 2025 | Multi-Granularity Aggregation Network for Remote Sensing Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) aims to segment a query image using a limited number of densely annotated support images from the same category. Most existing conventional FSS methods are tailored for coping with images from natural scenarios. Unlike natural images, remote sensing images usually have a similar background context among the support and query image pairs, and more severe intraclass inconsistency exists due to overhead shooting views. However, facing such realistic and challenging remote sensing FSS tasks, the existing methods seldom consider these intrinsic characteristics from a unified viewpoint, thus leading to inferior results. To solve the above dilemma, we propose a multi-granularity aggregation network (MGANet) to progressively capture multi-granularity discriminative information, for tackling the remote sensing FSS task. Specifically, MGANet consists of a multi-granularity similarity (MGS) module and an adaptive multiprototype aggregation (AMPA) module. To fully utilize background context, MGS extracts multi-granularity support and query feature maps from the backbone network to calculate a holistic correlation by incorporating the background information. Next, to alleviate the intraclass inconsistency of remote sensing images, AMPA decomposes the support foreground region into mainstay and auxiliary subregions by the guidance of reverse prediction on support features, thus generating three types of prototypes by masked average pooling (MAP) on these paired features and masks. Furthermore, these multiprototypes are collaboratively interacted with the query features to pursue reinforced discriminative features, relying on prototype-aware slot attention (PASA). Extensive experiments on iSAID-$5^{i}$and LoveDA-$2^{i}$demonstrate well the superiority of the proposed MGANet. The source code is available athttps://github.com/CVL-hub/MGANet/. Shi-Feng Peng, Guosen Xie, Fang Zhao 0006, Xiangbo Shu, Qingshan Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Uncertainty-Aware Transformer for Referring Camouflaged Object DetectionabstractReferring camouflaged object detection (Ref-COD) is a recently proposed task, aiming to segment specified camouflaged objects by leveraging visual reference, i.e., a small set of referring images with salient target objects. Ref-COD poses a considerable challenge due to the difficulty of discerning camouflaged objects from their highly similar backgrounds, as well as the significant feature differences between the camouflaged objects and the provided visual reference. To tackle the above dilemma, we propose a novel uncertainty-aware transformer for the Ref-COD task, termed UAT. UAT first utilizes a cross-attention mechanism to align and integrate visual reference to guide camouflaged feature learning, and then models dependencies between patches in a probabilistic manner to learn predictive uncertainty and excavate discriminative camouflaged features. Specifically, we first design a referring feature aggregation (RFA) module to align and incorporate referring features with camouflaged features, guiding targeted specific feature learning within the feature space of camouflaged images. Then, to enhance multi-level feature extraction, we develop a cross-attention encoder (CAE) to integrate global information and multi-scale semantics between adjacent layers to excavate critical camouflage cues. More importantly, we propose a transformer probabilistic decoder (TPD) to model the dependencies between patches as Gaussian random variables to capture uncertainty-aware camouflaged features. Extensive experiments on the golden Ref-COD benchmark demonstrate the superiority of UAT over existing state-of-the-art competitors. The proposed UAT also achieves competitive performance on several conventional COD datasets, further demonstrating its scalability. The source code is available at https://github.com/CVL-hub/UAT. Ranwan Wu, Tian-Zhu Xiang, Guosen Xie, Rongrong Gao, Xiangbo Shu, Fang Zhao 0006, Ling Shao 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | AFANet: Adaptive Frequency-Aware Network for Weakly-Supervised Few-Shot Semantic SegmentationabstractFew-shot learning aims to recognize novel concepts by leveraging prior knowledge learned from a few samples. However, for visually intensive tasks such as few-shot semantic segmentation, pixel-level annotations are time-consuming and costly. Therefore, in this paper, we utilize the more challenging image-level annotations and propose an adaptive frequency-aware network (AFANet) for weakly-supervised few-shot semantic segmentation (WFSS). Specifically, we first propose a cross-granularity frequency-aware module (CFM) that decouples RGB images into high-frequency and low-frequency distributions and further optimizes semantic structural information by realigning them. Unlike most existing WFSS methods using the textual information from the multi-modal language-vision model, e.g., CLIP, in an offline learning manner, we further propose a CLIP-guided spatial-adapter module (CSM), which performs spatial domain adaptive transformation on textual information through online learning, thus providing enriched cross-modal semantic information for CFM. Extensive experiments on the Pascal-5iand COCO-20idatasets demonstrate that AFANet has achieved state-of-the-art performance. Jiaqi Ma 0006, Guosen Xie, Fang Zhao 0006, Zechao Li |
IEEE Trans. Multim. | 3 |
| 2025 | Biphasic Face Photo-Sketch Synthesis via Semantic-Driven Generative Adversarial Network With Graph Representation LearningabstractBiphasic face photo-sketch synthesis has significant practical value in wide-ranging fields such as digital entertainment and law enforcement. Previous approaches directly generate the photo-sketch in a global view, they always suffer from the low quality of sketches and complex photograph variations, leading to unnatural and low-fidelity results. In this article, we propose a novel semantic-driven generative adversarial network to address the above issues, cooperating with graph representation learning. Considering that human faces have distinct spatial structures, we first inject class-wise semantic layouts into the generator to provide style-based spatial information for synthesized face photographs and sketches. In addition, to enhance the authenticity of details in generated faces, we construct two types of representational graphs via semantic parsing maps upon input faces, dubbed the intraclass semantic graph (IASG) and the interclass structure graph (IRSG). Specifically, the IASG effectively models the intraclass semantic correlations of each facial semantic component, thus producing realistic facial details. To preserve the generated faces being more structure-coordinated, the IRSG models interclass structural relations among every facial component by graph representation learning. To further enhance the perceptual quality of synthesized images, we present a biphasic interactive cycle training strategy by fully taking advantage of the multilevel feature consistency between the photograph and sketch. Extensive experiments demonstrate that our method outperforms the state-of-the-art competitors on the CUHK Face Sketch (CUFS) and CUHK Face Sketch FERET (CUFSF) datasets. Xingqun Qi, Muyi Sun, Zijian Wang 0009, Jiaming Liu 0003, Qi Li 0005, Fang Zhao 0006, Shanghang Zhang, Caifeng Shan |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Attribute Prompt Alignment Network for Zero-Shot LearningabstractIn the vanilla zero-shot learning (ZSL) paradigm, category attributes is the key for knowledge generalizable transfer from seen to unseen classes. By contrast, the current contrastive language-image pretraining (CLIP) model relies on the category names to achieve a more general ZSL-like prediction. When vanilla ZSL meets general CLIP, however, most existing methods on both sides struggle to benefit from each other. In this brief, we resort to attribute prompt tuning (APT) for improving the knowledge transferability from the pretrained CLIP model to the downstream ZSL framework for pursuing desirable feature representations. Our approach, termed as attribute prompt alignment network (APAN), leverages APT for cross-network feature alignment (CFA). In this way, we can investigate the effects of CLIP to vanilla ZSL task in the era of large model by the two branch APAN architecture. Specifically, APT takes as an input the templates of class attribute descriptions to produce attribute prompts, which are further used to both guide the localizations of visual regions across two frozen feature extraction networks, through a visual-semantic interaction attention. This enables APAN to progressively refine and align these cross-network features, thus resulting in generalizable feature representations that can capture fine-grained attribute information. For CFA, we simply introduce prediction alignment loss that constrains the predictions from these two cross-network visual features. Experimental results on three benchmark datasets well demonstrate that APAN outperforms the state-of-the-art methods by absorbing generalizable knowledge from CLIP models. Guosen Xie, Ting Guo 0004, Xiangbo Shu, Fang Zhao 0006, Zheng Zhang 0006, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | SynSP: Synergy of Smoothness and Precision in Pose Sequences RefinementabstractPredicting human pose sequences via existing pose estimators often encounters various estimation errors. Motion refinement methods aim to optimize the predicted human pose sequences from pose estimators while ensuring minimal computational overhead and latency. Prior investigations have primarily concentrated on striking a balance between the two objectives, i.e., smoothness and precision, while optimizing the predicted pose sequences. However, it has come to our attention that the tension between these two objectives can provide additional quality cues about the predicted pose sequences. These cues, in turn, are able to aid the network in optimizing lower-quality poses. To leverage this quality information, we propose a motion refinement network, termed SynSP, to achieve a Synergy of Smoothness and Precision in the sequence refinement tasks. Moreover, SynSP can also address multi-view poses of one person simultaneously, fixing inaccuracies in predicted poses through heightened attention to similar poses from other views, thereby amplifying the resultant quality cues and overall performance. Compared with previous methods, SynSP benefits from both pose quality and multi-view information with a much shorter input sequence length, achieving state-of-the-art results among four challenging datasets involving 2D, 3D, and SMPL pose representations in both single-view and multi-view scenes. Github code: https://github.com/InvertedForest/SynSP. Tao Wang 0011, Lei Jin 0003, Zheng Wang 0007, Jianshu Li, Liang Li 0003, Fang Zhao 0006, Yu Cheng 0009, Li Yuan 0007, Junliang Xing, Jian Zhao 0006 |
CVPR | 6 |
| 2024 | MMR: Multi-scale Motion Retargeting between Skeleton-agnostic CharactersabstractWe present a simple yet effective method for skeleton-agnostic motion retargeting. Previous methods transfer motion between high-resolution meshes, failing to preserve the inherent local-part motions in the mesh. Addressing this issue, our proposed method learns the correspondence in a coarse-to-fine fashion by disentangling the retargeting process within multi-scale meshes. First, we propose a mesh-pooling module that pools the mesh representations for better motion transfer. This module improves the ability to handle small-part motion and preserves the local motion interdependence between neighboring mesh vertices. Furthermore, we leverage a multi-scale refinement procedure to complement missing mesh details by gradually refining the low-resolution mesh output with a higher-resolution one. We evaluate our method on several well-known 3D character datasets, and it yields an average improvement of 25% on point-wise mesh Euclidean distance (PMD) against the start-of-art method. Qualitative results show that our method is significantly helpful in preserving the moving consistency of different body parts on the target character due to disentangling body-part structures and mesh details in a multi-scale way. Haoyu Wang 0018, Shaoli Huang, Fang Zhao 0006, Chun Yuan 0003 |
IJCNN | 3 |
| 2024 | SIAM: A parameter-free, Spatial Intersection Attention Module
Gaoge Han, Shaoli Huang, Fang Zhao 0006, Jinglei Tang |
Pattern Recognit. | 3 |
| 2023 | Skinned Motion Retargeting with Residual Perception of Motion Semantics & GeometryabstractA good motion retargeting cannot be reached without reasonable consideration of source-target differences on both the skeleton and shape geometry levels. In this work, we propose a novel Residual RETargeting network (R2ET) structure, which relies on two neural modification modules, to adjust the source motions to fit the target skeletons and shapes progressively. In particular, a skeleton-aware module is introduced to preserve the source motion semantics. A shape-aware module is designed to perceive the geometries of target characters to reduce interpenetration and contact-missing. Driven by our explored distance-based losses that explicitly model the motion semantics and geometry, these two modules can learn residual motion modifications on the source motion to generate plausible retargeted motion in a single inference without postprocessing. To balance these two modifications, we further present a balancing gate to conduct linear interpolation between them. Extensive experiments on the public dataset Mixamo demonstrate that our R2ET achieves the state-of-the-art performance, and provides a good balance between the preservation of motion semantics as well as the attenuation of interpenetration and contact-missing. Code is available at https://github.com/Kebii/R2ET. Junwu Weng, Fang Zhao 0006, Shaoli Huang, Xuefei Zhe, Linchao Bao, Ying Shan, Jue Wang 0001, Zhigang Tu 0001 |
CVPR | 4 |
| 2023 | Learning Anchor Transformations for 3D Garment AnimationabstractThis paper proposes an anchor-based deformation model, namely AnchorDEF, to predict 3D garment animation from a body motion sequence. It deforms a garment mesh template by a mixture of rigid transformations with extra nonlinear displacements. A set of anchors around the mesh surface is introduced to guide the learning of rigid transformation matrices. Once the anchor transformations are found, per-vertex nonlinear displacements of the garment template can be regressed in a canonical space, which reduces the complexity of deformation space learning. By explicitly constraining the transformed anchors to satisfy the consistencies of position, normal and direction, the physical meaning of learned anchor transformations in space is guaranteed for better generalization. Furthermore, an adaptive anchor updating is proposed to optimize the anchor position by being aware of local mesh topology for learning representative anchor transformations. Qualitative and quantitative experiments on different types of garments demonstrate that AnchorDEF achieves the state-of-the-art performance on 3D garment deformation prediction in motion, especially for loose-fitting garments. Fang Zhao 0006, Zekun Li 0002, Shaoli Huang, Junwu Weng, Tianfei Zhou, Guosen Xie, Jue Wang 0001, Ying Shan |
CVPR | 1 |
| 2023 | Leveraging Balanced Semantic Embedding for Generative Zero-Shot LearningabstractGenerative (generalized) zero-shot learning [(G)ZSL] models aim to synthesize unseen class features by using only seen class feature and attribute pairs as training data. However, the generated fake unseen features tend to be dominated by the seen class features and thus classified as seen classes, which can lead to inferior performances under zero-shot learning (ZSL), and unbalanced results under generalized ZSL (GZSL). To address this challenge, we tailor a novel balanced semantic embedding generative network (BSeGN), which incorporates balanced semantic embedding learning into generative learning scenarios in the pursuit of unbiased GZSL. Specifically, we first design a feature-to-semantic embedding module (FEM) to distinguish real seen and fake unseen features collaboratively with the generator in an online manner. We introduce the bidirectional contrastive and balance losses for the FEM learning, which can guarantee a balanced prediction for the interdomain features. In turn, the updated FEM can boost the learning of the generator. Next, we propose a multilevel feature integration module (mFIM) from the cycle-consistency branch of BSeGN, which can mitigate the domain bias through feature enhancement. To the best of our knowledge, this is the first work to explore embedding and generative learning jointly within the field of ZSL. Extensive evaluations on four benchmarks demonstrate the superiority of BSeGN over its state-of-the-art counterparts. Guosen Xie, Xu-Yao Zhang, Tian-Zhu Xiang, Fang Zhao 0006, Zheng Zhang 0006, Ling Shao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic SegmentationabstractLearning semantic segmentation from weakly-labeled (e.g., image tags only) data is challenging since it is hard to infer dense object regions from sparse semantic tags. Despite being broadly studied, most current efforts directly learn from limited semantic annotations carried by individual image or image pairs, and struggle to obtain integral localization maps. Our work alleviates this from a novel perspective, by exploring rich semantic contexts synergistically among abundant weakly-labeled training data for network learning and inference. In particular, we propose regional semantic contrast and aggregation (RCA). RCA is equipped with a regional memory bank to store massive, diverse object patterns appearing in training data, which acts as strong support for exploration of dataset-level semantic structure. Particularly, we propose i) semantic contrast to drive network learning by contrasting massive categorical object regions, leading to a more holistic object pattern understanding, and ii) semantic aggregation to gather diverse relational contexts in the memory to enrich semantic repre-sentations. In this manner, RCA earns a strong capability of fine-grained semantic understanding, and eventually establishes new state-of-the-art results on two popular benchmarks, i.e., PASCAL VOC 2012 and COCO 2014. Tianfei Zhou, Meijie Zhang, Fang Zhao 0006, Jianwu Li |
CVPR | 3 |
| 2022 | Beyond Monocular Deraining: Parallel Stereo Deraining Network Via Semantic Prior
Kaihao Zhang, Wenhan Luo, Yanjiang Yu, Wenqi Ren, Fang Zhao 0006, Lin Ma 0002, Wei Liu 0005, Hongdong Li |
Int. J. Comput. Vis. | 5 |
| 2022 | Attentive WaveBlock: Complementarity-Enhanced Mutual Networks for Unsupervised Domain Adaptation in Person Re-Identification and BeyondabstractUnsupervised domain adaptation (UDA) for person re-identification is challenging because of the huge gap between the source and target domain. A typical self-training method is to use pseudo-labels generated by clustering algorithms to iteratively optimize the model on the target domain. However, a drawback to this is that noisy pseudo-labels generally cause trouble in learning. To address this problem, a mutual learning method by dual networks has been developed to produce reliable soft labels. However, as the two neural networks gradually converge, their complementarity is weakened and they likely become biased towards the same kind of noise. This paper proposes a novel light-weight module, the Attentive WaveBlock (AWB), which can be integrated into the dual networks of mutual learning to enhance the complementarity and further depress noise in the pseudo-labels. Specifically, we first introduce a parameter-free module, the WaveBlock, which creates a difference between features learned by two networks by waving blocks of feature maps differently. Then, an attention mechanism is leveraged to enlarge the difference created and discover more complementary features. Furthermore, two kinds of combination strategies, i.e. pre-attention and post-attention, are explored. Experiments demonstrate that the proposed method achieves state-of-the-art performance with significant improvements on multiple UDA person re-identification tasks. We also prove the generality of the proposed method by applying it to vehicle re-identification and image classification tasks. Our codes and models are available at: AWB. Fang Zhao 0006, Shengcai Liao, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | DomainMix: Learning Generalizable Person Re-Identification Without Human Annotations
Shengcai Liao, Fang Zhao 0006, Cuicui Kang, Ling Shao 0001 |
BMVC | 3 |
| 2021 | Learning Anchored Unsigned Distance Functions with Gradient Direction Alignment for Single-view Garment ReconstructionabstractWhile single-view 3D reconstruction has made significant progress benefiting from deep shape representations in recent years, garment reconstruction is still not solved well due to open surfaces, diverse topologies and complex geometric details. In this paper, we propose a novel learn-able Anchored Unsigned Distance Function (AnchorUDF) representation for 3D garment reconstruction from a single image. AnchorUDF represents 3D shapes by predicting unsigned distance fields (UDFs) to enable open garment surface modeling at arbitrary resolution. To capture diverse garment topologies, AnchorUDF not only computes pixel-aligned local image features of query points, but also leverages a set of anchor points located around the surface to enrich 3D position features for query points, which provides stronger 3D space context for the distance function. Furthermore, in order to obtain more accurate point projection direction at inference, we explicitly align the spatial gradient direction of AnchorUDF with the ground-truth direction to the surface during training. Extensive experiments on two public 3D garment datasets, i.e., MGN and Deep Fashion3D, demonstrate that AnchorUDF achieves the state-of-the-art performance on single-view garment reconstruction. Code is available at https://github.com/zhaofang0627/AnchorUDF. Fang Zhao 0006, Shengcai Liao, Ling Shao 0001 |
ICCV | 1 |
| 2021 | VMAN: A Virtual Mainstay Alignment Network for Transductive Zero-Shot LearningabstractTransductive zero-shot learning (TZSL) extends conventional ZSL by leveraging (unlabeled) unseen images for model training. A typical method for ZSL involves learning embedding weights from the feature space to the semantic space. However, the learned weights in most existing methods are dominated by seen images, and can thus not be adapted to unseen images very well. In this paper, to align the (embedding) weights for better knowledge transfer between seen/unseen classes, we propose the virtual mainstay alignment network (VMAN), which is tailored for the transductive ZSL task. Specifically, VMAN is casted as a tied encoder-decoder net, thus only one linear mapping weights need to be learned. To explicitly learn the weights in VMAN, for the first time in ZSL, we propose to generate virtual mainstay (VM) samples for each seen class, which serve as new training data and can prevent the weights from being shifted to seen images, to some extent. Moreover, a weighted reconstruction scheme is proposed and incorporated into the model training phase, in both the semantic/feature spaces. In this way, the manifold relationships of the VM samples are well preserved. To further align the weights to adapt to more unseen images, a novel instance-category matching regularization is proposed for model re-training. VMAN is thus modeled as a nested minimization problem and is solved by a Taylor approximate optimization paradigm. In comprehensive evaluations on four benchmark datasets, VMAN achieves superior performances under the (Generalized) TZSL setting. Guosen Xie, Xu-Yao Zhang, Yazhou Yao, Zheng Zhang 0006, Fang Zhao 0006, Ling Shao 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Region Graph Embedding Network for Zero-Shot Learning
Guosen Xie, Li Liu 0004, Fan Zhu 0001, Fang Zhao 0006, Zheng Zhang 0006, Yazhou Yao, Jie Qin 0004, Ling Shao 0001 |
ECCV (4) | 4 |
| 2020 | Beyond Monocular Deraining: Stereo Image Deraining via Semantic Understanding
Kaihao Zhang, Wenhan Luo, Wenqi Ren, Jingwen Wang 0003, Fang Zhao 0006, Lin Ma 0002, Hongdong Li |
ECCV (27) | 5 |
| 2020 | Unsupervised Domain Adaptation with Noise Resistible Mutual-Training for Person Re-identification
Fang Zhao 0006, Shengcai Liao, Guosen Xie, Jian Zhao 0006, Kaihao Zhang, Ling Shao 0001 |
ECCV (11) | 1 |
| 2020 | Human Parsing Based Texture Transfer from Single Image to 3D Human via Cross-View ConsistencyabstractThis paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation of human image and preserve the spatial distribution of semantic parts. Meanwhile, in order to improve the prediction for textures of invisible parts, we explicitly enforce the consistency across different views of the same subject by exchanging the textures predicted by two views to render images during training. The perception loss and total variation regularization are optimized to maximize the similarity between rendered and input images, which does not necessitate extra 3D texture supervision. Experimental results on pedestrian images and fashion photos demonstrate that our method can produce higher quality textures with convincing details than other texture generation methods. Fang Zhao 0006, Shengcai Liao, Kaihao Zhang, Ling Shao 0001 |
NeurIPS | 1 |
| 2019 | Look across Elapse: Disentangled Representation Learning and Photorealistic Cross-Age Face Synthesis for Age-Invariant Face RecognitionabstractDespite the remarkable progress in face recognition related technologies, reliably recognizing faces across ages still remains a big challenge. The appearance of a human face changes substantially over time, resulting in significant intraclass variations. As opposed to current techniques for ageinvariant face recognition, which either directly extract ageinvariant features for recognition, or first synthesize a face that matches target age before feature extraction, we argue that it is more desirable to perform both tasks jointly so that they can leverage each other. To this end, we propose a deep Age-Invariant Model (AIM) for face recognition in the wild with three distinct novelties. First, AIM presents a novel unified deep architecture jointly performing cross-age face synthesis and recognition in a mutual boosting way. Second, AIM achieves continuous face rejuvenation/aging with remarkable photorealistic and identity-preserving properties, avoiding the requirement of paired data and the true age of testing samples. Third, we develop effective and novel training strategies for end-to-end learning the whole deep architecture, which generates powerful age-invariant face representations explicitly disentangled from the age variation. Extensive experiments on several cross-age datasets (MORPH, CACD and FG-NET) demonstrate the superiority of the proposed AIM model over the state-of-the-arts. Benchmarking our model on one of the most popular unconstrained face recognition datasets IJB-C additionally verifies the promising generalizability of AIM in recognizing faces in the wild. Jian Zhao 0006, Yu Cheng 0009, Yang Yang 0002, Fang Zhao 0006, Jianshu Li, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
AAAI | 5 |
| 2019 | Multi-Prototype Networks for Unconstrained Set-based Face RecognitionabstractIn this paper, we address the challenging unconstrained set-based face recognition problem where each subject face is instantiated by a set of media (images and videos) instead of a single image. Naively aggregating information from all the media within a set would suffer from the large intra-set variance caused by heterogeneous factors (e.g., varying media modalities, poses and illumination) and fail to learn discriminative face representations. A novel Multi-Prototype Network (MP- Net) model is thus proposed to learn multiple prototype face representations adaptively from the media sets. Each learned prototype is representative for the subject face under certain condition in terms of pose, illumination and media modality. Instead of handcrafting the set partition for prototype learn- ing, MPNet introduces a Dense SubGraph (DSG) learning sub-net that implicitly untangles inconsistent media and learns a number of representative prototypes. Qualitative and quantitative experiments clearly demonstrate the superiority of the proposed model over state-of-the-arts. Jian Zhao 0006, Jianshu Li, Xiaoguang Tu, Fang Zhao 0006, Yuan Xin, Junliang Xing, Hengzhu Liu, Shuicheng Yan, Jiashi Feng |
IJCAI | 4 |
| 2018 | Weakly Supervised Phrase Localization With Multi-Scale Anchored Transformer NetworkabstractIn this paper, we propose a novel weakly supervised model, Multi-scale Anchored Transformer Network (MATN), to accurately localize free-form textual phrases with only image-level supervision. The proposed MATN takes region proposals as localization anchors, and learns a multiscale correspondence network to continuously search for phrase regions referring to the anchors. In this way, MATN can exploit useful cues from these anchors to reliably reason about locations of the regions described by the phrases given only image-level supervision. Through differentiable sampling on image spatial feature maps, MATN introduces a novel training objective to simultaneously minimize a contrastive reconstruction loss between different phrases from a single image and a set of triplet losses among multiple images with similar phrases. Superior to existing region proposal based methods, MATN searches for the optimal bounding box over the entire feature map instead of selecting a sub-optimal one from discrete region proposals. We evaluate MATN on the Flickr30K Entities and ReferItGame datasets. The experimental results show that MATN significantly outperforms the state-of-the-art methods. Fang Zhao 0006, Jianshu Li, Jian Zhao 0006, Jiashi Feng |
CVPR | 1 |
| 2018 | Towards Pose Invariant Face Recognition in the WildabstractPose variation is one key challenge in face recognition. As opposed to current techniques for pose invariant face recognition, which either directly extract pose invariant features for recognition, or first normalize profile face images to frontal pose before feature extraction, we argue that it is more desirable to perform both tasks jointly to allow them to benefit from each other. To this end, we propose a Pose Invariant Model (PIM) for face recognition in the wild, with three distinct novelties. First, PIM is a novel and unified deep architecture, containing a Face Frontalization sub-Net (FFN) and a Discriminative Learning sub-Net (DLN), which are jointly learned from end to end. Second, FFN is a well-designed dual-path Generative Adversarial Network (GAN) which simultaneously perceives global structures and local details, incorporated with an unsupervised cross-domain adversarial training and a "learning to learn" strategy for high-fidelity and identity-preserving frontal view synthesis. Third, DLN is a generic Convolutional Neural Network (CNN) for face recognition with our enforced cross-entropy optimization strategy for learning discriminative yet generalized feature representation. Qualitative and quantitative experiments on both controlled and in-the-wild benchmarks demonstrate the superiority of the proposed model over the state-of-the-arts. Jian Zhao 0006, Yu Cheng 0009, Yan Xu 0009, Jianshu Li, Fang Zhao 0006, Jayashree Karlekar, Sugiri Pranata, Shengmei Shen, Junliang Xing, Shuicheng Yan, Jiashi Feng |
CVPR | 6 |
| 2018 | Dynamic Conditional Networks for Few-Shot Learning
Fang Zhao 0006, Jian Zhao 0006, Shuicheng Yan, Jiashi Feng |
ECCV (15) | 1 |
| 2018 | Landmark Free Face Attribute PredictionabstractFace attribute prediction in the wild is important for many facial analysis applications yet it is very challenging due to ubiquitous face variations. In this paper, we address face attribute prediction in the wild by proposing a novel method, lAndmark Free Face AttrIbute pRediction (AFFAIR). Unlike traditional face attribute prediction methods that require facial landmark detection and face alignment, AFFAIR uses an endto- end learning pipeline to jointly learn a hierarchy of spatial transformations that optimize facial attribute prediction with no reliance on landmark annotations or pre-trained landmark detectors. AFFAIR achieves this through simultaneously 1) learning a global transformation which effectively alleviates negative effect of global face variation for the following attribute prediction tailored for each face, 2) locating the most relevant facial part for attribute prediction and 3) aggregating the global and local features for robust attribute prediction. Within AFFAIR, a new competitive learning strategy is developed that effectively enhances global transformation learning for better attribute prediction. We show that with zero information about landmarks, AFFAIR achieves state-of-the-art performance on three face attribute prediction benchmarks, which simultaneously learns the face-level transformation and attribute-level localization within a unified framework. Jianshu Li, Fang Zhao 0006, Jiashi Feng, Sujoy Roy, Shuicheng Yan, Terence Sim |
IEEE Trans. Image Process. | 2 |
| 2018 | Robust LSTM-Autoencoders for Face De-Occlusion in the WildabstractFace recognition techniques have been developed significantly in recent years. However, recognizing faces with partial occlusion is still challenging for existing face recognizers, which is heavily desired in real-world applications concerning surveillance and security. Although much research effort has been devoted to developing face de-occlusion methods, most of them can only work well under constrained conditions, such as all of faces are from a pre-defined closed set of subjects. In this paper, we propose a robust LSTM-Autoencoders (RLA) model to effectively restore partially occluded faces even in the wild. The RLA model consists of two LSTM components, which aims at occlusion-robust face encoding and recurrent occlusion removal respectively. The first one, named multi-scale spatial LSTM encoder, reads facial patches of various scales sequentially to output a latent representation, and occlusion-robustness is achieved owing to the fact that the influence of occlusion is only upon some of the patches. Receiving the representation learned by the encoder, the LSTM decoder with a dual channel architecture reconstructs the overall face and detects occlusion simultaneously, and by feat of LSTM, the decoder breaks down the task of face de-occlusion into restoring the occluded part step by step. Moreover, to minimize identify information loss and guarantee face recognition accuracy over recovered faces, we introduce an identity-preserving adversarial training scheme to further improve RLA. Extensive experiments on both synthetic and real data sets of faces with occlusion clearly demonstrate the effectiveness of our proposed RLA in removing different types of facial occlusion at various locations. The proposed method also provides significantly larger performance gain than other de-occlusion methods in promoting recognition performance over partially-occluded faces. Fang Zhao 0006, Jiashi Feng, Jian Zhao 0006, Wenhan Yang, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2017 | Marginalized CNN: Learning Deep Invariant Representations
Jian Zhao 0006, Jianshu Li, Fang Zhao 0006, Xuecheng Nie, Yunpeng Chen, Shuicheng Yan, Jiashi Feng |
BMVC | 3 |
| 2017 | Deep Attribute-preserving Metric Learning for Natural Language Object RetrievalabstractRetrieving image content with a natural language expression is an emerging interdisciplinary problem at the intersection of multimedia, natural language processing and artificial intelligence. Existing methods tackle this challenging problem by learning features from the visual and linguistic domains independently while the critical semantic correlations bridging two domains have been under-explored in the feature learning process. In this paper, we propose to exploit sharable semantic attributes as "anchors" to ensure the learned features are well aligned across domains for better object retrieval. We define "attributes" as the common concepts that are informative for object retrieval and can be easily learned from both visual content and language expression. In particular, diverse and complex attributes (e.g., location, color, category, interaction between object and context) are modeled and incorporated to promote cross-domain alignment for feature learning from multiple perspectives. Based on the sharable attributes, we propose a deep Attribute-Preserving Metric learning (AP-Metric) framework that jointly generates unique query-sensitive region proposals and conducts novel cross-modal feature learning that explicitly pursues consistency over semantic attribute abstraction within both domains for deep metric learning. Benefiting from the cross-modal semantic correlations, our proposed framework can localize challenging visual objects to match complex query expressions within cluttered background accurately. The overall framework is end-to-end trainable. Extensive evaluations on popular datasets including ReferItGame, RefCOCO, and RefCOCO+ well demonstrate its superiority. Notably, it achieves state-of-the-art performance on the challenging ReferItGame dataset. Jianan Li 0001, Yunchao Wei, Xiaodan Liang, Fang Zhao 0006, Jianshu Li, Tingfa Xu, Jiashi Feng |
ACM Multimedia | 4 |
| 2017 | Integrated Face Analytics Networks through Cross-Dataset Hybrid TrainingabstractFace analytics benefits many multimedia applications. It consists of a number of tasks, such as facial emotion recognition and face parsing, and most existing approaches generally treat these tasks independently, which limits their deployment in real scenarios. In this paper we propose an integrated Face Analytics Network (iFAN), which is able to perform multiple tasks jointly for face analytics with a novel carefully designed network architecture to fully facilitate the informative interaction among different tasks. The proposed integrated network explicitly models the interactions between tasks so that the correlations between tasks can be fully exploited for performance boost. In addition, to solve the bottleneck of the absence of datasets with comprehensive training data for various tasks, we propose a novel cross-dataset hybrid training strategy. It allows "plug-in and play'' of multiple datasets annotated for different tasks without the requirement of a fully labeled common dataset for all the tasks. We experimentally show that the proposed iFAN achieves state-of-the-art performance on multiple face analytics tasks using a single integrated model. Specifically, iFAN achieves an overall F-score of 91.15% on the Helen dataset for face parsing, a normalized mean error of 5.81% on the MTFL dataset for facial landmark localization and an accuracy of 45.73% on the BNU dataset for emotion recognition with a single model. Jianshu Li, Shengtao Xiao, Fang Zhao 0006, Jian Zhao 0006, Jianan Li 0001, Jiashi Feng, Shuicheng Yan, Terence Sim |
ACM Multimedia | 3 |
| 2017 | Dual-Agent GANs for Photorealistic and Identity Preserving Profile Face SynthesisabstractSynthesizing realistic profile faces is promising for more efficiently training deep pose-invariant models for large-scale unconstrained face recognition, by populating samples with extreme poses and avoiding tedious annotations. However, learning from synthetic faces may not achieve the desired performance due to the discrepancy between distributions of the synthetic and real face images. To narrow this gap, we propose a Dual-Agent Generative Adversarial Network (DA-GAN) model, which can improve the realism of a face simulator's output using unlabeled real faces, while preserving the identity information during the realism refinement. The dual agents are specifically designed for distinguishing real v.s. fake and identities simultaneously. In particular, we employ an off-the-shelf 3D face model as a simulator to generate profile face images with varying poses. DA-GAN leverages a fully convolutional network as the generator to generate high-resolution images and an auto-encoder as the discriminator with the dual agents. Besides the novel architecture, we make several key modifications to the standard GAN to preserve pose and texture, preserve identity and stabilize training process: (i) a pose perception loss; (ii) an identity perception loss; (iii) an adversarial loss with a boundary equilibrium regularization term. Experimental results show that DA-GAN not only presents compelling perceptual results but also significantly outperforms state-of-the-arts on the large-scale and challenging NIST IJB-A unconstrained face recognition benchmark. In addition, the proposed DA-GAN is also promising as a new approach for solving generic transfer learning problems more effectively. Jian Zhao 0006, Jayashree Karlekar, Jianshu Li, Fang Zhao 0006, Zhecan Wang, Sugiri Pranata, Shengmei Shen, Shuicheng Yan, Jiashi Feng |
NIPS | 5 |
| 2017 | Deep Edge Guided Recurrent Residual Learning for Image Super-ResolutionabstractIn this paper, we consider the image super-resolution (SR) problem. The main challenge of image SR is to recover high-frequency details of a low-resolution (LR) image that are important for human perception. To address this essentially ill-posed problem, we introduce a Deep Edge Guided REcurrent rEsidual (DEGREE) network to progressively recover the high-frequency details. Different from most of the existing methods that aim at predicting high-resolution (HR) images directly, the DEGREE investigates an alternative route to recover the difference between a pair of LR and HR images by recurrent residual learning. DEGREE further augments the SR process with edge-preserving capability, namely the LR image and its edge map can jointly infer the sharp edge details of the HR image during the recurrent recovery process. To speed up its training convergence rate, by-pass connections across the multiple layers of DEGREE are constructed. In addition, we offer an understanding on DEGREE from the view-point of sub-band frequency decomposition on image signal and experimentally demonstrate how the DEGREE can recover different frequency bands separately. Extensive experiments on three benchmark data sets clearly demonstrate the superiority of DEGREE over the well-established baselines and DEGREE also provides new state-of-the-arts on these data sets. We also present addition experiments for JPEG artifacts reduction to demonstrate the good generality and flexibility of our proposed DEGREE network to handle other image processing tasks. Wenhan Yang, Jiashi Feng, Jianchao Yang, Fang Zhao 0006, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
IEEE Trans. Image Process. | 4 |
| 2016 | Robust Face Recognition with Deep Multi-View Representation LearningabstractThis paper describes our proposed method targeting at the MSR Image Recognition Challenge MS-Celeb-1M. The challenge is to recognize one million celebrities from their face images captured in the real world. The challenge provides a large scale dataset crawled from the Web, which contains a large number of celebrities with many images for each subject. Given a new testing image, the challenge requires an identify for the image and the corresponding confidence score. To complete the challenge, we propose a two-stage approach consisting of data cleaning and multi-view deep representation learning. The data cleaning can effectively reduce the noise level of training data and thus improves the performance of deep learning based face recognition models. The multi-view representation learning enables the learned face representations to be more specific and discriminative. Thus the difficulties of recognizing faces out of a huge number of subjects are substantially relieved. Our proposed method achieves a coverage of 46.1% at 95% precision on the random set and a coverage of 33.0% at 95% precision on the hard set of this challenge. Jianshu Li, Jian Zhao 0006, Fang Zhao 0006, Hao Liu 0003, Jing Li 0050, Shengmei Shen, Jiashi Feng, Terence Sim |
ACM Multimedia | 3 |
| 2016 | Learning Relevance Restricted Boltzmann Machine for Unstructured Group Activity and Event Understanding
Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tao Xiang 0002, Tieniu Tan |
Int. J. Comput. Vis. | 1 |
| 2015 | Deep semantic ranking based hashing for multi-label image retrievalabstractWith the rapid growth of web images, hashing has received increasing interests in large scale image retrieval. Research efforts have been devoted to learning compact binary codes that preserve semantic similarity based on labels. However, most of these hashing methods are designed to handle simple binary similarity. The complex multi-level semantic structure of images associated with multiple labels have not yet been well explored. Here we propose a deep semantic ranking based method for learning hash functions that preserve multilevel semantic similarity between multi-label images. In our approach, deep convolutional neural network is incorporated into hash functions to jointly learn feature representations and mappings from them to hash codes, which avoids the limitation of semantic representation power of hand-crafted features. Meanwhile, a ranking list that encodes the multilevel similarity information is employed to guide the learning of such deep hash functions. An effective scheme based on surrogate loss is used to solve the intractable optimization problem of nonsmooth and multivariate ranking measures involved in the learning procedure. Experimental results show the superiority of our proposed approach over several state-of-the-art hashing methods in term of ranking evaluation metrics when tested on multi-label image datasets. Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
CVPR | 1 |
| 2013 | Discovering compact topical descriptors for web video retrievalabstractDescribing videos efficiently is an important task for content based web video retrieval. To solve this problem, we propose an unsupervised approach based on an undirected topic model to learn a compact topical descriptor upon the bag-of-words (BoW) video representation. In our method, words in a BoW are assumed to have different topic features, and the topical descriptor of an entire video is obtained by aggregating those features, which makes the descriptor contain information about relative strength of topics. To improve the descriptor interpretability, an L1penalty is used to control the topical sparsity. Furthermore, efficient learning and inference algorithms are presented. We evaluate the proposed descriptor on the Columbia Consumer Video dataset. Experimental results demonstrate that compared with the BoW and other topical representations, the proposed compact descriptor has better performance in web video retrieval. Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
ICIP | 1 |
| 2013 | Relevance Topic Model for Unstructured Social Group Activity RecognitionabstractUnstructured social group activity recognition in web videos is a challenging task due to 1) the semantic gap between class labels and low-level visual features and 2) the lack of labeled training data. To tackle this problem, we propose a relevance topic model" for jointly learning meaningful mid-level representations upon bag-of-words (BoW) video representations and a classifier with sparse weights. In our approach, sparse Bayesian learning is incorporated into an undirected topic model (i.e., Replicated Softmax) to discover topics which are relevant to video classes and suitable for prediction. Rectified linear units are utilized to increase the expressive power of topics so as to explain better video data containing complex contents and make variational inference tractable for the proposed model. An efficient variational EM algorithm is presented for model parameter estimation and inference. Experimental results on the Unstructured Social Activity Attribute dataset show that our model achieves state of the art performance and outperforms other supervised topic model in terms of classification accuracy, particularly in the case of a very small number of labeled training videos." Fang Zhao 0006, Yongzhen Huang, Liang Wang 0001, Tieniu Tan |
NIPS | 1 |