VLDB 2026 Research / reviewers in the wild / expert
Zhanyu Ma
dblp:56/8107
· DBLP profile ↗
192ranked-venue papers
24as first author
138since 2021 · last 2026
0000-0003-2950-2488ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 116 · 9 first-author · 89 since 2021Artificial intelligence and machine learning · 103 · 18 first-author · 72 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 2 since 2021Computer networks · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion GuidanceabstractClean images are crucial for visual tasks such as small object detection, especially at high resolutions. However, real-world images are often degraded by adverse weather, and weather restoration methods may sacrifice high-frequency details critical for analyzing small objects. A natural solution is to apply super-resolution (SR) after weather removal to recover both clarity and fine structures. However, simply cascading restoration and SR struggle to bridge their inherent conflict: removal aims to remove high-frequency weather-induced noise, while SR aims to hallucinate high-frequency textures from existing details, leading to inconsistent restoration contents. In this paper, we take deraining as a case study and propose DHGM, a Diffusion-based High-frequency Guided Model for generating clean and high-resolution images. DHGM integrates pre-trained diffusion priors with high-pass filters to simultaneously remove rain artifacts and enhance structural details. Extensive experiments demonstrate that DHGM achieves superior performance over existing methods, with lower costs. Jinglei Shi, Heng Guo 0003, Zhanyu Ma |
AAAI | 5 |
| 2026 | MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level PrecisionabstractAccurately grounding regions of interest (ROIs) is critical for diagnosis and treatment planning in medical imaging. While multimodal large language models (MLLMs) combine visual perception with natural language, current medical-grounding pipelines still rely on supervised fine-tuning with explicit spatial hints, making them ill-equipped to handle the implicit queries common in clinical practice. This work makes three core contributions. We first define Unified Medical Reasoning Grounding (UMRG), a novel vision–language task that demands clinical reasoning and pixel-level grounding. Second, we release U-MRG-14K, a dataset of 14K samples featuring pixel-level masks alongside implicit clinical queries and reasoning traces, spanning 10 modalities, 15 super-categories, and 108 specific categories. Finally, we introduce MedReasoner, a modular framework that distinctly separates reasoning from segmentation: an MLLM reasoner is optimized with reinforcement learning, while a frozen segmentation expert converts spatial prompts into masks, with alignment achieved through format and accuracy rewards. MedReasoner achieves state-of-the-art performance on U-MRG-14K and demonstrates strong generalization to unseen clinical queries, underscoring the significant promise of reinforcement learning for interpretable medical grounding. Zhonghao Yan, Muxi Diao, Ruoyan Jing, Jiayuan Xu, Kaizhou Zhang, Lele Yang, Yanxi Liu 0006, Kongming Liang, Zhanyu Ma |
AAAI | 10 |
| 2026 | Efficient Paths and Dense Rewards: Probabilistic Flow Reasoning for Large Language ModelsabstractYan Liu, Feng Zhang, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Han Liu, Yangdong Deng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhanyu Ma, Jun Xu 0001, Jiuchong Gao, Jinghua Hao, Renqing He, Yangdong Deng |
ACL (1) | 3 |
| 2026 | Fine-Mem: Fine-Grained Feedback Alignment for Long-Horizon Memory ManagementabstractWeitao Ma, Xiaocheng Feng, Lei Huang, Xiachong Feng, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weitao Ma, Lei Huang 0021, Xiachong Feng, Zhanyu Ma, Jun Xu 0001, Jiuchong Gao, Jinghua Hao, Renqing He, Bing Qin 0001 |
ACL (1) | 5 |
| 2026 | SpecGen: Neural Spectral BRDF Generation via Spectral-Spatial Tri-plane AggregationabstractSynthesizing spectral images across different wavelengths is essential for photorealistic rendering. Unlike conventional spectral uplifting methods that convert RGB images into spectral ones, we introduce SpecGen, a method that generates spectral bidirectional reflectance distribution functions (BRDFs) from a single RGB image of a sphere. This enables spectral image rendering under arbitrary illuminations and shapes covered by the corresponding material. A key challenge in spectral BRDF generation is the scarcity of measured spectral BRDF data. To address this, we propose the Spectral-Spatial Tri-plane Aggregation (SSTA) network, which models reflectance responses across wavelengths and incident-outgoing directions, allowing the training strategy to leverage abundant RGB BRDF data to enhance spectral BRDF generation. Experiments show that our method accurately reconstructs spectral BRDFs from limited spectral data and surpasses existing methods in hyperspectral image reconstruction, achieving an improvement of 8dB in PSNR. Codes link: https://github.com/sosjzy/SpecGen. Zhenyu Jin, Zhanyu Ma, Heng Guo 0003 |
WACV | 3 |
| 2026 | MIX-based Foreground and Background Patch Augmentation Guided by Physics and Material Properties for X-ray DetectionabstractThe performance of deep learning-based models for X-ray prohibited item detection heavily relies on large-scale, diverse datasets, which are often unavailable. While data augmentation offers a promising solution, prevalent methods ignore the fundamental principles of X-ray imaging, leading to artifacts such as distorted material properties and unnatural thickness perturbations. To bridge this gap, we present MIX, a physics-grounded data augmentation pipeline. The core idea of MIX is to manipulate image attributes in a way that reflects real-world physical variations. Our contributions are twofold: (1) To address material ambiguity, MIX modulates foreground pseudo-colors by directly manipulating hue and saturation, informed by the relationship between color and effective atomic number. This forces the model to learn more robust material representations. (2) To simulate variations in object density and thickness, MIX introduces a novel thickness perturbation technique based on X-ray attenuation principles. This significantly improves the model’s adaptability to geometric changes. Our proposed method seamlessly integrates with existing detectors and yields substantial performance gains across multiple benchmarks. Our work not only provides an effective augmentation solution but also highlights the critical need for domain-specific approaches in X-ray computer vision. Dongliang Chang, Yujun Tong, Zhanyu Ma |
WACV | 4 |
| 2026 | FCNet: Extracting undistorted images for fine-grained image classification
Junhan Chen, Dongliang Chang, Yujun Tong, Ruoyi Du, Yingqing Wang, Zhanyu Ma, Yi-Zhe Song |
Neurocomputing | 6 |
| 2026 | Attention-Enhanced Cross-Modality Alignment for Adapting Vision-Language ModelsabstractAbstract Prompt learning is an effective approach for adapting pre-trained vision-language models (VLMs) to a variety of downstream tasks. However, prompts designed manually or generated by large language models may not effectively capture key discriminative visual features. In addition, pre-trained VLMs may not align images and text well at a fine-grained level. To address these two issues, we propose an attention-enhanced cross-modality alignment network, which includes an adaptive channel attention (ACA) module and a cross-modal measurement (CMM) module. The ACA module adapts the existing efficient channel attention to highlight discriminative visual and textual features. The CMM module leverages four pairs of image-text similarities across both frozen and learnable branches, improving the alignment of fine-grained discriminative visual and textual features. Experiments show that the proposed method outperforms state-of-the-art methods on two representative tasks: base-to-novel generalization and cross-dataset evaluation. Our code is available at https://github.com/xueshaoying/XSY_AECA.git . Shaoying Xue, Jie Cao 0014, Zhanyu Ma |
Mach. Learn. | 5 |
| 2026 | Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute EditingabstractSemantic segmentation takes a pivotal role in various applications such as autonomous driving and medical image analysis. When deploying segmentation models in practice, it is critical to test their behaviors in varied and complex scenes in advance. In this paper, we construct an automatic data generation pipeline Gen4Seg to stress-test semantic segmentation models by generating various challenging samples with different attribute changes. Beyond previous evaluation paradigms focusing solely on global weather and style transfer, we investigate variations in both appearance and geometry attributes at the object and image level. These include object color, material, size, and position, as well as image-level variations such as weather and style. To achieve this, we propose to edit visual attributes of existing real images with precise control of structural information, empowered by diffusion models. In this way, the existing segmentation labels can be reused for the edited images, which greatly reduces the labor costs of constructing datasets. Using our pipeline, we construct two new benchmarks, Pascal-EA and COCO-EA. We benchmark a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models. We have several key findings: 1) advanced open-vocabulary models do not exhibit greater robustness compared to closed-set methods under geometric variations; 2) traditional data augmentation techniques, such as CutOut and CutMix, are limited in enhancing robustness against appearance variations; 3) our generation pipeline can also be employed as a data augmentation tool and improve both in-distribution and out-of-distribution performances. Our work suggests the potential of generative models as effective tools for automatically analyzing segmentation models, and we hope our findings will assist practitioners and researchers in developing more robust and reliable segmentation models. Zijin Yin, Bing Li 0015, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | A fine-grained entity understanding network for weakly supervised phrase grounding
Pengyue Lin, Ruifan Li, Fangxiang Feng, Lun Ke, Zhanyu Ma, Xiaojie Wang 0006 |
Pattern Recognit. | 5 |
| 2026 | HumanRecon: Neural reconstruction of dynamic human using geometric cues and physical priors
Junhui Yin, Wei Yin 0006, Hao Chen 0041, Xuqian Ren, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001 |
Pattern Recognit. | 5 |
| 2026 | BARE: Towards Bias-Aware and Reasoning-Enhanced One-Tower Visual GroundingabstractVisual Grounding (VG), which aims to locate a specific region referred to by expressions, is a fundamental yet challenging task in the multimodal understanding fields. While recent grounding transfer works have advanced the field through one-tower architectures, they still suffer from two primary limitations: (1) over-entangled multimodal representations that exacerbate deceptive modality biases, and (2) insufficient semantic reasoning that hinders the comprehension of referential cues. In this paper, we propose BARE, a bias-aware and reasoning-enhanced framework for one-tower visual grounding. BARE introduces a mechanism that preserves modality-specific features and constructs referential semantics through three novel modules: (i) language salience modulator, (ii) visual bias correction and (iii) referential relationship enhancement, which jointly mitigate multimodal distractions and enhance referential comprehension. Extensive experimental results on five benchmarks demonstrate that BARE not only achieves state-of-the-art performance but also delivers superior computational efficiency compared to existing approaches. The code is publicly accessible at https://github.com/Marloweeee/BARE. Hongbing Li, Linhui Xiao, Bo Xiao 0006, Zhanyu Ma |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Fine-Tuning via Linked Domains: A Closed-Form Dual Alignment Mechanism for Transferring Vision-Language ModelsabstractAdapters and prompt learning have become two de facto strategies to fine-tune pre-trained vision-language models, mitigating the high computational cost of fine-tuning an entire model for downstream tasks. They can align the prediction from the fine-tuned model with that from the pre-trained model. However, the existing methods of these strategies primarily focus on aligning within a single modality, and the exploration of bidirectional interactions between modalities remains limited. To address this issue, we propose a closed-form dual alignment mechanism (DAM) thatnot only ensures the consistency in predictions within a single modality but also achieves the alignment of features across different modalities. In DAM, all alignments are achieved by closed-form solutions to ridge regression, without inducing a massive number of learnable parameters. Experimental results demonstrate that DAM outperforms the state-of-the-art methods on 11 benchmarks over various evaluation metrics. Our codes are available at https://github.com/Peiy-Lu/DAM. Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Toward Generalizable Forgery Detection and ReasoningabstractAccurate and interpretable detection of AI-generated images is essential for mitigating risks associated with AI misuse. However, the substantial domain gap among generative models makes it challenging to develop a generalizable forgery detection model. Moreover, since every pixel in an AI-generated image is synthesized, traditional saliency-based forgery explanation methods are not well suited for this task. To address these challenges, we formulate detection and explanation as a unified Forgery Detection and Reasoning task (FDR-Task), leveraging Multi-Modal Large Language Models (MLLMs) to provide accurate detection through reliable reasoning over forgery attributes. To facilitate this task, we introduce the Multi-Modal Forgery Reasoning dataset (MMFR-Dataset), a large-scale dataset containing 120K images across 10 generative models, with 378K reasoning annotations on forgery attributes, enabling comprehensive evaluation of the FDR-Task. Furthermore, we propose FakeReasoning, a forgery detection and reasoning framework with three key components: 1) a dual-branch visual encoder that integrates CLIP and DINO to capture both high-level semantics and low-level artifacts; 2) a Forgery-Aware Feature Fusion Module that leverages DINO's attention maps and cross-attention mechanisms to guide MLLMs toward forgery-related clues; 3) a Classification Probability Mapper that couples language modeling and forgery detection, enhancing overall performance. Experiments across multiple generative models demonstrate that FakeReasoning not only achieves robust generalization but also outperforms state-of-the-art methods on both detection and reasoning tasks. The code is available at: https://github.com/PRIS-CV/FakeReasoning. Yueying Gao, Dongliang Chang, Bingyao Yu, Haotian Qin, Muxi Diao, Lei Chen 0069, Kongming Liang, Zhanyu Ma |
IEEE Trans. Image Process. | 8 |
| 2026 | FourierSR: A Fourier Token-Based Plugin for Efficient Image Super-ResolutionabstractImage super-resolution (SR) aims to recover low-resolution images to high-resolution images, where improving SR efficiency is a high-profile challenge. However, commonly used units in SR, like convolutions and window-based Transformers, have limited receptive fields, making it challenging to apply them to improve SR under extremely limited computational cost. To address this issue, inspired by modeling convolution theorem through token mix, we propose a Fourier token-based plugin called FourierSR to improve SR uniformly, which avoids the instability or inefficiency of existing token mix technologies when applied as plug-ins. Furthermore, compared to convolutions and windows-based Transformers, our FourierSR only utilizes Fourier transform and multiplication operations, greatly reducing complexity while having global receptive fields. Experimental results show that our FourierSR as a plug-and-play unit brings an average PSNR gain of 0.34dB for existing efficient SR methods on Manga109 test set at the scale of $\times 4$ , while the average increase in the number of Params and FLOPs is only 0.6% and 1.5% of original sizes. We will release our codes upon acceptance. Heng Guo 0003, Yuefeng Hou, Zhanyu Ma |
IEEE Trans. Image Process. | 4 |
| 2026 | From Sight to Insight: Enhancing Confusable Structure Segmentation via Vision-Language Mutual PromptingabstractConfusable structure segmentation (CSS) is a type of semantic segmentation applied in remote sensing sea fog detection, medical image segmentation, camouflaged object detection, etc. Structural similarity and visual ambiguity are two critical issues in CSS that pose difficulties in distinguishing foreground objects from the background. Current methods focus primarily on enhancing visual representations and do not often incorporate multimodal information, which leads to performance bottlenecks. Inspired by recent achievements in vision-language models, we proposeVision-LanguageMutualPrompting (VLMP), a novel and unified language-guided framework that leverages text prompts to enhance CSS. Specifically, VLMP consists of vision-to-language prompting and language-to-vision prompting, which bidirectionally model the interactions between visual and linguistic features, thereby facilitating cross-modal complementary information flow. To prevent the predominance of one modality over another, we design a feature integration modulator that modulates and balances feature weights for adaptive multimodal fusion. Our framework is designed to be modular and flexible, allowing for integration with any backbone, including CNNs and transformers. We evaluate VLMP with three diverse datasets: SFDD-H8, QaTa-COV19, and CAMO-COD10K. Extensive experiments demonstrate the effectiveness and superiority of the proposed framework over those of state-of-the-art methods across these datasets. This shift from basicsightto deeperinsightin CSS through vision-language integration represents a significant advancement in the field. Yihao Zuo, Mengqiu Xu, Kaixin Chen 0001, Ming Wu 0001, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Multim. | 7 |
| 2026 | Dual-Domain Modulation Network for Lightweight Image Super-ResolutionabstractLightweight image super-resolution (SR) aims to reconstruct high-resolution images from low-resolution images under limited computational costs. We find existing frequencybased SR methods cannot balance the reconstruction of overall structures and high-frequency parts. Meanwhile, these methods are inefficient for handling frequency features and unsuitable for lightweight SR. In this paper, we show introducing both wavelet and Fourier information allows our model to consider both highfrequency features and overall SR structure reconstruction while reducing costs. Specifically, we propose a Dual-domain Modulation Network that integrates both wavelet and Fourier information for enhanced frequency modeling. Unlike existing methods that rely on a single frequency representation, our design combines wavelet-domain modulation via a Wavelet-domain Modulation Transformer (WMT) with global Fourier supervision, enabling complementary spectral learning well-suited for lightweight SR. Experimental results show that our method achieves a comparable PSNR of SRFormer [1] and MambaIR [2] while with less than 50% and 60% of their FLOPs and achieving inference speeds 15.4× and 5.4× faster, respectively, demonstrating the effectiveness of our method on SR quality and lightweight. Heng Guo 0003, Yuefeng Hou, Guangwei Gao, Zhanyu Ma |
IEEE Trans. Multim. | 5 |
| 2025 | ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion TransferabstractThe development of Text-to-Video (T2V) generation has made motion transfer possible, enabling the control of video motion based on existing footage. However, current methods have two limitations: 1) struggle to handle multi-subjects videos, failing to transfer specific subject motion; 2) struggle to preserve the diversity and accuracy of motion as transferring to subjects with varying shapes. To overcome these, we introduce ConMo, a zero-shot framework that disentangle and recompose the motions of subjects and camera movements. ConMo isolates individual subject and background motion cues from complex trajectories in source videos using only subject masks, and reassembles them for target video generation. This approach enables more accurate motion control across diverse subjects and improves performance in multi-subject scenarios. Additionally, we propose soft guidance in the recomposition stage which controls the retention of original motion to adjust shape constraints, aiding subject shape adaptation and semantic transformation. Unlike previous methods, ConMo unlocks a wide range of applications, including subject size and position editing, subject removal, semantic modifications, and camera motion simulation. Extensive experiments demonstrate that ConMo significantly outperforms state-of-the-art methods in motion fidelity and semantic consistency. The code is available at https://github.com/Andyplus1/ConMo. Jiayi Gao, Zijin Yin, Changcheng Hua, Yuxin Peng 0001, Kongming Liang, Zhanyu Ma, Jun Guo 0002, Yang Liu 0105 |
CVPR | 6 |
| 2025 | PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface ReconstructionabstractReflective and textureless surfaces remain a challenge in multi-view 3D reconstruction. Both camera pose calibration and shape reconstruction often fail due to insufficient or unreliable cross-view visual features. To address these issues, we present PMNI (Pose-free Multi-view Normal Integration), a neural surface reconstruction method that incorporates rich geometric information by leveraging surface normal maps instead of RGB images. By enforcing geometric constraints from surface normals and multi-view shape consistency within a neural signed distance function (SDF) optimization framework, PMNI simultaneously recovers accurate camera poses and high-fidelity surface geometry. Experimental results on synthetic and real-world datasets show that our method achieves state-of-the-art performance in the reconstruction of reflective surfaces, even without reliable initial camera poses. Mingzhi Pei, Heng Guo 0003, Zhanyu Ma |
CVPR | 5 |
| 2025 | PIDSR: Complementary Polarized Image Demosaicing and Super-ResolutionabstractPolarization cameras can capture multiple polarized images with different polarizer angles in a single shot, bringing convenience to polarization-based downstream tasks. However, their direct outputs are color-polarization filter array (CPFA) raw images, requiring demosaicing to reconstruct full-resolution, full-color polarized images; unfortunately, this necessary step introduces artifacts that make polarization-related parameters such as the degree of polarization (DoP) and angle of polarization (AoP) prone to error. Besides, limited by the hardware design, the resolution of a polarization camera is often much lower than that of a conventional RGB camera. Existing polarized image demosaicing (PID) methods are limited in that they cannot enhance resolution, while polarized image super-resolution (PISR) methods, though designed to obtain high-resolution (HR) polarized images from the demosaicing results, tend to retain or even amplify errors in the DoP and AoP introduced by demosaicing artifacts. In this paper, we propose PIDSR, a joint framework that performs complementary Polarized Image Demosaicing and Super-Resolution, showing the ability to robustly obtain high-quality HR polarized images with more accurate DoP and AoP from a CPFA raw image in a direct manner. Experiments show our PIDSR not only achieves state-of-the-art performance on both synthetic and real data, but also facilitates downstream tasks. Shuangfan Zhou, Chu Zhou, Youwei Lyu, Heng Guo 0003, Zhanyu Ma, Boxin Shi, Imari Sato |
CVPR | 5 |
| 2025 | PolGS: Polarimetric Gaussian Splatting for Fast Reflective Surface Reconstruction
Yufei Han 0002, Bowen Tie, Heng Guo 0003, Youwei Lyu, Si Li 0001, Boxin Shi, Zhanyu Ma |
ICCV | 8 |
| 2025 | VisualCloze: A Universal Image Generation Framework via Visual in-Context Learning
Ruoyi Du, Juncheng Yan, Le Zhuo, Peng Gao 0007, Zhanyu Ma, Ming-Ming Cheng |
ICCV | 7 |
| 2025 | FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
Tianwei Cao, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ICCV | 6 |
| 2025 | PolarAnything: Diffusion-based Polarimetric Image SynthesisabstractPolarization images facilitate image enhancement and 3D reconstruction tasks, but the limited accessibility of polarization cameras hinders their broader application. This gap drives the need for synthesizing photorealistic polarization images. The existing polarization simulator Mitsuba relies on a parametric polarization image formation model and requires extensive 3D assets covering shape and PBR materials, preventing it from generating large-scale photorealistic images. To address this problem, we propose PolarAnything, capable of synthesizing polarization images from a single RGB input with both photorealism and physical accuracy, eliminating the dependency on 3D asset collections. Drawing inspiration from the zero-shot performance of pretrained diffusion models, we introduce a diffusion-based generative framework with an effective representation strategy that preserves the fidelity of polarization properties. Experiments show that our model generates high-quality polarization images and supports downstream tasks like shape from polarization. Kailong Zhang, Youwei Lyu, Heng Guo 0003, Si Li 0001, Zhanyu Ma, Boxin Shi |
ICCV | 5 |
| 2025 | SelectVision: Adaptive Vision Resolution Selection for Visual Document Understanding
Zhongjiang He, Han Fang 0002, Hao Sun 0015, Kongming Liang, Zhanyu Ma |
ICDAR (4) | 7 |
| 2025 | A Debiasing Framework For Attribute Binding In Diffusion-Based Text-To-Image GenerationabstractDespite the impressive generative capabilities of text-to-image (T2I) models, accurate binding of objects and attributes specified in input prompts remains a significant challenge. Existing approaches often fail to address inherent biases in text embeddings, where objects tend to associate with their frequent attributes in the training data. This work identifies these biases and introduces a novel debiasing framework centered on alternating prompt vector binding, dynamically improving semantic alignment during image generation. First, we quantify bias strength in object-attribute pairs using large language models (LLMs), revealing problematic associations. Next, biased object terms are replaced with broader concepts during diffusion sampling to reduce bias. Finally, object embeddings are refined by pulling target attributes while suppressing irrelevant ones, avoiding attribute confusion. Extensive experiments on public datasets demonstrate the framework’s effectiveness in addressing attribute binding challenges through bias mitigation. Yueheng Luo, Tianwei Cao, Ling Jin 0004, Donghui Gao, Kongming Liang, Zhanyu Ma |
ICIP | 8 |
| 2025 | Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face RestorationabstractOld-photo face restoration poses significant challenges due to compounded degradations such as breakage, fading, and severe blur. Existing pre-trained diffusion-guided methods either rely on explicit degradation priors or global statistical guidance, which struggle with localized artifacts or face color. We propose Self-Supervised Selective-Guided Diffusion (SSDiff), which leverages pseudo-reference faces generated by a pre-trained diffusion model under weak guidance. These pseudo-labels exhibit structurally aligned contours and natural colors, enabling region-specific restoration via staged supervision: structural guidance applied throughout the denoising process and color refinement in later steps, aligned with the coarse-to-fine nature of diffusion. By incorporating face parsing maps and scratch masks, our method selectively restores breakage regions while avoiding identity mismatch. We further construct VintageFace, a 300-image benchmark of real old face photos with varying degradation levels. SSDiff outperforms existing GAN-based and diffusion-based methods in perceptual quality, fidelity, and regional controllability. Code link: https://github.com/PRIS-CV/SSDiff. Heng Guo 0003, Guangwei Gao, Zhanyu Ma |
NeurIPS | 5 |
| 2025 | CineTechBench: A Benchmark for Cinematographic Technique Understanding and GenerationabstractCinematography is a cornerstone of film production and appreciation, shaping mood, emotion, and narrative through visual elements such as camera movement, shot composition, and lighting. Despite recent progress in multimodal large language models (MLLMs) and video generation models, the capacity of current models to grasp and reproduce cinematographic techniques remains largely uncharted, hindered by the scarcity of expert-annotated data. To bridge this gap, we present CineTechBench, a pioneering benchmark founded on precise, manual annotation by seasoned cinematography experts across key cinematography dimensions. Our benchmark covers seven essential aspects—shot scale, shot angle, composition, camera movement, lighting, color, and focal length—and includes over 600 annotated movie images and 120 movie clips with clear cinematographic techniques. For the understanding task, we design question–answer pairs and annotated descriptions to assess MLLMs’ ability to interpret and explain cinematographic techniques. For the generation task, we assess advanced video generation models on their capacity to reconstruct cinema-quality camera movements given conditions such as textual prompts or keyframes. We conduct a large-scale evaluation on 15+ MLLMs and 5+ video generation models. Our results offer insights into the limitations of current models and future directions for cinematography understanding and generation in automatical film production and appreciation. The code and benchmark can be accessed at \url{https://github.com/PRIS-CV/CineTechBench}. Songyu Xu, Xiangxuan Shan, Muxi Diao, Xueyan Duan, Yanhua Huang, Kongming Liang, Zhanyu Ma |
NeurIPS | 9 |
| 2025 | Animal-CLIP: A Dual-Prompt Enhanced Vision-Language Model for Animal Action Recognition
Yinuo Jing, Kongming Liang, Ruxu Zhang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma |
Int. J. Comput. Vis. | 7 |
| 2025 | Interactive triplet attention for few-shot fine-grained image classificationabstractFew-shot fine-grained classification aims to identify novel fine-grained classes from extremely few examples with ultra-high semantic similarity between classes, hence a notoriously hard task. To extract discriminative features from few samples for recognizing subtle differences between fine-grained classes , it is pivotal to exploit comprehensive interactions across all dimensions in space and channel, which, however, is unexplored yet by state-of-the-art methods in this challenging area. To address this issue, in this paper we show that a simple adjustment to the existing triplet attention module (TAM) can be highly effective for few-shot fine-grained image classification. More specifically, building on TAM which comprises three parallel branches for pairwise interactions between height, width, and channel dimensions, we introduce an additional interaction between the output of these three branches, capable of modeling the dependency across all three dimensions; the revised method is dubbed interactive triplet attention module (ITAM). ITAM is a plug-and-play module, which can be inserted into any metric-based few-shot fine-grained image classifiers for performance enhancement. Extensive experiments, on CUB-200–2011, Flowers, Stanford-Cars, and Stanford-Dogs, showcase the superiority of ITAM against state-of-the-art few-shot fine-grained image classifiers. Shaoying Xue, Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue |
Neurocomputing | 5 |
| 2025 | Practically Unbiased Pairwise Loss for Recommendation With Implicit FeedbackabstractRecommender systems have been widely employed on various online platforms to improve user experience. In these systems, recommendation models are often learned from the users' historical behaviors that are automatically collected. Notably, recommender systems differ slightly from ordinary supervised learning tasks. In recommender systems, there is an exposure mechanism that decides which items could be presented to each specific user, which breaks the i.i.d assumption of supervised learning and brings biases into the recommendation models. In this paper, we focus on unbiased ranking loss weighted by inversed propensity scores (IPS), which are widely used in recommendations with implicit feedback labels. More specifically, we first highlight the fact that there is a gap between theory and practice in IPS-weighted unbiased loss. The existing pairwise loss could be theoretically unbiased by adopting an IPS weighting scheme. Unfortunately, the propensity scores are hard to estimate due to the inaccessibility of each user-item pair's true exposure status. In practical scenarios, we can only approximate the propensity scores. In this way, the theoretically unbiased loss would be still practically biased. To solve this problem, we first construct a theoretical framework to obtain a generalization upper bound of the current theoretically unbiased loss. The bound illustrates that we can ensure the theoretically unbiased loss's generalization ability if we lower its implementation loss and practical bias at the same time. To that aim, we suggest treating feedback label as a noisy proxy for exposure result for each user-item pair . Here we assume the noise rate meets the condition that . According to our analysis, this is a mild assumption that can be satisfied by many real-world applications. Based on this, we could train an accurate propensity model directly by leveraging a noise-resistant loss function. Then we could construct a practically unbiased recommendation model weighted by precise propensity scores. Lastly, experimental findings on public datasets demonstrate our suggested method's effectiveness. Tianwei Cao, Qianqian Xu 0001, Zhiyong Yang 0001, Zhanyu Ma, Qingming Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Understanding Episode Hardness in Few-Shot LearningabstractAchieving generalization for deep learning models has usually suffered from the bottleneck of annotated sample scarcity. As a common way of tackling this issue, few-shot learning focuses on "episodes", i.e., sampled tasks that help the model acquire generalizable knowledge onto unseen categories - better the episodes, the higher a model's generalisability. Despite extensive research, the characteristics of episodes and their potential effects are relatively less explored. A recent paper discussed that different episodes exhibit different prediction difficulties, and coined a new metric "hardness" to quantify episodes, which however is too wide-range for an arbitrary dataset and thus remains impractical for realistic applications. In this paper therefore, we for the first time conduct an algebraic analysis of the critical factors influencing episode hardness supported by experimental demonstrations, that reveal episode hardness to largely depend on classes within an episode, and importantly propose an efficient pre-sampling hardness assessment technique named Inverse-Fisher Discriminant Ratio (IFDR). This enables sampling hard episodes at the class level via class-level (CL) sampling scheme that drastically decreases quantification cost. Delving deeper, we also develop a variant called class-pair-level (CPL) sampling, which further reduces the sampling cost while guaranteeing the sampled distribution. Finally, comprehensive experiments conducted on benchmark datasets verify the efficacy of our proposed method. Yurong Guo 0001, Ruoyi Du, Aneeshan Sain, Kongming Liang, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2025 | Disentangling Before Composing: Learning Invariant Disentangled Features for Compositional Zero-Shot LearningabstractCompositional Zero-Shot Learning (CZSL) aims to recognize novel compositions using knowledge learned from seen attribute-object compositions in the training set. Previous works mainly project an image and its corresponding composition into a common embedding space to measure their compatibility score. However, both attributes and objects share the visual representations learned above, leading the model to exploit spurious correlations and bias towards seen compositions. Instead, we reconsider CZSL as an out-of-distribution generalization problem. If an object is treated as a domain, we can learn object-invariant features to recognize attributes attached to any object reliably, and vice versa. Specifically, we propose an invariant feature learning framework to align different domains at the representation and gradient levels to capture the intrinsic characteristics associated with the tasks. To further facilitate and encourage the disentanglement of attributes and objects, we propose an "encoding-reshuffling-decoding" process to help the model avoid spurious correlations by randomly regrouping the disentangled features into synthetic features. Ultimately, our method improves generalization by learning to disentangle features that represent two independent factors of attributes and objects. Experiments demonstrate that the proposed method achieves state-of-the-art or competitive performance in both closed-world and open-world scenarios. Tian Zhang 0029, Kongming Liang, Ruoyi Du, Wei Chen 0071, Zhanyu Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | CDN4: A cross-view Deep Nearest Neighbor Neural Network for fine-grained few-shot classificationabstractThe fine-grained few-shot classification is a challenging task in computer vision, aiming to classify images with subtle and detailed differences given scarce labeled samples. A promising avenue to tackle this challenge is to use spatially local features to densely measure the similarity between query and support samples. Compared with image-level global features, local features contain more low-level information that is rich and transferable across categories. However, methods based on spatially localized features have difficulty distinguishing subtle category differences due to the lack of sample diversity. To address this issue, we propose a novel method called Cross-view Deep Nearest Neighbor Neural Network (CDN4). CDN4 applies a random geometric transformation to augment a different view of support and query samples and subsequently exploits four similarities between the original and transformed views of query local features and those views of support local features. The geometric augmentation increases the diversity between samples of the same class, and the cross-view measurement encourages the model to focus more on discriminative local features for classification through the cross-measurements between the two branches. Extensive experiments validate the superiority of CDN4, which achieves new state-of-the-art results in few-shot classification across various fine-grained benchmarks. Code is available at . • A novel fine-grained FSL method for improving local feature discriminativeness. • Construct the Episodic Dual-Branch Structure to enhance sample diversity. • Exploit four cross-view metric pairs to enforce learning of discriminative features. • CDN4 achieves state-of-the-art performance on three fine-grained benchmark datasets. Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 5 |
| 2025 | Clarity in chaos: Boosting few-shot classification through information suppression and sparsificationabstractThe advance of deep learning has invigorated the research of few-shot classification. However, the interference of non-target information in feature representations hampers classification generalization. To tackle this issue, we propose an irrelevant information suppression (IIS) module, which is focused on suppressing the weight of unimportant information and elevating the sparsity of feature representations . An IIS network with three consecutive IIS modules is developed, to illustrate the progressive suppression of unimportant information and highlighting of key discriminative features of the target. Extensive experiments showcase the superior performance of our IIS network on five widely-used benchmark datasets. Furthermore, we show that the IIS module can be readily used as a plug-in module by state-of-the-art few-shot classifiers, and can clearly further improve their performance. Our code is available on GitHub at https://github.com/LC4188/IISNet . • We propose an IIS module to progressively suppress non-target information. • The IIS module can be readily used as a plug-in module. • The IIS module can clearly improve the performance of few-shot classifiers. Luchen Ji, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 4 |
| 2025 | Self-randomized focuses effectively boost metric-based few-shot classifiers
Zhen Li 0026, Zhongyuan Liu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jing-Hao Xue, Yi-Zhe Song |
Pattern Recognit. | 6 |
| 2025 | SRML: Structure-relation mutual learning network for few-shot image classificationabstractFew-shot image classification aims at tackling a challenging but practical classification setting, where only few labelled images are available for training. Metric-based methods are main-stream solutions for few-shot image classification, but many of them extract features that are either irrelevant to target objects in the query images or insufficient to describe the local shape or structural patterns within images, which can lead to mis-identification of the target objects, especially when the images are of multiple objects. To resolve this issue, we propose the structure-relation mutual learning (SRML) network, which first learns both the intra-image structural features and the inter-image relational features in a parallel fashion via two parallel branches, the structural feature extractor (SFE) and the relational feature extractor (RFE), and then harnesses mutual learning to enable knowledge exchange between them. In such a manner, the structural features learnt from the SFE branch not only contain the structural patterns within the images, but also focus more on the target objects, guided by the relational knowledge from the RFE branch. In return, the RFE branch can exploit the more-focused structural knowledge to better match the target objects in the support and query images. We conduct extensive experiments on four few-shot classification benchmark datasets to showcase the superior classification of the proposed SRML network, achieving a 3.17% improvement in classification accuracy over the leading competitor, RENet Kang et al. (2021). The code of this work can be found in https://github.com/Rilliant7/SRML . Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
Pattern Recognit. | 4 |
| 2025 | Rise by Lifting Others: Interacting Features to Uplift Few-Shot Fine-Grained ClassificationabstractFew-shot fine-grained classification entails notorious subtle inter-class variation. Recent works address this challenge by developing attention mechanisms, such as the task discrepancy maximization (TDM) that can highlight discriminative channels. This paper, however, aims to reveal that, besides designing sophisticated attention modules, a well-designed input scheme, which simply blends two types of features and their interactions capturing different properties of the target object, can also greatly promote the quality of the learnt weights. To illustrate, we design a bi-feature interactive TDM (BiFI-TDM) module to serve as a strong foundation for TDM to discover the most discriminative channels with ease. Specifically, we design a novel mixing strategy to produce four sets of channel weights with different focuses, reflecting the properties of the corresponding input features and their interactions, as well as a proper feature re-weighting scheme. Extensive experiments on four benchmark fine-grained image datasets showcase superior performance of BiFI-TDM in metric-based few-shot methods. Our codes are available athttps://github.com/Peiy-Lu/BiFI-TDM. Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Selectively Augmented Attention Network for Few-Shot Image ClassificationabstractFew-shot image classification is a challenging task that aims to learn from a limited number of labelled training images a classification model that can be generalised to unseen classes. Two strategies are usually taken to improve the classification performances of few-shot image classifiers: either applying data augmentation to enlarge the sample size of the training set and reduce overfitting, or involving attention mechanisms to highlight discriminative spatial regions or channels. However, naively applying them to few-shot classifiers directly and separately may lead to undesirable results; for example, some augmented images may focus majorly on the background rather than the object, which brings additional noises to the training process. In this paper, we propose a unified framework, the selectively augmented attention (SAA) network, that carefully integrates the best of the two approaches in an end-to-end fashion via a selective best match module to select the most representative images from the augmented training set. The selected images tend to concentrate on the objects with less irrelevant background, which can assist the subsequent calculation of attentions by alleviating the interference from background. Moreover, we design a joint attention module to jointly learn both the spatial and channel-wise attentions. Experimental results on four benchmark datasets showcase the superior classification performance of the proposed SAA network compared with the state-of-the-arts. Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Adaptive Multi-Resolution Feature Fusion for Fine-Grained Visual ClassificationabstractDespite significant progress, the shortage of labeled data and expert knowledge remains a challenge for Fine-grained Visual Classification (FGVC). Some multi-source approaches that incorporate additional modalities, such as sound or bounding boxes, show promise for data enrichment but introduce added complexity to data collection. In this paper, we pose the question: can multi-source capabilities be achieved solely with existing images? The answer, confirmed by a pilot study, is affirmative. By analyzing the probability distribution of model output with different resolutions image, we find that complementary information beneficial to FGVC exists among images of different resolutions. Although the classification accuracy of low-resolution images is lower than high-resolution images, it can provide additional information for high-resolution input images. We designed a naive baseline that uses mixed training of multi-resolution images. Through the experimental results of the baseline, we find that i) not all low-resolution images are beneficial, and ii) adaptively selecting low-resolution images is what we need. Therefore, we proposed a meta-learning-based adaptive “resolution” pooling layer. Through the pooling operation, the features of low-resolution images are obtained from high-resolution images, and the most appropriate complementary features are selected for the features of high-resolution images through the gating mechanism, which enables the model to fully and autonomously exploit the complementary information. Experimental results on three FGVC datasets validate the effectiveness of our proposed method. Our code is available athttps://github.com/PRIS-CV/Adaptive-Multi-Resolution-Feature-Fusion. Dongliang Chang, Ruoyi Du, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Query-Aware Cross-Mixup and Cross-Reconstruction for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained image classification is prominent but challenging in computer vision, aiming to distinguish sub-classes under the same parent class but with only a few labeled support samples. Data augmentation techniques were explored to address the few-shot issue, but they often fail to mitigate the bias between support and query samples. Therefore, in this paper we propose a query-aware cross-mixup and cross-reconstruction method to address both few-shot and fine-grained issues. Specifically, in the training phase, we randomly select query samples and mix them with the support samples from the same class to augment the support set. This first strategy ensures the augmented support set query-aware within each sub-class. Then, we reconstruct both query samples and support samples from both original and cross-mixed support samples, thus leveraging both cross-reconstruction and self-reconstruction to enhance classification. This second strategy, enabling the reconstruction also query-aware, further mitigates the bias between support and query samples, leading to more reliable generalization. We evaluate our proposed method on four widely used few-shot fine-grained image classification datasets, and experimental results demonstrate its effectiveness in achieving the state-of-the-art classification performance. Dongliang Chang, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Class-Customized Domain Adaptation: Unlock Each Customer-Specific Class With Single AnnotationabstractModel customization mitigates the issues of inadequate performance, resource wastage, and privacy risks associated with using general-purpose models in specialized domains and well-defined tasks. However, achieving customization at a low annotation cost still poses a challenge. Existing domain adaptation research has addressed cases where all customized classes are present in the labeled database, yet scenarios involving customer-specific classes are still unresolved. Therefore, this paper proposes a novel Class-Customized Domain Adaptation (CCDA) method, addressing the latter scenario with just one additional annotation for each customer-specific class. CCDA adopts the classic adaptation training framework and comprises two innovative techniques. Firstly, to ensure the shared class knowledge from the database and the private class knowledge from additional annotations are transferred and propagated to the correct regions within the target domain, we design the partial-feature alignment strategy, based on the mechanical properties of feature alignment. Second, we propose soft-balanced sampling to tackle the long-tail distribution problem in labeled data, preventing the model from overfitting to the labeled samples of customer-specific classes. The effectiveness of CCDA has been validated across 48 tasks simulated on domain adaptation benchmarks and two real-world customization scenarios, consistently showing excellent performance. Additionally, extensive analytical experiments illustrate the contributions of two innovative techniques. The code is available at https://github.com/CHEN-kx/ClassCustomizedDA. Kaixin Chen 0001, Huiying Chang, Mengqiu Xu, Ruoyi Du, Ming Wu 0001, Zhanyu Ma |
IEEE Trans. Image Process. | 6 |
| 2025 | Reserve to Adapt: Mining Inter-Class Relations for Open-Set Domain AdaptationabstractOpen-Set Domain Adaptation (OSDA) aims at adapting a model trained on a labelled source domain, to an unlabeled target domain that is corrupted with unknown classes. The key challenge inherent to this open-set setting is therefore how best to avoid the negative transfer incurred by unknown classes during model adaptation. Most existing works tackle this challenge by simply pushing the entire unknown classes away. In this paper, we take a different stance - instead of addressing these unknown classes as a single entity, we "reserve" in-between spaces for their subsets in the learned embedding. Our key finding is that the inter-class relations learned off the source domain, can help to enforce class separations in the target domain - thereby reserving spaces for unknown classes. More specifically, we first prep the "reservation" by tightening the known-class representations while enlarging their inter-class margin. We then learn soft-label prototypes in the source domain to facilitate the discrimination of known and unknown samples in the target domain. It follows that these two steps are iterated at each epoch in a mutually beneficial manner - better discrimination of unknown samples helps with space reservation, and vice versa. We show state-of-the-art results on four standard OSDA datasets, Office-31, Office-Home, VisDA and ImageCLEF, and conduct further analysis to help understand our method. Codes are available at: https://github.com/PRIS-CV/Reserve_to_Adapt. Yujun Tong, Dongliang Chang, Da Li 0001, Kongming Liang, Zhongjiang He, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 8 |
| 2025 | Multi-View Knowledge Guided Semantic Prototype Learning for Generalized Zero-Shot Action RecognitionabstractGeneralized zero-shot skeleton-based action recognition (GZSSAR) is an emerging and challenging problem in the computer vision community. It requires models to recognize human actions, including some classes that are unseen during training. Previous studies typically rely solely on action labels to bridge the gap between seen and unseen action classes. However, the limited action semantic information hinders the learning of comprehensive semantic prototypes, thereby restricting the model's ability to generalize to unseen classes. To address this issue, in addition to the original action labels, we explore four types of textual action descriptions (i.e., interpretive and motional descriptions derived from manual expert annotation and large language model) for each action class. In order to comprehensively utilize multi-view semantic information for zeroshot classification, an Attentional Multi-view Semantic Fusion (AMSF) model is proposed. It effectively integrates the multi-view semantic features and aligns visual and semantic features in a common space, subsequently realizing the recognition of unseen action classes. Furthermore, previous works typically evaluate models in settings that include specific unseen classes, which is insufficient for GZSSAR research. To thoroughly evaluate different models, we introduce two novel distinct experimental settings, termed the “easy setting” and the “hard setting”, based on the semantic similarities between action classes. Extensive experimental results on three large-scale skeleton-based action recognition benchmarks (PKU-MMD, NTU-60, and NTU-120) not only validate the advantages of the proposed multi-view action descriptions and the AMSF model but also demonstrate the rationality of the novel experimental settings. All the data and code of this paper are publicly available on GitHub. Ming-Zhe Li, Zhang Zhang 0001, Yaoning Li, Zhanyu Ma, Liang Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Detailed Object Description With Controllable DimensionsabstractObject description plays an important role for visually impaired individuals to understand and compare the differences between objects. Recent multimodal large language models (MLLMs) exhibit powerful perceptual abilities and demonstrate impressive potential for generating object-centric descriptions. However, the descriptions generated by such models may still usually contain a lot of content that is not relevant to the user intent or miss some important object dimension details. Under special scenarios, users may only need the details of certain dimensions of an object. In this paper, we propose a training-free object description refinement pipeline,Dimension Tailor, designed to enhance user-specified details in object descriptions. This pipeline includes three steps: dimension extracting, erasing, and supplementing, which decompose the description into user-specified dimensions. Dimension Tailor can not only improve the quality of object details but also offer flexibility in including or excluding specific dimensions based on user preferences. We conducted extensive experiments to demonstrate the effectiveness of Dimension Tailor on controllable object descriptions. Notably, the proposed pipeline can consistently improve the performance of the recent MLLMs. The code is currently accessible athttps://github.com/PRIS-CV/ControllableObjectDescription. Haiwen Zhang, Baoteng Li, Kongming Liang, Hao Sun 0015, Zhongjiang He, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Multim. | 7 |
| 2024 | Dual-Prior Augmented Decoding Network for Long Tail Distribution in HOI DetectionabstractHuman object interaction detection aims at localizing human-object pairs and recognizing their interactions. Trapped by the long-tailed distribution of the data, existing HOI detection methods often have difficulty recognizing the tail categories. Many approaches try to improve the recognition of HOI tasks by utilizing external knowledge (e.g. pre-trained visual-language models). However, these approaches mainly utilize external knowledge at the HOI combination level and achieve limited improvement in the tail categories. In this paper, we propose a dual-prior augmented decoding network by decomposing the HOI task into two sub-tasks: human-object pair detection and interaction recognition. For each subtask, we leverage external knowledge to enhance the model's ability at a finer granularity. Specifically, we acquire the prior candidates from an external classifier and embed them to assist the subsequent decoding process. Thus, the long-tail problem is mitigated from a coarse-to-fine level with the corresponding external knowledge. Our approach outperforms existing state-of-the-art models in various settings and significantly boosts the performance on the tail HOI categories. The source code is available at https://github.com/PRIS-CV/DP-ADN. Jiayi Gao, Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002 |
AAAI | 5 |
| 2024 | Class-Aware Contrastive Learning for Fine-Grained Skeleton-Based Action Recognition
Xinyu Bian, Dongliang Chang, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ACCV (1) | 6 |
| 2024 | Hierarchical Prompting for Diffusion Classifiers
Wenxin Ning, Dongliang Chang, Yujun Tong, Zhongjiang He, Kongming Liang, Zhanyu Ma |
ACCV (8) | 6 |
| 2024 | Polyp-E: Benchmarking the Robustness of Deep Segmentation Models via Polyp EditingabstractIn daily clinical practice, clinicians exhibit robustness in identifying polyps with both location and size variations. It is uncertain if deep segmentation models can achieve comparable robustness in automated colonoscopic analysis. To benchmark the model robustness, we focus on evaluating the segmentation models on the polyps with various attributes (e.g. location and size) and healthy samples. Based on the Latent Diffusion Model, we perform attribute editing on real polyps and build a new dataset named Polyp-E. Our synthetic dataset boasts exceptional realism, to the extent that clinical experts find it challenging to discern them from real data. We evaluate various existing polyp segmentation models on the proposed benchmark. The results reveal most of the models are highly sensitive to attribute variations. As a novel data augmentation technique, the proposed editing pipeline can improve both in-distribution and out-ofdistribution generalization ability. The code and datasets has been released at https://github.com/RunpuWei/Polyp-E-Benchmark. Runpu Wei, Zijin Yin, Kongming Liang, Min Min, Chengwei Pan, Haonan Huang, Zhanyu Ma |
BIBM | 9 |
| 2024 | DemoFusion: Democratising High-Resolution Image Generation With No $$$abstractHigh-resolution image generation with Generative Artificial Intelligence (GenAl) has immense potential but, due to the enormous capital investment required for training, it is increasimgly centralised to a few large corporations, and hidden behind paywalls. This paper aims to democratise high-resolution GenAl by advancing the frontier of high-resolution generation while remaining accessible to a broad audience. We demonstrate that existing Latent Diffusion Models (LDMs) possess untapped potential for higher-resolution image generation. Our novel DemoFusion framework seamlessly extends open-source GenAl models, employing Progressive Upscaling, Skip Residual, and Di-lated Sampling mechanisms to achieve higher-resolution image generation. The progressive nature of DemoFusion requires more passes, but the intermediate results can serve as “previews”, facilitating rapid prompt iteration. Ruoyi Du, Dongliang Chang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 5 |
| 2024 | NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized ImagesabstractWe present NeRSP, a Neural 3D reconstruction technique for Reflective surfaces with Sparse Polarized images. Reflective surface reconstruction is extremely challenging as specular reflections are view-dependent and thus violate the multiview consistency for multiview stereo. On the other hand, sparse image inputs, as a practical capture setting, commonly cause incomplete or distorted results due to the lack of correspondence matching. This paper jointly han-dles the challenges from sparse inputs and reflective surfaces by leveraging polarized images. We derive photomet-ric and geometric cues from the polarimetric image formation model and multiview azimuth consistency, which jointly optimize the surface geometry modeled via implicit neural representation. Based on the experiments on our synthetic and real datasets, we achieve the state-of-the-art surface reconstruction results with only 6 views as input. Yufei Han 0002, Heng Guo 0003, Koki Fukai, Hiroaki Santo, Boxin Shi, Fumio Okura, Zhanyu Ma |
CVPR | 7 |
| 2024 | Benchmarking Segmentation Models with Mask-Preserved Attribute EditingabstractWhen deploying segmentation models in practice, it is critical to evaluate their behaviors in varied and complex scenes. Different from the previous evaluation paradigms only in consideration of global attribute variations (e.g. adverse weather), we investigate both local and global attribute variations for robustness evaluation. To achieve this, we construct a mask-preserved attribute editing pipeline to edit visual attributes of real images with precise control of structural information. Therefore, the original segmentation labels can be reused for the edited images. Using our pipeline, we construct a benchmark covering both object and image attributes (e.g. color, material, pattern, style). We evaluate a broad variety of semantic segmentation models, spanning from conventional close-set models to recent open-vocabulary large models on their robustness to different types of variations. We find that both local and global attribute variations affect segmentation performances, and the sensitivity of models diverges across different variation types. We argue that local attributes have the same importance as global attributes, and should be considered in the robustness evaluation of segmentation models. Code: https://github.com/PRIS-CV/Pascal-EA. Zijin Yin, Kongming Liang, Bing Li 0015, Zhanyu Ma, Jun Guo 0002 |
CVPR | 4 |
| 2024 | A Benchmark of Zero-Shot Cross-Lingual Task-Oriented Dialogue Based on Adversarial Contrastive Representation LearningabstractIdentifying user intents and their corresponding slots is the first step in the utterance interpretation pipeline of many task-oriented conversational AI systems. A multilingual system that does not adequately address unbalanced issues may provide unsatisfactory experiences for users who communicate in low-resource languages, limiting the system’s usability. Since data collection of machine learning models for this task is time-consuming, it is desirable to make use of existing data in a high-resource language to train models in low-resource languages. However, the development of such models has largely been hindered by the lack of multilingual datasets annotated according to the same guidelines with enough languages. In this paper, we present a new Cross-Lingual Task-Oriented Dialogue (CLTOD) Dataset comprised of 19k annotated utterances in 10 low-resource languages across 12 intent types. And we propose An Adversarial Contrastive Zero-Shot Learning for Cross-Lingual (2ACL) training strategy to obtain better multi-lingual semantic representation. To the end, we utilize this dataset and other publicly available datasets to conduct a comprehensive benchmarking study. This show that our model significantly outperforms state-of-the-art baselines under both zero-shot and few-shot settings. In particular, other models can improve performance by up to 200% with fine-tuning of our data. Shuang Cheng, Zhanyu Ma |
ICME | 2 |
| 2024 | Learning Conditional Prompt for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) strives to learn attributes and objects from seen compositions and transfer the acquired knowledge to unseen compositions. Existing methods either learn primitive concepts in an entangled manner, leading to the model relying on spurious correlations between attributes and objects. Alternatively, they adopt a decoupled approach, causing the model to overlook relationships between attributes and objects. In this paper, we propose a conditional prompting (CoP) method to enhance the performance of vision-language models (e.g., CLIP) in CZSL. Specifically, we utilize two image-to-word mapping networks to learn pseudo attribute and object word embeddings that can represent the corresponding semantics of input images. Subsequently, the model recognizes one concept based on another generated pseudo word embeddings, enabling the recognition of individual sub-concepts while leveraging the image-specific correlations between attributes and objects. The experimental results on three CZSL benchmarks indicate that the proposed method achieves competitive performance compared to previous state-of-the-art methods. Tian Zhang 0029, Kongming Liang, Ke Zhang 0005, Zhanyu Ma |
ICME | 4 |
| 2024 | Efficient Face Super-Resolution via Wavelet-based Feature Enhancement NetworkabstractFace super-resolution aims to reconstruct a high-resolution face image from a low-resolution face image. Previous methods typically employ an encoder-decoder structure to extract facial structural features, where the direct downsampling inevitably introduces distortions, especially to high-frequency features such as edges. To address this issue, we propose a wavelet-based feature enhancement network, which mitigates feature distortion by losslessly decomposing the input feature into high and low-frequency components using the wavelet transform and processing them separately. To improve the efficiency of facial feature extraction, a full domain Transformer is further proposed to enhance local, regional, and global facial features. Such designs allow our method to perform better without stacking many modules as previous methods did. Experiments show that our method effectively balances performance, model size, and speed. Code link: https://github.com/PRIS-CV/WFEN. Heng Guo 0003, Xuannan Liu, Kongming Liang, Jiani Hu, Zhanyu Ma, Jun Guo 0002 |
ACM Multimedia | 6 |
| 2024 | Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak SupervisionabstractPhrase Grounding, i.e., PG aims to locate objects referred by noun phrases. Recently, PG under weak supervision (i.e., grounding without region-level annotations) and zero-shot PG (i.e., grounding from seen categories to unseen ones) are proposed, respectively. However, for real-world applications these two approaches are limited due to slight annotations and numerable categories during training. In this paper, we propose a framework of zero-shot PG under weak supervision. Specifically, our PG framework is built on triple alignment strategies. Firstly, we propose a region-text alignment (RTA) strategy to build region-level attribute associations via CLIP. Secondly, we propose a domain alignment (DomA) strategy by minimizing the difference between distributions of seen classes in the training and those of the pre-training. Thirdly, we propose a category alignment (CatA) strategy by considering both category semantics and region-category relations. Extensive experimental results show that our proposed PG framework outperforms previous zero-shot methods and weakly-supervised methods. Our code is available at https://github.com/LinPengyue/ZS-WSG. Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006 |
ACM Multimedia | 6 |
| 2024 | Channel-Spatial Support-Query Cross-Attention for Fine-Grained Few-Shot Image ClassificationabstractFew-shot fine-grained image classification aims to use only few labelled samples to successfully recognize subtle sub-classes within the same parent class. This task is extremely challenging, due to the co-occurrence of large inter-class similarity, low intra-class similarity, and only few labelled samples. In this paper, to address these challenges, we propose a new Channel-Spatial Cross-Attention Module (CSCAM), which can effectively drive a model to extract discriminative fine-grained feature representations with only few shots. CSCAM collaboratively integrates a channel cross-attention module and a spatial cross-attention module, for the attentions across support and query samples. In addition, to fit for the characteristics of fine-grained images, a support averaging method is proposed in CSCAM to reduce the intra-class distance and increase the inter-class distance. Extensive experiments on four few-shot fine-grained classification datasets validate the effectiveness of CSCAM. Furthermore, CSCAM is a plug-and-play module, conveniently enabling effective improvement of state-of-the-art methods for few-shot fine-grained image classification. Shicheng Yang, Dongliang Chang, Zhanyu Ma, Jing-Hao Xue |
ACM Multimedia | 4 |
| 2024 | Animal-Bench: Benchmarking Multimodal Video Models for Animal-centric Video UnderstandingabstractWith the emergence of large pre-trained multimodal video models, multiple benchmarks have been proposed to evaluate model capabilities. However, most of the benchmarks are human-centric, with evaluation data and tasks centered around human applications. Animals are an integral part of the natural world, and animal-centric video understanding is crucial for animal welfare and conservation efforts. Yet, existing benchmarks overlook evaluations focused on animals, limiting the application of the models. To address this limitation, our work established an animal-centric benchmark, namely Animal-Bench, to allow for a comprehensive evaluation of model capabilities in real-world contexts, overcoming agent-bias in previous benchmarks. Animal-Bench includes 13 tasks encompassing both common tasks shared with humans and special tasks relevant to animal conservation, spanning 7 major animal categories and 819 species, comprising a total of 41,839 data entries. To generate this benchmark, we defined a task system centered on animals and proposed an automated pipeline for animal-centric data processing. To further validate the robustness of models against real-world challenges, we utilized a video editing approach to simulate realistic scenarios like weather changes and shooting parameters due to animal movements. We evaluated 8 current multimodal video models on our benchmark and found considerable room for improvement. We hope our work provides insights for the community and opens up new avenues for research in multimodal video models. Our data and code will be released at https://github.com/PRIS-CV/Animal-Bench. Yinuo Jing, Ruxu Zhang, Kongming Liang, Zhongjiang He, Zhanyu Ma, Jun Guo 0002 |
NeurIPS | 6 |
| 2024 | Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiTabstractLumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduce a sigmoid time discretization schedule for diffusion sampling, which achieves high-quality generation in 5-10 steps combined with higher-order ODE solvers. Thanks to these improvements, Lumina-Next not only improves the basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities as well as multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-views, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights at https://github.com/Alpha-VLLM/Lumina-T2X, we aim to advance the development of next-generation generative AI capable of universal modeling. Le Zhuo, Ruoyi Du, Han Xiao 0010, Yangguang Li 0001, Rongjie Huang 0001, Wenze Liu, Fu-Yun Wang, Zhanyu Ma, Zehan Wang 0001, Kaipeng Zhang, Lirui Zhao, Si Liu 0001, Xiangyu Yue 0001, Wanli Ouyang, Yu Qiao 0001, Hongsheng Li 0001, Peng Gao 0007 |
NeurIPS | 10 |
| 2024 | Mixture-of-Hand-Experts: Repainting the Deformed Hand Images Generated by Diffusion Models
Tianwei Cao, Kongming Liang, Zhongjiang He, Hao Sun 0015, Zhanyu Ma |
PRCV (5) | 7 |
| 2024 | Evaluating Attribute Comprehension in Large Vision-Language Models
Haiwen Zhang, Zixi Yang, Zheqi He, Kongming Liang, Zhanyu Ma |
PRCV (5) | 7 |
| 2024 | Learning Dynamic Prototypes for Visual Pattern DebiasingabstractAbstract Deep learning has achieved great success in academic benchmarks but fails to work effectively in the real world due to the potential dataset bias. The current learning methods are prone to inheriting or even amplifying the bias present in a training dataset and under-represent specific demographic groups. More recently, some dataset debiasing methods have been developed to address the above challenges based on the awareness of protected or sensitive attribute labels. However, the number of protected or sensitive attributes may be considerably large, making it laborious and costly to acquire sufficient manual annotation. To this end, we propose a prototype-based network to dynamically balance the learning of different subgroups for a given dataset. First, an object pattern embedding mechanism is presented to make the network focus on the foreground region. Then we design a prototype learning method to discover and extract the visual patterns from the training data in an unsupervised way. The number of prototypes is dynamic depending on the pattern structure of the feature space. We evaluate the proposed prototype-based network on three widely used polyp segmentation datasets with abundant qualitative and quantitative experiments. Experimental results show that our proposed method outperforms the CNN-based and transformer-based state-of-the-art methods in terms of both effectiveness and fairness metrics. Moreover, extensive ablation studies are conducted to show the effectiveness of each proposed component and various parameter values. Lastly, we analyze how the number of prototypes grows during the training process and visualize the associated subgroups for each learned prototype. The code and data will be released at https://github.com/zijinY/dynamic-prototype-debiasing . Kongming Liang, Zijin Yin, Min Min, Zhanyu Ma, Jun Guo 0002 |
Int. J. Comput. Vis. | 5 |
| 2024 | Semi-Supervised Learning for FGVC With Out-of-Category DataabstractDespite great strides made on fine-grained visual classification (FGVC), current methods are still heavily reliant on fully-supervised paradigms where ample expert labels are called for. Semi-supervised learning (SSL) techniques, acquiring knowledge from unlabeled data, provide a considerable means forward and have shown great promise for coarse-grained problems. However, exiting SSL paradigms mostly assume in-category (i.e., category-aligned) unlabeled data, which hinders their effectiveness when re-proposed on FGVC. In this paper, we put forward a novel design specifically aimed at making out-of-category data work for semi-supervised FGVC. We work off an important assumption that all fine-grained categories naturally follow a hierarchical structure (e.g., the phylogenetic tree of "Aves" that covers all bird species). It follows that, instead of operating on individual samples, we can instead predict sample relations within this tree structure as the optimization goal of SSL. Beyond this, we further introduced two strategies uniquely brought by these tree structures to achieve inter-sample consistency regularization and reliable pseudo-relation. Our experimental results reveal that (i) the proposed method yields good robustness against out-of-category data, and (ii) it can be equipped with prior arts, boosting their performance thus yielding state-of-the-art results. Ruoyi Du, Dongliang Chang, Zhanyu Ma, Kongming Liang, Yi-Zhe Song, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Bi-Directional Ensemble Feature Reconstruction Network for Few-Shot Fine-Grained ClassificationabstractThe main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting - a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. We introduce the snapshot ensemble method in the episodic learning strategy - a simple trick to further improve model performance without increasing training costs. Experimental results on three widely used fine-grained image classification datasets, as well as general and cross-domain few-shot image datasets, consistently show considerable improvements compared with other methods. Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | A simple scheme to amplify inter-class discrepancy for improving few-shot fine-grained image classificationabstractFew-shot image classification is a challenging topic in pattern recognition and computer vision. Few-shot fine-grained image classification is even more challenging, due to not only the few shots of labelled samples but also the subtle differences to distinguish subcategories in fine-grained images. A recent method called task discrepancy maximisation (TDM) can be embedded into the feature map reconstruction network (FRN) to generate discriminative features, by preserving the appearance details through reconstructing the query image and then assigning higher weights to more discriminative channels, producing the state-of-the-art performance for few-shot fine-grained image classification. However, due to the small inter-class discrepancy in fine-grained images and the small training set in few-shot learning, the training of FRN+TDM can result in excessively flexible boundaries between subcategories and hence overfitting. To resolve this problem, we propose a simple scheme to amplify inter-class discrepancy and thus improve FRN+TDM. To achieve this aim, instead of developing new modules, our scheme only involves two simple amendments to FRN+TDM: relaxing the inter-class score in TDM, and adding a centre loss to FRN. Extensive experiments on five benchmark datasets showcase that, although embarrassingly simple, our scheme is quite effective to improve the performance of few-shot fine-grained image classification. The code is available at https://github.com/Airgods/AFRN.git. Zijie Guo, Rui Zhu 0006, Zhanyu Ma, Jun Guo 0002, Jing-Hao Xue |
Pattern Recognit. | 4 |
| 2024 | Self-reconstruction network for fine-grained few-shot classificationabstractMetric-based methods are one of the most common methods to solve the problem of few-shot image classification. However, traditional metric-based few-shot methods suffer from overfitting and local feature misalignment. The recently proposed feature reconstruction-based approach, which reconstructs query image features from the support set features of a given class and compares the distance between the original query features and the reconstructed query features as the classification criterion, effectively solves the feature misalignment problem. However, the issue of overfitting still has not been considered. To this end, we propose a self-reconstruction metric module for diversifying query features and a restrained cross-entropy loss for avoiding over-confident predictions. By introducing them, the proposed self-reconstruction network can effectively alleviate overfitting. Extensive experiments on five benchmark fine-grained datasets demonstrate that our proposed method achieves state-of-the-art performance on both 5-way 1-shot and 5-way 5-shot classification tasks. Code is available at https://github.com/liz-lut/SRM-main. Zhen Li 0026, Jiyang Xie 0001, Jing-Hao Xue, Zhanyu Ma |
Pattern Recognit. | 6 |
| 2024 | Generating Accurate and Diverse Audio Captions Through Variational Autoencoder FrameworkabstractGenerating both diverse and accurate descriptions is an essential goal in the audio captioning task. Traditional methods mainly focus on improving the accuracy of the generated captions but ignore their diversity. In contrast, recent methods have considered generating diverse captions for a given audio clip, but with the potential trade-off in caption accuracy. In this work, we propose a new diverse audio captioning method based on a variational autoencoder structure, dubbed AC-VAE, aiming to achieve a better trade-off between the diversity and accuracy of the generated captions. To improve diversity, AC-VAE learns the latent word distribution at each location based on contextual information. To uphold accuracy, AC-VAE incorporates an autoregressive prior module and a global constraint module, which enable precise modeling of word distribution and encourage semantic consistency of captions at the sentence level. We evaluate the proposed AC-VAE on the Clotho dataset. Experimental results show that AC-VAE achieves a better trade-off between diversity and accuracy compared to the state-of-the-art methods. The code is publicly available at https://github.com/XinMing0411/AC-VAE Yiming Zhang 0025, Ruoyi Du, Zheng-Hua Tan, Wenwu Wang 0001, Zhanyu Ma |
IEEE Signal Process. Lett. | 5 |
| 2024 | Mind the Gap: Open Set Domain Adaptation via Mutual-to-Separate FrameworkabstractUnsupervised domain adaptation aims to leverage labeled data from a source domain to learn a classifier for an unlabeled target domain. Amongst its many variants, open set domain adaptation (OSDA) is perhaps the most challenging one, as it further assumes the presence of unknown classes in the target domain. In this paper, we study OSDA with a particular focus on enriching its ability to traverse across larger domain gaps, and we show that existing state-of-the-art methods suffer a considerable performance drop in the presence of larger domain gaps, especially on a new dataset (PACS) that we re-purposed for OSDA. Exploring this is pivotal for OSDA as with increasing domain shift, identifying unknown samples in the target domain becomes harder for the model, thus making negative transfer between source and target domains more challenging. Accordingly, we propose a Mutual-to-Separate (MTS) framework to address the larger domain gaps. Essentially we design two networks – (a) Sample Separation Network (SSN): which is trained to learn a hyperplane for separating unknown samples from known ones, and (b) Distribution Matching Network (DMN): which is trained to maximise domain confusion between source and target domains without unknown samples under the guidance of the SSN. The key insight lies in how we exploit the mutually beneficial information between these two networks. On closer observation, we see that SSN can reveal which samples in the target domain belong to the unknown class by instance weighting whereas, DMN pushes apart the samples that most likely belong to the unknown class in the target domain, which in turn reduces the difficulty of SSN in identifying unknown samples. It follows that (a) and (b) will mutually supervise each other and alternate until convergence, which can better align the source and target domains in the shared label space. Extensive experiments on five datasets (Office-31, Office-Home, PACS, VisDA, andmini_DomainNet) demonstrate the efficiency of the proposed method. Detailed ablation experiments also validate the effectiveness of each component and the generality of the proposed framework. Codes are available at: https://github.com/PRIS-CV/Mutual-to-Separate. Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Ruiping Wang 0001, Jun Guo 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | LGR-NET: Language Guided Reasoning Network for Referring Expression ComprehensionabstractReferring Expression Comprehension(REC) is a fundamental task in the vision and language domain, which aims to locate an image region according to a natural language expression. REC requires the models to capture key clues in the text and perform accurate cross-modal reasoning. A recent trend employs transformer-based methods to address this problem. However, most of these methods typically treat image and text equally. They usually perform cross-modal reasoning in a crude way, and utilize textual features as a whole without detailed considerations (e.g., spatial information). This insufficient utilization of textual features will lead to sub-optimal results. In this paper, we propose aLanguage Guided Reasoning Network(LGR-NET) to fully utilize the guidance of the referring expression. To localize the referred object, we set a prediction token to capture cross-modal features. Furthermore, to sufficiently utilize the textual features, we extend them by our Textual Feature Extender (TFE) from three aspects.First, we design a novel coordinate embedding based on textual features. The coordinate embedding is incorporated to the prediction token to promote its capture of language-related visual features.Second, we employ the extracted textual features for Text-guided Cross-modal Alignment (TCA) and Fusion (TCF), alternately.Third, we devise a novel cross-modal loss to enhance cross-modal alignment between the referring expression and the learnable prediction token. We conduct extensive experiments on five benchmark datasets, and the experimental results show that our LGR-NET achieves a new state-of-the-art. Source code is available at https://github.com/lmc8133/LGR-NET. Mingcong Lu, Ruifan Li, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | CreativeSeg: Semantic Segmentation of Creative SketchesabstractThe problem of sketch semantic segmentation is far from being solved. Despite existing methods exhibiting near-saturating performances on simple sketches with high recognisability, they suffer serious setbacks when the target sketches are products of an imaginative process with high degree of creativity. We hypothesise that human creativity, being highly individualistic, induces a significant shift in distribution of sketches, leading to poor model generalisation. Such hypothesis, backed by empirical evidences, opens the door for a solution that explicitly disentangles creativity while learning sketch representations. We materialise this by crafting a learnable creativity estimator that assigns a scalar score of creativity to each sketch. It follows that we introduce CreativeSeg, a learning-to-learn framework that leverages the estimator in order to learn creativity-agnostic representation, and eventually the downstream semantic segmentation task. We empirically verify the superiority of CreativeSeg on the recent "Creative Birds" and "Creative Creatures" creative sketch datasets. Through a human study, we further strengthen the case that the learned creativity score does indeed have a positive correlation with the subjective creativity of human. Codes are available at https://github.com/PRIS-CV/Sketch-CS. Yixiao Zheng, Kaiyue Pang, Ayan Das 0003, Dongliang Chang, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 6 |
| 2024 | DualGCN: Exploring Syntactic and Semantic Information for Aspect-Based Sentiment AnalysisabstractThe task of aspect-based sentiment analysis aims to identify sentiment polarities of given aspects in a sentence. Recent advances have demonstrated the advantage of incorporating the syntactic dependency structure with graph convolutional networks (GCNs). However, their performance of these GCN-based methods largely depends on the dependency parsers, which would produce diverse parsing results for a sentence. In this article, we propose a dual GCN (DualGCN) that jointly considers the syntax structures and semantic correlations. Our DualGCN model mainly comprises four modules: 1) SynGCN: instead of explicitly encoding syntactic structure, the SynGCN module uses the dependency probability matrix as a graph structure to implicitly integrate the syntactic information; 2) SemGCN: we design the SemGCN module with multihead attention to enhance the performance of the syntactic structure with the semantic information; 3) Regularizers: we propose orthogonal and differential regularizers to precisely capture semantic correlations between words by constraining attention scores in the SemGCN module; and 4) Mutual BiAffine: we use the BiAffine module to bridge relevant information between the SynGCN and SemGCN modules. Extensive experiments are conducted compared with up-to-date pretrained language encoders on two groups of datasets, one including Restaurant14, Laptop14, and Twitter and the other including Restaurant15 and Restaurant16. The experimental results demonstrate that the parsing results of various dependency parsers affect their performance of the GCN-based models. Our DualGCN model achieves superior performance compared with the state-of-the-art approaches. The source code and preprocessed datasets are provided and publicly available on GitHub (see https://github.com/CCChenhao997/DualGCN-ABSA). Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image ClassificationabstractThe main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting -- a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we for the first time introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. Experimental results on three widely used fine-grained image classification datasets consistently show considerable improvements compared with other methods. Codes are available at: https://github.com/PRIS-CV/Bi-FRN. Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song |
AAAI | 5 |
| 2023 | LaDA: Latent Dialogue Action For Zero-shot Cross-lingual Neural Network Language Modeling
Zhanyu Ma, Shuang Cheng |
CogSci | 1 |
| 2023 | An Erudite Fine-Grained Visual Classification ModelabstractCurrent fine-grained visual classification (FGVC) models are isolated. In practice, we first need to identify the coarse-grained label of an object, then select the corresponding FGVC model for recognition. This hinders the application of FGVC algorithms in real-life scenarios. In this paper, we propose an erudite FGVC model jointly trained by several different datasets11In this paper, different datasets mean different fine-grained visual classification datasets., which can efficiently and accurately predict an object's fine-grained label across the combined label space. We found through a pilot study that positive and negative transfers co-occur when different datasets are mixed for training, i.e., the knowledge from other datasets is not always useful. Therefore, we first propose a feature disentanglement module and a feature re-fusion module to reduce negative transfer and boost positive transfer between different datasets. In detail, we reduce negative transfer by decoupling the deep features through many dataset-specific feature extractors. Subsequently, these are channel-wise re-fused to facilitate positive transfer. Finally, we propose a meta-learning based dataset-agnostic spatial attention layer to take full advantage of the multi-dataset training data, given that localisation is dataset-agnostic between different datasets. Experimental results across 11 different mixed-datasets built on four different FGVC datasets demonstrate the effectiveness of the proposed method. Furthermore, the proposed method can be easily combined with existing FGVC methods to obtain state-of-the-art results. Our code is available at https://github.com/PRIS-CV/An-Erudite-FGVC-Model. Dongliang Chang, Yujun Tong, Ruoyi Du, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 6 |
| 2023 | On-the-Fly Category DiscoveryabstractAlthough machines have surpassed humans on visual recognition problems, they are still limited to providing closed-set answers. Unlike machines, humans can cognize novel categories at the first observation. Novel category discovery (NCD) techniques, transferring knowledge from seen categories to distinguish unseen categories, aim to bridge the gap. However, current NCD methods assume a transductive learning and offline inference paradigm, which restricts them to a predefined query set and renders them unable to deliver instant feedback. In this paper, we study on-the-fly category discovery (OCD) aimed at making the model instantaneously aware of novel category samples (i.e., enabling inductive learning and streaming inference). We first design a hash coding-based expandable recognition model as a practical baseline. Afterwards, noticing the sensitivity of hash codes to intra-category variance, we further propose a novel Sign-Magnitude dIsentangLEment (SMILE) architecture to alleviate the disturbance it brings. Our experimental results demonstrate the superiority of SMILE against our baseline model and prior art. Our code is available at https://github.com/PRIS-CV/On-the-fly-Category-Discovery. Ruoyi Du, Dongliang Chang, Kongming Liang, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
CVPR | 6 |
| 2023 | Super-Resolution Information Enhancement for Crowd CountingabstractCrowd counting is a challenging task due to the heavy occlusions, scales, and density variations. Existing methods handle these challenges effectively while ignoring low-resolution (LR) circumstances. The LR circumstances weaken the counting performance deeply for two crucial reasons: 1) limited detail information; 2) overlapping head regions accumulate in density maps and result in extreme ground-truth values. An intuitive solution is to employ super-resolution (SR) pre-processes for the input LR images. However, it complicates the inference steps and thus limits application potentials when requiring real-time. We propose a more elegant method termed Multi-Scale Super-Resolution Module (MSSRM). It guides the network to estimate the lost details and enhances the detailed information in the feature space. Noteworthy that the MSSRM is plug-in plug-out and deals with the LR problems with no inference cost. As the proposed method requires SR labels, we further propose a Super-Resolution Crowd Counting dataset (SR-Crowd). Extensive experiments on three datasets demonstrate the superiority of our method. The code will be available at https://github.com/PRIS-CV/MSSRM.git. Wei Xu 0037, Dingkang Liang, Zhanyu Ma, Kongming Liang, Ling Jin 0004 |
ICASSP | 4 |
| 2023 | Semantic Memory Guided Image Representation for Polyp SegmentationabstractPolyp segmentation is important in the early diagnosis and treatment of colorectal cancer. Since polyps vary in shape, size, color, and texture, accurate polyp segmentation is very challenging. One promising solution is to model the contextual relation for each pixel. However, previous methods only focus on learning the dependencies between the position within an individual image and ignore the contextual relation across different images. In this paper, we propose a memory-based feature enhancement module to capture the cross-image contextual relations. Specifically, we first present a polyp-centric representation. Then a semantic memory is designed to extract the polyp prototypes across different images. The feature at one position can be further enhanced by the contextual embeddings stored in the semantic memory. The enhanced feature is propagated into the features of the previous levels as the multi-scale guidance. The experimental results show that our method achieves better performance than other state-of-the-art methods. Zijin Yin, Runpu Wei, Kongming Liang, Yiyang Lin, Zhanyu Ma, Min Min, Jun Guo 0002 |
ICASSP | 6 |
| 2023 | Multi-View Active Fine-Grained Visual RecognitionabstractDespite the remarkable progress of Fine-grained visual classification (FGVC) with years of history, it is still limited to recognizing 2D images. Recognizing objects in the physical world (i.e., 3D environment) poses a unique challenge – discriminative information is not only present in visible local regions but also in other unseen views. Therefore, in addition to finding the distinguishable part from the current view, efficient and accurate recognition requires inferring the critical perspective with minimal glances. E.g., a person might recognize a "Ford sedan" with a glance at its side and then know that looking at the front can help tell which model it is. In this paper, towards FGVC in the real physical world, we put forward the problem of multi-view active fine-grained visual recognition (MAFR) and complete this study in three steps: (i) a multi-view, fine-grained vehicle dataset is collected as the testbed, (ii) a pilot experiment is designed to validate the need and research value of MAFR, (iii) a policy-gradient-based framework along with a dynamic exiting strategy is proposed to achieve efficient recognition with active view selection. Our comprehensive experiments demonstrate that the proposed method outperforms previous multi-view recognition works and can extend existing state-of-the-art FGVC methods and advanced neural networks to become "FGVC experts" in the 3D environment. Our code is available at https://github.com/PRIS-CV/MAFR. Ruoyi Du, Wenqing Yu, Heqing Wang, Ting-En Lin, Dongliang Chang, Zhanyu Ma |
ICCV | 6 |
| 2023 | Task-aware Adaptive Learning for Cross-domain Few-shot LearningabstractAlthough existing few-shot learning works yield promising results for in-domain queries, they still suffer from weak cross-domain generalization. Limited support data requires effective knowledge transfer, but domain-shift makes this harder. Towards this emerging challenge, researchers improved adaptation by introducing task-specific parameters, which are directly optimized and estimated for each task. However, adding a fixed number of additional parameters fails to consider the diverse domain shifts between target tasks and the source domain, limiting efficacy. In this paper, we first observe the dependence of task-specific parameter configuration on the target task. Abundant task-specific parameters may over-fit, and insufficient task-specific parameters may result in under-adaptation – but the optimal task-specific configuration varies for different test tasks. Based on these findings, we propose the Task-aware Adaptive Network (TA2-Net), which is trained by reinforcement learning to adaptively estimate the optimal task-specific parameter configuration for each test task. It learns, for example, that tasks with significant domain-shift usually have a larger need for task-specific parameters for adaptation. We evaluate our model on Meta-dataset. Empirical results show that our model outperforms existing state-of-the-art methods. Our code is available at https://github.com/PRIS-CV/TA2-Net. Yurong Guo 0001, Ruoyi Du, Timothy M. Hospedales, Yi-Zhe Song, Zhanyu Ma |
ICCV | 6 |
| 2023 | Multi-semantic Fusion Model For Generalized Zero-Shot Skeleton-Based Action Recognition
Ming-Zhe Li, Zhang Zhang 0001, Zhanyu Ma, Liang Wang 0001 |
ICIG (1) | 4 |
| 2023 | Self-Enhanced Training Framework for Referring Expression GroundingabstractWeakly-supervised referring expression grounding (REG) aims at locating the image region described by a query sentence, where the mapping between the referential region and query is not available during the training stage. Noticing the significant gap between the fully- and weakly-supervised approaches, we develop a Self-Enhanced Training(SET) framework in this paper. Specifically, we first train the network under a weakly-supervised setting. Then, the model outputs are collected and filtered according to the confidence score and serve as pseudo-labels. Finally, with the help of these pseudo-labels, we tune the model under a fully-supervised setting. The SET framework provides a simple way of generating pseudo-labels that build a bridge between weak and full supervision. Experimental results demonstrate that model trained through our SET framework outperforms existing traditional methods on RefCOCO, RefCOCO+, and RefCOCOg datasets. The code is available at https://github.com/HTDL98/SET-framework. Ruoyi Du, Kongming Liang, Zhanyu Ma |
ICIP | 4 |
| 2023 | Attribute Learning with Knowledge Enhanced Partial AnnotationsabstractUnder limited annotation cost, large-scale attribute learning datasets only contain partial labels for each image. The conventional methods treat the un-annotated attributes as negative or ignore their loss without considering the associated knowledge. In this paper, we present a knowledge enhanced selective loss for partially labeled attribute learning. Given a visual instance, we investigate the object-attribute co-occurrence as internal knowledge to subdivide the unannotated attributes into feasible and infeasible sets. Based on that, we can enhance the model to focus on the learning of feasible un-annotated attributes and remove the distraction from the infeasible ones. Besides the internal knowledge, we adopt external knowledge to excavate the unseen object-attribute pairs. Experimental results show that our proposed loss can achieve state-of-the-art performance on the newly cleaned VAW2 dataset that contains 170,407 instances, 1763 objects, and 591 attributes. The code and VAW2 dataset are available at https://github.com/GriffinLiang/seal. Kongming Liang, Wei Chen 0071, Zhanyu Ma, Jun Guo 0002 |
ICIP | 5 |
| 2023 | KSRL: Knowledge Selection Based Reinforcement Learning for Knowledge-Grounded Dialogue
Zhanyu Ma, Shuang Cheng |
KSEM (4) | 1 |
| 2023 | Category-Specific Prompts for Animal Action Recognition with Pretrained Vision-Language ModelsabstractAnimal action recognition has a wide range of applications. However, the field largely remains unexplored due to the greater challenges compared to human action recognition, such as lack of annotated training data, large intra-class variation, and interference of cluttered background. Most of the existing methods directly apply human action recognition techniques, which essentially require a large amount of annotated data. In recent years, contrastive vision-language pretraining has demonstrated strong zero-shot generalization ability and has been used for human action recognition. Inspired by the success, we develop a highly performant action recognition framework based on the CLIP model. Our model addresses the above challenges via a novel category-specific prompting module to generate adaptive prompts for both text and video based on the animal category detected in input videos. On one hand, it can generate more precise and customized textual descriptions for each action and animal category pair, being helpful in the alignment of textual and visual space. On the other hand, it allows the model to focus on video features of the target animal in the video and reduce the interference of video background noise. Experimental results demonstrate that our method outperforms five previous action recognition methods on the Animal Kingdom dataset and has shown best generalization ability on unseen animals. Yinuo Jing, Chunyu Wang 0001, Ruxu Zhang, Kongming Liang, Zhanyu Ma |
ACM Multimedia | 5 |
| 2023 | Hierarchical Visual Attribute Learning in the WildabstractObserving objects' attributes at different levels of detail is a fundamental aspect of how humans perceive and understand the world around them. Existing studies focused on attribute prediction in a flat way, but they overlook the underlying attribute hierarchy, e.g., navy blue is a subcategory of blue. In recent years, large language models, e.g., ChatGPT, have emerged with the ability to perform an extensive range of natural language processing tasks like text generation and classification. The factual knowledge learned by LLM can assist us build the hierarchical relations of visual attributes in the wild. Based on that, we propose a model called the object-specific attribute relation net, which takes advantage of three types of relations among attributes - positive, negative, and hierarchical - to better facilitate attribute recognition in images. Guided by the extracted hierarchical relations, our model can predict attributes from coarse to fine. Additionally, we introduce several evaluation metrics for attribute hierarchy to comprehensively assess the model's ability to comprehend hierarchical relations. Our extensive experiments demonstrate that our proposed hierarchical annotation brings improvements to the model's understanding of hierarchical relations of attributes, and the object-specific attribute relation net can recognize visual attributes more accurately. Kongming Liang, Haiwen Zhang, Zhanyu Ma, Jun Guo 0002 |
ACM Multimedia | 4 |
| 2023 | Focus the Overlapping Problem on Few-Shot Object Detection via Multiple Predictions
Mandan Guan, Wenqing Yu, Yurong Guo 0001, Keyan Huang, Jiaxun Zhang, Kongming Liang, Zhanyu Ma |
PRCV (2) | 7 |
| 2023 | Plugging Stylized Controls in Open-Stylized Image Captioning
Yixiao Zheng, Ruoyi Du, Yiming Zhang 0025, Kongming Liang, Zhanyu Ma |
PRCV (1) | 6 |
| 2023 | Image Generation Based Intra-class Variance Smoothing for Fine-Grained Visual Classification
Ruoyi Du, Kongming Liang, Wei Chen 0071, Zhanyu Ma |
PRCV (6) | 6 |
| 2023 | ReNAP: Relation network with adaptiveprototypical learning for few-shot classification
Yalan Li, Yixiao Zheng, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014 |
Neurocomputing | 5 |
| 2023 | Making a Bird AI Expert Work for You and MeabstractAs powerful as fine-grained visual classification (FGVC) is, responding your query with a bird name of “Whip-poor-will” or “Mallard” probably does not make much sense. This however commonly accepted in the literature, underlines a fundamental question interfacing AI and human – what constitutes transferable knowledge for human to learn from AI? This paper sets out to answer this very question using FGVC as a test bed. Specifically, we envisage a scenario where a trained FGVC model (the AI expert) functions as a knowledge provider in enabling average people (you and me) to become better domain experts ourselves,i.e.,those capable in distinguishing between “Whip-poor-will” and “Mallard”. Fig. 1 lays out our approach in answering this question. Assuming an AI expert trained using expert human labels, we ask (i) what is the best transferable knowledge we can extract from AI, and (ii) what is the most practical means to measure the gains in expertise given that knowledge? On the former, we propose to represent knowledge as highly discriminative visual regions that are expert-exclusive. For that, we devise a multi-stage learning framework, which starts with modelling visual attention of domain experts and novices separately, before discriminatively distilling their differences to acquire those exclusive to experts. For the latter, we simulate the evaluation process as a book guide to best accommodate the learning practice of that is accustomed to humans. A comprehensive human study of 15,000 trials shows our method is able to consistently improve people of divergent bird expertise to recognise once unrecognisable birds. To counter the lack of reproducibility of perceptual studies, and in turn to make a sustainable direction out of our “AI for Human” effort, we further propose a quantitative metric, namely Transferable Effective Model Attention (TEMI). TEMI acts as a crude but benchmarkable metric to replace large-scale human studies, and therefore allows future efforts in this direction to be comparable to ours. We attest to the integrity of TEMI by (i) empirically showing a strong correlation between TEMI scores and raw human study data, and (ii) its expected behaviour holds for a large body of attention models. Last but not least, our approach also leads to improved FGVC performance in the conventional benchmarking sense, when the extracted knowledge defined is utilised as means to achieve discriminative localisation. Codes and all details on the human study are available at:https://github.com/PRIS-CV/Making-a-Bird-AI-Expert-Work-for-You-and-Me. Dongliang Chang, Kaiyue Pang, Ruoyi Du, Yujun Tong, Yi-Zhe Song, Zhanyu Ma, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Deep metric learning for few-shot image classification: A Review of recent developments
Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 3 |
| 2023 | ACTUAL: Audio Captioning With Caption Feature Space RegularizationabstractAudio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio content, different people may perceive the same audio clip differently, resulting in caption disparities (i.e., the same audio clip may be described by several captions with diverse semantics). In the literature, the one-to-many strategy is often employed to train the audio captioning models, where a related caption is randomly selected as the optimization target for each audio clip at each training iteration. However, we observe that this can lead to significant variations during the optimization process and adversely affect the performance of the model. In this paper, we address this issue by proposing an audio captioning method, named ACTUAL (Audio Captioning with capTion featUre spAce reguLarization). ACTUAL involves a two-stage training process: (i) in the first stage, we use contrastive learning to construct a proxy feature space where the similarities between captions at the audio level are explored, and (ii) in the second stage, the proxy feature space is utilized as additional supervision to improve the optimization of the model in a more stable direction. We conduct extensive experiments to demonstrate the effectiveness of the proposed ACTUAL method. The results show that proxy caption embedding can significantly improve the performance of the baseline model and the proposed ACTUAL method offers competitive performance on two datasets compared to state-of-the-art methods. The code is publicly available athttps://github.com/PRIS-CV/Caption-Feature-Space-Regularization. Yiming Zhang 0025, Hong Yu 0006, Ruoyi Du, Zheng-Hua Tan, Wenwu Wang 0001, Zhanyu Ma |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Locally-Enriched Cross-Reconstruction for Few-Shot Fine-Grained Image ClassificationabstractFew-shot fine-grained image classification has attracted considerable attention in recent years for its realistic setting to imitate how humans conduct recognition tasks. Metric-based few-shot classifiers have achieved high accuracies. However, their metric function usually requires two arguments of vectors, while transforming or reshaping three-dimensional feature maps to vectors can result in loss of spatial information. Image reconstruction is thus involved to retain more appearance details: the test images are reconstructed by different classes and then classified to the one with the smallest reconstruction error. However, discriminative local information, vital to distinguish sub-categories in fine-grained images with high similarities, is not well elaborated when only the base features from a usual embedding module are adopted for reconstruction. Hence, we propose the novel local content-enriched cross-reconstruction network (LCCRN) for few-shot fine-grained classification. In LCCRN, we design two new modules: the local content-enriched module (LCEM) to learn the discriminative local features, and the cross-reconstruction module (CRM) to fully engage the local features with the appearance details obtained from a separate embedding module. The classification score is calculated based on the weighted sum of reconstruction errors of the cross-reconstruction tasks, with weights learnt from the training process. Extensive experiments on four fine-grained datasets showcase the superior classification performance of LCCRN compared with the state-of-the-art few-shot classification methods. Codes are available at:https://github.com/lutsong/LCCRN. Jijie Wu, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | A Real-Time Memory Updating Strategy for Unsupervised Person Re-IdentificationabstractRecently, clustering-based methods have been the dominant solution for unsupervised person re-identification (ReID). Memory-based contrastive learning is widely used for its effectiveness in unsupervised representation learning. However, we find that the inaccurate cluster proxies and the momentum updating strategy do harm to the contrastive learning system. In this paper, we propose a real-time memory updating strategy (RTMem) to update the cluster centroid with a randomly sampled instance feature in the current mini-batch without momentum. Compared to the method that calculates the mean feature vectors as the cluster centroid and updating it with momentum, RTMem enables the features to be up-to-date for each cluster. Based on RTMem, we propose two contrastive losses, i.e., sample-to-instance and sample-to-cluster, to align the relationships between samples to each cluster and to all outliers not belonging to any other clusters. On the one hand, sample-to-instance loss explores the sample relationships of the whole dataset to enhance the capability of density-based clustering algorithm, which relies on similarity measurement for the instance-level images. On the other hand, with pseudo-labels generated by the density-based clustering algorithm, sample-to-cluster loss enforces the sample to be close to its cluster proxy while being far from other proxies. With the simple RTMem contrastive learning strategy, the performance of the corresponding baseline is improved by 9.3% on Market-1501 dataset. Our method consistently outperforms state-of-the-art unsupervised learning person ReID methods on three benchmark datasets. Code is made available at:https://github.com/PRIS-CV/RTMem. Junhui Yin, Xinyu Zhang 0015, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2023 | Sketch-Segformer: Transformer-Based Segmentation for Figurative and Creative SketchesabstractSketch is a well-researched topic in the vision community by now. Sketch semantic segmentation in particular, serves as a fundamental step towards finer-level sketch interpretation. Recent works use various means of extracting discriminative features from sketches and have achieved considerable improvements on segmentation accuracy. Common approaches for this include attending to the sketch-image as a whole, its stroke-level representation or the sequence information embedded in it. However, they mostly focus on only a part of such multi-facet information. In this paper, we for the first time demonstrate that there is complementary information to be explored across all these three facets of sketch data, and that segmentation performance consequently benefits as a result of such exploration of sketch-specific information. Specifically, we propose the Sketch-Segformer, a transformer-based framework for sketch semantic segmentation that inherently treats sketches as stroke sequences other than pixel-maps. In particular, Sketch-Segformer introduces two types of self-attention modules having similar structures that work with different receptive fields (i.e., whole sketch or individual stroke). The order embedding is then further synergized with spatial embeddings learned from the entire sketch as well as localized stroke-level information. Extensive experiments show that our sketch-specific design is not only able to obtain state-of-the-art performance on traditional figurative sketches (such as SPG, SketchSeg-150K datasets), but also performs well on creative sketches that do not conform to conventional object semantics (CreativeSketch dataset) thanks for our usage of multi-facet sketch information. Ablation studies, visualizations, and invariance tests further justifies our design choice and the effectiveness of Sketch-Segformer. Codes are available at https://github.com/PRIS-CV/Sketch-SF. Yixiao Zheng, Jiyang Xie 0001, Aneeshan Sain, Yi-Zhe Song, Zhanyu Ma |
IEEE Trans. Image Process. | 5 |
| 2023 | On the Comparisons of Decorrelation Approaches for Non-Gaussian Neutral Vector Variablesabstract-norm equals one. In addition, its neutral properties make it significantly different from the commonly studied vector variables (e.g., the Gaussian vector variables). Due to the aforementioned properties, the conventionally applied linear transformation approaches [e.g., principal component analysis (PCA) and independent component analysis (ICA)] are not suitable for neutral vector variables, as PCA cannot transform a neutral vector variable, which is highly negatively correlated, into a set of mutually independent scalar variables and ICA cannot preserve the bounded property after transformation. In recent work, we proposed an efficient nonlinear transformation approach, i.e., the parallel nonlinear transformation (PNT), for decorrelating neutral vector variables. In this article, we extensively compare PNT with PCA and ICA through both theoretical analysis and experimental evaluations. The results of our investigations demonstrate the superiority of PNT for decorrelating the neutral vector variables. Zhanyu Ma, Xiaoou Lu, Jiyang Xie 0001, Zhen Yang 0004, Jing-Hao Xue, Zheng-Hua Tan, Bo Xiao 0006, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | HCLD: A Hierarchical Framework for Zero-shot Cross-lingual Dialogue SystemabstractRecently, many task-oriented dialogue systems need to serve users in different languages. However, it is time-consuming to collect enough data of each language for training. Thus, zero-shot adaptation of cross-lingual task-oriented dialog systems has been studied. Most of existing methods consider the word-level alignments to conduct two main tasks for task-oriented dialogue system, i.e., intent detection and slot filling, and they rarely explore the dependency relations among these two tasks. In this paper, we propose a hierarchical framework to classify the pre-defined intents in the high-level and fulfill slot filling under the guidance of intent in the low-level. Particularly, we incorporate sentence-level alignment among different languages to enhance the performance of intent detection. The extensive experiments report that our proposed method achieves the SOTA performance on a public task-oriented dialog dataset. Zhanyu Ma, Xurui Yang |
COLING | 1 |
| 2022 | ScaleNet: Searching for the Model to Scale
Jiyang Xie 0001, Xiu Su, Shan You, Zhanyu Ma, Fei Wang 0032, Chen Qian 0006 |
ECCV (21) | 4 |
| 2022 | Learning Invariant Visual Representations for Compositional Zero-Shot Learning
Tian Zhang 0029, Kongming Liang, Ruoyi Du, Zhanyu Ma, Jun Guo 0002 |
ECCV (24) | 5 |
| 2022 | Structured Dropconnect for Uncertainty Inference in Image ClassificationabstractUncertainty inference has become an important task to prove the reliability for deep neural networks. For image classification tasks, we propose a structured DropConnect (SDC) framework to model the output of a deep neural network into a distribution. We introduce a DropConnect strategy in a structured manner in the fully connected layers, split the network into several sub-networks during testing, and choose the Dirichlet distribution to model the outputs of these sub-networks. The entropy of the parameterized Dirichlet distribution is finally utilized for uncertainty inference. In this paper, this framework is implemented on VGG16, and ResNet18 models for misclassification detection and open-set out-of-domain detection on CIFAR-10 and CIFAR-100 datasets. Experimental results show that the performance of the proposed SDC can be comparable to other uncertainty inference methods. Wenqing Zheng, Jiyang Xie 0001, Zhanyu Ma |
ICIP | 4 |
| 2022 | Improving Image Paragraph Captioning with Dual RelationsabstractImage paragraph captioning aims to generate multiple de-scriptive sentences for an image. However, most previous methods ignore the explicit relations among objects resulting in unsatisfactory performance. In this paper, we propose a novel model (i.e., DualRel) to capture spatial and seman-tic relations among objects. Specifically, the spatial relation embedding is obtained solely from images using a predefined geometry pattern. With the help of captions, the semantic relation embedding is learned in a weakly supervised man-ner. These two relation embeddings are then interacted with regional features of objects through a relation-aware attention interaction. It first obtains a visual context vector using regional features. Then with the visual context vector, we obtain the corresponding spatial and semantic relation-aware vectors using attentions. These three vectors are fused with two gates for language decoding to further generate a para-graph. Experimental results on Stanford benchmark dataset show that DualRel achieves remarkable improvements11Code released at https://github.com/fuyunll07/DualRel. Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006 |
ICME | 5 |
| 2022 | Domain Generalization via Frequency-domain-based Feature Disentanglement and InteractionabstractAdaptation to out-of-distribution data is a meta-challenge for all statistical learning algorithms that strongly rely on the i.i.d. assumption. It leads to unavoidable labor costs and confidence crises in realistic applications. For that, domain generalization aims at mining domain-irrelevant knowledge from multiple source domains that can generalize to unseen target domains. In this paper, by leveraging the frequency domain of an image, we uniquely work with two key observations: (i) the high-frequency information of an image depicts object edge structure, which preserves high-level semantic information of the object is naturally consistent across different domains, and (ii) the low-frequency component retains object smooth structure, while this information is susceptible to domain shifts. Motivated by the above observations, we introduce (i) an encoder-decoder structure to disentangle high- and low-frequency features of an image, (ii) an information interaction mechanism to ensure the helpful knowledge from both two parts can cooperate effectively, and (iii) a novel data augmentation technique that works on the frequency domain to encourage the robustness of frequency-wise feature disentangling. The proposed method obtains state-of-the-art performance on three widely used domain generalization benchmarks (Digit-DG, Office-Home, and PACS). Jingye Wang, Ruoyi Du, Dongliang Chang, Kongming Liang, Zhanyu Ma |
ACM Multimedia | 5 |
| 2022 | Complex Scenario-Oriented Fine-Grained Visual Classification PlatformabstractIn recent years, fine-grained visual classification (FGVC) algorithms have achieved excellent performance across a variety of datasets. However, it is still rare to see these algorithms applied in daily life. The main reasons for this are i) the algorithms are developed based on different design guidelines and cannot be deployed in the same environment; ii) there is not a simple and efficient platform to present the algorithm's results to the user - the accuracy is meaningless to the users. To address the above problem, we built a complex scenario-oriented fine-grained visual classification platform. The platform consists of a PyTorch-based fine-grained visual recognition algorithm library (FGL) and a WeChat applet-based user interaction module (WEM). We can quickly develop new algorithms or readily apply existing algorithms in the same environment through FGL. Driven by FGL, the WEM enables users to achieve fine-grained recognition of complex scenes interactively. In addition to showing the user the fine-grained labels of objects, we will also show how the model makes decisions to help the user master the ability to recognise the fine-grained object so that everyone can become a domain expert. A video demo shows an example of the proposed platform in a real-world scenario: https://reurl.cc/rRZE7O. Dongliang Chang, Junhan Chen, Ruoyi Du, Wenqing Yu, Yujun Tong, Kongming Liang, Yi-Zhe Song, Zhanyu Ma |
MMSP | 10 |
| 2022 | Multi-modal Human-machine Conversation System for Real Physical WorldabstractEnabling machines to process multi-modal information and understand the real physical world is an important step to achieving free human-machine conversation. However, previous human-machine conversation systems are mostly limited to single-modal (e.g., chat robot), single-round (e.g., visual Q & A), and static visual information (e.g., visual dialogue). To address the above problem, we develop a multi-modal human-machine conversation system for specific application scenarios. The system includes two modules: (i) an interactive visual grounding module that can actively disambiguate user's queries, and (ii) an interactive fine-grained recognition module that can model objects in the 3D environment and actively ask for missing visual information. A video demo of our system under the automobile sales scenario can be found here11https://drive.google.com/file/d/1IfBsMKq55ryLOZchIT6G-R5CyWnC3Skiew?usp=sharing. Shibo Nie, Mandan Guan, Ruoyi Du, Dongliang Chang, Kongming Liang, Zhanyu Ma |
MMSP | 8 |
| 2022 | Cross-Layer Feature based Multi-Granularity Visual ClassificationabstractIn contrast to traditional fine-grained visual clas-sification, multi-granularity visual classification is no longer limited to identifying the different sub-classes belonging to the same super-class (e.g., bird species, cars, and aircraft models). Instead, it gives a sequence of labels from coarse to fine (e.g., Passeriformes → Corvidae → Fish Crow), which is more convenient in practice. The key to solving this task is how to use the relationships between the different levels of labels to learn feature representations that contain different levels of granularity. Interestingly, the feature pyramid structure naturally implies different granularity of feature representation, with the shallow layers representing coarse-grained features and the deep layers representing fine-grained features. Therefore, in this paper, we exploit this property of the feature pyramid structure to decouple features and obtain feature representations corre-sponding to different granularities. Specifically, we use shallow features for coarse-grained classification and deep features for fine-grained classification. In addition, to enable fine-grained features to enhance the coarse-grained classification, we propose a feature reinforcement module based on the feature pyramid structure, where deep features are first upsampled and then combined with shallow features to make decisions. Experimental results on three widely used fine-grained image classification datasets such as CUB-200-2011, Stanford Cars, and FGVC-Aircraft validate the method's effectiveness. Code available at https://github.com/PRIS-CV/CGVC. Junhan Chen, Dongliang Chang, Jiyang Xie 0001, Ruoyi Du, Zhanyu Ma |
VCIP | 5 |
| 2022 | ENDE-GNN: An Encoder-decoder GNN Framework for Sketch Semantic SegmentationabstractSketch semantic segmentation serves as an important part of sketch interpretation. Recently, some researchers have obtained significant results using graph neural networks (GNN) for this task. However, existing GNN-based methods usually neglect the drawing order of sketches thus missing out the sequence information inherent to sketches. Towards solving this problem to achieve better performance on sketch semantic segmentation, we propose an encoder-decoder GNN framework named ENDE-GNN. Working with an auxiliary decoder, our ENDE-GNN guides the GNN backbone network to not only extract the inter-stroke and intra-stroke features, but also pays attention to the drawing order of sketches. This decoder acts during training only, preventing any additional overhead during testing. The proposed ENDE-GNN obtains state-of-the-art per-formances on three public sketch semantic segmentation datasets, namely SPG, SketchSeg-150K, and CreativeSketch. We further evaluate the effectiveness of ENDE-GNN via ablation studies and visualizations. Codes are available at https://github.com/PRIS-CV/ENDE_For_SSS. Yixiao Zheng, Jiyang Xie 0001, Aneeshan Sain, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
VCIP | 4 |
| 2022 | Dual-granularity feature alignment for cross-modality person re-identification
Junhui Yin, Zhanyu Ma, Jiyang Xie 0001, Shibo Nie, Kongming Liang, Jun Guo 0002 |
Neurocomputing | 2 |
| 2022 | Progressive Learning of Category-Consistent Multi-Granularity Features for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) is much more challenging than traditional classification tasks due to the inherently subtle intra-class object variations. Recent works are mainly part-driven (either explicitly or implicitly), with the assumption that fine-grained information naturally rests within the parts. In this paper, we take a different stance, and show that part operations are not strictly necessary - the key lies with encouraging the network to learn at different granularities and progressively fusing multi-granularity features together. In particular, we propose: (i) a progressive training strategy that effectively fuses features from different granularities, and (ii) a consistent block convolution that encourages the network to learn the category-consistent features at specific granularities. We evaluate on several standard FGVC benchmark datasets, and demonstrate the proposed method consistently outperforms existing alternatives or delivers competitive results. Codes are available at https://github.com/PRIS-CV/PMG-V2. Ruoyi Du, Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Yi-Zhe Song, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | GPCA: A Probabilistic Framework for Gaussian Process Embedded Channel AttentionabstractChannel attention mechanisms have been commonly applied in many visual tasks for effective performance improvement. It is able to reinforce the informative channels as well as to suppress the useless channels. Recently, different channel attention modules have been proposed and implemented in various ways. Generally speaking, they are mainly based on convolution and pooling operations. In this paper, we propose Gaussian process embedded channel attention (GPCA) module and further interpret the channel attention schemes in a probabilistic way. The GPCA module intends to model the correlations among the channels, which are assumed to be captured by beta distributed variables. As the beta distribution cannot be integrated into the end-to-end training of convolutional neural networks (CNNs) with a mathematically tractable solution, we utilize an approximation of the beta distribution to solve this problem. To specify, we adapt a Sigmoid-Gaussian approximation, in which the Gaussian distributed variables are transferred into the interval [0,1]. The Gaussian process is then utilized to model the correlations among different channels. In this case, a mathematically tractable solution is derived. The GPCA module can be efficiently implemented and integrated into the end-to-end training of the CNNs. Experimental results demonstrate the promising performance of the proposed GPCA module. Codes are available at https://github.com/PRIS-CV/GPCA. Jiyang Xie 0001, Zhanyu Ma, Dongliang Chang, Guoqiang Zhang 0003, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Advanced Dropout: A Model-Free Methodology for Bayesian Dropout OptimizationabstractDue to lack of data, overfitting ubiquitously exists in real-world applications of deep neural networks (DNNs). We propose advanced dropout, a model-free methodology, to mitigate overfitting and improve the performance of DNNs. The advanced dropout technique applies a model-free and easily implemented distribution with parametric prior, and adaptively adjusts dropout rate. Specifically, the distribution parameters are optimized by stochastic gradient variational Bayes in order to carry out an end-to-end training. We evaluate the effectiveness of the advanced dropout against nine dropout techniques on seven computer vision datasets (five small-scale datasets and two large-scale datasets) with various base models. The advanced dropout outperforms all the referred techniques on all the datasets. We further compare the effectiveness ratios and find that advanced dropout achieves the highest one on most cases. Next, we conduct a set of analysis of dropout rate characteristics, including convergence of the adaptive dropout rate, the learned distributions of dropout masks, and a comparison with dropout rate generation without an explicit distribution. In addition, the ability of overfitting prevention is evaluated and confirmed. Finally, we extend the application of the advanced dropout to uncertainty inference, network pruning, text classification, and regression. The proposed advanced dropout is also superior to the corresponding referred methods. Codes are available at https://github.com/PRIS-CV/AdvancedDropout. Jiyang Xie 0001, Zhanyu Ma, Jianjun Lei 0001, Guoqiang Zhang 0003, Jing-Hao Xue, Zheng-Hua Tan, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | MPCCL: Multiview predictive coding with contrastive learning for person re-identification
Junhui Yin, Jiyang Xie 0001, Zhanyu Ma, Jun Guo 0002 |
Pattern Recognit. | 3 |
| 2022 | Unsupervised person re-identification via simultaneous clustering and mask prediction
Junhui Yin, Siqing Zhang 0001, Jiyang Xie 0001, Zhanyu Ma, Jun Guo 0002 |
Pattern Recognit. | 4 |
| 2022 | Exploring Local Detail Perception for Scene Sketch Semantic SegmentationabstractIn this paper, we aim to explore the fine-grained perception ability of deep models for the newly proposed scene sketch semantic segmentation task. Scene sketches are abstract drawings containing multiple related objects. It plays a vital role in daily communication and human-computer interaction. The study has only recently started due to a main obstacle of the absence of large-scale datasets. The currently available dataset SketchyScene is composed of clip art-style edge maps, which lacks abstractness and diversity. To drive further research, we contribute two new large-scale datasets based on real hand-drawn object sketches. A general automatic scene sketch synthesis process is developed to assist with new dataset composition. Furthermore, we propose to enhancing local detail perception in deep models to realize accurate stroke-oriented scene sketch segmentation. Due to the inherent differences between hand-drawn sketches and natural images, extreme low-level local features of strokes are incorporated to improve detail discrimination. Stroke masks are also integrated into model training to guide the learning attention. Extensive experiments are conducted on three large-scale scene sketch datasets. Our method achieves state-of-the-art performance under four evaluation metrics and yields meaningful interpretability via visual analytics. Ce Ge, Haifeng Sun 0001, Yi-Zhe Song, Zhanyu Ma, Jianxin Liao |
IEEE Trans. Image Process. | 4 |
| 2022 | Learning Calibrated Class Centers for Few-Shot Classification by Pair-Wise SimilarityabstractMetric-based methods achieve promising performance on few-shot classification by learning clusters on support samples and generating shared decision boundaries for query samples. However, existing methods ignore the inaccurate class center approximation introduced by the limited number of support samples, which consequently leads to biased inference. Therefore, in this paper, we propose to reduce the approximation error by class center calibration. Specifically, we introduce the so-called Pair-wise Similarity Module (PSM) to generate calibrated class centers adapted to the query sample by capturing the semantic correlations between the support and the query samples, as well as enhancing the discriminative regions on support representation. It is worth noting that the proposed PSM is a simple plug-and-play module and can be inserted into most metric-based few-shot learning models. Through extensive experiments in metric-based models, we demonstrate that the module significantly improves the performance of conventional few-shot classification methods on four few-shot image classification benchmark datasets. Codes are available at: https://github.com/PRIS-CV/Pair-wise-Similarity-module. Yurong Guo 0001, Ruoyi Du, Jiyang Xie 0001, Zhanyu Ma |
IEEE Trans. Image Process. | 5 |
| 2022 | Actor and Action Modular Network for Text-Based Video SegmentationabstractText-based video segmentation aims to segment an actor in video sequences by specifying the actor and its performing action with a textual query. Previous methods fail to explicitly align the video content with the textual query in a fine-grained manner according to the actor and its action, due to the problem of semantic asymmetry. The semantic asymmetry implies that two modalities contain different amounts of semantic information during the multi-modal fusion process. To alleviate this problem, we propose a novel actor and action modular network that individually localizes the actor and its action in two separate modules. Specifically, we first learn the actor-/action-related content from the video and textual query, and then match them in a symmetrical manner to localize the target tube. The target tube contains the desired actor and action which is then fed into a fully convolutional network to predict segmentation masks of the actor. Our method also establishes the association of objects cross multiple frames with the proposed temporal proposal aggregation mechanism. This enables our method to segment the video effectively and keep the temporal consistency of predictions. The whole model is allowed for joint learning of the actor-action matching and segmentation, as well as achieves the state-of-the-art performance for both single-frame segmentation and full video segmentation on A2D Sentences and J-HMDB Sentences datasets. Yan Huang 0008, Kai Niu 0002, Linjiang Huang, Zhanyu Ma, Liang Wang 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Dirichlet Process Mixture of Generalized Inverted Dirichlet Distributions for Positive Vector Data With Extended Variational InferenceabstractA Bayesian nonparametric approach for estimation of a Dirichlet process (DP) mixture of generalized inverted Dirichlet distributions [i.e., an infinite generalized inverted Dirichlet mixture model (InGIDMM)] has been proposed. The generalized inverted Dirichlet distribution has been proven to be efficient in modeling the vectors that contain only positive elements. Under the classical variational inference (VI) framework, the key challenge in the Bayesian estimation of InGIDMM is that the expectation of the joint distribution of data and variables cannot be explicitly calculated. Therefore, numerical methods are usually applied to simulate the optimal posterior distributions. With the recently proposed extended VI (EVI) framework, we introduce lower bound approximations to the original variational objective function in the VI framework such that an analytically tractable solution can be derived. Hence, the problem in numerical simulation has been overcome. By applying the DP mixture technique, an InGIDMM can automatically determine the number of mixture components from the observed data. Moreover, the DP mixture model with an infinite number of mixture components also avoids the problems of underfitting and overfitting. The performance of the proposed approach is demonstrated with both synthesized data and real-life data applications. Zhanyu Ma, Yuping Lai, Jiyang Xie 0001, Deyu Meng, W. Bastiaan Kleijn, Jun Guo 0002, Jingyi Yu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2021 | Dual Graph Convolutional Networks for Aspect-based Sentiment AnalysisabstractRuifan Li, Hao Chen, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang, Eduard Hovy. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Ruifan Li, Hao Chen 0041, Fangxiang Feng, Zhanyu Ma, Xiaojie Wang 0006, Eduard H. Hovy |
ACL/IJCNLP (1) | 4 |
| 2021 | Channel DropBlock: An Improved Regularization Method for Fine-Grained Visual Classification
Shuwei Dong, Yujun Tong, Zhanyu Ma, Haibin Ling |
BMVC | 4 |
| 2021 | Your "Flamingo" is My "Bird": Fine-Grained, or NotabstractWhether what you see in Figure 1 is a "flamingo" or a "bird", is the question we ask in this paper. While fine-grained visual classification (FGVC) strives to arrive at the former, for the majority of us non-experts just "bird" would probably suffice. The real question is therefore – how can we tailor for different fine-grained definitions under divergent levels of expertise. For that, we re-envisage the traditional setting of FGVC, from single-label classification, to that of top-down traversal of a pre-defined coarse-to-fine label hierarchy – so that our answer becomes "bird" ⇒ "Phoenicopteriformes" ⇒ "Phoenicopteridae" ⇒ "flamingo".To approach this new problem, we first conduct a comprehensive human study where we confirm that most participants prefer multi-granularity labels, regardless whether they consider themselves experts. We then discover the key intuition that: coarse-level label prediction exacerbates fine-grained feature learning, yet fine-level feature betters the learning of coarse-level classifier. This discovery enables us to design a very simple albeit surprisingly effective solution to our new problem, where we (i) leverage level-specific classification heads to disentangle coarse-level features with fine-grained ones, and (ii) allow finer-grained features to participate in coarser-grained label predictions, which in turn helps with better disentanglement. Experiments show that our method achieves superior performance in the new FGVC setting, and performs better than state-of-the-art on the traditional single-label FGVC problem as well. Thanks to its simplicity, our method can be easily implemented on top of any existing FGVC frameworks and is parameter-free. Dongliang Chang, Kaiyue Pang, Yixiao Zheng, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
CVPR | 4 |
| 2021 | Part Uncertainty Estimation Convolutional Neural Network For Person Re-IdentificationabstractDue to the large amount of noisy data in person re-identification (ReID) task, the ReID models are usually affected by the data uncertainty. Therefore, the deep uncertainty estimation method is important for improving the model robustness and matching accuracy. To this end, we propose a part-based uncertainty convolutional neural network (PUCNN), which introduces the part-based uncertainty estimation into the baseline model. On the one hand, PUCNN improves the model robustness to noisy data by distributilizing the feature embedding and constraining the part-based uncertainty. On the other hand, PUCNN improves the cumulative matching characteristics (CMC) performance of the model by filtering out low-quality training samples according to the estimated uncertainty score. The experiments on both non-video datasets, the noised Market-1501 and DukeMTMC, and video datasets, PRID2011, iLiDS-VID and MARS, demonstrate that our proposed method achieves encouraging and promising performance. Wenyu Sun, Jiyang Xie 0001, Jiayan Qiu, Zhanyu Ma |
ICIP | 4 |
| 2021 | CMF: Cascaded Multi-Model Fusion For Referring Image SegmentationabstractIn this work, we address the task of referring image segmentation (RIS), which aims at predicting a segmentation mask for the object described by a natural language expression. Most existing methods focus on establishing unidirectional or directional relationships between visual and linguistic features to associate two modalities together, while the multi-scale context is ignored or insufficiently modeled. Multi-scale context is crucial to localize and segment those objects that have large scale variations during the multi-modal fusion process. To solve this problem, we propose a simple yet effective Cascaded Multi-modal Fusion (CMF) module, which stacks multiple atrous convolutional layers in parallel and further introduces a cascaded branch to fuse visual and linguistic features. The cascaded branch can progressively integrate multi-scale contextual information and facilitate the alignment of two modalities during the multi-modal fusion process. Experimental results on four benchmark datasets demonstrate that our method outperforms most state-of-the-art methods. Code is available at https://github.com/jianhua2022/CMF-Refseg. Yan Huang 0008, Zhanyu Ma, Liang Wang 0001 |
ICIP | 3 |
| 2021 | Knowledge Transfer Based Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) aims to distinguish the sub-classes of the same category and its essential solution is to mine the subtle and discriminative regions. Convolution neural networks (CNNs), which employ the cross entropy loss (CE-loss) as the loss function, show poor performance since the model can only learn the most discriminative part and ignore other meaningful regions. Some existing works try to solve this problem by mining more discriminative regions by some detection techniques or attention mechanisms. However, most of them will meet the background noise problem when trying to find more discriminative regions. In this paper, we address it in a knowledge transfer learning manner. Multiple models are trained one by one, and all previously trained models are regarded as teacher models to supervise the training of the current one. Specifically, a orthogonal loss (OR-loss) is proposed to encourage the network to find diverse and meaningful regions. In addition, the first model is trained with only CE-Loss. Finally, all models’ outputs with complementary knowledge are combined together for the final prediction result. We demonstrate the superiority of the proposed method and obtain state-of-the-art (SOTA) performances on three popular FGVC datasets. Siqing Zhang 0001, Ruoyi Du, Dongliang Chang, Zhanyu Ma, Jun Guo 0002 |
ICME | 4 |
| 2021 | Cross-layer Navigation Convolutional Neural Network for Fine-grained Visual ClassificationabstractFine-grained visual classification (FGVC) aims to classify sub-classes of objects in the same super-class (e.g., species of birds, models of cars). For the FGVC tasks, the essential solution is to find discriminative subtle information of the target from local regions. Traditional FGVC models preferred to use the refined features, i.e., high-level semantic information for recognition and rarely use low-level information. However, it turns out that low-level information which contains rich detail information also has effect on improving performance. Therefore, in this paper, we propose cross-layer navigation convolutional neural network for feature fusion. First, the feature maps extracted by the backbone network are fed into a convolutional long short-term memory model sequentially from high-level to low-level to perform feature aggregation. Then, attention mechanisms are used after feature fusion to extract spatial and channel information while linking the high-level semantic information and the low-level texture features, which can better locate the discriminative regions for the FGVC. In the experiments, three commonly used FGVC datasets, including CUB-200-2011, Stanford-Cars, and FGVC-Aircraft datasets, are used for evaluation and we demonstrate the superiority of the proposed method by comparing it with other referred FGVC methods to show that this method achieves superior results. https://github.com/PRIS-CV/CN-CNN.git Chenyu Guo, Jiyang Xie 0001, Kongming Liang, Zhanyu Ma |
MMAsia | 5 |
| 2021 | S2TD: A Tree-Structured Decoder for Image Paragraph CaptioningabstractImage paragraph captioning, a task to generate the paragraph description for a given image, usually requires mining and organizing linguistic counterparts from abundant visual clues. Limited by sequential decoding perspective, previous methods have difficulty in organizing the visual clues holistically or capturing the structural nature of linguistic descriptions. In this paper, we propose a novel tree-structured visual paragraph decoder network, called Splitting to Tree Decoder (S2TD) to address this problem. The key idea is to model the paragraph decoding process as a top-down binary tree expansion. S2TD consists of three modules: a split module, a score module, and a word-level RNN. The split module iteratively splits ancestral visual representations into two parts through a gating mechanism. To determine the tree topology, the score module uses cosine similarity to evaluate the nodes splitting. A novel tree structure loss is proposed to enable end-to-end learning. After the tree expansion, the word-level RNN decodes leaf nodes into sentences forming a coherent paragraph. Extensive experiments are conducted on the Stanford benchmark dataset. The experimental results show promising performance of our proposed S2TD. Yihui Shi, Fangxiang Feng, Ruifan Li, Zhanyu Ma, Xiaojie Wang 0006 |
MMAsia | 5 |
| 2021 | Exploring Category-Shared and Category-Specific Features for Fine-Grained Image Classification
Dongliang Chang, Bo Xiao 0002, Zhanyu Ma, Jun Guo 0002, Yaning Chang |
PRCV (1) | 5 |
| 2021 | TLRM: Task-level Relation Module for GNN-based Few-Shot LearningabstractRecently, graph neural networks (GNNs) have shown powerful ability to handle few-shot classification problem, which aims at classifying unseen samples when trained with limited labeled samples per class. GNN-based few-shot learning architectures mostly replace traditional metric with a learnable GNN. In the GNN, the nodes are set as the samples' embedding, and the relationship between two connected nodes can be obtained by a network, the input of which is the difference of their embedding features. We consider this method of measuring relation of samples only models the sample-to-sample relation, while neglects the specificity of different tasks. That is, this method of measuring relation does not take the task-level information into account. To this end, we propose a new relation measure method, namely the task-level relation module (TLRM), to explicitly model the task-level relation of one sample to all the others. The proposed module captures the relation representations between nodes by considering the sample-to-task instead of sample-to-sample embedding features. We conducted extensive experiments on four benchmark datasets: mini-ImageNet, tiered-ImageNet, CUB-200-2011, and CIFAR-FS. Experimental results demonstrate that the proposed module is effective for GNN-based few-shot learning. Yurong Guo 0001, Zhanyu Ma |
VCIP | 2 |
| 2021 | Progressive Co-Attention Network for Fine-Grained Visual ClassificationabstractFine-grained visual classification aims to recognize images belonging to multiple sub-categories within a same category. It is a challenging task due to the inherently subtle variations among highly-confused categories. Most existing methods only take an individual image as input, which may limit the ability of models to recognize contrastive clues from different images. In this paper, we propose an effective method called progressive co-attention network (PCA-Net) to tackle this problem. Specifically, we calculate the channel-wise similarity by encouraging interaction between the feature channels within same-category image pairs to capture the common discriminative features. Considering that complementary information is also crucial for recognition, we erase the prominent areas enhanced by the channel interaction to force the network to focus on other discriminative regions. The proposed model has achieved competitive results on three fine-grained visual classification benchmark datasets: CUB-200-2011, Stanford Cars, and FGVC Aircraft. Tian Zhang 0029, Dongliang Chang, Zhanyu Ma, Jun Guo 0002 |
VCIP | 3 |
| 2021 | Deep InterBoost networks for small-sample image classification
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002 |
Neurocomputing | 3 |
| 2021 | A concise review of recent few-shot meta-learning methods
Jing-Hao Xue, Zhanyu Ma |
Neurocomputing | 4 |
| 2021 | Guest Editorial: Special issue on deep learning with small samples
Jing-Hao Xue, Jufeng Yang, Yan Yan 0001, Yujiu Yang 0001, Zongqing Lu 0001, Zhanyu Ma |
Neurocomputing | 7 |
| 2021 | Competing ratio loss for discriminative multi-class image classification
Ke Zhang 0005, Yurong Guo 0001, Dongliang Chang, Zhenbing Zhao, Zhanyu Ma, Tony X. Han |
Neurocomputing | 6 |
| 2021 | Generalisations of stochastic supervision models
Xiaoou Lu, Yangqi Qiao, Rui Zhu 0006, Guijin Wang, Zhanyu Ma, Jing-Hao Xue |
Pattern Recognit. | 5 |
| 2021 | Dilated-Scale-Aware Category-Attention ConvNet for Multi-Class Object CountingabstractObject counting aims to estimate the number of objects in images. The leading counting approaches focus on single-category counting tasks and achieve impressive performance. Nevertheless, there are multiple categories of objects in real scenes. Multi-class object counting expands the scope of application of object counting tasks. The multi-target detection task can achieve multi-class object counting in some scenarios. However, it requires the dataset annotated with bounding boxes. Compared with the point-level annotations used in mainstream object counting issues, the box-level annotations are more difficult to be obtained. In this paper, we propose a simple yet efficient counting network based on point-level annotations. Specifically, we first change the traditional estimated density map from one to the number of categories to achieve multi-class object counting. Since all categories of objects use the same feature extractor, their features will interfere mutually in the shared feature space. We further design a multi-mask structure to suppress the negative interaction among objects. Extensive experiments on the challenging benchmarks demonstrate that the proposed method achieves state-of-the-art counting performance.The code is available athttps://github.com/PRIS-CV/DSACA. Wei Xu 0037, Dingkang Liang, Yixiao Zheng, Zhanyu Ma |
IEEE Signal Process. Lett. | 5 |
| 2021 | ReMarNet: Conjoint Relation and Margin Learning for Small-Sample Image ClassificationabstractDespite achieving state-of-the-art performance, deep learning methods generally require a large amount of labeled data during training and may suffer from overfitting when the sample size is small. To ensure good generalizability of deep networks under small sample sizes, learning discriminative features is crucial. To this end, several loss functions have been proposed to encourage large intra-class compactness and inter-class separability. In this paper, we propose to enhance the discriminative power of features from a new perspective by introducing a novel neural network termed Relation-and-Margin learning Network (ReMarNet). Our method assembles two networks of different backbones so as to learn the features that can perform excellently in both of the aforementioned two classification mechanisms. Specifically, a relation network is used to learn the features that can support classification based on the similarity between a sample and a class prototype; at the meantime, a fully connected network with the cross entropy loss is used for classification via the decision boundary. Experiments on four image datasets demonstrate that our approach is effective in learning discriminative features from a small set of labeled samples and achieves competitive performance against state-of-the-art methods. Code is available at https://github.com/liyunyu08/ReMarNet. Liyun Yu, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Fine-Grained Instance-Level Sketch-Based Video RetrievalabstractExisting sketch-analysis work studies sketches depicting static objects or scenes. In this work, we propose a novel cross-modal retrieval problem of fine-grained instance-level sketch-based video retrieval (FG-SBVR), where a sketch sequence is used as a query to retrieve a specific target video instance. Compared with sketch-based still image retrieval, and coarse-grained category-level video retrieval, this is more challenging as both visual appearance and motion need to be simultaneously matched at a fine-grained level. We contribute the first FG-SBVR dataset with rich annotations. We then introduce a novel multi-stream multi-modality deep network to perform FG-SBVR under both strong and weakly supervised settings. The key component of the network is a relation module, designed to prevent model overfitting given scarce training data. We show that this model significantly outperforms a number of existing state-of-the-art models designed for video analysis. Peng Xu 0005, Kun Liu 0016, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | AP-CNN: Weakly Supervised Attention Pyramid Convolutional Neural Network for Fine-Grained Visual ClassificationabstractClassifying the sub-categories of an object from the same super-category (e.g., bird species and cars) in fine-grained visual classification (FGVC) highly relies on discriminative feature representation and accurate region localization. Existing approaches mainly focus on distilling information from high-level features. In this article, by contrast, we show that by integrating low-level information (e.g., color, edge junctions, texture patterns), performance can be improved with enhanced feature representation and accurately located discriminative regions. Our solution, named Attention Pyramid Convolutional Neural Network (AP-CNN), consists of 1) a dual pathway hierarchy structure with a top-down feature pathway and a bottom-up attention pathway, hence learning both high-level semantic and low-level detailed feature representation, and 2) an ROI-guided refinement strategy with ROI-guided dropblock and ROI-guided zoom-in operation, which refines features with discriminative local regions enhanced and background noises eliminated. The proposed AP-CNN can be trained end-to-end, without the need of any additional bounding box/part annotation. Extensive experiments on three popularly tested FGVC datasets (CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that our approach achieves state-of-the-art performance. Models and code are available at https://github.com/PRIS-CV/AP-CNN_Pytorch-master. Zhanyu Ma, Shaoguo Wen, Jiyang Xie 0001, Dongliang Chang, Zhongwei Si, Ming Wu 0001, Haibin Ling |
IEEE Trans. Image Process. | 2 |
| 2021 | BSNet: Bi-Similarity Network for Few-shot Fine-grained Image ClassificationabstractFew-shot learning for fine-grained image classification has gained recent attention in computer vision. Among the approaches for few-shot learning, due to the simplicity and effectiveness, metric-based methods are favorably state-of-the-art on many tasks. Most of the metric-based methods assume a single similarity measure and thus obtain a single feature space. However, if samples can simultaneously be well classified via two distinct similarity measures, the samples within a class can distribute more compactly in a smaller feature space, producing more discriminative feature maps. Motivated by this, we propose a so-called Bi-Similarity Network (BSNet) that consists of a single embedding module and a bi-similarity module of two similarity measures. After the support images and the query images pass through the convolution-based embedding module, the bi-similarity module learns feature maps according to two similarity measures of diverse characteristics. In this way, the model is enabled to learn more discriminative and less similarity-biased features from few shots of fine-grained images, such that the model generalization ability can be significantly improved. Through extensive experiments by slightly modifying established metric/similarity based networks, we show that the proposed approach produces a substantial improvement on several fine-grained image benchmark datasets. Codes are available at: https://github.com/PRIS-CV/BSNet. Jijie Wu, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Image Process. | 4 |
| 2021 | DS-UI: Dual-Supervised Mixture of Gaussian Mixture Models for Uncertainty Inference in Image RecognitionabstractThis paper proposes a dual-supervised uncertainty inference (DS-UI) framework for improving Bayesian estimation-based UI in DNN-based image recognition. In the DS-UI, we combine the classifier of a DNN, i.e., the last fully-connected (FC) layer, with a mixture of Gaussian mixture models (MoGMM) to obtain an MoGMM-FC layer. Unlike existing UI methods for DNNs, which only calculate the means or modes of the DNN outputs' distributions, the proposed MoGMM-FC layer acts as a probabilistic interpreter for the features that are inputs of the classifier to directly calculate the probabilities of them for the DS-UI. In addition, we propose a dual-supervised stochastic gradient-based variational Bayes (DS-SGVB) algorithm for the MoGMM-FC layer optimization. Unlike conventional SGVB and optimization algorithms in other UI methods, the DS-SGVB not only models the samples in the specific class for each Gaussian mixture model (GMM) in the MoGMM, but also considers the negative samples from other classes for the GMM to reduce the intra-class distances and enlarge the inter-class margins simultaneously for enhancing the learning ability of the MoGMM-FC layer in the DS-UI. Experimental results show the DS-UI outperforms the state-of-the-art UI methods in misclassification detection. We further evaluate the DS-UI in open-set out-of-domain/-distribution detection and find statistically significant improvements. Visualizations of the feature spaces demonstrate the superiority of the DS-UI. Codes are available at https://github.com/PRIS-CV/DS-UI. Jiyang Xie 0001, Zhanyu Ma, Jing-Hao Xue, Guoqiang Zhang 0003, Yinhe Zheng, Jun Guo 0002 |
IEEE Trans. Image Process. | 2 |
| 2020 | Fine-Grained Visual Classification via Progressive Multi-granularity Training of Jigsaw Patches
Ruoyi Du, Dongliang Chang, Ayan Kumar Bhunia, Jiyang Xie 0001, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
ECCV (20) | 5 |
| 2020 | GINet: Graph Interaction Network for Scene Parsing
Yu Zhu 0006, Ming Wu 0001, Zhanyu Ma, Guodong Guo |
ECCV (17) | 6 |
| 2020 | IU-Module: Intersection and Union Module for Fine-Grained Visual ClassificationabstractA predominant viewpoint in previous works of fine-grained visual classification (FGVC) is to the localize discriminative parts by auxiliary networks and extract the part-based finegrained features for classification. In this paper, we propose a simple yet effective approach by introducing an intersection and union module (IU-Module). The IU-Module aims to capture more discriminative features by 1) dividing features into distinct groups, 2) sharing parts of interests within each group, and 3) adding a differentiation loss to reduce the similarity among those grouped feature channels. Without adding any new learnable parameters, the proposed approach imposes two straightforward operations, namely channel intersection (CI) and channel union (CU) operations, on the convolutional features and achieves competitive results compared with the state-of-the-art methods. Experimental results on three publicly available FGVC datasets show the effectiveness of the IU-Module. Ablation studies and visualizations are also provided to make further demonstrations. Yixiao Zheng, Dongliang Chang, Jiyang Xie 0001, Zhanyu Ma |
ICME | 4 |
| 2020 | CC-Loss: Channel Correlation Loss for Image ClassificationabstractThe loss function is a key component in deep learning models. A commonly used loss function for classification is the cross entropy loss, which is a simple yet effective application of information theory for classification problems. Based on this loss, many other loss functions have been proposed, e.g., by adding intra-class and inter-class constraints to enhance the discriminative ability of the learned features. However, these loss functions fail to consider the connections between the feature distribution and the model structure. Aiming at addressing this problem, we propose a channel correlation loss (CC-Loss) that is able to constrain the specific relations between classes and channels as well as maintain the intra-class and the inter-class separability. CC-Loss uses a channel attention module to generate channel attention of features for each sample in the training stage. Next, an Euclidean distance matrix is calculated to make the channel attention vectors associated with the same class become identical and to increase the difference between different classes. Finally, we obtain a feature embedding with good intra-class compactness and inter-class separability. Experimental results show that two different backbone models trained with the proposed CC-Loss outperform the state-of-the-art loss functions on three image classification datasets. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan |
ICPR | 3 |
| 2020 | Dual-attention Guided Dropblock Module for Weakly Supervised Object LocalizationabstractAttention mechanisms is frequently used to learn the discriminative features for better feature representations. In this paper, we extend the attention mechanism to the task of weakly supervised object localization (WSOL) and propose the dual-attention guided dropblock module (DGDM), which aims at learning the informative and complementary visual patterns for WSOL. This module contains two key components, the channel attention guided dropout (CAGD) and the spatial attention guided dropblock (SAGD). To model channel interdependencies, the CAGD ranks the channel attentions and treats the top-k attentions with the largest magnitudes as the important ones. It also keeps some low-valued elements to increase their value if they become important during training. The SAGD can efficiently remove the most discriminative information by erasing the contiguous regions of feature maps rather than individual pixels. This guides the model to capture the less discriminative parts for classification. Furthermore, it can also distinguish the foreground objects from the background regions to alleviate the attention misdirection. Experimental results demonstrate that the proposed method achieves new state-of-the-art localization performance. Junhui Yin, Siqing Zhang 0001, Dongliang Chang, Zhanyu Ma, Jun Guo 0002 |
ICPR | 4 |
| 2020 | Global Context Enhanced Multi-modal Fusion for Referring Image Segmentation
Yan Huang 0008, Linjiang Huang, Yunbo Wang, Zhanyu Ma, Liang Wang 0001 |
PRCV (1) | 5 |
| 2020 | Clustering Analysis in the Wireless Propagation Channel with a Variational Gaussian Mixture ModelabstractIn this paper, the Gaussian mixture model (GMM) is introduced to implement channel multipath clustering. The GMM incorporates the covariance structure and the mean information of the channel multipaths, thus it can effectively reveal the similarity of the channel multipaths. First, the expectation-maximization (EM) algorithm is utilized to search for the posterior estimation of the GMM parameters. Then, the variational Bayesian (VB) algorithm is employed to optimize the GMM parameters to enhance the searching ability of EM and further to determine the optimal number of Gaussian distributions without resorting to cross-validation. Finally, a compact index (CI) is proposed to validate the clustering results reasonably. Thanks to the proposed CI, it is possible to find a close relationship among the GMM clustering mechanism, the multipath propagation characteristics and the CI evaluation index. Experiments with synthetic data and outdoor-to-indoor (O2I) channel data are presented to demonstrate the effectiveness of the proposed method. Jianhua Zhang 0001, Zhanyu Ma, Yu Zhang 0054 |
IEEE Trans. Big Data | 3 |
| 2020 | Deep Neural Network-Based Impacts Analysis of Multimodal Factors on Heat Demand PredictionabstractPrediction of heat demand using artificial neural networks has attracted enormous research attention. Weather conditions, such as direct solar irradiance and wind speed, have been identified as key parameters affecting heat demand. This paper employs an Elman neural network to investigate the impacts of direct solar irradiance and wind speed on the heat demand from the perspective of the entire district heating network. Results of the overall mean absolute percentage error (MAPE) show that direct solar irradiance and wind speed have quite similar impacts. However, the involvement of direct solar irradiance can clearly reduce the maximum absolute deviation when only involving direct solar irradiance and wind speed, respectively. In addition, the simultaneous involvement of both wind speed and direct solar irradiance does not show an obvious improvement of MAPE. Moreover, the prediction accuracy can also be affected by other factors like data discontinuity and outliers. Zhanyu Ma, Jiyang Xie 0001, Qie Sun, Fredrik Wallin, Zhongwei Si, Jun Guo 0002 |
IEEE Trans. Big Data | 1 |
| 2020 | Semi-Heterogeneous Three-Way Joint Embedding Network for Sketch-Based Image RetrievalabstractSketch-based image retrieval (SBIR) is a challenging task due to the large cross-domain gap between sketches and natural images. How to align abstract sketches and natural images into a common high-level semantic space remains a key problem in SBIR. In this paper, we propose a novel semi-heterogeneous three-way joint embedding network (Semi3-Net), which integrates three branches (a sketch branch, a natural image branch, and an edgemap branch) to learn more discriminative cross-domain feature representations for the SBIR task. The key insight lies with how we cultivate the mutual and subtle relationships amongst the sketches, natural images, and edgemaps. A semi-heterogeneous feature mapping is designed to extract bottom features from each domain, where the sketch and edgemap branches are shared while the natural image branch is heterogeneous to the other branches. In addition, a joint semantic embedding is introduced to embed the features from different domains into a common high-level semantic space, where all of the three branches are shared. To further capture informative features common to both natural images and the corresponding edgemaps, a co-attention model is introduced to conduct common channel-wise feature recalibration between different domains. A hybrid-loss mechanism is designed to align the three branches, where an alignment loss and a sketch-edgemap contrastive loss are presented to encourage the network to learn invariant cross-domain representations. Experimental results on two widely used category-level datasets (Sketchy and TU-Berlin Extension) demonstrate that the proposed method outperforms state-of-the-art methods. Jianjun Lei 0001, Bo Peng 0007, Zhanyu Ma, Ling Shao 0001, Yi-Zhe Song |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Fine-Grained Age Estimation in the Wild With Attention LSTM NetworksabstractAge estimation from a single face image has been an essential task in the field of human-computer interaction and computer vision, which has a wide range of practical application values. Accuracy of age estimation of face images in the wild is relatively low for existing methods, because they only take into account the global features, while neglecting the fine-grained features of age-sensitive areas. We propose a novel method based on our attention long short-term memory (AL) network for fine-grained age estimation in the wild, inspired by the fine-grained categories and the visual attention mechanism. This method combines the residual networks (ResNets) or the residual network of residual network (RoR) models with LSTM units to construct AL-ResNets or AL-RoR networks to extract local features of age-sensitive regions, which effectively improves the age estimation accuracy. First, a ResNets or a RoR model pretrained on ImageNet dataset is selected as the basic model, which is then fine-tuned on the IMDB-WIKI-101 dataset for age estimation. Then, we fine-tune the ResNets or the RoR on the target age datasets to extract the global features of face images. To extract the local features of age-sensitive regions, the LSTM unit is then presented to obtain the coordinates of the age-sensitive region automatically. Finally, the age group classification is conducted directly on the Adience dataset, and age-regression experiments are performed by the Deep EXpectation algorithm (DEX) on MORPH Album 2, FG-NET and 15/16LAP datasets. By combining the global and the local features, we obtain our final prediction results. Experimental results illustrate the effectiveness and robustness of the proposed AL-ResNets or AL-RoR for age estimation in the wild, where it achieves better state-of-the-art performance than all other convolutional neural network (CNN) methods on the Adience, MORPH Album 2, FG-NET and 15/16LAP datasets. Ke Zhang 0005, Xingfang Yuan, Xinyao Guo, Ce Gao, Zhenbing Zhao, Zhanyu Ma |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2020 | The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image ClassificationabstractThe key to solving fine-grained image categorization is finding discriminate and local regions that correspond to subtle visual traits. Great strides have been made, with complex networks designed specifically to learn part-level discriminate feature representations. In this paper, we show that it is possible to cultivate subtle details without the need for overly complicated network designs or training mechanisms - a single loss is all it takes. The main trick lies with how we delve into individual feature channels early on, as opposed to the convention of starting from a consolidated feature map. The proposed loss function, termed as mutual-channel loss (MC-Loss), consists of two channel-specific components: a discriminality component and a diversity component. The discriminality component forces all feature channels belonging to the same class to be discriminative, through a novel channel-wise attention mechanism. The diversity component additionally constraints channels so that they become mutually exclusive across the spatial dimension. The end result is therefore a set of feature channels, each of which reflects different locally discriminative regions for a specific class. The MC-Loss can be trained end-to-end, without the need for any bounding-box/part annotations, and yields highly discriminative regions during inference. Experimental results show our MC-Loss when implemented on top of common base networks can achieve state-of-the-art performance on all four fine-grained categorization datasets (CUB-Birds, FGVC-Aircraft, Flowers-102, and Stanford Cars). Ablative studies further demonstrate the superiority of the MC-Loss when compared with other recently proposed general-purpose losses for visual classification, on two different base networks. Dongliang Chang, Jiyang Xie 0001, Ayan Kumar Bhunia, Zhanyu Ma, Ming Wu 0001, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Image Process. | 6 |
| 2020 | OSLNet: Deep Small-Sample Classification With an Orthogonal Softmax LayerabstractA deep neural network of multiple nonlinear layers forms a large function space, which can easily lead to overfitting when it encounters small-sample data. To mitigate overfitting in small-sample classification, learning more discriminative features from small-sample data is becoming a new trend. To this end, this paper aims to find a subspace of neural networks that can facilitate a large decision margin. Specifically, we propose the Orthogonal Softmax Layer (OSL), which makes the weight vectors in the classification layer remain orthogonal during both the training and test processes. The Rademacher complexity of a network using the OSL is only 1/K, where K is the number of classes, of that of a network using the fully connected classification layer, leading to a tighter generalization error bound. Experimental results demonstrate that the proposed OSL has better performance than the methods used for comparison on four small-sample benchmark datasets, as well as its applicability to large-sample datasets. Codes are available at: https://github.com/dongliangchang/OSLNet. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jingyi Yu 0001, Jun Guo 0002 |
IEEE Trans. Image Process. | 3 |
| 2020 | Insights Into Multiple/Single Lower Bound Approximation for Extended Variational Inference in Non-Gaussian Structured Data ModelingabstractFor most of the non-Gaussian statistical models, the data being modeled represent strongly structured properties, such as scalar data with bounded support (e.g., beta distribution), vector data with unit length (e.g., Dirichlet distribution), and vector data with positive elements (e.g., generalized inverted Dirichlet distribution). In practical implementations of non-Gaussian statistical models, it is infeasible to find an analytically tractable solution to estimating the posterior distributions of the parameters. Variational inference (VI) is a widely used framework in Bayesian estimation. Recently, an improved framework, namely, the extended VI (EVI), has been introduced and applied successfully to a number of non-Gaussian statistical models. EVI derives analytically tractable solutions by introducing lower bound approximations to the variational objective function. In this paper, we compare two approximation strategies, namely, the multiple lower bounds (MLBs) approximation and the single lower bound (SLB) approximation, which can be applied to carry out the EVI. For implementation, two different conditions, the weak and the strong conditions, are discussed. Convergence of the EVI depends on the selection of the lower bound, regardless of the choice of weak or strong condition. We also discuss the convergence properties to clarify the differences between MLB and SLB. Extensive comparisons are made based on some EVI-based non-Gaussian statistical models. Theoretical analysis is conducted to demonstrate the differences between the weak and strong conditions. Experimental results based on real data show advantages of the SLB approximation over the MLB approximation. Zhanyu Ma, Jiyang Xie 0001, Yuping Lai, Jalil Taghia, Jing-Hao Xue, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Guest Editorial Special Issue on Recent Advances in Theory, Methodology, and Applications of Imbalanced Learning
Jing-Hao Xue, Zhanyu Ma, Manuel Roveri, Nathalie Japkowicz |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Channel-Wise and Feature-Points Reweights Densenet for Image ClassificationabstractRecent network research has demonstrated that the performance of convolutional neural networks can be improved by introducing a learning block that capture spatial correlations and channel-wise correlations . In this work, we propose a novel Channel-wise and Feature-points Reweights DenseNet (CAPR-DenseNet) architecture. The CAPR-DenseNet improves the representation power of the DenseNet by adaptively recalibrating the channel-wise feature responses and explicitly modeling the interdependencies between feature-points. First, in order to perform dynamic channel-wise feature recalibration, we construct the Channel-wise Feature Reweight DenseNet (CFR-DenseNet) by introducing the Squeeze-and-Excitation Module (SEM) to DenseNet. Then, we present a novel CAPR-DenseNet by adding a Feature-points Reweight Module (FPRM) to the CFR-DenseNet. Through massive experiments, we demonstrate that by recalibrating the channel-wise feature and the feature-points responses. Our CAPR-DenseNet performs better than DenseNet across challenging datasets CIFAR-10 and CIFAR-100. Ke Zhang 0005, Yurong Guo 0001, Jinsha Yuan, Zhanyu Ma, Zhenbing Zhao |
ICIP | 5 |
| 2019 | Deep Zero-Shot Learning for Scene SketchabstractWe introduce a novel problem of scene sketch zero-shot learning (SSZSL), which is a challenging task, since (i) different from photo, the gap between common semantic domain (e.g., word vector) and sketch is too huge to exploit common semantic knowledge as the bridge for knowledge transfer, and (ii) compared with single-object sketch, more expressive feature representation for scene sketch is required to accommodate its high-level of abstraction and complexity. To overcome these challenges, we propose a deep embedding model for scene sketch zero-shot learning. In particular, we propose the augmented semantic vector to conduct domain alignment by fusing multi-modal semantic knowledge (e.g., cartoon image, natural image, text description), and adopt attention-based network for scene sketch feature learning. Moreover, we propose a novel distance metric to improve the similarity measure during testing. Extensive experiments and ablation studies demonstrate the benefit of our sketch-specific design. Peng Xu 0005, Zhanyu Ma |
ICIP | 3 |
| 2019 | Competing Ratio Loss for Multi-class image ClassificationabstractCross-entropy loss function (CEL) is widely used for training a multi-class classification deep convolutional neural network (DCNN). While CEL has been successfully implemented in image classification tasks, it only focuses on the posterior probability of correct class when the labels of training images are one-hot. It cannot be discriminated against the classes not belong to correct class (wrong classes) directly. Negative Log Likelihood Ratio Loss (NLLR) is proposed to better discriminate the correct class from competing wrong classes. But optimization of the loss function is normally presented as a minimization problem. In training DCNN, the value of NLLR is not constantly positive or negative, which affects the convergence of NLLR adversely. So, we propose competing ratio loss (CRL), which calculates the posterior probability ratio between the correct class and competing wrong classes to better widen the difference between the probability of the correct class and the probabilities of wrong classes, which also assures the value of CRL is constantly positive. Through massive experiments, we demonstrate the effectiveness and robustness of CRL on deep convolutional neural networks, our CRL outperforms CEL and NLLR on CIFAR-10/100 datasets. Ke Zhang 0005, Yurong Guo 0001, Zhenbing Zhao, Zhanyu Ma |
VCIP | 5 |
| 2019 | FICAL: Focal Inter-Class Angular Loss for Image ClassificationabstractConvolutional Neural Networks (CNNs) have been successfully applied in various image analysis tasks and gradually become one of the most powerful machine learning approaches. In order to improve the capability of the model generalization and performance in image classification, a new trend is to learn more discriminative features via CNNs. The main contribution of this paper is to increase the angles between the categories to extract discriminative features and enlarge the inter-class variance. To this end, we propose a loss function named focal inter-class angular loss (FICAL) which introduces the confusion rate-weighted cosine distance as the similarity measurement between categories. This measurement is dynamically evaluated during each iteration to adapt the model. Compared with other loss functions, experimental results demonstrate that the proposed FICAL achieved best performance among the referred loss functions on two image classificaton datasets. Xinran Wei, Dongliang Chang, Jiyang Xie 0001, Yixiao Zheng, Chen Gong 0002, Zhanyu Ma |
VCIP | 7 |
| 2019 | Image-text dual neural network with decision strategy for small-sample image classification
Fangyi Zhu, Zhanyu Ma, Guang Chen 0003, Jen-Tzung Chien, Jing-Hao Xue, Jun Guo 0002 |
Neurocomputing | 2 |
| 2019 | Variational Bayesian Learning for Dirichlet Process Mixture of Inverted Dirichlet Distributions in Non-Gaussian Image Feature ModelingabstractIn this paper, we develop a novel variational Bayesian learning method for the Dirichlet process (DP) mixture of the inverted Dirichlet distributions, which has been shown to be very flexible for modeling vectors with positive elements. The recently proposed extended variational inference (EVI) framework is adopted to derive an analytically tractable solution. The convergency of the proposed algorithm is theoretically guaranteed by introducing single lower bound approximation to the original objective function in the EVI framework. In principle, the proposed model can be viewed as an infinite inverted Dirichlet mixture model that allows the automatic determination of the number of mixture components from data. Therefore, the problem of predetermining the optimal number of mixing components has been overcome. Moreover, the problems of overfitting and underfitting are avoided by the Bayesian estimation approach. Compared with several recently proposed DP-related methods and conventional applied methods, the good performance and effectiveness of the proposed method have been demonstrated with both synthesized data and real data evaluations. Zhanyu Ma, Yuping Lai, W. Bastiaan Kleijn, Yi-Zhe Song, Liang Wang 0001, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | SketchMate: Deep Hashing for Million-Scale Human Sketch RetrievalabstractWe propose a deep hashing framework for sketch retrieval that, for the first time, works on a multi-million scale human sketch dataset. Leveraging on this large dataset, we explore a few sketch-specific traits that were otherwise under-studied in prior literature. Instead of following the conventional sketch recognition task, we introduce the novel problem of sketch hashing retrieval which is not only more challenging, but also offers a better testbed for large-scale sketch analysis, since: (i) more fine-grained sketch feature learning is required to accommodate the large variations in style and Abstraction, and (ii) a compact binary code needs to be learned at the same time to enable efficient retrieval. Key to our network design is the embedding of unique characteristics of human sketch, where (i) a two-branch CNN-RNN architecture is adapted to explore the temporal ordering of strokes, and (ii) a novel hashing loss is specifically designed to accommodate both the temporal and Abstract traits of sketches. By working with a 3.8M sketch dataset, we show that state-of-the-art hashing models specifically engineered for static images fail to perform well on temporal sketch data. Our network on the other hand not only offers the best retrieval performance on various code sizes, but also yields the best generalization performance under a zero-shot setting and when re-purposed for sketch recognition. Such superior performances effectively demonstrate the benefit of our sketch-specific design. Peng Xu 0005, Yongye Huang, Tongtong Yuan, Kaiyue Pang, Yi-Zhe Song, Tao Xiang 0002, Timothy M. Hospedales, Zhanyu Ma, Jun Guo 0002 |
CVPR | 8 |
| 2018 | Clustering in wireless propagation channel with a statistics-based frameworkabstractIn this article, we introduce a statistics-based clustering framework that is able to model the clustering problem corresponding to the channel propagation characteristics and evaluate the clustering results effectively. In the framework a Gaussian mixture model (GMM) is employed to model the channel multipaths. Then, we optimize the GMM parameters with the expectation-maximization (EM) algorithm. To evaluate the clustering results effectively, a compact index (CI) is devised, in which both the mean and variance of the clusters are considered. In the simulation, outdoor-to-indoor (O2I) channel measurement data is presented to demonstrate the effectiveness of the proposed framework. Jianhua Zhang 0001, Zhanyu Ma |
WCNC | 3 |
| 2018 | Supervised latent Dirichlet allocation with a mixture of sparse softmax
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002 |
Neurocomputing | 2 |
| 2018 | Corrigendum to "Supervised latent Dirichlet allocation with a mixture of sparse softmax" [Neurocomputing, volume 312, 27 October 2018, Pages 324-335]
Zhanyu Ma, Feiyue Huang, Xiaojie Wang 0006, Jun Guo 0002 |
Neurocomputing | 2 |
| 2018 | Recent advances in machine learning for non-Gaussian data processing
Zhanyu Ma, Jen-Tzung Chien, Zheng-Hua Tan, Yi-Zhe Song, Jalil Taghia, Ming Xiao 0001 |
Neurocomputing | 1 |
| 2018 | Cross-modal subspace learning for fine-grained sketch-based image retrieval
Peng Xu 0005, Qiyue Yin, Yongye Huang, Yi-Zhe Song, Zhanyu Ma, Liang Wang 0001, Tao Xiang 0002, W. Bastiaan Kleijn, Jun Guo 0002 |
Neurocomputing | 5 |
| 2018 | LRID: A new metric of multi-class imbalance degree based on likelihood-ratio testabstractIn this paper, we introduce a new likelihood ratio imbalance degree (LRID) to measure the class-imbalance extent of multi-class data. Imbalance ratio (IR) is usually used to measure class-imbalance extent in imbalanced learning problems. However, IR cannot capture the detailed information in the class distribution of multi-class data, because it only utilises the information of the largest majority class and the smallest minority class. Imbalance degree (ID) has been proposed to solve the problem of IR for multi-class data. However, we note that improper use of distance metric in ID can have harmful effect on the results. In addition, ID assumes that data with more minority classes are more imbalanced than data with less minority classes, which is not always true in practice. Thus ID cannot provide reliable measurement when the assumption is violated. In this paper, we propose a new metric based on the likelihood-ratio test, LRID, to provide a more reliable measurement of class-imbalance extent for multi-class data. Experiments on both simulated and real data show that LRID is competitive with IR and ID, and can reduce the negative correlation with F1 scores by up to 0.55. Rui Zhu 0006, Ziyu Wang 0003, Zhanyu Ma, Guijin Wang, Jing-Hao Xue |
Pattern Recognit. Lett. | 3 |
| 2018 | Decorrelation of Neutral Vector Variables: Theory and ApplicationsabstractIn this paper, we propose novel strategies for neutral vector variable decorrelation. Two fundamental invertible transformations, namely, serial nonlinear transformation and parallel nonlinear transformation, are proposed to carry out the decorrelation. For a neutral vector variable, which is not multivariate-Gaussian distributed, the conventional principal component analysis cannot yield mutually independent scalar variables. With the two proposed transformations, a highly negatively correlated neutral vector can be transformed to a set of mutually independent scalar variables with the same degrees of freedom. We also evaluate the decorrelation performances for the vectors generated from a single Dirichlet distribution and a mixture of Dirichlet distributions. The mutual independence is verified with the distance correlation measurement. The advantages of the proposed decorrelation strategies are intensively studied and demonstrated with synthesized data and practical application evaluations. Zhanyu Ma, Jing-Hao Xue, Arne Leijon, Zheng-Hua Tan, Zhen Yang 0004, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic FeaturesabstractWith the development of speech synthesis technology, automatic speaker verification (ASV) systems have encountered the serious challenge of spoofing attacks. In order to improve the security of ASV systems, many antispoofing countermeasures have been developed. In the front-end domain, much research has been conducted on finding effective features which can distinguish spoofed speech from genuine speech and the published results show that dynamic acoustic features work more effectively than static ones. In the back-end domain, Gaussian mixture model (GMM) and deep neural networks (DNNs) are the two most popular types of classifiers used for spoofing detection. The log-likelihood ratios (LLRs) generated by the difference of human and spoofing log-likelihoods are used as spoofing detection scores. In this paper, we train a five-layer DNN spoofing detection classifier using dynamic acoustic features and propose a novel, simple scoring method only using human log-likelihoods (HLLs) for spoofing detection. We mathematically prove that the new HLL scoring method is more suitable for the spoofing detection task than the classical LLR scoring method, especially when the spoofing speech is very similar to the human speech. We extensively investigate the performance of five different dynamic filter bank-based cepstral features and constant Q cepstral coefficients (CQCC) in conjunction with the DNN-HLL method. The experimental results show that, compared to the GMM-LLR method, the DNN-HLL method is able to significantly improve the spoofing detection accuracy. Compared with the CQCC-based GMM-LLR baseline, the proposed DNN-HLL model reduces the average equal error rate of all attack types to 0.045%, thus exceeding the performance of previously published approaches for the ASVspoof 2015 Challenge task. Fusing the CQCC-based DNN-HLL spoofing detection system with ASV systems, the false acceptance rate on spoofing attacks can be reduced significantly. Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | A Survey on Machine Learning-Based Mobile Big Data Analysis: Challenges and ApplicationsabstractThis paper attempts to identify the requirement and the development of machine learning‐based mobile big data (MBD) analysis through discussing the insights of challenges in the mobile big data. Furthermore, it reviews the state‐of‐the‐art applications of data analysis in the area of MBD. Firstly, we introduce the development of MBD. Secondly, the frequently applied data analysis methods are reviewed. Three typical applications of MBD analysis, namely, wireless channel modeling, human online and offline behavior analysis, and speech recognition in the Internet of Vehicles, are introduced, respectively. Finally, we summarize the main challenges and future development directions of mobile big data analysis. Jiyang Xie 0001, Yanting Zhang 0001, Hong Yu 0006, Jinnan Zhan, Zhanyu Ma, Yuanyuan Qiao 0002, Jianhua Zhang 0001, Jun Guo 0002 |
Wirel. Commun. Mob. Comput. | 7 |
| 2017 | Adversarial Network Bottleneck Features for Noise Robust Speaker VerificationabstractIn this paper, we propose a noise robust bottleneck feature representation which is generated by an adversarial network (AN).The AN includes two cascade connected networks, an encoding network (EN) and a discriminative network (DN).Melfrequency cepstral coefficients (MFCCs) of clean and noisy speech are used as input to the EN and the output of the EN is used as the noise robust feature.The EN and DN are trained in turn, namely, when training the DN, noise types are selected as the training labels and when training the EN, all labels are set as the same, i.e., the clean speech label, which aims to make the AN features invariant to noise and thus achieve noise robustness.We evaluate the performance of the proposed feature on a Gaussian Mixture Model-Universal Background Model based speaker verification system, and make comparison to MFCC features of speech enhanced by short-time spectral amplitude minimum mean square error (STSA-MMSE) and deep neural network-based speech enhancement (DNN-SE) methods.Experimental results on the RSR2015 database show that the proposed AN bottleneck feature (AN-BN) dramatically outperforms the STSA-MMSE and DNN-SE based MFCCs for different noise types and signal-to-noise ratios.Furthermore, the AN-BN feature is able to improve the speaker verification performance under the clean condition. Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Jun Guo 0002 |
INTERSPEECH | 3 |
| 2016 | Feature selection for neutral vector in EEG signal classification
Zhanyu Ma, Zheng-Hua Tan, Jun Guo 0002 |
Neurocomputing | 1 |
| 2016 | A Comprehensive Review of Smart Energy Meters in Intelligent Energy NetworksabstractThe significant increase in energy consumption and the rapid development of renewable energy, such as solar power and wind power, have brought huge challenges to energy security and the environment, which, in the meantime, stimulate the development of energy networks toward a more intelligent direction. Smart meters are the most fundamental components in the intelligent energy networks (IENs). In addition to measuring energy flows, smart energy meters can exchange the information on energy consumption and the status of energy networks between utility companies and consumers. Furthermore, smart energy meters can also be used to monitor and control home appliances and other devices according to the individual consumer's instruction. This paper systematically reviews the development and deployment of smart energy meters, including smart electricity meters, smart heat meters, and smart gas meters. By examining various functions and applications of smart energy meters, as well as associated benefits and costs, this paper provides insights and guidelines regarding the future development of smart meters. Qie Sun, Zhanyu Ma, Chao Wang 0015, Javier Campillo, Qi Zhang 0019, Fredrik Wallin, Jun Guo 0002 |
IEEE Internet Things J. | 3 |
| 2015 | Activation force-based air pollution observation station clustering
Di Huang 0006, Hong Yu 0006, Huanyu Zhou, Zhanyu Ma, Weisong Hu, Jun Guo 0002 |
QSHINE | 5 |
| 2015 | A multi-level system for sequential update summarization
Chunyun Zhang, Zhanyu Ma, Jiayue Zhang, Weiran Xu, Jun Guo 0002 |
QSHINE | 2 |
| 2015 | Multi-label learning with prior knowledge for facial expression analysis
Kaili Zhao, Honggang Zhang 0002, Zhanyu Ma, Yi-Zhe Song, Jun Guo 0002 |
Neurocomputing | 3 |
| 2015 | Construction of semantic bootstrapping models for relation extraction
Chunyun Zhang, Weiran Xu, Zhanyu Ma, Sheng Gao 0001, Qun Li 0002, Jun Guo 0002 |
Knowl. Based Syst. | 3 |
| 2015 | Mining activation force defined dependency patterns for relation extraction
Chunyun Zhang, Yichang Zhang, Weiran Xu, Zhanyu Ma, Yan Leng, Jun Guo 0002 |
Knowl. Based Syst. | 4 |
| 2015 | Variational Bayesian Matrix Factorization for Bounded Support DataabstractA novel Bayesian matrix factorization method for bounded support data is presented. Each entry in the observation matrix is assumed to be beta distributed. As the beta distribution has two parameters, two parameter matrices can be obtained, which matrices contain only nonnegative values. In order to provide low-rank matrix factorization, the nonnegative matrix factorization (NMF) technique is applied. Furthermore, each entry in the factorized matrices, i.e., the basis and excitation matrices, is assigned with gamma prior. Therefore, we name this method as beta-gamma NMF (BG-NMF). Due to the integral expression of the gamma function, estimation of the posterior distribution in the BG-NMF model can not be presented by an analytically tractable solution. With the variational inference framework and the relative convexity property of the log-inverse-beta function, we propose a new lower-bound to approximate the objective function. With this new lower-bound, we derive an analytically tractable solution to approximately calculate the posterior distributions. Each of the approximated posterior distributions is also gamma distributed, which retains the conjugacy of the Bayesian estimation. In addition, a sparse BG-NMF can be obtained by including a sparseness constraint to the gamma prior. Evaluations with synthetic data and real life data demonstrate the good performance of the proposed method. Zhanyu Ma, Andrew E. Teschendorff, Arne Leijon, Yuanyuan Qiao 0002, Honggang Zhang 0002, Jun Guo 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Line spectral frequencies modeling by a mixture of von Mises-Fisher distributions
Zhanyu Ma, Jalil Taghia, W. Bastiaan Kleijn, Arne Leijon, Jun Guo 0002 |
Signal Process. | 1 |
| 2014 | Nonlinear estimation of missing ΔLSF parameters by a mixture of Dirichlet distributionsabstractIn packet networks, a reliable scheme to handle packet loss during speech transmission is of great importance. As a common representation of the linear predictive coding (LPC) model, the line spectral frequency (LSF) parameters are widely used in speech quantization and transmission. In this paper, we propose a novel scheme to estimate the missing values occurring during LPC model transmission. In order to exploit the boundary and ordering properties of the LSF parameters, we utilize the ΔLSF representation and apply the Dirichlet mixture model (DMM) to capture the correlations among the elements in the ΔLSF vector. With the conditional distribution of the missing part given the received part, an optimal nonlinear minimum mean square error estimator for the missing values is proposed. Compared to the previously presented Gaussian mixture model based method, the proposed DMM based nonlinear estimator shows a convincing improvement. Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002, Honggang Zhang 0002 |
ICASSP | 1 |
| 2014 | Bayesian Estimation of the von-Mises Fisher Mixture Model with Variational InferenceabstractThis paper addresses the Bayesian estimation of the von-Mises Fisher (vMF) mixture model with variational inference (VI). The learning task in VI consists of optimization of the variational posterior distribution. However, the exact solution by VI does not lead to an analytically tractable solution due to the evaluation of intractable moments involving functional forms of the Bessel function in their arguments. To derive a closed-form solution, we further lower bound the evidence lower bound where the bound is tight at one point in the parameter distribution. While having the value of the bound guaranteed to increase during maximization, we derive an analytically tractable approximation to the posterior distribution which has the same functional form as the assigned prior distribution. The proposed algorithm requires no iterative numerical calculation in the re-estimation procedure, and it can potentially determine the model complexity and avoid the over-fitting problem associated with conventional approaches based on the expectation maximization. Moreover, we derive an analytically tractable approximation to the predictive density of the Bayesian mixture model of vMF distributions. The performance of the proposed approach is verified by experiments with both synthetic and real data. Jalil Taghia, Zhanyu Ma, Arne Leijon |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Bayesian estimation of Dirichlet mixture model with variational inference
Zhanyu Ma, Pravin Kumar Rana, Jalil Taghia, Markus Flierl, Arne Leijon |
Pattern Recognit. | 1 |
| 2014 | Dirichlet mixture modeling to estimate an empirical lower bound for LSF quantization
Zhanyu Ma, Saikat Chatterjee, W. Bastiaan Kleijn, Jun Guo 0002 |
Signal Process. | 1 |
| 2013 | Multiview depth map enhancement by variational bayes inference estimation of Dirichlet mixture modelsabstractHigh quality view synthesis is a prerequisite for future free-viewpoint television. It will enable viewers to move freely in a dynamic real world scene. Depth image based rendering algorithms will play a pivotal role when synthesizing an arbitrary number of novel views by using a subset of captured views and corresponding depth maps only. Usually, each depth map is estimated individually by stereo-matching algorithms and, hence, shows lack of inter-view consistency. This inconsistency affects the quality of view synthesis negatively. This paper enhances the inter-view consistency of multiview depth imagery. First, our approach classifies the color information in the multiview color imagery by modeling color with a mixture of Dirichlet distributions where the model parameters are estimated in a Bayesian framework with variational inference. Second, using the resulting color clusters, we classify the corresponding depth values in the multiview depth imagery. Each clustered depth image is subject to further sub-clustering. Finally, the resulting mean of each sub-cluster is used to enhance the depth imagery at multiple viewpoints. Experiments show that our approach improves the average quality of virtual views by up to 0.8 dB when compared to views synthesized by using conventionally estimated depth maps. Pravin Kumar Rana, Zhanyu Ma, Jalil Taghia, Markus Flierl |
ICASSP | 2 |
| 2013 | On von-mises fisher mixture model in text-independent speaker identificationabstractThis paper addresses text-independent speaker identification (SI) based on line spectral frequencies (LSFs). The LSFs are transformed to differential LSFs (MLSF) in order to exploit their boundary and ordering properties. We show that the square root of MLSF has interesting directional characteristics implying that their distribution can be modeled by a mixture of von-Mises Fisher (vMF) distributions. We analytically estimate the mixture model parameters in a fully Bayesian treatment by using variational inference. In the Bayesian inference, we can potentially determine the model complexity and avoid overfitting problem associated with conventional approaches based on the expectation maximization. The experimental results confirm the effectiveness of the proposed SI system. Jalil Taghia, Zhanyu Ma, Arne Leijon |
INTERSPEECH | 2 |
| 2013 | Vector quantization of LSF parameters with a mixture of dirichlet distributionsabstractQuantization of the linear predictive coding parameters is an important part in speech coding. Probability density function (PDF)-optimized vector quantization (VQ) has been previously shown to be more efficient than VQ based only on training data. For data with bounded support, some well-defined bounded-support distributions (e.g., the Dirichlet distribution) have been proven to outperform the conventional Gaussian mixture model (GMM), with the same number of free parameters required to describe the model. When exploiting both the boundary and the order properties of the line spectral frequency (LSF) parameters, the distribution of LSF differences LSF can be modelled with a Dirichlet mixture model (DMM). We propose a corresponding DMM based VQ. The elements in a Dirichlet vector variable are highly mutually correlated. Motivated by the Dirichlet vector variable's neutrality property, a practical non-linear transformation scheme for the Dirichlet vector variable can be obtained. Similar to the Karhunen-Loève transform for Gaussian variables, this non-linear transformation decomposes the Dirichlet vector variable into a set of independent beta-distributed variables. Using high rate quantization theory and by the entropy constraint, the optimal inter- and intra-component bit allocation strategies are proposed. In the implementation of scalar quantizers, we use the constrained-resolution coding to approximate the derived constrained-entropy coding. A practical coding scheme for DVQ is designed for the purpose of reducing the quantization error accumulation. The theoretical and practical quantization performance of DVQ is evaluated. Compared to the state-of-the-art GMM-based VQ and recently proposed beta mixture model (BMM) based VQ, DVQ performs better, with even fewer free parameters and lower computational cost Zhanyu Ma, Arne Leijon, W. Bastiaan Kleijn |
IEEE Trans. Speech Audio Process. | 1 |
| 2011 | Super-Dirichlet Mixture Models Using Differential Line Spectral Frequencies for Text-Independent Speaker IdentificationabstractA new text-independent speaker identification (SI) system is proposed. This system utilizes the line spectral frequencies (LSFs) as alternative feature set for capturing the speaker characteristics ... Zhanyu Ma, Arne Leijon |
INTERSPEECH | 1 |
| 2011 | Bayesian Estimation of Beta Mixture Models with Variational InferenceabstractBayesian estimation of the parameters in beta mixture models (BMM) is analytically intractable. The numerical solutions to simulate the posterior distribution are available, but incur high computational cost. In this paper, we introduce an approximation to the prior/posterior distribution of the parameters in the beta distribution and propose an analytically tractable (closed form) Bayesian approach to the parameter estimation. The approach is based on the variational inference (VI) framework. Following the principles of the VI framework and utilizing the relative convexity bound, the extended factorized approximation method is applied to approximate the distribution of the parameters in BMM. In a fully Bayesian model where all of the parameters of the BMM are considered as variables and assigned proper distributions, our approach can asymptotically find the optimal estimate of the parameters posterior distribution. Also, the model complexity can be determined based on the data. The closed-form solution is proposed so that no iterative numerical calculation is required. Meanwhile, our approach avoids the drawback of overfitting in the conventional expectation maximization algorithm. The good performance of this approach is verified by experiments with both synthetic and real data. Zhanyu Ma, Arne Leijon |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | Expectation propagation for estimating the parameters of the beta distributionabstractParameter estimation for the beta distribution is analytically intractable due to the integration expression in the normalization constant. For maximum likelihood estimation, numerical methods can be used to calculate the parameters. For Bayesian estimation, we can utilize different approximations to the posterior parameter distribution. A method based on the variational inference (VI) framework reported the posterior mean of the parameters analytically but the approximating distribution violated the correlation between the parameters. We now propose a method via the expectation propagation (EP) framework to approximate the posterior distribution analytically and capture the correlation between the parameters. Compared to the method based on VI, the EP based algorithm performs better with small amounts of data and is more stable. Zhanyu Ma, Arne Leijon |
ICASSP | 1 |
| 2010 | Modelling speech line spectral frequencies with dirichlet mixture modelsabstractStatistical modeling plays an important role in various research areas. It provides away to connect the data with the statistics. Based on the statistical properties of theobserved data, an appropriate model can be chosen that leads to a promising practicalperformance. The Gaussian distribution is the most popular and dominant probabilitydistribution used in statistics, since it has an analytically tractable Probability DensityFunction (PDF) and analysis based on it can be derived in an explicit form. However,various data in real applications have bounded support or semi-bounded support. As the support of the Gaussian distribution is unbounded, such type of data is obviously notGaussian distributed. Thus we can apply some non-Gaussian distributions, e.g., the betadistribution, the Dirichlet distribution, to model the distribution of this type of data.The choice of a suitable distribution is favorable for modeling efficiency. Furthermore,the practical performance based on the statistical model can also be improved by a bettermodeling. An essential part in statistical modeling is to estimate the values of the parametersin the distribution or to estimate the distribution of the parameters, if we consider themas random variables. Unlike the Gaussian distribution or the corresponding GaussianMixture Model (GMM), a non-Gaussian distribution or a mixture of non-Gaussian dis-tributions does not have an analytically tractable solution, in general. In this dissertation,we study several estimation methods for the non-Gaussian distributions. For the Maxi-mum Likelihood (ML) estimation, a numerical method is utilized to search for the optimalsolution in the estimation of Dirichlet Mixture Model (DMM). For the Bayesian analysis,we utilize some approximations to derive an analytically tractable solution to approxi-mate the distribution of the parameters. The Variational Inference (VI) framework basedmethod has been shown to be efficient for approximating the parameter distribution byseveral researchers. Under this framework, we adapt the conventional Factorized Approx-imation (FA) method to the Extended Factorized Approximation (EFA) method and useit to approximate the parameter distribution in the beta distribution. Also, the LocalVariational Inference (LVI) method is applied to approximate the predictive distributionof the beta distribution. Finally, by assigning a beta distribution to each element in thematrix, we proposed a variational Bayesian Nonnegative Matrix Factorization (NMF) forbounded support data. The performances of the proposed non-Gaussian model based methods are evaluatedby several experiments. The beta distribution and the Dirichlet distribution are appliedto model the Line Spectral Frequency (LSF) representation of the Linear Prediction (LP)model for statistical model based speech coding. For some image processing applications,the beta distribution is also applied. The proposed beta distribution based variationalBayesian NMF is applied for image restoration and collaborative filtering. Comparedto some conventional statistical model based methods, the non-Gaussian model basedmethods show a promising improvement. Zhanyu Ma, Arne Leijon |
INTERSPEECH | 1 |
| 2010 | PDF-optimized LSF vector quantization based on beta mixture modelsabstractThe line spectral frequencies (LSF) are known to be the mostefficient representation of the linear predictive coding (LPC) parametersfrom both the distortion and perceptual point of view.By considering the bounded property of the LSF parameters,we apply beta mixture models (BMM) to model the distributionof the LSF parameters. Meanwhile, by following the principlesof probability density function (PDF) optimized vector quantization(VQ), we derive the bit allocation strategy for the BMM.The LSF parameters are obtained from the TIMIT database anda practical VQ is designed. By taking the Bayesian informationcriterion (BIC), the square error (SE) and the spectral distortion(SD) as the criteria, the BMM based VQ outperforms theGaussian mixture model based VQ with uncorrelated Gaussiancomponent (UGMVQ) by about 1-2 bits/vector. Zhanyu Ma, Arne Leijon |
INTERSPEECH | 1 |
| 2009 | Beta mixture models and the application to image classificationabstractStatistical pattern recognition is one of the most studied and applied approaches in the area of pattern recognition. Mixture modelling of densities is an efficient statistical pattern recognition method for continuous data. We propose a classifier based on the beta mixture models for strictly bounded and asymmetrically distributed data. Due to the property of the mixture modelling, the statistical dependence in a multi-dimensional variable is captured, even with the conditional independence assumption in each mixture component. A synthetic example and the USPS handwriting digit data was used to verify the effectiveness of this approach. Compared to the conventional Gaussian mixture models (GMM), the beta mixture models has a better performance on data which has strictly bounded value and asymmetric distribution. The performance of beta mixture models is about equivalent to that of GMM applied to data transformed via a strictly increasing link function. Zhanyu Ma, Arne Leijon |
ICIP | 1 |
| 2009 | Human audio-visual consonant recognition analyzed with three bimodal integration modelsabstractWith A-V recordings. ten normal hearing people took recognition tests at different signal-to-noise ratios (SNR). The AV recognition results are predicted by the fuzzy logical model of perception (FLMP) and the post-labelling integration model (POSTL). We also applied hidden Markov models (HMMs) and multi-stream HMMs (MSHMMs) for the recognition. As expected, all the models agree qualitatively with the results that the benefit gained from the visual signal is larger at lower acoustic SNRs. However, the FLMP severely overestimates the AV integration result, while the POSTL model underestimates it. Our automatic speech recognizers integrated the audio and visual stream efficiently. The visual automatic speech recognizer could be adjusted to correspond to human visual performance. The MSHMMs combine the audio and visual streams efficiently, but the audio automatic speech recognizer must be further improved to allow precise quantitative comparisons with human audio-visual performance. Zhanyu Ma, Arne Leijon |
INTERSPEECH | 1 |