VLDB 2026 Research / reviewers in the wild / expert
Shiliang Zhang
dblp:52/6186
· DBLP profile ↗
205ranked-venue papers
36as first author
122since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 158 · 27 first-author · 88 since 2021Artificial intelligence and machine learning · 108 · 15 first-author · 69 since 2021Databases, data management, data science and information retrieval · 7 · 1 since 2021Computer networks · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCAN: Self-Calibrated AutoregressioN for High-Quality Visual GenerationabstractHuman artists can continuously refine their coarse sketches during artistic creation. This is quite different from existing autoregressive generation, where a token is determined once sampled. Aiming to flexibly refine the generated contents, this paper presents a Self-Calibrated AutoregressioN (SCAN) model capable of self-evaluating and refining generation quality without regenerating the entire image. We unify image token generation and quality evaluation into a single autoregressive model, formulating both tasks as categorical prediction problems. During inference, the model first generates a coarse initial image, then iteratively refines the lowest-quality patches until satisfactory image quality is achieved. Experimental results demonstrate that SCAN effectively handles diverse real-world generation errors and achieves a promising balance between image quality and speed. For example, SCAN-XL achieves an FID of 2.10 and an IS of 326.1, surpassing the LlamaGen-XL by 1.29 (+38%) in FID and 99.0 (+43.6%) in IS, with a 5.6× speedup (19.76s to 3.56s). Compared to recent works, SCAN improves FID and speed by +18.3% and +23% over VAR-d20, and by +7% and +46% over RandAR-XL. Zhanzhou Feng, Qingpei Guo, Jingdong Chen, Ming Yang 0007, Shiliang Zhang |
AAAI | 6 |
| 2026 | When Person Re-Identification Meets Event Camera: A Benchmark Dataset and an Attribute-Guided Re-Identification FrameworkabstractRecent researchers have proposed using event cameras for person re-identification (ReID) due to their promising performance and better balance in terms of privacy protection, event camera-based person ReID has attracted significant attention. Currently, mainstream event-based person ReID algorithms primarily focus on fusing visible light and event stream, as well as preserving privacy. Although significant progress has been made, these methods are typically trained and evaluated on small-scale or simulated event camera datasets, making it difficult to assess their real identification performance and generalization ability. To address the issue of data scarcity, this paper introduces a large-scale RGB-event based person ReID dataset, called EvReID. The dataset contains 118,988 image pairs and covers 1200 pedestrian identities, with data collected across multiple seasons, scenes, and lighting conditions. We also evaluate 15 state-of-the-art person ReID algorithms, laying a solid foundation for future research in terms of both data and benchmarking. Based on our newly constructed dataset, this paper further proposes a pedestrian attribute-guided contrastive learning framework to enhance feature learning for person re-identification, termed TriPro-ReID. This framework not only effectively explores the visual features from both RGB frames and event streams, but also fully utilizes pedestrian attributes as mid-level semantic features. Extensive experiments on the EvReID dataset and MARS datasets fully validated the effectiveness of our proposed RGB-Event person ReID framework. Xiao Wang 0014, Shujuan Wu, Bo Jiang 0002, Shiliang Zhang |
AAAI | 5 |
| 2026 | Binary Descriptor Learning via Diffusion-based Feature Disentanglement
Shunan Mao, Xuefei Lv, Shiliang Zhang |
ISCAS | 5 |
| 2026 | VideoAligner: Text-driven feature decomposition for precise video-text alignment
Zhanzhou Feng, Shunan Mao, Yaowei Wang 0001, Shiliang Zhang |
Pattern Recognit. | 4 |
| 2026 | SA-BCT: Self-Adapting Backward-Compatible TrainingabstractBackward-compatible training enables the deployment of advanced models without requiring updates to old gallery databases. However, existing methods, including old-prototype-based (i.e., those relying on prototypes from the old model) and instance-based approaches, often overlook the impact of the old model's quality. High-quality old models exhibit compact intra-class feature distributions, which facilitate effective alignment between old and new models across various methods. In contrast, low-quality old models produce dispersed features, making it difficult for old-prototype-based methods to extract sufficient information. Additionally, instance-based methods are overly restrictive, limiting the flexibility of new models. In this work, we propose SA-BCT, an extremely simple yet effective backward-compatible training method that offers a unified framework for accommodating old models of varying quality. SA-BCT employs a single loss function applied to both old and new features, self-adaptively adjusting the constraint space for new features based on the distribution of old features. Extensive experiments in diverse settings demonstrate the effectiveness of SA-BCT. Code is available athttps://github.com/yuleung/SA-BCT. Yufeng Zhang 0001, Shiliang Zhang, Sheng Xiao, Rong Xiao 0003, Xiaoyu Wang 0002, Kenli Li 0001 |
IEEE Trans. Multim. | 3 |
| 2026 | Efficient Human Feature Refinement for Weakly Supervised Group Activity RecognitionabstractWeakly supervised group activity recognition (WSGAR) aims to identify the joint activity of a group of people without relying on hand-annotated human bounding boxes. Existing WSGAR methods typically acquire coarse human-level features by pooling from detected bounding boxes or applying human queries with cross attentions. These approaches focus on learning human relations from the acquired features. However, discriminative person-specific clues might be confused with irrelevant backgrounds, hindering the effectiveness of downstream human relation learning. To address this limitation, we propose a Human Feature Refinement framework that enhances human-level information with graph convolutional networks and self-attention. We define in-box regions as tokens and learn their spatial correspondence through GCN and self-attention. By explicitly extracting in-box details and suppressing irrelevant regions, our method acquires more discriminative human-level features for relation learning and group activity prediction. We further propose a Graph-based Token Merging algorithm to reduce the computation cost of Human Feature Refinement, while minimizing information loss and overfitting risk. Experiments show that our method outperforms previous WSGAR methods on Volleyball, NBA and JRDB-PAR benchmarks, with reduced computation cost. Yihui Zhou, Hao Chen 0061, Zhanzhou Feng, Shiliang Zhang |
IEEE Trans. Multim. | 4 |
| 2026 | Adaptive Ensemble Control for Stochastic Systems With Mixed Asymmetric Laplace NoisesabstractThis article presents an adaptive ensemble control for stochastic systems subject to asymmetric noises and outliers. Asymmetric noises skew system observations, and outliers with large amplitude deteriorate the observations even further. Such disturbances induce poor system estimation and degraded stochastic system control. In this work, we model the asymmetric noises and outliers by mixed asymmetric Laplace distributions (ALDs) and propose an optimal control for stochastic systems with mixed ALD noises. Particularly, we segregate the system disturbed by mixed ALD noises into subsystems, each of which is subject to a specific ALD noise. For each subsystem, we design an iterative quantile filter (IQF) to estimate the system parameters using system observations. With the estimated parameters by the IQF, we derive the certainty equivalence (CE) control law for each subsystem. Then we use the Bayesian approach to ensemble the subsystem CE controllers, with each of the controllers weighted by its posterior probability. We finalize our control law as the weighted sum of the control signals by the subsystem CE controllers. To demonstrate our approach, we conduct three numerical simulations and Monte Carlo analyses. The results show improved tracking performance by our approach for skew noises and its robustness to outliers, compared with the RLS-based control policy. Yajie Yu, Xuehui Ma, Shiliang Zhang, Xubing Shi, Yushuai Li, Tingwen Huang |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2025 | Speech Recognition Meets Large Language Model: Benchmarking, Models, and ExplorationabstractIn this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decisions in LLM-based ASR systems are often inadequately justified. This lack of clarity impedes the field's progress, making it challenging to pinpoint which design choices truly improve model performance. To address these challenges, we conduct a comprehensive series of experiments that explore various aspects, leading to the optimal LLM-based ASR system. We found that delicate designs are not necessary, while a clean setup with little task-specific design is competent. The models achieve strong performance on the Librispeech and Gigaspeech datasets, compared to both LLM-based models and non-LLM-based models. Finally, we explore the capability emergence of LLM-based ASR in the process of modal alignment. We hope that our study can facilitate the research on extending LLM with cross-modality capacity and shed light on the LLM-based ASR community. Ziyang Ma 0001, Guanrou Yang, Yifan Yang 0005, Zhifu Gao, Jiaming Wang 0004, Zhihao Du, Fan Yu 0002, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
AAAI | 10 |
| 2025 | MV-VTON: Multi-View Virtual Try-On with Diffusion ModelsabstractThe goal of image-based virtual try-on is to generate an image of the target person naturally wearing the given clothing. However, existing methods solely focus on the frontal try-on using the frontal clothing. When the views of the clothing and person are significantly inconsistent, particularly when the person's view is non-frontal, the results are unsatisfactory. To address this challenge, we introduce Multi-View Virtual Try-ON (MV-VTON), which aims to reconstruct the dressing results from multiple views using the given clothes. Given that single-view clothes provide insufficient information for MV-VTON, we instead employ two images, i.e., the frontal and back views of the clothing, to encompass the complete view as much as possible. Moreover, we adopt diffusion models that have demonstrated superior abilities to perform our MV-VTON. In particular, we propose a view-adaptive selection method where hard-selection and soft-selection are applied to the global and local clothing feature extraction, respectively. This ensures that the clothing features are roughly fit to the person's view. Subsequently, we suggest joint attention blocks to align and fuse clothing features with person features. Additionally, we collect a MV-VTON dataset MVG, in which each person has multiple photos with diverse views and poses. Experiments show that the proposed method not only achieves state-of-the-art results on MV-VTON task using our MVG dataset, but also has superiority on frontal-view virtual try-on task using VITON-HD and DressCode datasets. Zhilu Zhang 0001, Donglin Di, Shiliang Zhang, Wangmeng Zuo |
AAAI | 4 |
| 2025 | UniCodec: Unified Audio Codec with Single Domain-Adaptive CodebookabstractYidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yidi Jiang, Qian Chen 0003, Shengpeng Ji, Yu Xi, Wen Wang 0001, Chong Zhang 0003, Xianghu Yue, Shiliang Zhang, Haizhou Li 0001 |
ACL (1) | 8 |
| 2025 | OmniFlatten: An End-to-end GPT Model for Seamless Voice ConversationabstractQinglin Zhang, Luyao Cheng, Chong Deng, Qian Chen, Wen Wang, Siqi Zheng, Jiaqing Liu, Hai Yu, Chao-Hong Tan, Zhihao Du, ShiLiang Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Luyao Cheng, Chong Deng, Qian Chen 0003, Wen Wang 0001, Jiaqing Liu, Chao-Hong Tan, Zhihao Du, Shiliang Zhang |
ACL (1) | 11 |
| 2025 | NN-Former: Rethinking Graph Structure in Neural Architecture RepresentationabstractThe growing use of deep learning necessitates efficient network design and deployment, making neural predictors vital for estimating attributes such as accuracy and latency. Recently, Graph Neural Networks (GNNs) and transformers have shown promising performance in representing neural architectures. However, each of both methods has its disadvantages. GNNs lack the capabilities to represent complicated features, while transformers face poor generalization when the depth of architecture grows. To mitigate the above issues, we rethink neural architecture topology and show that sibling nodes are pivotal while overlooked in previous research. We thus propose a novel predictor leveraging the strengths of GNNs and transformers to learn the enhanced topology. We introduce a novel token mixer that considers siblings, and a new channel mixer named bidirectional graph isomorphism feed-forward network. Our approach consistently achieves promising performance in both accuracy and latency prediction, providing valuable insights for learning Directed Acyclic Graph (DAG) topology. The code is available at https://github.com/XuRuihan/NNFormer. Ruihan Xu 0002, Haokui Zhang, Yaowei Wang 0001, Wei Zeng 0006, Shiliang Zhang |
CVPR | 5 |
| 2025 | Generalizable Object Keypoint Localization from Generative PriorsabstractGeneralizable object keypoint localization is a fundamental computer vision task in understanding the object structure. It is challenging for existing keypoint localization methods because their limited training data cannot provide generalizable shape and semantic cues, leading to inferior performance and generalization capability. Instead of relying on large scale training data, this work tackles this challenge by exploiting the rich priors from large generative models. We propose a data-efficient generalizable localization method named GenLoc. GenLoc extracts the generative priors from a pre-trained image generation model by calculating the correlation map between image latent feature and condition embedding. Those priors are hence optimized with our proposed heatmap expectation loss to perform object keypoint localization. Benefited by the rich knowledge of generative priors in understanding of object semantics and structures, GenLoc achieves superior performance on various object keypoint localization benchmarks. It shows more substantial performance enhancements in cross-domain, few-shot and zero-shot evaluation settings, e.g., getting 20%+ AP enhancement over CLAMP [43] in various zero-shot settings. Dongkai Wang, Jiang Duan, Liangjian Wen, Shiyu Xuan, Hao Chen 0061, Shiliang Zhang |
CVPR | 6 |
| 2025 | Self-Distillation Prototypes Network: Learning Robust Speaker Representations without SupervisionabstractTraining speaker-discriminative and robust speaker verification systems without explicit speaker labels remains a persistent challenge. In this paper, we propose a novel self-supervised speaker verification approach, Self-Distillation Prototypes Network (SDPN), which effectively facilitates self-supervised speaker representation learning. SDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the augmented and original views. Due to lack of negative pairs in the SDPN training process, the network tends to align positive pairs quite closely in the embedding space, a phenomenon known as model collapse. To mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN. Comprehensive experiments on the VoxCeleb datasets demonstrate the superiority of SDPN among self-supervised speaker verification approaches. SDPN sets a new state-of-the-art on the VoxCeleb1 speaker verification evaluation benchmark, achieving Equal Error Rate 1.80%, 1.99%, and 3.62% for trial VoxCeleb1-O, VoxCeleb1-E and VoxCeleb1H respectively1, without using any speaker labels in training. Ablation studies show that both proposed learnable prototypes in self-distillation network and diversity regularization contribute to the verification performance. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Chong Deng, Shiliang Zhang, Wen Wang 0001 |
ICASSP | 7 |
| 2025 | 3D-Speaker-Toolkit: An Open-Source Toolkit for Multimodal Speaker Verification and DiarizationabstractWe introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to comprehend the substance and context of spoken language, thereby augmenting the system’s proficiency in distinguishing speakers through linguistic patterns. The visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to achieve substantially improved accuracy and reliability in speaker-related tasks. With 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis. The toolkit also includes a handful of open-source state-of-the-art models and a large-scale dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/modelscope/3D-Speaker. Yafeng Chen, Hui Wang 0030, Luyao Cheng, Tinglong Zhu, Rongjie Huang 0001, Chong Deng, Qian Chen 0003, Shiliang Zhang, Wen Wang 0001, Xihao Li |
ICASSP | 9 |
| 2025 | Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data GapabstractWhile automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompassing dialects, accents, minority languages, and long-tail hotwords, domains with significant practical relevance. With the advent of versatile and powerful text-to-speech (TTS) models, capable of generating speech with human-level naturalness, expressiveness, and diverse speaker profiles, leveraging TTS for ASR data augmentation provides a cost-effective and practical approach to enhancing ASR performance. Comprehensive experiments on an unprecedentedly rich variety of low-resource datasets demonstrate consistent and substantial performance improvements, proving that the proposed method of enhancing low-resource ASR through a versatile TTS model is highly effective and has broad application prospects. Furthermore, we delve deeper into key characteristics of synthesized speech data that contribute to ASR improvement, examining factors such as text diversity, speaker diversity, and the volume of synthesized data, with text diversity being studied for the first time in this work. We hope our findings provide helpful guidance and reference for the practical application of TTS-based data augmentation and push the advancement of low-resource ASR one step further. Guanrou Yang, Fan Yu 0002, Ziyang Ma 0001, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 6 |
| 2025 | Unified Visual Generation via Next-Set Prediction in Continuous Domain
Zhanzhou Feng, Qingpei Guo, Xinyu Xiao, Ruihan Xu 0002, Ming Yang 0007, Shiliang Zhang |
ICCV | 6 |
| 2025 | OmniAudio: Generating Spatial Audio from 360-Degree VideoabstractTraditional video-to-audio generation techniques primarily focus on perspective video and non-spatial audio, often missing the spatial cues necessary for accurately representing sound sources in 3D environments. To address this limitation, we introduce a novel task, 360V2SA, to generate spatial audio from 360-degree videos, specifically producing First-order Ambisonics (FOA) audio - a standard format for representing 3D spatial audio that captures sound directionality and enables realistic 3D audio reproduction. We first create Sphere360, a novel dataset tailored for this task that is curated from real-world data. We also design an efficient semi-automated pipeline for collecting and cleaning paired video-audio data. To generate spatial audio from 360-degree video, we propose a novel framework OmniAudio, which leverages self-supervised pre-training using both spatial audio data (in FOA format) and large-scale non-spatial data. Furthermore, OmniAudio features a dual-branch framework that utilizes both panoramic and perspective video inputs to capture comprehensive local and global information from 360-degree videos. Experimental results demonstrate that OmniAudio achieves state-of-the-art performance across both objective and subjective metrics on Sphere360. Code and datasets are available at https://github.com/liuhuadai/OmniAudio. The project website is available at https://OmniAudio-360V2SA.github.io. Huadai Liu, Tianyi Luo, Kaicheng Luo, Qikai Jiang, Peiwen Sun, Rongjie Huang 0001, Qian Chen 0003, Wen Wang 0001, Xiangtai Li, Shiliang Zhang, Zhijie Yan, Zhou Zhao 0001, Wei Xue 0002 |
ICML | 11 |
| 2025 | Efficient Multi-modal Long Context Learning for Training-free AdaptationabstractTraditional approaches to adapting multi-modal large language models (MLLMs) to new tasks have relied heavily on fine-tuning. This paper introduces Efficient Multi-Modal Long Context Learning (EMLoC), a novel training-free alternative that embeds demonstration examples directly into the model input. EMLoC offers a more efficient, flexible, and scalable solution for task adaptation. Because extremely lengthy inputs introduce prohibitive computational and memory overhead, EMLoC contributes a chunk-wise compression mechanism combined with layer-wise adaptive pruning. It condenses long-context multimodal inputs into compact, task-specific memory representations. By adaptively pruning tokens at each layer under a Jensen-Shannon divergence constraint, our method achieves a dramatic reduction in inference complexity without sacrificing performance. This approach is the first to seamlessly integrate compression and pruning techniques for multi-modal long-context learning, offering a scalable and efficient solution for real-world applications. Extensive experiments on diverse vision-language benchmarks demonstrate that EMLoC achieves performance on par with or superior to naive long-context approaches. Our results highlight the potential of EMLoC as a groundbreaking framework for efficient and flexible adaptation of multi-modal models in resource-constrained environments. Codes are publicly available at https://github.com/Zehong-Ma/EMLoC. Zehong Ma, Shiliang Zhang, Longhui Wei, Qi Tian 0001 |
ICML | 2 |
| 2025 | Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
Yifan Yang 0005, Jiajun Deng, Jiawen Kang 0002, Shujie Hu, Tianzi Wang, Zhaoqing Li, Shiliang Zhang, Xie Chen 0001, Xunying Liu |
INTERSPEECH | 8 |
| 2025 | Differentiable Reward Optimization for LLM based TTS system
Changfeng Gao, Zhihao Du, Shiliang Zhang |
INTERSPEECH | 3 |
| 2025 | MoGE: A Benchmark for Comprehensive Evaluation of Molecular Generation Models in De Novo Drug Design
Shiliang Zhang, Renyi Zhou |
ISBRA (1) | 1 |
| 2025 | EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text PromptingabstractHuman speech goes beyond the mere transfer of information; it is a profound exchange of emotions and a connection between individuals. While Text-to-Speech (TTS) models have made huge progress, they still face challenges in controlling the emotional expression in the generated speech. In this work, we propose EmoVoice, a novel emotion-controllable TTS model that exploits large language models (LLMs) to enable fine-grained freestyle natural language emotion control, and a phoneme boost variant design that makes the model output phoneme tokens and audio tokens in parallel to enhance content consistency, inspired by chain-of-thought (CoT) and modality-of-thought (CoM) techniques. Besides, we introduce EmoVoice-DB, a high-quality 40-hour English emotion dataset featuring expressive speech and fine-grained emotion labels with natural language descriptions. EmoVoice achieves state-of-the-art performance on the English EmoVoice-DB test set using only synthetic training data, and on the Chinese Secap test set using our in-house data. We further investigate the reliability of existing emotion evaluation metrics and their alignment with human perceptual preferences, and explore using SOTA multimodal LLMs GPT-4o-audio and Gemini to assess emotional speech. Dataset, code, checkpoints and demo samples are available at https://github.com/yanghaha0908/EmoVoice. Guanrou Yang, Qian Chen 0003, Ziyang Ma 0001, Wen Wang 0019, Tianrui Wang, Yifan Yang 0005, Zhikang Niu, Wenrui Liu 0003, Fan Yu 0002, Zhihao Du, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
ACM Multimedia | 14 |
| 2025 | HDCFN: Haze Distribution-aware Cross-modal Fusion Network for Infrared-guided Dense Haze Removal in UAVsabstractIn UAV applications, dense haze severely obscures small ground-level objects, hindering the recovery of fine details. Existing visible-only dehazing methods struggle with such dense occlusions, while infrared imaging lacks color and fine texture information. To address these limitations, we propose the Haze Distribution-aware Cross-modal Fusion Network (HDCFN). HDCFN features two key components: (i) an infrared-guided multiscale feature enhancement framework that integrates haze-resistant structural cues from infrared modality with visible features across coarse to fine, improving the recovery of small objects, and (ii) a haze distribution-aware cross-modal fusion module that adaptively prioritizes relevant information from each modality according to haze density. This framework effectively combines the complementary strengths of visible and infrared imaging for dense haze removal. Extensive experiments on multiple public datasets show that HDCFN outperforms state-of-the-art dehazing and fusion methods, yielding higher-quality and more detailed images. Junwei Zhao 0003, Qianchun Luo, Shiliang Zhang, Shen Gao, Jie Wu 0001 |
ACM Multimedia | 3 |
| 2025 | MagCache: Fast Video Generation with Magnitude-Aware CacheabstractExisting acceleration techniques for video diffusion models often rely on uniform heuristics or time-embedding variants to skip timesteps and reuse cached features. These approaches typically require extensive calibration with curated prompts and risk inconsistent outputs due to prompt-specific overfitting. In this paper, we introduce a novel and robust discovery: a unified magnitude law observed across different models and prompts. Specifically, the magnitude ratio of successive residual outputs decreases monotonically, steadily in most timesteps while rapidly in the last several steps. Leveraging this insight, we introduce a Magnitude-aware Cache (MagCache) that adaptively skips unimportant timesteps using an error modeling mechanism and adaptive caching strategy. Unlike existing methods requiring dozens of curated samples for calibration, MagCache only requires a single sample for calibration. Experimental results show that MagCache achieves 2.10×-2.68× speedups on Open-Sora, CogVideoX, Wan 2.1, and HunyuanVideo, while preserving superior visual fidelity. It significantly outperforms existing methods in LPIPS, SSIM, and PSNR, under similar computational budgets. Zehong Ma, Longhui Wei, Shiliang Zhang, Qi Tian 0001 |
NeurIPS | 4 |
| 2025 | Extended Vehicle Energy Dataset (eVED): An Enhanced Large-Scale Dataset for Vehicle Energy Consumption AnalysisabstractThis work presents an extended version of the Vehicle Energy Dataset (VED), which is a openly released large-scale dataset of vehicle trip energy consumption records. Compared with its original version, the extended VED (eVED) dataset11A description of our eVED dataset can be found at https://github.com/zhangs12013/eVED, and users can download our data and code via git clone https://bitbucket.org/datarepo/eved-dataset.git is enhanced with accurate vehicle trip GPS coordinates. Based on the accurate trip trajectories, we associate the VED trip records with external information that is essential in analyzing vehicle energy consumption e.g., road speed limit and intersections, from open-source map services. Particularly, we calibrate all the GPS trace records in the original VED data, upon which we associated the VED data with external attributes extracted from Geographic Information System (QGIS), the Overpass API, the Open Street Map API, and Google Maps API. The extracted attributes include$12,609,170$records of road elevation, 12,203,044 of speed limit, 12,281,719 of bi-directional speed limit, 584,551 of intersections, 429,638 of bus stop, 312,196 of crossings, 195,856 of traffic signals, 29,397 of stop signs, 5,848 of turning loops, 4,053 of railway crossings, 3,554 of turning circles, and 2,938 of motorway junctions. With the accurate GPS traces and enriched features of the vehicle trip records, the obtained eVED dataset can facilitate research on vehicle energy consumption and energy-efficient approaches, especially machine/deep learning approaches that are demanding on data volume and richness. Moreover, our software work of data calibration and enrichment can be reused to generate further vehicle trip datasets for specific user cases that empower vehicle behavior and traffic dynamic analyses. We anticipate that the eVED dataset and our data enrichment software can serve the academia and industry as an apparatus in developing future vehicle technologies. Shiliang Zhang, Dyako Fatih, Fahmi Abdulqadir Ahmed, Tobias Schwarz, Xuehui Ma |
VTC2025-Spring | 1 |
| 2025 | Generalization-preserving adaptation of vision-language models for open-vocabulary segmentation
Zhen Chen 0020, Hao Tang 0005, Shiliang Zhang |
Comput. Vis. Image Underst. | 3 |
| 2025 | Incremental Model Enhancement via Memory-based Contrastive Learning
Shiyu Xuan, Ming Yang 0007, Shiliang Zhang |
Int. J. Comput. Vis. | 3 |
| 2025 | Anti-Forgetting Adaptation for Unsupervised Person Re-IdentificationabstractRegular unsupervised domain adaptive person re-identification (ReID) focuses on adapting a model from a source domain to a fixed target domain. However, an adapted ReID model can hardly retain previously-acquired knowledge and generalize to unseen data. In this paper, we propose a Dual-level Joint Adaptation and Anti-forgetting (DJAA) framework, which incrementally adapts a model to new domains without forgetting source domain and each adapted target domain. We explore the possibility of using prototype and instance-level consistency to mitigate the forgetting during the adaptation. Specifically, we store a small number of representative image samples and corresponding cluster prototypes in a memory buffer, which is updated at each adaptation step. With the buffered images and prototypes, we regularize the image-to-image similarity and image-to-prototype similarity to rehearse old knowledge. After the multi-step adaptation, the model is tested on all seen domains and several unseen domains to validate the generalization ability of our method. Extensive experiments demonstrate that our proposed method significantly improves the anti-forgetting, generalization and backward-compatible ability of an unsupervised person ReID model. Hao Chen 0061, François Brémond, Nicu Sebe, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Evolved Hierarchical Masking for Self-Supervised LearningabstractExisting Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those mask patterns resort to different criteria to depict image contents, sticking to a fixed pattern leads to a limited vision cues modeling capability. This paper introduces an evolved hierarchical masking method to pursue general visual cues modeling in self-supervised learning. The proposed method leverages the vision model being trained to parse the input visual cues into a hierarchy structure, which is hence adopted to generate masks accordingly. The accuracy of hierarchy is on par with the capability of the model being trained, leading to evolved mask patterns at different training stages. Initially, generated masks focus on low-level visual cues to grasp basic textures, then gradually evolve to depict higher-level cues to reinforce the learning of more complicated object semantics and contexts. Our method does not require extra pre-trained models or annotations and ensures training efficiency by evolving the training difficulty. We conduct extensive experiments on seven downstream tasks including partial-duplicate image retrieval relying on low-level details, as well as image classification and semantic segmentation that require semantic parsing capability. Experimental results demonstrate that it substantially boosts performance across these tasks. For instance, it surpasses the recent MAE by 1.1% in imageNet-1K classification and 1.4% in ADE20K segmentation with the same training epochs. We also align the proposed method with the current research focus on LLMs. The proposed approach bridges the gap with large-scale pre-training on semantic demanding tasks and enhances intricate detail perception in tasks requiring low-level feature recognition. Zhanzhou Feng, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Multi-Modal Reference Learning for Fine-Grained Text-to-Image RetrievalabstractFine-grained text-to-image retrieval aims to retrieve a fine-grained target image with a given text query. Existing methods typically assume that each training image is accurately depicted by its textual descriptions. However, textual descriptions can be ambiguous and fail to depict discriminative visual details in images, leading to inaccurate representation learning. To alleviate the effects of text ambiguity, we propose a Multi-Modal Reference learning framework to learn robust representations. We first propose a multi-modal reference construction module to aggregate all visual and textual details of the same object into a comprehensive multi-modal reference. The multi-modal reference hence facilitates the subsequent representation learning and retrieval similarity computation. Specifically, a reference-guided representation learning module is proposed to use multi-modal references to learn more accurate visual and textual representations. Additionally, we introduce a reference-based refinement method that employs the object references to compute a reference-based similarity that refines the initial retrieval results. Extensive experiments are conducted on five fine-grained text-to-image retrieval datasets for different text-to-image retrieval tasks. The proposed method has achieved superior performance over state-of-the-art methods. For instance, on the text-to-person image retrieval dataset RSTPReid, our method achieves the Rank1 accuracy of 56.2%, surpassing the recent CFine by 5.6%. Zehong Ma, Hao Chen 0061, Wei Zeng 0006, Limin Su, Shiliang Zhang |
IEEE Trans. Multim. | 5 |
| 2025 | Context-Assisted Active Learning for Weakly Supervised Person SearchabstractPerson search is a challenging task that aims to jointly detect and identify a target person from a large-scale scene image dataset. Fully supervised person search requires both bounding boxes and person identity annotations, making it hard to deploy in real-world applications. Although recent weakly supervised person search methods can alleviate annotation workloads, they often result in severe performance degradation when compared to supervised methods. To pursue better performance with a lower annotation budget, we propose to integrate active learning into weakly supervised person search, where a small number of pairwise identity annotations are actively acquired from oracles. Specifically, we propose a context-assisted active learning framework that selects informative instance pairs for labeling and refines pseudo labels for representation learning. The proposed framework consists of a split module and a merge module, which leverage two types of contextual cues for label refinement. Besides, a pairwise relationship predictor is introduced to estimate relations between instances so that annotation cost can be further reduced. Extensive experiments demonstrate that the proposed method could achieve comparable or even better performance than recent fully supervised methods at a much lower annotation cost. Notably, our method achieves 61.4% mAP on PRW dataset, which outperforms recent fully supervised methods at a much lower annotation cost. Rinyoichi Takezoe, Hao Chen 0061, Xuefei Lv, Yaowei Wang 0001, Shiliang Zhang, Xiaoyu Wang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Decoupled Optimisation for Long-Tailed Visual RecognitionabstractWhen training on a long-tailed dataset, conventional learning algorithms tend to exhibit a bias towards classes with a larger sample size. Our investigation has revealed that this biased learning tendency originates from the model parameters, which are trained to disproportionately contribute to the classes characterised by their sample size (e.g., many, medium, and few classes). To balance the overall parameter contribution across all classes, we investigate the importance of each model parameter to the learning of different class groups, and propose a multistage parameter Decouple and Optimisation (DO) framework that decouples parameters into different groups with each group learning a specific portion of classes. To optimise the parameter learning, we apply different training objectives with a collaborative optimisation step to learn complementary information about each class group. Extensive experiments on long-tailed datasets, including CIFAR100, Places-LT, ImageNet-LT, and iNaturaList 2018, show that our framework achieves competitive performance compared to the state-of-the-art. Cong Cong 0001, Shiyu Xuan, Sidong Liu, Shiliang Zhang, Maurice Pagnucco, Yang Song 0001 |
AAAI | 4 |
| 2024 | Decoupled Contrastive Learning for Long-Tailed RecognitionabstractSupervised Contrastive Loss (SCL) is popular in visual representation learning. Given an anchor image, SCL pulls two types of positive samples, i.e., its augmentation and other images from the same class together, while pushes negative images apart to optimize the learned embedding. In the scenario of long-tailed recognition, where the number of samples in each class is imbalanced, treating two types of positive samples equally leads to the biased optimization for intra-category distance. In addition, similarity relationship among negative samples, that are ignored by SCL, also presents meaningful semantic cues. To improve the performance on long-tailed recognition, this paper addresses those two issues of SCL by decoupling the training objective. Specifically, it decouples two types of positives in SCL and optimizes their relations toward different objectives to alleviate the influence of the imbalanced dataset. We further propose a patch-based self distillation to transfer knowledge from head to tail classes to relieve the under-representation of tail classes. It uses patch-based features to mine shared visual patterns among different instances and leverages a self distillation procedure to transfer such knowledge. Experiments on different long-tailed classification benchmarks demonstrate the superiority of our method. For instance, it achieves the 57.7% top-1 accuracy on the ImageNet-LT dataset. Combined with the ensemble-based method, the performance can be further boosted to 59.7%, which substantially outperforms many recent works. Our code will be released. Shiyu Xuan, Shiliang Zhang |
AAAI | 2 |
| 2024 | Recognizing Ultra-High-Speed Moving Objects with Bio-Inspired Spike CameraabstractBio-inspired spike camera mimics the sampling principle of primate fovea. It presents high temporal resolution and dynamic range, showing great promise in fast-moving object recognition. However, the physical limit of CMOS technology in spike cameras still hinders their capability of recognizing ultra-high-speed moving objects, e.g., extremely fast motions cause blur during the imaging process of spike cameras. This paper presents the first theoretical analysis for the causes of spiking motion blur and proposes a robust representation that addresses this issue through temporal-spatial context learning. The proposed method leverages multi-span feature aggregation to capture temporal cues and employs residual deformable convolution to model spatial correlation among neighbouring pixels. Additionally, this paper contributes an original real-captured spiking recognition dataset consisting of 12,000 ultra-high-speed (equivalent speed > 500 km/h) moving objects. Experimental results show that the proposed method achieves 73.2% accuracy in recognizing 10 classes of ultra-high-speed moving objects, outperforming all existing spike-based recognition methods. Resources will be available at https://github.com/Evin-X/UHSR. Junwei Zhao 0003, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001 |
AAAI | 2 |
| 2024 | OVMR: Open-Vocabulary Recognition with Multi-Modal ReferencesabstractThe challenge of open-vocabulary recognition lies in the model has no clue of new categories it is applied to. Existing works have proposed different methods to embed category cues into the model, e.g., through few-shot fine-tuning, providing category names or textual descriptions to Vision-Language Models. Fine-tuning is time-consuming and degrades the generalization capability. Textual descriptions could be ambiguous and fail to depict visual details. This paper tackles open-vocabulary recognition from a different perspective by referring to multi-modal clues composed of textual descriptions and exemplar images. Our method, named OVMR, adopts two innovative components to pursue a more robust category cues embedding. A multi-modal classifier is first generated by dynamically complementing textual descriptions with image exemplars. A preference-based refinement module is hence applied to fuse uni-modal and multi-modal classifiers, with the aim to alleviate issues of low-quality exemplar images or textual descriptions. The proposed OVMR is a plug-and-play module, and works well with exemplar images randomly crawled from the Internet. Extensive experiments have demonstrated the promising performance of OVMR, e.g., it outperforms existing methods across various scenarios and setups. Codes are publicly available at https://github.com/Zehong-Ma/OVMR. Zehong Ma, Shiliang Zhang, Longhui Wei, Qi Tian 0001 |
CVPR | 2 |
| 2024 | LocLLM: Exploiting Generalizable Human Keypoint Localization via Large Language ModelabstractThe capacity of existing human keypoint localization models is limited by keypoint priors provided by the training data. To alleviate this restriction and pursue more gen-eral model, this work studies keypoint localization from a different perspective by reasoning locations based on key-piont clues in text descriptions. We propose LocLLM, the first Large-Language Model (LLM) based keypoint local-ization model that takes images and text instructions as in-puts and outputs the desired keypoint coordinates. LocLLM leverages the strong reasoning capability of LLM and clues of keypoint type, location, and relationship in textual de-scriptions for keypoint localization. To effectively tune Lo-cLLM, we construct localization-based instruction conver-sations to connect keypoint description with corresponding coordinates in input image, and fine-tune the whole model in a parameter-efficient training pipeline. LocLLM shows remarkable performance on standard 2D/3D keypoint lo-calization benchmarks. Moreover, incorporating language clues into the localization makes LocLLM show superior flexibility and generalizable capability in cross dataset key-point localization, and even detecting novel type of key-points unseen during training††Project page: https://github.com/kennethwdk/LocLLM. Dongkai Wang, Shiyu Xuan, Shiliang Zhang |
CVPR | 3 |
| 2024 | Spatial-Aware Regression for Keypoint LocalizationabstractRegression-based keypoint localization shows advan-tages of high efficiency and better robustness to quantization errors than heatmap-based methods. However, existing regression-based methods discard the spatial location prior in input image with a global pooling, leading to in-ferior accuracy and are limited to single instance localization tasks. We study the regression-based keypoint localization from a new perspective by leveraging the spatiallocation prior. Instead of regressing on the pooled feature, the proposed Spatial-Aware Regression (SAR) maintains the spatial location map and outputs spatial coordinates and confidence score for each grid, which are optimized with a unified objective. Benefited by the location prior, these spatial-aware outputs can be efficiently optimized, resulting in better localization performance. Moreover, incorporating spatial prior makes SAR more general and can be applied into various keypoint localization tasks. We test the proposed method in 4 keypoint localization tasks including single/multi-person 2D/3D pose estimation, and the whole-body pose estimation. Extensive experiments demonstrate its promising performance, e.g., consistently outperforming recent regressions-based methods††project pagn: https://github.com/kennethwdk/SAR. Dongkai Wang, Shiliang Zhang |
CVPR | 2 |
| 2024 | Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMsabstractMulti-modal Large Language Models (MLLMs) have shown remarkable capabilities in various multi-modal tasks. Nevertheless, their performance in fine-grained image understanding tasks is still limited. To address this issue, this paper proposes a new framework to enhance the fine-grained image understanding abilities of MLLMs. Specifically, we present a new method for constructing the instruction tuning dataset at a low cost by leveraging annotations in existing datasets. A self-consistent bootstrapping method is also introduced to extend existing dense object annotations into high-quality referring-expression-bounding-box pairs. These methods enable the generation of high-quality instruction data which includes a wide range of fundamental abilities essential for fine-grained image perception. Moreover, we argue that the visual encoder should be tuned during instruction tuning to mitigate the gap between full image perception and fine-grained image perception. Experimental results demonstrate the superior performance of our method. For instance, our model exhibits a 5.2% accuracy improvement over Qwen-VL on GQA and surpasses the accuracy of Kosmos-2 by 24.7% on RefCOCO_val. We have also attained the top rank on the leaderboard of MM-Bench. This promising performance is achieved by training on only publicly available data, making it easily reproducible. The models, datasets, and codes are publicly available at https://github.com/SY-Xuan/Pink. Shiyu Xuan, Qingpei Guo, Ming Yang 0007, Shiliang Zhang |
CVPR | 4 |
| 2024 | Loss Masking Is Not Needed In Decoder-Only Transformer For Discrete-Token-Based ASRabstractRecently, unified speech-text models, such as SpeechGPT, VioLA, and AudioPaLM, have achieved remarkable performance on various speech tasks. These models discretize speech signals into tokens (speech discretization) and use a shared vocabulary for both text and speech tokens. Then they train a single decoder-only Transformer on a mixture of speech tasks. However, these models rely on the Loss Masking strategy for the ASR task, which ignores the dependency among speech tokens. In this paper, we propose to model speech tokens in an autoregressive way, similar to text. We find that applying the conventional cross-entropy loss on input speech tokens does not consistently improve the ASR performance over the Loss Masking approach. To address this issue, we propose a novel approach denoted Smoothed Label Distillation (SLD), which applies a KL divergence loss with smoothed labels on speech tokens. Our experiments show that SLD effectively models speech tokens and outperforms Loss Masking for decoder-only Transformers in ASR tasks with different speech discretization methods1. Qian Chen 0003, Wen Wang 0001, Shiliang Zhang, Chong Deng, Jiaqing Liu, Chong Zhang 0003 |
ICASSP | 5 |
| 2024 | FunCodec: A Fundamental, Reproducible and Integrable Open-Source Toolkit for Neural Speech CodecabstractThis paper presents FunCodec, a fundamental neural speech codec toolkit, which is an extension of the open-source speech processing toolkit FunASR. FunCodec provides reproducible training recipes and inference scripts for the latest neural speech codec models, such as SoundStream and Encodec. Thanks to the unified design with FunASR, FunCodec can be easily integrated into downstream tasks, such as speech recognition. Along with FunCodec, pretrained models are also provided, which can be used for academic or generalized purposes. Based on the toolkit, we further propose the frequency-domain codec models, FreqCodec, which can achieve comparable speech quality with much lower computation and parameter complexity. Experimental results show that, under the same compression ratio, FunCodec can achieve better reconstruction quality compared with other toolkits and released models. We also demonstrate that the pre-trained models are suitable for downstream tasks, including automatic speech recognition and personalized text-to-speech synthesis. This toolkit is publicly available at https://github.com/alibaba-damo-academy/FunCodec. Zhihao Du, Shiliang Zhang |
ICASSP | 2 |
| 2024 | Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion RecognitionabstractIn this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervised pre-trained models, and we found that data2vec has a good representation ability on the SER task. Second, we employed a powerful large language model (LLM), GPT-4, and emotional text-to-speech (TTS) model, Azure TTS, to generate emotionally congruent text and speech. We carefully designed the text prompt and dataset construction, to obtain the synthetic emotional speech data with high quality. Third, we studied different ways of data augmentation to promote the SER task with synthetic speech, including random mixing, adversarial training, transfer learning, and curriculum learning. Experiments and ablation studies on the IEMOCAP dataset demonstrate the effectiveness of our method, compared with other data augmentation methods, and data augmentation with other synthetic data. Ziyang Ma 0001, Wen Wu 0007, Zhisheng Zheng, Qian Chen 0003, Shiliang Zhang, Xie Chen 0001 |
ICASSP | 6 |
| 2024 | SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization AbilityabstractHotword customization is one of the concerned issues remained in ASR field - it is of value to enable users of ASR systems to customize names of entities, persons and other phrases to obtain better experience. The past few years have seen effective modeling strategies for ASR contextualization developed, but they still exhibit space for improvement about training stability and the invisible activation process. In this paper we propose Semantic-Augmented Contextual-Paraformer (SeACo-Paraformer) a novel NAR based ASR system with flexible and effective hotword customization ability. It possesses the advantages of AED-based model’s accuracy, NAR model’s efficiency, and explicit customization capacity of superior performance. Through extensive experiments with 50,000 hours of industrial big data, our proposed model outperforms strong baselines in customization. Besides, we explore an efficient way to filter large-scale incoming hotwords for further improvement. The industrial models compared, source codes and two hotword test sets are all open source. Xian Shi, Yexin Yang, Yanni Chen, Zhifu Gao, Shiliang Zhang |
ICASSP | 6 |
| 2024 | SlideSpeech: A Large Scale Slide-Enriched Audio-Visual CorpusabstractMulti-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the utilization of extra supplementary textual information has been overlooked. Recognizing the abundance of online conference videos with slides, which provide rich domain-specific information in the form of text and images, we release SlideSpeech, a large-scale audio-visual corpus enriched with slides. The corpus contains 1,705 videos, 1,000+ hours, with 473 hours of high-quality transcribed speech. Moreover, the corpus contains a significant amount of real-time synchronized slides. In this work, we present the pipeline for constructing the corpus and propose baseline methods for utilizing text information in the visual slide context. Through the application of keyword extraction and contextual ASR methods in the benchmark system, we demonstrate the potential of improving speech recognition performance by incorporating textual information from supplementary video slides. Haoxu Wang, Fan Yu 0002, Xian Shi, Yuezhang Wang, Shiliang Zhang, Ming Li 0026 |
ICASSP | 5 |
| 2024 | Hourglass-AVSR: Down-Up Sampling-Based Computational Efficiency Model for Audio-Visual Speech RecognitionabstractRecently audio-visual speech recognition (AVSR), which better leverages video modality as additional information to extend automatic speech recognition (ASR), has shown promising results in complex acoustic environments. However, there is still substantial space to improve as complex computation of visual modules and ineffective fusion of audio-visual modalities. To eliminate these drawbacks, we propose a down-up sampling-based AVSR model (Hourglass-AVSR) to enjoy high efficiency and performance, whose time length is scaled during the intermediate processing, resembling an hourglass. Firstly, we propose a context and residual aware video upsampling approach to improve the recognition performance, which utilizes contextual information from visual representations and captures residual information between adjacent video frames. Secondly, we introduce a visual-audio alignment approach during the upsampling by explicitly incorporating boundary constraint loss. Besides, we propose a cross-layer attention fusion to capture the modality dependencies within each visual encoder layer. Experiments conducted on the MISP-AVSR dataset reveal that our proposed Hourglass-AVSR model outperforms ASR model by 12.9% and 20.8% relative concatenated minimum permutation character error rate (cpCER) reduction on far-field and middle-field test sets, respectively. Moreover, compared to other state-of-the-art AVSR models, our model exhibits the highest improvement in cpCER for the visual module. Furthermore, on the benefit of our down-up sampling approach, Hourglass-AVSR model reduces 54.2% overall computation costs with minor performance degradation. Fan Yu 0002, Haoxu Wang, Ziyang Ma 0001, Shiliang Zhang |
ICASSP | 4 |
| 2024 | LCB-Net: Long-Context Biasing for Audio-Visual Speech RecognitionabstractThe growing prevalence of online conferences and courses presents a new challenge in improving automatic speech recognition (ASR) with enriched textual information from video slides. In contrast to rare phrase lists, the slides within videos are synchronized in real-time with the speech, enabling the extraction of long contextual bias. Therefore, we propose a novel long-context biasing network (LCB-net) for audio-visual speech recognition (AVSR) to leverage the long-context information available in videos effectively. Specifically, we adopt a bi-encoder architecture to simultaneously model audio and long-context biasing. Besides, we also propose a biasing prediction module that utilizes binary cross entropy (BCE) loss to explicitly determine biased phrases in the long-context biasing. Furthermore, we introduce a dynamic contextual phrases simulation to enhance the generalization and robustness of our LCB-net. Experiments on the SlideSpeech, a large-scale audio-visual corpus enriched with slides, reveal that our proposed LCB-net outperforms general ASR model by 9.4%/9.1%/10.9% relative WER/U-WER/B-WER reduction on test set, which enjoys high unbiased and biased performance. Moreover, we also evaluate our model on LibriSpeech corpus, leading to 23.8%/19.2%/35.4% relative WER/U-WER/B-WER reduction over the ASR model. Fan Yu 0002, Haoxu Wang, Xian Shi, Shiliang Zhang |
ICASSP | 4 |
| 2024 | ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
Yafeng Chen, Hui Wang 0030, Luyao Cheng, Qian Chen 0003, Shiliang Zhang |
INTERSPEECH | 6 |
| 2024 | Personality-memory Gated Adaptation: An Efficient Speaker Adaptation for Personalized End-to-end Automatic Speech Recognition
Zhihao Du, Shiliang Zhang, Jiqing Han 0001, Yongjun He 0002 |
INTERSPEECH | 3 |
| 2024 | Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
Yifan Yang 0005, Tian Tan 0002, Shiliang Zhang, Xie Chen 0001 |
INTERSPEECH | 5 |
| 2024 | MaLa-ASR: Multimedia-Assisted LLM-Based ASR
Guanrou Yang, Ziyang Ma 0001, Fan Yu 0002, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
INTERSPEECH | 5 |
| 2024 | CoTuning: A Large-Small Model Collaborating Distillation Framework for Better Model GeneralizationabstractModel compression and distillation techniques have become essential for deploying deep learning models efficiently. However, existing methods often encounter challenges related to model generalization and scalability for harnessing the expertise of pre-trained large models. This paper introduces CoTuning, a novel framework designed to enhance the generalization ability of neural networks by leveraging collaborative learning between large and small models. CoTuning overcomes the limitations of traditional compression and distillation techniques by introducing strategies for knowledge exchange and simultaneous optimization. Our framework comprises an adapter-based co-tuning mechanism between cloud and edge models, a scale-shift projection for feature alignment, and a novel collaborative knowledge distillation mechanism for domain-agnostic tasks. Extensive experiments conducted on various benchmark datasets demonstrate the effectiveness of CoTuning in improving model generalization while maintaining computational efficiency and scalability. The proposed framework exhibits a significant advancement in model compression and distillation, with broad implications for research in the collaborative evolution of large-small models. Zimo Liu, Kangjun Liu, Mingyue Guo 0001, Shiliang Zhang, Yaowei Wang 0001 |
ACM Multimedia | 4 |
| 2024 | CTC-Assisted LLM-Based Contextual ASRabstractContextual ASR or hotword customization holds substantial practical value. Despite the impressive performance of current end-to-end (E2E) automatic speech recognition (ASR) systems, they often face challenges in accurately recognizing rare words. Typical E2E contextual ASR models commonly feature complex architectures and decoding mechanisms, limited in performance and susceptible to interference from distractor words. With large language model (LLM)-based ASR models emerging as the new mainstream, we propose a CTC-Assisted LLM-Based Contextual ASR model with an efficient filtering algorithm. By using coarse CTC decoding results to filter potential relevant hotwords and incorporating them into LLM prompt input, our model attains WER/B-WER of $1.27 \% / 3.67 \%$ and $2.72 \% / 8.02 \%$ on the Librispeech test-clean and test-other sets targeting on recognizing rare long-tail words, demonstrating significant improvements compared to the baseline LLM-based ASR model, and substantially surpassing other related work. More remarkably, with the help of the large language model and proposed filtering algorithm, our contextual ASR model still performs well with 2000 biasing words.1 Guanrou Yang, Ziyang Ma 0001, Zhifu Gao, Shiliang Zhang, Xie Chen 0001 |
SLT | 4 |
| 2024 | Open Set Recognition in Real World
Zhen Yang 0026, Jun Yue 0004, Pedram Ghamisi, Shiliang Zhang, Jiayi Ma 0001, Leyuan Fang |
Int. J. Comput. Vis. | 4 |
| 2024 | Switched Surplus-Based Distributed Security Dispatch for Smart Grid With Persistent Packet LossabstractCommunication network failure, e.g., persistent packet loss, may considerably affect the safe and stable operation of smart grids. This may degrade the performance of various components and applications, including energy management and economic dispatch. We propose a switched surplus-based distributed security dispatch approach to cope with the persistent packet loss under an unreliable communication network environment. First, we jointly consider the packet loss sequence and the dynamic triggering sequence to define actual affected periods caused by the persistent packet loss. Then, we outline an incentive scheme, integrate primal-dual analysis and eigenvalue perturbation theory to design the switched surplus-based distributed security dispatch algorithm. Further, we design a dynamic triggering mechanism that enables the proposed algorithm to dynamically switch to different modes according to the change in network state. With those components, the proposed method offers strong robustness against persistent packet loss. In addition, we provide the convergence and optimality proofs of the algorithm. Finally, simulation results are provided to validate the proposed method and to demonstrate its effectiveness. Rufei Ren, Yushuai Li, Qiuye Sun, Shiliang Zhang, David Wenzhong Gao, Sabita Maharjan |
IEEE Internet Things J. | 4 |
| 2024 | Graph-based social relation inference with multi-level conditional attention
Xiaotian Yu, Hanling Yi, Qie Tang, Wenze Hu, Shiliang Zhang, Xiaoyu Wang 0002 |
Neural Networks | 6 |
| 2024 | Intra-Inter Domain Similarity for Unsupervised Person Re-IdentificationabstractMost of unsupervised person Re-Identification (ReID) works produce pseudo-labels by measuring the feature similarity without considering the domain discrepancy among cameras, leading to degraded accuracy in pseudo-label computation across cameras. This paper targets to address this challenge by decomposing the similarity computation into two stages, i.e., the intra-domain and inter-domain computations, respectively. The intra-domain similarity directly leverages CNN features learned within each camera, hence generates pseudo-labels on different cameras to train the ReID model in a multi-branch network. The inter-domain similarity considers the classification scores of each sample on different cameras as a new feature vector. This new feature effectively alleviates the domain discrepancy among cameras and generates more reliable pseudo-labels. We further propose the Instance and Camera Style Normalization (ICSN) to enhance the robustness to domain discrepancy. ICSN alleviates the intra-camera variations by adaptively learning a combination of instance and batch normalization. ICSN also boosts the robustness to inter-camera variations through TNorm which converts the original style of features into target styles. The proposed method achieves competitive performance on multiple datasets under fully unsupervised, intra-camera supervised and domain generalization settings, e.g., it achieves rank-1 accuracy of 64.4% on the MSMT17 dataset, outperforming the recent unsupervised methods by 20+%. Shiyu Xuan, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | SpiReco: Fast and Efficient Recognition of High-Speed Moving Objects With Spike CameraabstractBenefited from the high temporal resolution and high dynamic range, spike cameras have shown great potential in recognizing high-speed moving objects. However, the computer vision community has not explored this task due to the lack of spike data and annotations of high-speed moving objects. This paper contributes a novel dataset, namedSpiReco(Spiking datasets forRecognition), by recording high-speed moving objects using a spike camera. To annotate the dataset, image labels from established datasets such as MNIST, CIFAR10, and CALTECH101 are utilized. Based on this new dataset, this paper proposes the first spike-based object recognition framework. The proposed framework includes a denoise module, which is designed to suppress spike noise by learning spatio-temporal correlation from neighbouring pixels. Additionally, a motion enhancement module is introduced to address high-speed and random motions. Afterward, binarized neural networks are adopted to save computation costs. These efforts result in a fast and efficient processing framework for spiking data. Experimental results demonstrate the effectiveness of the proposed methods. For example, the proposed spike-based recognition framework achieves 80.2% accuracy in recognizing 101 classes of high-speed moving objects using only 2.2ms of spike streams. The SpiReco is available at https://github.com/Evin-X/SpiReco. Junwei Zhao 0003, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Robust Fine-Grained Visual Recognition With Neighbor-Attention Label CorrectionabstractExisting deep learning methods for fine-grained visual recognition often rely on large-scale, well-annotated training data. Obtaining fine-grained annotations in the wild typically requires concentration and expertise, such as fine category annotation for species recognition, instance annotation for person re-identification (re-id) and dense annotation for segmentation, which inevitably leads to label noise. This paper aims to tackle label noise in deep model training for fine-grained visual recognition. We propose a Neighbor-Attention Label Correction (NALC) model to correct labels during the training stage. NALC samples a training batch and a validation batch from the training set. It hence leverages a meta-learning framework to correct labels in the training batch based on the validation batch. To enhance the optimization efficiency, we introduce a novel nested optimization algorithm for the meta-learning framework. The proposed training procedure consistently improves label accuracy in the training batch, consequently enhancing the learned image representation. Experimental results demonstrate that our method significantly increases label accuracy from 70% to over 98% and outperforms recent approaches by up to 13.4% in mean Average Precision (mAP) on various fine-grained image retrieval (FGIR) tasks, including instance retrieval on CUB200 and person re-id on Market1501. We also demonstrate the efficacy of NALC on noisy semantic segmentation datasets generated from Cityscapes, where it achieves a significant 7.8% improvement in mIOU score. NALC also exhibits robustness to different types of noise, including simulated noise such as Asymmetric, Pair-Flip, and Pattern noise, as well as practical noisy labels generated by tracklets and clustering. Shunan Mao, Shiliang Zhang |
IEEE Trans. Image Process. | 2 |
| 2024 | Adapting Vision-Language Models via Learning to Inject KnowledgeabstractPre-trained vision-language models (VLM) such as CLIP, have demonstrated impressive zero-shot performance on various vision tasks. Trained on millions or even billions of image-text pairs, the text encoder has memorized a substantial amount of appearance knowledge. Such knowledge in VLM is usually leveraged by learning specific task-oriented prompts, which may limit its performance in unseen tasks. This paper proposes a new knowledge injection framework to pursue a generalizable adaption of VLM to downstream vision tasks. Instead of learning task-specific prompts, we extract task-agnostic knowledge features, and insert them into features of input images or texts. The fused features hence gain better discriminative capability and robustness to intra-category variances. Those knowledge features are generated by inputting learnable prompt sentences into text encoder of VLM, and extracting its multi-layer features. A new knowledge injection module (KIM) is proposed to refine text features or visual features using knowledge features. This knowledge injection framework enables both modalities to benefit from the rich knowledge memorized in the text encoder. Experiments show that our method outperforms recently proposed methods under few-shot learning, base-to-new classes generalization, cross-dataset transfer, and domain generalization settings. For instance, it outperforms CoOp by 4.5% under the few-shot learning setting, and CoCoOp by 4.4% under the base-to-new classes generalization setting. Our code will be released. Shiyu Xuan, Ming Yang 0007, Shiliang Zhang |
IEEE Trans. Image Process. | 3 |
| 2024 | AAformer: Auto-Aligned Transformer for Person Re-IdentificationabstractIn person re-identification (re-ID), extracting part-level features from person images has been verified to be crucial to offer fine-grained information. Most of the existing CNN-based methods only locate the human parts coarsely, or rely on pretrained human parsing models and fail in locating the identifiable nonhuman parts (e.g., knapsack). In this article, we introduce an alignment scheme in transformer architecture for the first time and propose the auto-aligned transformer (AAformer) to automatically locate both the human parts and nonhuman ones at patch level. We introduce the "Part tokens ([PART]s)," which are learnable vectors, to extract part features in the transformer. A [PART] only interacts with a local subset of patches in self-attention and learns to be the part representation. To adaptively group the image patches into different subsets, we design the auto-alignment. Auto-alignment employs a fast variant of optimal transport (OT) algorithm to online cluster the patch embeddings into several groups with the [PART]s as their prototypes. AAformer integrates the part alignment into the self-attention and the output [PART]s can be directly used as part features for retrieval. Extensive experiments validate the effectiveness of [PART]s and the superiority of AAformer over various state-of-the-art methods. Kuan Zhu, Haiyun Guo, Shiliang Zhang, Yaowei Wang 0001, Jing Liu 0001, Jinqiao Wang, Ming Tang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Efficient Video Transformers via Spatial-temporal Token Merging for Action RecognitionabstractTransformer has exhibited promising performance in various video recognition tasks but brings a huge computational cost in modeling spatial-temporal cues. This work aims to boost the efficiency of existing video transformers for action recognition through eliminating redundancies in their tokens and efficiently learning motion cues of moving objects. We propose a lightweight and plug-and-play module, namely Spatial-temporal Token Merger (STTM), to merge the tokens belonging to the same object into a more compact representation. STTM first adaptively identifies crucial object clues underlying the video as meta tokens. Similarity scores between input tokens and meta tokens are hence computed and used to guide the fusion of similar tokens in both spatial and temporal domains, respectively. To compensate for motion cues lost in the merging procedure, we compute the linear aggregation of spatial-temporal positions of tokens as motion features. STTM hence outputs a compact set of tokens fusing both appearance and motion features of moving objects. This procedure substantially decreases the number of tokens that need to be processed by each Transformer block and boosts the efficiency. As a general module, STTM can be applied to different layers of various video Transformers. Extensive experiments on the action recognition datasets Kinectics-400 and SSv2 demonstrate its promising performance. For example, it reduces the computation complexity of ViT by 38% while maintaining a similar performance on Kinectics-400. It also brings 1.7% gains of top-1 accuracy on SSv2 under the same computational cost. Zhanzhou Feng, Lei Ma 0008, Shiliang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | TPTE: Text-Guided Patch Token Exploitation for Unsupervised Fine-Grained Representation LearningabstractRecent advances in pre-trained vision-language models have successfully boosted the performance of unsupervised image representation in many vision tasks. Most of existing works focus on learning global visual features with Transformers and neglect detailed local cues, leading to suboptimal performance in fine-grained vision tasks. In this article, we propose a text-guided patch token exploitation framework to enhance the discriminative power of unsupervised representation by exploiting more detailed local features. Our text-guided decoder extracts local features with the guidance of texts or learned prompts describing discriminative object parts. We hence introduce a local-global relation distillation loss to promote the joint optimization of local and global features. The proposed method allows to flexibly extract either global or global-local features as the image representation. It significantly outperforms previous methods in fine-grained image retrieval and base-to-new fine-grained classification tasks. For instance, our Recall@1 metric surpasses the recent unsupervised retrieval method STML by 6.0% on the SOP dataset. The code is publicly available at https://github.com/maosnhehe/TPTE . Shunan Mao, Hao Chen 0061, Yaowei Wang 0001, Wei Zeng 0006, Shiliang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Sa-Paraformer: Non-Autoregressive End-To-End Speaker-Attributed ASRabstractJoint modeling of multi-speaker ASR and speaker diarization has recently shown promising results in speaker-attributed automatic speech recognition (SA-ASR). Although being able to obtain state-of-the-art (SOTA) performance, most of the studies are based on an autoregressive (AR) decoder which generates tokens one-by-one and results in a large real-time factor (RTF). To speed up inference, we introduce a recently proposed non-autoregressive model Paraformer as an acoustic model in the SA-ASR model. Paraformer uses a single-step decoder to enable parallel generation, obtaining comparable performance to the SOTA AR transformer models. Besides, we propose a speaker-filling strategy to reduce speaker identification errors and adopt an inter-CTC strategy to enhance the encoder’s ability in acoustic modeling. Experiments on the AliMeeting corpus show that our model outperforms the cascaded SA-ASR model by a 6.1% relative speaker-dependent character error rate (SD-CER) reduction on the test set. Moreover, our model achieves a comparable SD-CER of 34.8% with only 1/10 RTF compared with the SOTA joint AR SA-ASR model. Yangze Li, Fan Yu 0002, Yuhao Liang, Mohan Shi, Zhihao Du, Shiliang Zhang, Lei Xie 0001 |
ASRU | 7 |
| 2023 | The Second Multi-Channel Multi-Party Meeting Transcription Challenge (M2MeT 2.0): A Benchmark for Speaker-Attributed ASRabstractWith the success of the first Multi-channel Multi-party Meeting Transcription challenge (M2MeT), the second M2MeT challenge (M2MeT 2.0) held in ASRU2023 particularly aims to tackle the complex task of speaker-attributed ASR (SAASR), which directly addresses the practical and challenging problem of “who spoke what at when” at typical meeting scenario. We particularly established two sub-tracks. The fixed training condition sub-track, where the training data is constrained to predetermined datasets, but participants can use any open-source pre-trained model. The open training condition sub-track, which allows for the use of all available data and models without limitation. In addition, we release a new 10-hour test set for challenge ranking. This paper provides an overview of the dataset, track settings, results, and analysis of submitted systems, as a benchmark to show the current state of speaker-attributed ASR. Yuhao Liang, Mohan Shi, Fan Yu 0002, Yangze Li, Shiliang Zhang, Zhihao Du, Qian Chen 0003, Lei Xie 0001, Yanmin Qian, Jian Wu 0027, Zhuo Chen 0006, Kong-Aik Lee, Zhijie Yan, Hui Bu |
ASRU | 5 |
| 2023 | Evolved Part Masking for Self-Supervised LearningabstractExisting Masked Image Modeling methods apply fixed mask patterns to guide the self-supervised training. As those patterns resort to different criteria to mask local regions, sticking to a fixed pattern leads to limited vision cues modeling capability. This paper proposes an evolved part-based masking to pursue more general visual cues modeling in self-supervised learning. Our method is based on an adaptive part partition module, which leverages the vision model being trained to construct a part graph, and partitions parts with graph cut. The accuracy of partitioned parts is on par with the capability of the pretrained model, leading to evolved mask patterns at different training stages. It generates simple patterns at the initial training stage to learn low-level visual cues, which hence evolves to eliminate accurate object parts to reinforce the learning of object semantics and contexts. Our method does not require extra pretrained models or annotations, and effectively ensures the training efficiency by evolving the training difficulty. Experiment results show that it substantially boosts the performance on various tasks including image classification, object detection, and semantic segmentation. For example, it outperforms the recent MAE by 0.69% on imageNet-1K classification and 1.61% on ADE20K segmentation with the same training epochs. Zhanzhou Feng, Shiliang Zhang |
CVPR | 2 |
| 2023 | Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech RecognitionabstractIn recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech as input for the backend. However, it is difficult for speech enhancement systems to directly separate speech from input due to the diverse types of noise with different intensities. Furthermore, speech distortion and residual noise are often observed in enhanced speech, and the distortion of speech and noise is different. Most existing methods focus on fusing enhanced and noisy features to address this issue. In this paper, we propose a dual-stream spectrogram refine network to simultaneously refine the speech and noise and decouple the noise from the noisy input. Our proposed method can achieve better performance with a relative 8.6% CER reduction. Haoyu Lu, Tongtong Song, Longbiao Wang, Jianwu Dang 0001, Xiaobao Wang, Shiliang Zhang |
ICASSP | 7 |
| 2023 | TOLD: a Novel Two-Stage Overlap-Aware Framework for Speaker DiarizationabstractRecently, end-to-end neural diarization (EEND) is introduced and achieves promising results in speaker-overlapped scenarios. In EEND, speaker diarization is formulated as a multi-label prediction problem, where speaker activities are estimated independently and their dependency are not well considered. To overcome these disadvantages, we employ the power set encoding to reformulate speaker diarization as a single-label classification problem and propose the overlap-aware EEND (EEND-OLA) model, in which speaker overlaps and dependency can be modeled explicitly. Inspired by the success of two-stage hybrid systems, we further propose a novel Two-stage OverLap-aware Diarization framework (TOLD) by involving a speaker overlap-aware post-processing (SOAP) model to iteratively refine the diarization results of EEND-OLA. Experimental results show that, compared with the original EEND, the proposed EEND-OLA achieves a 14.39% relative improvement in terms of diarization error rates (DER), and utilizing SOAP provides another 19.33% relative improvement. As a result, our method TOLD achieves a DER of 10.14% on the CALLHOME dataset, which is a new state-of-the-art result on this benchmark to the best of our knowledge. Jiaming Wang 0004, Zhihao Du, Shiliang Zhang |
ICASSP | 3 |
| 2023 | ParCNetV2: Oversized Kernel with Enhanced Attention*abstractTransformers have shown great potential in various computer vision tasks. By borrowing design concepts from transformers, many studies revolutionized CNNs and showed remarkable results. This paper falls in this line of studies. Specifically, we propose a new convolutional neural network, ParCNetV2, that extends the research line of ParCNetV1 by bridging the gap between CNN and ViT. It introduces two key designs: 1) Oversized Convolution (OC) with twice the size of the input, and 2) Bifurcate Gate Unit (BGU) to ensure that the model is input adaptive. Fusing OC and BGU in a unified CNN, ParCNetV2 is capable of flexibly extracting global features like ViT, while maintaining lower latency and better accuracy. Extensive experiments demonstrate the superiority of our method over other convolutional neural networks and hybrid models that combine CNNs and transformers. The code are publicly available at https://github.com/XuRuihan/ParCNetV2. Ruihan Xu 0002, Haokui Zhang, Wenze Hu, Shiliang Zhang, Xiaoyu Wang 0002 |
ICCV | 4 |
| 2023 | 3D Human Mesh Recovery with Sequentially Global Rotation EstimationabstractModel-based 3D human mesh recovery aims to reconstruct a 3D human body mesh by estimating its parameters from monocular RGB images. Most of recent works adopt the Skinned Multi-Person Linear (SMPL) model to regress relative rotations for each body joint along the kinematics chain. This pipeline needs to transform each relative rotation matrix into a global rotation matrix to articulate the canonical mesh, and suffers from accumulated errors along the kinematics chain. This paper proposes to directly estimate the global rotation of each joint to avoid error accumulation and pursue better accuracy. The proposed Sequentially Global Rotation Estimation (SGRE) directly predicts the global rotation matrix of each joint on the kinematics chain. SGRE features a residual learning module to leverage complementary features and previously predicted rotations of parent joints to guide the estimation of subsequent child joints. Thanks to this global estimation pipeline and residual learning module, SGRE alleviates error accumulation and produces more accurate 3D human mesh. It can be flexibly integrated into existing regression-based methods and achieves superior performance on various benchmarks. For example, it improves the latest method 3DCrowdNet by 3.3 mm MPJPE and 5.0 mm PVE on 3DPW dataset and 3.0 AP on COCO dataset, respectively†. Dongkai Wang, Shiliang Zhang |
ICCV | 2 |
| 2023 | BAT: Boundary aware transducer for memory-efficient and low-latency ASR
Keyu An, Xian Shi, Shiliang Zhang |
INTERSPEECH | 3 |
| 2023 | FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Zhifu Gao, Jiaming Wang 0004, Haoneng Luo, Xian Shi, Mengzhe Chen, Yabin Li, Lingyun Zuo, Zhihao Du, Shiliang Zhang |
INTERSPEECH | 10 |
| 2023 | Personality-aware Training based Speaker Adaptation for End-to-end Speech Recognition
Zhihao Du, Shiliang Zhang, Qian Chen 0003, Jiqing Han 0001 |
INTERSPEECH | 3 |
| 2023 | Rethinking the Visual Cues in Audio-Visual Speaker Extraction
Meng Ge, Zexu Pan, Longbiao Wang, Jianwu Dang 0001, Shiliang Zhang |
INTERSPEECH | 7 |
| 2023 | BA-SOT: Boundary-Aware Serialized Output Training for Multi-Talker ASR
Yuhao Liang, Fan Yu 0002, Yangze Li, Shiliang Zhang, Qian Chen 0003, Lei Xie 0001 |
INTERSPEECH | 5 |
| 2023 | CASA-ASR: Context-Aware Speaker-Attributed ASR
Mohan Shi, Zhihao Du, Qian Chen 0003, Fan Yu 0002, Yangze Li, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 6 |
| 2023 | Accurate and Reliable Confidence Estimation Based on Non-Autoregressive End-to-End Speech Recognition System
Xian Shi, Haoneng Luo, Zhifu Gao, Shiliang Zhang, Zhijie Yan |
INTERSPEECH | 4 |
| 2023 | Semantic VAD: Low-Latency Voice Activity Detection for Speech Interaction
Mohan Shi, Yuchun Shu, Lingyun Zuo, Qian Chen 0003, Shiliang Zhang, Jie Zhang 0042, Li-Rong Dai 0001 |
INTERSPEECH | 5 |
| 2023 | MMSpeech: Multi-modal Multi-task Encoder-Decoder Pre-training for speech recognition
Xiaohuan Zhou, Jiaming Wang 0004, Zeyu Cui, Shiliang Zhang, Zhijie Yan, Jingren Zhou 0001, Chang Zhou 0005 |
INTERSPEECH | 4 |
| 2023 | HumVis: Human-Centric Visual Analysis SystemabstractHuman-centric visual analysis is a fundamental task for many multimedia and computer vision applications, such as self-driving, multimedia retrieval, and augmented reality, etc. Based on our recent research efforts on fine-grained human visual analysis, we develop a robust and efficient human-centric visual analysis system named as HumVis. HumVis is built on a simple yet efficient contextual instance decoupling (CID) module, which can effectively separate different persons in an input image and output corresponding person structure information for visual analysis. Based on CID, HumVis achieves accurate multi-person pose estimation, multi-person foreground segmentation, multi-person part segmentation and 3D human mesh recovery for user-uploaded images/videos and support live stream presentation. Dongkai Wang, Shiliang Zhang, Yaowei Wang 0001, Yonghong Tian 0001, Tiejun Huang 0001, Wen Gao 0001 |
ACM Multimedia | 2 |
| 2023 | Recognizing High-Speed Moving Objects with Spike CameraabstractSpike camera is a novel bio-inspired vision sensor that mimics the sampling mechanism of the primate fovea. It presents high temporal resolution and dynamic range, showing great potentials in the high-speed moving object recognition task, which has not been fully explored in the Multimedia community due to the lack of data and annotations. This paper contributes the first large-scale High-Speed Spiking Recognition (HSSR) dataset, by recording high-speed moving objects using a spike camera. The HSSR dataset contains 135,000 indoor objects annotated using ImageNet labels and 3,100 outdoor objects collected from real-world scenarios. Furthermore, we propose an original spiking recognition framework, which employs long-term spike stream features to supervise the feature learning from short-term spike streams. This framework improves the recognition accuracy, meanwhile substantially decreasing the recognition latency, making our method can accurately recognize moving objects at an equivalent speed of 514 km/h, using only 1 ms of spike stream. Experimental results show that, the proposed method achieves 76.5% accuracy for recognizing 100 fine-grained indoor objects and 84.3% accuracy for recognizing 8 outdoor objects using 1 ms of spike streams. Resources will be available at https://github.com/Evin-X/HSSR. Junwei Zhao 0003, Jianming Ye, Shiliang Zhang, Zhaofei Yu, Tiejun Huang 0001 |
ACM Multimedia | 3 |
| 2023 | Unleashing the Full Potential of Product Quantization for Large-Scale Image RetrievalabstractDue to its promising performance, deep hashing has become a prevalent method for approximate nearest neighbors search (ANNs). However, most of current deep hashing methods are validated on relatively small-scale datasets, leaving potential threats when are applied to large-scale real-world scenarios. Specifically, they can be constrained either by the computational cost due to the large number of training categories and samples, or unsatisfactory accuracy. To tackle those issues, we propose a novel deep hashing framework based on product quantization (PQ). It uses a softmax-based differentiable PQ branch to learn a set of predefined PQ codes of the classes. Our method is easy to implement, does not involve large-scale matrix operations, and learns highly discriminate compact codes. We validate our method on multiple large-scaled datasets, including ImageNet100, ImageNet1K, and Glint360K, where the category size scales from 100 to 360K and sample number scales from 10K to 17 million, respectively. Extensive experiments demonstrate the superiority of our method. Code is available at https://github.com/yuleung/FPPQ. Shiliang Zhang, Li Ken Li |
NeurIPS | 2 |
| 2023 | Contextual Instance Decoupling for Instance-Level Human AnalysisabstractOne fundamental challenge of instance-level human analysis is to decouple instances in crowded scenes, where multiple persons are overlapped with each other. This paper proposes the Contextual Instance Decoupling (CID), which presents a new pipeline of decoupling persons for multi-person instance-level analysis. Instead of relying on person bounding boxes to spatially differentiate persons, CID decouples persons in an image into multiple instance-aware feature maps. Each of those feature maps is hence adopted to infer instance-level cues for a specific person, e.g., keypoints, instance mask or part segmentation masks. Compared with bounding box detection, CID is differentiable and robust to detection errors. Decoupling persons into different feature maps also allows to isolate distractions from other persons, and explore context cues at scales larger than the bounding box size. Extensive experiments on various tasks including multi-person pose estimation, person foreground segmentation, and part segmentation, show that CID consistently outperforms previous methods in both accuracy and efficiency. For instance, it achieves 71.3% AP on CrowdPose in multi-person pose estimation, outperforming the recent single-stage DEKR by 5.6%, the bottom-up CenterAttention by 3.7%, and the top-down JC-SPPE by 5.3%. This advantage sustains on multi-person segmentation and part segmentation tasks. Dongkai Wang, Shiliang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Multi-proxy feature learning for robust fine-grained visual recognition
Shunan Mao, Yaowei Wang 0001, Xiaoyu Wang 0002, Shiliang Zhang |
Pattern Recognit. | 4 |
| 2023 | A CIF-Based Speech Segmentation Method for Streaming E2E ASRabstractLong utterances segmentation is crucial in end-to-end (E2E) streaming automatic speech recognition (ASR). However, commonly used voice activity detection(VAD)-based and fixed-length segmentation methods may lead to long segments and semantic incompleteness, affecting the user experience and ASR performance. In this paper, we propose a speech segmentation method for streaming E2E ASR to solve the above issues. Both the decoder's dependence on acoustic information and the human average breath frequency are used for judging segment boundaries. Frame-level decoder's dependence information is provided by the Continuous Integrate-and-Fire (CIF) predictor, which optimizes jointly with ASR to guarantee a more suitable segmentation for ASR. Besides, the proposed method does not increase the model parameters and real-time factor (RTF). The experimental results show that our method can accurately detect the pauses in speech, and the segment usually contains relatively complete semantic information. Compared with VAD-based segmentation, 53.5% latency reduction and 3.7% CER reduction relatively are achieved. Yuchun Shu, Haoneng Luo, Shiliang Zhang, Longbiao Wang, Jianwu Dang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | DCR-ReID: Deep Component Reconstruction for Cloth-Changing Person Re-IdentificationabstractPerson re-identification (Re-ID) plays an important role in many areas such as robotics, multimedia and forensics. However, it becomes difficult when considering long-term scenarios, due to changing clothes irregularly for people. Therefore, cloth-changing person re-identification (CC-ReID) has attracted more attention recently. CC-ReID aims to identify the same person but with different clothes. Its main challenge is how to disentangle clothes-irrelevant features, such as face, shape, body, etc. Most existing methods force the model to learn clothes-irrelevant features by changing the colour of clothes or reconstructing people dressed in different colours. However, due to the lack of the ground truth for supervision, these methods inevitably introduce noises which spoil the discriminativeness of features and lead to uncontrollable disentanglement. In this paper, we propose a novel disentanglement framework, called Deep Component Reconstruction Re-ID (DCR-ReID), which can disentangle the clothes-irrelevant features and the clothes-relevant features in a controllable manner. Specifically, we propose a Component Reconstruction Disentanglement (CRD) module to disentangle the clothes-irrelevant features and the clothes-relevant features based on the reconstruction of human component regions. In addition, we propose a Deep Assembled Disentanglement (DAD) module, which further improves the discriminativeness of these disentangled features. Extensive experiments on three real-world benchmark CC-ReID datasets, LTCC, PRCC, and CCVID, are conducted to demonstrate the effectiveness of the proposed DCR-ReID. Empirical studies show that our DCR-ReID achieves the state-of-the-art performance against the other CC-ReID methods. The source code of this paper is available athttps://github.com/PKU-ICST-MIPL/DCR-ReID_TCSVT2023. Zhenyu Cui, Jiahuan Zhou, Yuxin Peng 0001, Shiliang Zhang, Yaowei Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Efficient Vision Transformer via Token MergerabstractVision Transformers (ViTs) split an image into fixed-size patches as tokens. This strategy has succeeded in computer vision tasks, but introduces considerable tokens similar in semantics and appearances. This work proposes Token Merger to spot redundant tokens and merge them into a compact representation to accelerate ViTs. For each forward inference, the Token Merger first identifies meta tokens to represent meaningful cues of the image content, then adaptively merges similar tokens into a uniform one referring to meta tokens. To pursue a reasonable tradeoff between accuracy and efficiency, we further introduce learnable gates to adaptively decide the token merge ratios of different layers. As a generalizable module, Token Merger can be easily plugged into different layers of ViTs to boost their efficiency. Visualizations show that Token Merger progressively merges tokens and finally learns a compact set of tokens representing clear semantics. Compared with token pruning methods, Token Merger is more effective in preserving meaning contextual cues, thus performs and generalizes substantially better in different vision tasks. Extensive experiments and comparisons with other state-of-the-art downsampling methods also demonstrate its promising performance. For instance, it reduces 95% tokens and accelerates the inference speed by 62%. Meanwhile, the ImageNet classification accuracy only drops by 0.4%. The code will be available. Zhanzhou Feng, Shiliang Zhang |
IEEE Trans. Image Process. | 2 |
| 2023 | PolarPose: Single-Stage Multi-Person Pose Estimation in Polar CoordinatesabstractRegression based multi-person pose estimation receives increasing attention because of its promising potential in achieving realtime inference. However, the challenges in long-range 2D offset regression have restricted the regression accuracy, leading to a considerable performance gap compared with heatmap based methods. This paper tackles the challenge of long-range regression through simplifying the 2D offset regression to a classification task. We present a simple yet effective method, named PolarPose, to perform 2D regression in Polar coordinate. Through transforming the 2D offset regression in Cartesian coordinate to quantized orientation classification and 1D length estimation in the Polar coordinate, PolarPose effectively simplifies the regression task, making the framework easier to optimize. Moreover, to further boost the keypoint localization accuracy in PolarPose, we propose a multi-center regression to relieve the quantization error during orientation quantization. The resulting PolarPose framework is able to regress the keypoint offsets in a more reliable way, and achieves more accurate keypoint localization. Tested with the single-model and single-scale setting, PolarPose achieves the AP of 70.2% on COCO test-dev dataset, outperforming the state-of-the-art regression based methods. PolarPose also achieves promising efficiency, e.g., 71.5% AP at 21.5FPS and 68.5%AP at 24.2FPS and 65.5%AP at 27.2FPS on COCO val2017 dataset, faster than current state-of-the-art. Jianing Li 0001, Yaowei Wang 0001, Shiliang Zhang |
IEEE Trans. Image Process. | 3 |
| 2023 | MFGNet: Dynamic Modality-Aware Filter Generation for RGB-T TrackingabstractMany RGB-T trackers attempt to attain robust feature representation by utilizing an adaptive weighting scheme (or attention mechanism). Different from these works, we propose a new dynamic modality-aware filter generation module (named MFGNet) to boost the message communication between visible and thermal data by adaptively adjusting the convolutional kernels for various input images in practical tracking. Given the image pairs as input, we first encode their features with the backbone network. Then, we concatenate these feature maps and generate dynamic modality-aware filters with two independent networks. The visible and thermal filters will be used to conduct a dynamic convolutional operation on their corresponding input feature maps respectively. Inspired by residual connection, both the generated visible and thermal feature maps will be summarized with input feature maps. The augmented feature maps will be fed into the RoI align module to generate instance-level features for subsequent classification. To address issues caused by heavy occlusion, fast motion and out-of-view, we propose to conduct a joint local and global search by exploiting a new direction-aware target driven attention mechanism. The spatial and temporal recurrent neural network is used to capture the direction-aware context for accurate global attention prediction. Extensive experiments on three large-scale RGB-T tracking benchmark datasets validated the effectiveness of our proposed algorithm. Xiao Wang 0014, Xiujun Shu, Shiliang Zhang, Bo Jiang 0002, Yaowei Wang 0001, Yonghong Tian 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Complementary Coarse-to-Fine Matching for Video Object SegmentationabstractSemi-supervised Video Object Segmentation (VOS) needs to establish pixel-level correspondences between a video frame and preceding segmented frames to leverage their segmentation clues. Most works rely on features at a single scale to establish those correspondences, e.g., perform dense matching with Convolutional Neural Network (CNN) features from a deep layer. Differently, this work explores complementary features at different scales to pursue more robust feature matching. A coarse feature from a deep layer is first adopted to get coarse pixel-level correspondences. We hence evaluate the quality of those correspondences, and select pixels with low-quality correspondences for fine-scale feature matching. Segmentation clues of previous frames are propagated by both coarse and fine-scale correspondences, which are fused with appearance features for object segmentation. Compared with previous works, this coarse-to-fine matching scheme is more robust to distractions by similar objects and better preserves object details. The sparse fine-scale matching also ensures a fast inference speed. On popular VOS datasets including DAVIS and YouTube-VOS, the proposed method shows promising performance compared with recent works. Zhen Chen 0020, Ming Yang 0007, Shiliang Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | Domain Generalization Capability Enhancement for Binary Neural Networks
Jianming Ye, Shunan Mao, Shiliang Zhang |
BMVC | 3 |
| 2022 | Contextual Instance Decoupling for Robust Multi-Person Pose EstimationabstractCrowded scenes make it challenging to differentiate persons and locate their pose keypoints. This paper proposes the Contextual Instance Decoupling (CID), which presents a new pipeline for multi-person pose estimation. Instead of relying on person bounding boxes to spatially differentiate persons, CID decouples persons in an image into multiple instance-aware feature maps. Each of those feature maps is hence adopted to infer keypoints for a specific person. Compared with bounding box detection, CID is differentiable and robust to detection errors. Decoupling persons into different feature maps allows to isolate distractions from other persons, and explore context cues at scales larger than the bounding box size. Experiments show that CID outperforms previous multi-person pose estimation pipelines on crowded scenes pose estimation benchmarks in both accuracy and efficiency. For instance, it achieves 71.3% AP on CrowdPose, outperforming the recent single-stage DEKR by 5.6%, the bottom-up CenterAttention by 3.7%, and the top-down JC-SPPE by 5.3%. This advantage sustains on the commonly used COCO benchmark††Code is available at https://github.com/kennethwdk/CID. Dongkai Wang, Shiliang Zhang |
CVPR | 2 |
| 2022 | Speaker Overlap-aware Neural Diarization for Multi-party Meeting AnalysisabstractRecently, hybrid systems of clustering and neural diarization models have been successfully applied in multi-party meeting analysis.However, current models always treat overlapped speaker diarization as a multi-label classification problem, where speaker dependency and overlaps are not well considered.To overcome the disadvantages, we reformulate overlapped speaker diarization task as a single-label prediction problem via the proposed power set encoding (PSE).Through this formulation, speaker dependency and overlaps can be explicitly modeled.To fully leverage this formulation, we further propose the speaker overlap-aware neural diarization (SOND) model, which consists of a contextindependent (CI) scorer to model global speaker discriminability, a context-dependent scorer (CD) to model local discriminability, and a speaker combining network (SCN) to combine and reassign speaker activities.Experimental results show that using the proposed formulation can outperform the state-ofthe-art methods based on target speaker voice activity detection, and the performance can be further improved with SOND, resulting in a 6.30% relative diarization error reduction. Zhihao Du, Shiliang Zhang, Zhijie Yan |
EMNLP | 2 |
| 2022 | Prosospeech: Enhancing Prosody with Quantized Vector Pre-Training in Text-To-SpeechabstractExpressive text-to-speech (TTS) has become a hot research topic recently, mainly focusing on modeling prosody in speech. Prosody modeling has several challenges: 1) the extracted pitch used in previous prosody modeling works have inevitable errors, which hurts the prosody modeling; 2) different attributes of prosody (e.g., pitch, duration and energy) are dependent on each other and produce the natural prosody together; and 3) due to high variability of prosody and the limited amount of high-quality data for TTS training, the distribution of prosody cannot be fully shaped. To tackle these issues, we propose ProsoSpeech, which enhances the prosody using quantized latent vectors pre-trained on large-scale unpaired and low-quality text and speech data. Specifically, we first introduce a word-level prosody encoder, which quantizes the low-frequency band of the speech and compresses prosody at-tributes in the latent prosody vector (LPV). Then we introduce an LPV predictor, which predicts LPV given word sequence. We pre-train the LPV predictor on large-scale text and low-quality speech data and fine-tune it on the high-quality TTS dataset. Finally, our model can generate expressive speech conditioned on the predicted LPV. Experimental results show that ProsoSpeech can generate speech with richer prosody compared with baseline methods. Yi Ren 0006, Zhiying Huang, Shiliang Zhang, Qian Chen 0003, Zhijie Yan, Zhou Zhao 0001 |
ICASSP | 4 |
| 2022 | M2Met: The Icassp 2022 Multi-Channel Multi-Party Meeting Transcription ChallengeabstractRecent development of speech signal processing, such as speech recognition, speaker diarization, etc., has inspired numerous applications of speech technologies. The meeting scenario is one of the most valuable and, at the same time, most challenging scenarios for the deployment of speech technologies. Speaker diarization and multi-speaker automatic speech recognition in meeting scenarios have attracted much attention recently. However, the lack of large public meeting data has been a major obstacle for advancement of the field. Therefore, we make available the AliMeeting corpus, which consists of 120 hours of recorded Mandarin meeting data, including far-field data collected by 8-channel microphone array as well as near-field data collected by headset microphone. Each meeting session is composed of 2-4 speakers with different speaker overlap ratio, recorded in meeting rooms with different size. Along with the dataset, we launch the ICASSP 2022 Multi-channel Multi-party Meeting Transcription Challenge (M2MeT) with two tracks, namely speaker diarization and multi-speaker ASR, aiming to provide a common testbed for meeting rich transcription and promote reproducible research in this field. In this paper we provide a detailed introduction of the AliMeeting dateset, challenge rules, evaluation methods and baseline systems. Fan Yu 0002, Shiliang Zhang, Yihui Fu, Lei Xie 0001, Zhihao Du, Weilong Huang, Zhijie Yan, Bin Ma 0001, Hui Bu |
ICASSP | 2 |
| 2022 | Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand ChallengeabstractThe ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic speech recognition (ASR) (track 2). Along with the challenge, we released 120 hours of real-recorded Mandarin meeting speech data with manual annotation, including far-field data collected by 8-channel micro-phone array as well as near-field data collected by each participants’ headset microphone. We briefly describe the released dataset, track setups, baselines and summarize the challenge results and major techniques used in the submissions. Fan Yu 0002, Shiliang Zhang, Yihui Fu, Zhihao Du, Weilong Huang, Lei Xie 0001, Zheng-Hua Tan, DeLiang Wang, Yanmin Qian, Kong-Aik Lee, Zhijie Yan, Bin Ma 0001, Hui Bu |
ICASSP | 2 |
| 2022 | Modeling The Detection Capability Of High-Speed Spiking CamerasabstractThe novel working principle enables spiking cameras to capture high-speed moving objects. However, the applications of spiking cameras can be affected by many factors, such as brightness intensity, detectable distance, and the maximum speed of moving targets. Improper settings such as weak ambient brightness and too short object-camera distance, will lead to failure in the application of such cameras. To address the issue, this paper proposes a modeling algorithm that studies the detection capability of spiking cameras. The algorithm deduces the maximum detectable speed of spiking cameras corresponding to different scenario settings (e.g., brightness intensity, camera lens, and object-camera distance) based on the basic technical parameters of cameras (e.g., pixel size, spatial and temporal resolution). Thereby, the proper camera settings for various applications can be determined. Extensive experiments verify the effectiveness of the modeling algorithm. To our best knowledge, it is the first work to investigate the detection capability of spiking cameras. Junwei Zhao 0003, Zhaofei Yu, Lei Ma 0008, Ziluo Ding, Shiliang Zhang, Yonghong Tian 0001, Tiejun Huang 0001 |
ICASSP | 5 |
| 2022 | Transformer-Based Domain Adaptation for Event Data ClassificationabstractEvent cameras encode the change of brightness into events, differing from conventional frame cameras. The novel working principle makes them to have stronger potential in high-speed applications. However, the lack of labeled event annotations limits the applications of such cameras in deep learning frameworks, making it appealing to study more efficient deep learning algorithms and architectures. This paper devises the Convolutional Transformer Network (CTN) for processing event data. The CTN enjoys the advantages of convolution networks and transformers, presenting stronger capability in event-based classification tasks compared with existing models. To address the insufficiency issue of annotated event data, we propose to train the CTN via the source-free Unsupervised Domain Adaptation (UDA) algorithm leveraging large-scale labeled image data. Extensive experiments verify the effectiveness of the UDA algorithm. And our CTN outperforms recent state-of-the-art methods on event-based classification tasks, suggesting that it is an effective model for this task. To our best acknowledge, it is an early attempt of employing vision transformers with the source-free UDA algorithm to process event data. Junwei Zhao 0003, Shiliang Zhang, Tiejun Huang 0001 |
ICASSP | 2 |
| 2022 | Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech RecognitionabstractTransformers have recently dominated the ASR field. Although able to yield good performance, they involve an autoregressive (AR) decoder to generate tokens one by one, which is computationally inefficient. To speed up inference, non-autoregressive (NAR) methods, e.g. single-step NAR, were designed, to enable parallel generation. However, due to an independence assumption within the output tokens, performance of single-step NAR is inferior to that of AR models, especially with a large-scale corpus. There are two challenges to improving single-step NAR: Firstly to accurately predict the number of output tokens and extract hidden variables; secondly, to enhance modeling of interdependence between output tokens. To tackle both challenges, we propose a fast and accurate parallel transformer, termed Paraformer. This utilizes a continuous integrate-and-fire based predictor to predict the number of tokens and generate hidden variables. A glancing language model sampler then generates semantic embeddings to enhance the NAR decoder's ability to model context interdependence. Finally, we design a strategy to generate negative samples for minimum word error rate training to further improve performance. Experiments using the AISHELL-1, AISHELL-2 benchmark, and an industrial-level 20,000 hour task demonstrate that the proposed Paraformer can attain comparable performance to the state-of-the-art AR transformer, with over 10x speedup. Zhifu Gao, Shiliang Zhang, Ian McLoughlin 0001, Zhijie Yan |
INTERSPEECH | 2 |
| 2022 | A Comparative Study on Speaker-attributed Automatic Speech Recognition in Multi-party MeetingsabstractIn this paper, we conduct a comparative study on speaker-attributed automatic speech recognition (SA-ASR) in the multi-party meeting scenario, a topic with increasing attention in meeting rich transcription. Specifically, three approaches are evaluated in this study. The first approach, FD-SOT, consists of a frame-level diarization model to identify speakers and a multi-talker ASR to recognize utterances. The speaker-attributed transcriptions are obtained by aligning the diarization results and recognized hypotheses. However, such an alignment strategy may suffer from erroneous timestamps due to the modular independence, severely hindering the model performance. Therefore, we propose the second approach, WD-SOT, to address alignment errors by introducing a word-level diarization model, which can get rid of such timestamp alignment dependency. To further mitigate the alignment issues, we propose the third approach, TS-ASR, which trains a target-speaker separation module and an ASR module jointly. By comparing various strategies for each SA-ASR approach, experimental results on a real meeting scenario corpus, AliMeeting, reveal that the WD-SOT approach achieves 10.7% relative reduction on averaged speaker-dependent character error rate (SD-CER), compared with the FD-SOT approach. In addition, the TS-ASR approach also outperforms the FD-SOT approach and brings 16.5% relative average SD-CER reduction. Fan Yu 0002, Zhihao Du, Shiliang Zhang, Yuxiao Lin, Lei Xie 0001 |
INTERSPEECH | 3 |
| 2022 | SpikingSIM: A Bio-Inspired Spiking SimulatorabstractLarge-scale neuromorphic dataset is costly to construct and difficult to annotate because of the unique high-speed asynchronous imaging principle of bio-inspired cameras. Lacking of large-scale annotated neuromorphic datasets has significantly hindered the applications of bio-inspired cameras in deep neural networks. Synthesizing neuromorphic data from annotated RGB images can be considered to alleviate this challenge. This paper proposes a simulator to generate simulated spiking data from images recorded by frame cameras. To minimize the deviations between synthetic data and real data, the proposed simulator named SpikingSIM considers the sensing principle of spiking cameras, and generates high-quality simulated spiking data, e.g., the noises in real data are also simulated. Experimental results show that, our simulator generates more realistic spiking data than existing methods. We hence train deep neural networks with synthesized spiking data. Experiments show that, the net- work trained by our simulated data generalizes well on real spiking data. The source code of SpikingSIM is available at http://github.com/Evin-X/SpikingSIM. Junwei Zhao 0003, Shiliang Zhang, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001 |
ISCAS | 2 |
| 2022 | Asymmetric Label Propagation for Video Object SegmentationabstractSemi-supervised video object segmentation aims to segment foreground objects across a video sequence based on their masks given at the first frame. The motion in adjacent frames tends to be smooth, yet object appearances could change substantially in subsequent frames due to clutters or occlusions. Most existing works segment a video frame by equally referring to segmentation masks of its previous frame and the first frame, and are prone to unreliable matching and accumulated segmentation errors. In order to alleviate this issue, this paper proposes to treat the first and previous frames differently to leverage the motion and appearance clues reliably, and presents an Asymmetric Label Propagation (ALP) method. ALP consists of a Confidence-guided Local Propagation (CLP) module and a Global Label Matching (GLM) module, respectively. CLP propagates labels from the previous frame to the current frame based on local affinity and appearance matching uncertainty. To further recover potential missing objects and alleviate error accumulation, GLM matches the current frame to both the foreground and background of the first frame, and adaptively fuses their matching results. The CLP and GLM outputs are fused to generate object-specific feature maps to perform multi-object segmentation. Extensive experiments on DAVIS and Youtube-VOS datasets demonstrate the effectiveness of the proposed method. Zhen Chen 0020, Ming Yang 0007, Shiliang Zhang |
MMAsia | 3 |
| 2022 | MFCCA:Multi-Frame Cross-Channel Attention for Multi-Speaker ASR in Multi-Party Meeting ScenarioabstractRecently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel attention focuses on either learning global correlations between sequences of different channels or exploiting fine-grained channel-wise information effectively at each time step. Considering the delay of microphone array receiving sound, we propose a multi-frame cross-channel attention, which models cross-channel information between adjacent frames to exploit the complementarity of both frame-wise and channel-wise knowledge. Besides, we also propose a multi-layer convolutional mechanism to fuse the multi -channel output and a channel masking strategy to combat the channel number mismatch problem between training and inference. Experiments on the AliMeeting, a real-world corpus, reveal that our proposed model outperforms single-channel model by 31.7% and 37.0% CER reduction on Eval and Test sets. Moreover, with comparable model parameters and training data, our proposed model achieves a new SOTA performance on the AliMeeting corpus, as compared with the top ranking systems in the ICASSP2022 M2MeT challenge, a recently held multi-channel multi-speaker ASR challenge. Fan Yu 0002, Shiliang Zhang, Yuhao Liang, Zhihao Du, Yuxiao Lin, Lei Xie 0001 |
SLT | 2 |
| 2022 | Unsupervised Person Re-Identification via Multi-Label Classification
Dongkai Wang, Shiliang Zhang |
Int. J. Comput. Vis. | 2 |
| 2022 | BDCN: Bi-Directional Cascade Network for Perceptual Edge DetectionabstractExploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a bi-directional cascade network (BDCN) architecture, where an individual layer is supervised by labeled edges at its specific scale, rather than directly applying the same supervision to different layers. Furthermore, to enrich multi-scale representations learned by each layer of BDCN, we introduce a scale enhancement module (SEM), which utilizes dilated convolution to generate multi-scale features, instead of using deeper CNNs. These new approaches encourage the learning of multi-scale representations in different layers and detect edges that are well delineated by their scales. Learning scale dedicated layers also results in a compact network with a fraction of parameters. We evaluate our method on three datasets, i.e., BSDS500, NYUDv2, and Multicue, and achieve ODS F-measure of 0.832, 2.7 percent higher than current state-of-the-art on the BSDS500 dataset. We also applied our edge detection result to other vision tasks. Experimental results show that, our method further boosts the performance of image segmentation, optical flow estimation, and object proposal generation. Shiliang Zhang, Ming Yang 0007, Yanhu Shan, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Pose-Guided Representation Learning for Person Re-IdentificationabstractThe large pose variations and misalignment errors exhibited by person images significantly increase the difficulty of person Re-Identification (ReID). Existing works commonly apply extra operations like pose estimation, part segmentation, etc., to alleviate those issues and improve the robustness of pedestrian representations. While boosting the ReID accuracy, those operations introduce considerable computational overheads and make the deep models complex and hard to tune. To chase a more efficient solution, we propose a Part-Guided Representation (PGR) composed of Pose Invariant Feature (PIF) and Local Descriptive Feature (LDF), respectively. We call PGR "Part-Guided" because it is trained and supervised by local part cues. Specifically, PIF approximates a pose invariant representation inferred by pose estimation and pose normalization. LDF focuses on discriminative body parts by approximating a representation learned with body region segmentation. In this way, extra pose extraction is only introduced during the training stage to supervise the learning of PGR, but is not required during the testing stage for feature extraction. Extensive comparisons with recent works on five widely used datasets demonstrate the competitive accuracy and efficiency of PGR. Jianing Li 0001, Shiliang Zhang, Qi Tian 0001, Meng Wang 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Who is closer: A computational method for domain gap evaluation
Shiliang Zhang |
Pattern Recognit. | 2 |
| 2022 | Large-Scale Spatio-Temporal Person Re-Identification: Algorithms and BenchmarkabstractPerson re-identification (re-ID) in the scenario with large spatial and temporal spans has not been fully explored. This fact partially occurs because existing benchmark datasets were mainly collected with limited spatial and temporal ranges,e.g.,using videos recorded in a few days by cameras in a specific region of the campus. Such limited spatial and temporal ranges make it hard to simulate the difficulties of person re-ID in real scenarios. In this work, we contribute a novel Large-scale Spatio-Temporal (LaST) person re-ID dataset, including 10,862 identities with more than 228k images. Compared with existing datasets, LaST presents more challenging and high-diversity re-ID settings and significantly larger spatial and temporal ranges. For instance, each person can appear in different cities or countries, and in various time slots from day to evening, and in different seasons from spring to winter. To our best knowledge, LaST is a novel person re-ID dataset with the largest spatio-temporal ranges. Based on LaST, we verified its challenge by conducting a comprehensive performance evaluation of 14 re-ID algorithms. We further propose an easy-to-implement baseline that works well in such challenging re-ID settings. We also verified that models pre-trained on LaST can generalize well on existing datasets with short-term and cloth-changing scenarios. We expect LaST to inspire future works toward more realistic and challenging re-ID tasks. More information about the dataset is available athttps://github.com/shuxjweb/last.git. Xiujun Shu, Xiao Wang 0014, Xianghao Zang, Shiliang Zhang, Yuanqi Chen, Ge Li 0002, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Bidirectional Posture-Appearance Interaction Network for Driver Behavior RecognitionabstractDriver behavior recognition has become one of the most important tasks for intelligent vehicles. This task, however, is very challenging since the background contents in real-world driving scenarios are often very complex. More critically, the difference between driving behaviors is often very minor, making it extremely difficult to distinguish them. Existing methods often rely only on RGB frames (or skeleton data), which may fail to capture the minor differences between behaviors and appearance information of objects simultaneously and thus fail to achieve promising performance. To address the above issues, in this paper, we propose a bidirectional posture-appearance interaction network (BPAI-Net), which simultaneously considers RGB frames and skeleton (i.e., posture) data for driver behavior recognition. Specifically, we propose a posture-guided convolutional neural network (PG-CNN) and an appearance-guided graph convolutional network (AG-GCN) to extract appearance and posture features, respectively. To exploit the complementary information between appearance and posture, we use the appearance features from PG-CNN for guiding AG-GCN to exploit the contextual information (e.g., nearby objects) to enhance posture features. Then, we use the enhanced posture features from AG-GCN to help PG-CNN focus on critical local areas of video frames that are related to driver behaviors. In this sense, we are able to use the interaction between two modalities to extract more discriminative features and thus improve the recognition accuracy. Experimental results on Drive&Act dataset show that our method outperforms state-of-the-art methods by a large margin (67.83% vs. 63.64%). Furthermore, we collect a bus driver behavior recognition dataset and yield consistent performance gain against baseline methods, demonstrating the effectiveness of our method in real-world applications. The source code and trained models are available at github.com/SCUT-AILab/BPAI-Net/. Mingkui Tan, Gengqin Ni, Xu Liu 0022, Shiliang Zhang, Xiangmiao Wu, Yaowei Wang 0001, Runhao Zeng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | Distillation-Guided Residual Learning for Binary Convolutional Neural NetworksabstractIt is challenging to bridge the performance gap between binary convolutional neural network (BCNN) and floating-point CNN (FCNN). This performance gap is mainly caused by the inferior modeling capability and training strategy of BCNN, which leads to substantial residuals in intermediate feature maps between BCNN and FCNN. To minimize the performance gap, we enforce BCNN to produce similar intermediate feature maps with the ones of FCNN. This intuition leads to a more effective training strategy for BCNN, i.e., optimizing each binary convolutional block with blockwise distillation loss derived from FCNN. The goal of minimizing the residuals in intermediate feature maps also motivates us to update the binary convolutional block architecture to facilitate the optimization of blockwise distillation loss. Specifically, a lightweight shortcut branch is inserted into each binary convolutional block to complement residuals at each block. Benefited from its squeeze-and-interaction (SI) structure, this shortcut branch introduces a fraction of parameters, e.g., less than 10% overheads, but effectively boosts the modeling capability of binary convolution blocks in BCNN. Extensive experiments on ImageNet demonstrate the superior performance of our method in both classification efficiency and accuracy, e.g., BCNN trained with our methods achieves the accuracy of 60.45% on ImageNet, better than many state-of-the-art ones. Jianming Ye, Jingdong Wang 0001, Shiliang Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identificationabstractintroduction Share on Introduction to the Special Issue on Fine-Grained Visual Recognition and Re-Identification Authors: Shiliang Zhang Peking University Peking UniversityView Profile , Guorong Li University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Weigang Zhang Harbin Institute of Technology Harbin Institute of TechnologyView Profile , Qingming Huang University of Chinese Academy of Sciences University of Chinese Academy of SciencesView Profile , Tiejun Huang Peking University Peking UniversityView Profile , Mubarak Shah University of Central Florida University of Central FloridaView Profile , Nicu Sebe University of Trento University of TrentoView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 18Issue 1sFebruary 2022 Article No.: 24pp 1–3https://doi.org/10.1145/3505280Online:25 January 2022Publication History 0citation169DownloadsMetricsTotal Citations0Total Downloads169Last 12 Months169Last 6 weeks22 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Shiliang Zhang, Guorong Li, Weigang Zhang, Qingming Huang, Tiejun Huang 0001, Mubarak Shah, Nicu Sebe |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2021 | Intra-Inter Camera Similarity for Unsupervised Person Re-IdentificationabstractMost of unsupervised person Re-Identification (Re-ID) works produce pseudo-labels by measuring the feature similarity without considering the distribution discrepancy among cameras, leading to degraded accuracy in label computation across cameras. This paper targets to address this challenge by studying a novel intra-inter camera similarity for pseudo-label generation. We decompose the sample similarity computation into two stage, i.e., the intra-camera and inter-camera computations, respectively. The intra-camera computation directly leverages the CNN features for similarity computation within each camera. Pseudo-labels generated on different cameras train the re-id model in a multi-branch network. The second stage considers the classification scores of each sample on different cameras as a new feature vector. This new feature effectively alleviates the distribution discrepancy among cameras and generates more reliable pseudo-labels. We hence train our re-id model in two stages with intra-camera and inter-camera pseudo-labels, respectively. This simple intra-inter camera similarity produces surprisingly good performance on multiple datasets, e.g., achieves rank-1 accuracy of 89.5% on the Market1501 dataset, outperforming the recent unsupervised works by 9+%, and is comparable with the latest transfer learning works that leverage extra annotations. Shiyu Xuan, Shiliang Zhang |
CVPR | 2 |
| 2021 | Graph Consistency Based Mean-Teaching for Unsupervised Domain Adaptive Person Re-IdentificationabstractRecent works show that mean-teaching is an effective framework for unsupervised domain adaptive person re-identification. However, existing methods perform contrastive learning on selected samples between teacher and student networks, which is sensitive to noises in pseudo labels and neglects the relationship among most samples. Moreover, these methods are not effective in cooperation of different teacher networks. To handle these issues, this paper proposes a Graph Consistency based Mean-Teaching (GCMT) method with constructing the Graph Consistency Constraint (GCC) between teacher and student networks. Specifically, given unlabeled training images, we apply teacher networks to extract corresponding features and further construct a teacher graph for each teacher network to describe the similarity relationships among training images. To boost the representation learning, different teacher graphs are fused to provide the supervise signal for optimizing student networks. GCMT fuses similarity relationships predicted by different teacher networks as supervision and effectively optimizes student networks with more sample relationships involved. Experiments on three datasets, i.e., Market-1501, DukeMTMCreID, and MSMT17, show that proposed GCMT outperforms state-of-the-art methods by clear margin. Specially, GCMT even outperforms the previous method that uses a deeper backbone. Experimental results also show that GCMT can effectively boost the performance with multiple teacher and student networks. Our code is available at https://github.com/liu-xb/GCMT . Shiliang Zhang |
IJCAI | 2 |
| 2021 | Extremely Low Footprint End-to-End ASR System for Smart DeviceabstractRecently, end-to-end (E2E) speech recognition has become popular, since it can integrate the acoustic, pronunciation and language models into a single neural network, which outperforms conventional models. Among E2E approaches, attention-based models, e.g. Transformer, have emerged as being superior. Such models have opened the door to deployment of ASR on smart devices, however they still suffer from requiring a large number of model parameters. We propose an extremely low footprint E2E ASR system for smart devices, to achieve the goal of satisfying resource constraints without sacrificing recognition accuracy. We design cross-layer weight sharing to improve parameter efficiency and further exploit model compression methods including sparsification and quantization, to reduce memory storage and boost decoding efficiency. We evaluate our approaches on the public AISHELL-1 and AISHELL-2 benchmarks. On the AISHELL-2 task, the proposed method achieves more than 10× compression (model size reduces from 248 to 24MB), at the cost of only minor performance loss (CER reduces from 6.49% to 6.92%). Zhifu Gao, Yiwu Yao, Shiliang Zhang, Ian McLoughlin 0001 |
Interspeech | 3 |
| 2021 | Investigation of Spatial-Acoustic Features for Overlapping Speech Detection in Multiparty Meetings
Shiliang Zhang, Weilong Huang, Hongbin Suo, Jinwei Feng, Zhijie Yan |
Interspeech | 1 |
| 2021 | An Energy Consumption Model for Electrical Vehicle Networks via Extended Federated-learningabstractElectrical vehicle (EV) raises to promote an eco-sustainable society. Nevertheless, the “range anxiety” of EV hinders its wider acceptance among customers. This paper proposes a novel solution to range anxiety based on a federated-learning model, which is capable of estimating battery consumption and providing energy-efficient route planning for vehicle networks. Specifically, the new approach extends the federated-learning structure with two components: anomaly detection and sharing policy. The first component identifies preventing factors in model learning, while the second component offers guidelines for information sharing amongst vehicle networks when the sharing is necessary to preserve learning efficiency. The two components collaborate to enhance learning robustness against data heterogeneities in networks. Numerical experiments are conducted, and the results show that compared with considered solutions, the proposed approach could provide higher accuracy of battery-consumption estimation for vehicles under heterogeneous data distributions, without increasing the time complexity or transmitting raw data among vehicle networks. Shiliang Zhang |
IV | 1 |
| 2021 | Hybrid Network Compression via Meta-LearningabstractNeural network pruning and quantization are two major lines of network compression. This raises a natural question that whether we can find the optimal compression by considering multiple network compression criteria in a unified framework. This paper incorporates two criteria and seeks layer-wise compression by leveraging the meta-learning framework. A regularization loss is applied to unify the constraint of input and output channel numbers, bit-width of network activations and weights, so that the compressed network can satisfy a given Bit-OPerations counts (BOPs) constraint. We further propose an iterative compression constraint for optimizing the compression procedure, which effectively achieves a high compression rate and maintains the original network performance. Extensive experiments on various networks and vision tasks show that the proposed method yields better performance and compression rates than recent methods. For instance, our method achieves better image classification accuracy and compactness than the recent DJPQ. It achieves similar performance with the recent DHP in image super-resolution, meanwhile saves about 50% computation. Jianming Ye, Shiliang Zhang, Jingdong Wang 0001 |
ACM Multimedia | 2 |
| 2021 | Robust Pose Estimation in Crowded Scenes with Direct Pose-Level InferenceabstractMulti-person pose estimation in crowded scenes is challenging because overlapping and occlusions make it difficult to detect person bounding boxes and infer pose cues from individual keypoints. To address those issues, this paper proposes a direct pose-level inference strategy that is free of bounding box detection and keypoint grouping. Instead of inferring individual keypoints, the Pose-level Inference Network (PINet) directly infers the complete pose cues for a person from his/her visible body parts. PINet first applies the Part-based Pose Generation (PPG) to infer multiple coarse poses for each person from his/her body parts. Those coarse poses are refined by the Pose Refinement module through incorporating pose priors, and finally are fused in the Pose Fusion module. PINet relies on discriminative body parts to differentiate overlapped persons, and applies visual body cues to infer the global pose cues. Experiments on several crowded scenes pose estimation benchmarks demonstrate the superiority of PINet. For instance, it achieves 59.8% AP on the OCHuman dataset, outperforming the recent works by a large margin. Dongkai Wang, Shiliang Zhang, Gang Hua 0001 |
NeurIPS | 2 |
| 2021 | Simplified Self-Attention for Transformer-Based end-to-end Speech RecognitionabstractTransformer models have been introduced into end-to-end speech recognition with state-of-the-art performance on various tasks owing to their superiority in modeling long-term dependencies. However, such improvements are usually obtained through the use of very large neural networks. Transformer models mainly include two submodules - position-wise feedforward layers and self-attention (SAN) layers. In this paper, to reduce the model complexity while maintaining good performance, we propose a simplified self-attention (SSAN) layer which employs FSMN memory blocks instead of projection layers to form query and key vectors for transformer-based end-to-end speech recognition. We evaluate the SSAN-based and the conventional SAN-based transformers on the public AISHELL-1, internal 1000-hour and 20,000-hour large-scale Mandarin tasks. Results show that our proposed SSAN-based transformer model can achieve over 20% reduction in model parameters and 6.7% relative CER reduction on the AISHELL-1 task. With impressively 20% parameter reduction, our model shows no loss of recognition performance on the 20,000-hour large-scale task. Haoneng Luo, Shiliang Zhang, Lei Xie 0001 |
SLT | 2 |
| 2021 | Viewpoint and Scale Consistency Reinforcement for UAV Vehicle Re-Identification
Shangzhi Teng, Shiliang Zhang, Qingming Huang, Nicu Sebe |
Int. J. Comput. Vis. | 2 |
| 2021 | Diverse part attentive network for video-based person re-identification
Xiujun Shu, Ge Li 0002, Longhui Wei, Jia-Xing Zhong, Xianghao Zang, Shiliang Zhang, Yaowei Wang 0001, Yongsheng Liang 0001, Qi Tian 0001 |
Pattern Recognit. Lett. | 6 |
| 2021 | Multi-View Spatial Attention Embedding for Vehicle Re-IdentificationabstractVehicle Re-Identification (Re-ID) is a challenging vision task mainly because the appearance of a vehicle varies dramatically under different viewpoints. Moreover, different vehicles with the same model and color commonly show similar appearance, thus are hard to be distinguished. To alleviate negative effects of viewpoint variance, we design a multi-view branch network where each branch learns a viewpoint-specific feature without parameter sharing. Being able to focus on a limited range of viewpoints, this viewpoint-specific feature performs substantially better than the general feature learned by an uniform network. To further differentiate visually similar vehicles, we strengthen the discriminative power on their subtle local differences by introducing a spatial attention model into each feature learning branch. The multi-view feature learning and spatial attention learning compose our neural network architecture, which is trained end to end with the softmax loss and triplet loss, respectively. We evaluate our methods on two large vehicle Re-ID datasets, i.e., VehicleID and VeRi-776, respectively. Extensive experiments show that our methods achieve promising performance. For example, we achieve mAP accuracy of 76.78% and 72.53% on VehicleID and VeRi-776 dataset respectively, substantially better than current state-of-the art. Shangzhi Teng, Shiliang Zhang, Qingming Huang, Nicu Sebe |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Progressive Feature Enhancement for Person Re-IdentificationabstractMost of person Re-Identification (ReID) works extract features from the top CNN layer for person image matching. The top CNN layer commonly corresponds to large receptive fields, thus is not effective in depicting visual cues at multiple scales, e.g., both global appearance and local details. This work proposes a Progressive Feature Enhancement (PFE) algorithm to spot and fuse multi-scale discriminative cues from different CNN layers into a single feature vector. The basic idea is to progressively learn complementary features with a layer-specific supervision from deep to shallow layers. The layer-specific supervision is inferred by the proposed Masked Feature Augmentation (MFA) module. For each CNN layer, MFA indicates cues that have been captured in its deeper layers. MFA hence supervises each layer to depict additional visual cues missed by its deeper layers. This framework effectively learns multi-scale features without requiring extra part annotations or dividing body parts. To further facilitate the layer-specific feature generation, a Two-Stage Attention Module (TSAM) is proposed to filter pixel-wise and channel-wise noises on intermediate feature maps. Extensive experiments on four ReID datasets show that our approach achieves competitive performance, e.g., with ResNet50 backbone, it achieves rank1 accuracy of 95.1%, 88.2%, 79.1% and 71.6% on Market-1501, DukeMTMC-ReID, MSMT17 and CUHK03 Detected, respectively, outperforming many state-of-the-art works. Yingji Zhong, Yaowei Wang 0001, Shiliang Zhang |
IEEE Trans. Image Process. | 3 |
| 2020 | Unsupervised Person Re-Identification via Multi-Label ClassificationabstractThe challenge of unsupervised person re-identification (ReID) lies in learning discriminative features without true labels. This paper formulates unsupervised person ReID as a multi-label classification task to progressively seek true labels. Our method starts by assigning each person image with a single-class label, then evolves to multi-label classification by leveraging the updated ReID model for label prediction. The label prediction comprises similarity computation and cycle consistency to ensure the quality of predicted labels. To boost the ReID model training efficiency in multi-label classification, we further propose the memory-based multi-label classification loss (MMCL). MMCL works with memory-based non-parametric classifier and integrates multi-label classification and single-label classification in an unified framework. Our label prediction and MMCL work iteratively and substantially boost the ReID performance. Experiments on several large-scale person ReID datasets demonstrate the superiority of our method in unsupervised person ReID. Our method also allows to use labeled person images in other domains. Under this transfer learning setting, our method also achieves state-of-the-art performance. Dongkai Wang, Shiliang Zhang |
CVPR | 2 |
| 2020 | Robust Partial Matching for Person Search in the WildabstractVarious factors like occlusions, backgrounds, etc., would lead to misaligned detected bounding boxes , e.g., ones covering only portions of human body. This issue is common but overlooked by previous person search works. To alleviate this issue, this paper proposes an Align-to-Part Network (APNet) for person detection and re-Identification (reID). APNet refines detected bounding boxes to cover the estimated holistic body regions, from which discriminative part features can be extracted and aligned. Aligned part features naturally formulate reID as a partial feature matching procedure, where valid part features are selected for similarity computation, while part features on occluded or noisy regions are discarded. This design enhances the robustness of person search to real-world challenges with marginal computation overhead. This paper also contributes a Large-Scale dataset for Person Search in the wild (LSPS), which is by far the largest and the most challenging dataset for person search. Experiments show that APNet brings considerable performance improvement on LSPS. Meanwhile, it achieves competitive performance on existing person search benchmarks like CUHK-SYSU and PRW. Yingji Zhong, Xiaoyu Wang 0002, Shiliang Zhang |
CVPR | 3 |
| 2020 | Joint Visual and Temporal Consistency for Unsupervised Domain Adaptive Person Re-identification
Jianing Li 0001, Shiliang Zhang |
ECCV (24) | 2 |
| 2020 | Pan: Phoneme-Aware Network for Monaural Speech EnhancementabstractCurrent methods for monaural speech enhancement only utilize acoustic information but seldom consider the phonetic information of an utterance. In the voice conversion community, significant progress has been achieved by using the phonetic information via the phonetic posteriorgrams (PPGs). Inspired by the progress, we propose a phoneme-aware network (PAN) to utilize the noisy PPGs for speech enhancement. Since the PPG prediction and speech enhancement benefit from each other, a PPG predictor is involved into the PAN and an iterative training algorithm is proposed for PAN. Experimental results show that the enhancement performance is improved by using the phonetic information in terms of speech intelligibility, perceptual quality and character error rate. To the best of our knowledge, this is the first time to introduce the PPG into speech enhancement. Zhihao Du, Jiqing Han 0001, Shiliang Zhang |
ICASSP | 4 |
| 2020 | Self-Supervised Adversarial Multi-Task Learning for Vocoder-Based Monaural Speech Enhancement
Zhihao Du, Jiqing Han 0001, Shiliang Zhang |
INTERSPEECH | 4 |
| 2020 | Neural Zero-Inflated Quality Estimation Model for Automatic Speech Recognition SystemabstractThe performances of automatic speech recognition (ASR) systems are usually evaluated by the metric word error rate (WER) when the manually transcribed data are provided, which are, however, expensively available in the real scenario.In addition, the empirical distribution of WER for most ASR systems usually tends to put a significant mass near zero, making it difficult to simulate with a single continuous distribution.In order to address the two issues of ASR quality estimation (QE), we propose a novel neural zero-inflated model to predict the WER of the ASR result without transcripts.We design a neural zeroinflated beta regression on top of a bidirectional transformer language model conditional on speech features (speech-BERT).We adopt the pre-training strategy of token level masked language modeling for speech-BERT as well, and further fine-tune with our zero-inflated layer for the mixture of discrete and continuous outputs.The experimental results show that our approach achieves better performance on WER prediction compared with strong baselines. Kai Fan 0002, Bo Li 0121, Jiayi Wang 0010, Shiliang Zhang, Boxing Chen, Niyu Ge, Zhijie Yan |
INTERSPEECH | 4 |
| 2020 | SAN-M: Memory Equipped Self-Attention for End-to-End Speech RecognitionabstractEnd-to-end speech recognition has become popular in recent years, since it can integrate the acoustic, pronunciation and language models into a single neural network. Among end-to-end approaches, attention-based methods have emerged as being superior. For example, Transformer, which adopts an encoder-decoder architecture. The key improvement introduced by Transformer is the utilization of self-attention instead of recurrent mechanisms, enabling both encoder and decoder to capture long-range dependencies with lower computational complexity. In this work, we propose boosting the self-attention ability with a DFSMN memory block, forming the proposed memory equipped self-attention (SAN-M) mechanism. Theoretical and empirical comparisons have been made to demonstrate the relevancy and complementarity between self-attention and the DFSMN memory block. Furthermore, the proposed SAN-M provides an efficient mechanism to integrate these two modules. We have evaluated our approach on the public AISHELL-1 benchmark and an industrial-level 20,000-hour Mandarin speech recognition task. On both tasks, SAN-M systems achieved much better performance than the self-attention based Transformer baseline system. Specially, it can achieve a CER of 6.46% on the AISHELL-1 task even without using any external LM, comfortably outperforming other state-of-the-art systems. Zhifu Gao, Shiliang Zhang, Ian McLoughlin 0001 |
INTERSPEECH | 2 |
| 2020 | Streaming Chunk-Aware Multihead Attention for Online End-to-End Speech RecognitionabstractRecently, streaming end-to-end automatic speech recognition (E2E-ASR) has gained more and more attention.Many efforts have been paid to turn the non-streaming attention-based E2E-ASR system into streaming architecture.In this work, we propose a novel online E2E-ASR system by using Streaming Chunk-Aware Multihead Attention (SCAMA) and a latency control memory equipped self-attention network (LC-SAN-M).LC-SAN-M uses chunk-level input to control the latency of encoder.As to SCAMA, a jointly trained predictor is used to control the output of encoder when feeding to decoder, which enables decoder to generate output in streaming manner.Experimental results on the open 170-hour AISHELL-1 and an industrial-level 20000-hour Mandarin speech recognition tasks show that our approach can significantly outperform the MoChA-based baseline system under comparable setup.On the AISHELL-1 task, our proposed method achieves a character error rate (CER) of 7.39%, to the best of our knowledge, which is the best published performance for online ASR. Shiliang Zhang, Zhifu Gao, Haoneng Luo, Zhijie Yan, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2020 | Domain Adaptive Person Re-Identification via Coupling OptimizationabstractDomain adaptive person Re-Identification (ReID) is challenging owing to the domain gap and shortage of annotations on target scenarios. To handle those two challenges, this paper proposes a coupling optimization method including the Domain-Invariant Mapping (DIM) method and the Global-Local distance Optimization (GLO), respectively. Different from previous methods that transfer knowledge in two stages, the DIM achieves a more efficient one-stage knowledge transfer by mapping images in labeled and unlabeled datasets to a shared feature space. GLO is designed to train the ReID model with unsupervised setting on the target domain. Instead of relying on existing optimization strategies designed for supervised training, GLO involves more images in distance optimization, and achieves better robustness to noisy label prediction. GLO also integrates distance optimizations in both the global dataset and local training batch, thus exhibits better training efficiency. Extensive experiments on three large-scale datasets,i.e., Market-1501, DukeMTMC-reID, andMSMT17, show that our coupling optimization outperforms state-of-the-art methods by a large margin. Our method also works well in unsupervised training, and even outperforms several recent domain adaptive methods. Shiliang Zhang |
ACM Multimedia | 2 |
| 2020 | E2BoWs: An end-to-end Bag-of-Words model via deep convolutional neural network for image retrieval
Shiliang Zhang, Tiejun Huang 0001, Qi Tian 0001 |
Neurocomputing | 2 |
| 2020 | CDbin: Compact Discriminative Binary Descriptor Learned With Efficient Neural NetworkabstractAs an important computer vision task, image matching requires efficient and discriminative local descriptors. Most of the existing descriptors like SIFT and ORB are hand-crafted; therefore it is necessary to study more optimized descriptors through end-to-end learning. This paper proposes the compact binary descriptors learned with a lightweight Convolutional Neural Network (CNN), which is efficient for training and testing. Specifically, we propose a CNN with no larger than five layers for descriptor learning. The resulting descriptors, i.e., Compact Discriminative binary descriptors (CDbin) are optimized with four complementary loss functions, i.e., 1) triplet loss to ensure the discriminative power; 2) quantization loss to decrease the quantization error; 3) correlation loss to ensure the feature compactness; and 4) even-distribution loss to enrich the embedded information. The extensive experiments on two image patch datasets and three image retrieval datasets show that the CDbin exhibits competitive performance compared with the existing descriptors. For example, the 64-bit CDbin substantially outperforms the 256-bit ORB and 1024-bit SIFT on Hpatches dataset. Although generated by a shallow CNN, CDbin also outperforms several recent deep descriptors. Jianming Ye, Shiliang Zhang, Tiejun Huang 0001, Yong Rui |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Multi-Scale Temporal Cues Learning for Video Person Re-IdentificationabstractTemporal cues embedded in videos provide important clues for person Re-Identification (ReID). To efficiently exploit temporal cues with a compact neural network, this work proposes a novel 3D convolution layer called Multi-scale 3D (M3D) convolution layer. The M3D layer is easy to implement and could be inserted into traditional 2D convolution networks to learn multi-scale temporal cues by end-to-end training. According to its inserted location, the M3D layer has two variants, i.e., local M3D layer and global M3D layer, respectively. The local M3D layer is inserted between 2D convolution layers to learn spatial-temporal cues among adjacent 2D feature maps. The global M3D layer is computed on adjacent frame feature vectors to learn their global temporal relations. The local and global M3D layers hence learn complementary temporal cues. Their combination introduces a fraction of parameters to traditional 2D CNN, but leads to the strong multi-scale temporal feature learning capability. The learned temporal feature is fused with a spatial feature to compose the final spatial-temporal representation for video person ReID. Evaluations on four widely used video person ReID datasets, i.e., MARS, DukeMTMC-VideoReID, PRID2011, and iLIDS-VID demonstrate the substantial advantages of our method over the state-of-the art. For example, it achieves rank1 accuracy of 88.63% on MARS without re-ranking. Our method also achieves a reasonable trade-off between ReID accuracy and model size, e.g., it saves about 40% parameters of I3D CNN. Jianing Li 0001, Shiliang Zhang, Tiejun Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Group-Group Loss-Based Global-Regional Feature Learning for Vehicle Re-IdentificationabstractVehicle Re-Identification (Re-ID) is challenging because vehicles of the same model commonly show similar appearance. We tackle this challenge by proposing a Global-Regional Feature (GRF) that depicts extra local details to enhance discrimination power in addition to the global context. It is motivated by the observation that, vehicles of same color, maker, and model can be distinguished by their regional difference, e.g., the decorations on the windshields. To accelerate the GRF learning and promote its discrimination power, we propose a Group-Group Loss (GGL) to optimize the distance within and across vehicle image groups. Different from the siamese or triplet loss, GGL is directly computed on image groups rather than individual sample pairs or triplets. By avoiding traversing numerous sample combinations, GGL makes the model training easier and more efficient. Those two contributions highlight this work from previous methods on vehicle Re-ID task, which commonly learn global features with triplet loss or its variants. We evaluate our methods on two large-scale vehicle Re-ID datasets, i.e., VeRi and VehicleID. Experimental results show our methods achieve promising performance in comparison with recent works. Shiliang Zhang, Xiaoyu Wang 0002, Richang Hong, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | Multi-Scale 3D Convolution Network for Video Based Person Re-IdentificationabstractThis paper proposes a two-stream convolution network to extract spatial and temporal cues for video based person ReIdentification (ReID). A temporal stream in this network is constructed by inserting several Multi-scale 3D (M3D) convolution layers into a 2D CNN network. The resulting M3D convolution network introduces a fraction of parameters into the 2D CNN, but gains the ability of multi-scale temporal feature learning. With this compact architecture, M3D convolution network is also more efficient and easier to optimize than existing 3D convolution networks. The temporal stream further involves Residual Attention Layers (RAL) to refine the temporal features. By jointly learning spatial-temporal attention masks in a residual manner, RAL identifies the discriminative spatial regions and temporal cues. The other stream in our network is implemented with a 2D CNN for spatial feature extraction. The spatial and temporal features from two streams are finally fused for the video based person ReID. Evaluations on three widely used benchmarks datasets, i.e.,MARS, PRID2011, and iLIDS-VID demonstrate the substantial advantages of our method over existing 3D convolution networks and state-of-art methods. Jianing Li 0001, Shiliang Zhang, Tiejun Huang 0001 |
AAAI | 2 |
| 2019 | Bi-Directional Cascade Network for Perceptual Edge DetectionabstractExploiting multi-scale representations is critical to improve edge detection for objects at different scales. To extract edges at dramatically different scales, we propose a Bi-Directional Cascade Network (BDCN) structure, where an individual layer is supervised by labeled edges at its specific scale, rather than directly applying the same supervision to all CNN outputs. Furthermore, to enrich multi-scale representations learned by BDCN, we introduce a Scale Enhancement Module (SEM) which utilizes dilated convolution to generate multi-scale features, instead of using deeper CNNs or explicitly fusing multi-scale edge maps. These new approaches encourage the learning of multi-scale representations in different layers and detect edges that are well delineated by their scales. Learning scale dedicated layers also results in compact network with a fraction of parameters. We evaluate our method on three datasets, i.e., BSDS500, NYUDv2, and Multicue, and achieve ODS Fmeasure of 0.828, 1.3% higher than current state-of-the art on BSDS500. Shiliang Zhang, Ming Yang 0007, Yanhu Shan, Tiejun Huang 0001 |
CVPR | 2 |
| 2019 | Investigation of Modeling Units for Mandarin Speech Recognition Using Dfsmn-ctc-smbrabstractThe choice of acoustic modeling units is critical to acoustic modeling in large vocabulary continuous speech recognition (LVCSR) tasks. The recent connectionist temporal classification (CTC) based acoustic models have more options for the choice of modeling units. In this work, we propose a DFSMN-CTC-sMBR acoustic model and investigate various modeling units for Mandarin speech recognition. In addition to the commonly used context-independent Initial/Finals (CI-IF), context-dependent Initial/Finals (CD-IF) and Syllable, we also propose a hybrid Character-Syllable modeling units by mixing high frequency Chinese characters and syllables. Experimental results show that DFSMN-CTC-sMBR models with all these types of modeling units can significantly outperform the well-trained conventional hybrid models. Moreover, we find that the proposed hybrid Character-Syllable modeling units is the best choice for CTC based acoustic modeling for Mandarin speech recognition in our work since it can dramatically reduce substitution errors in recognition results. In a 20,000 hours Mandarin speech recognition task, the DFSMN-CTC-sMBR system with hybrid Character-Syllable achieves a character error rate (CER) of 7.45% while performance of the well-trained DFSMN-CE-sMBR system is 9.49%. Shiliang Zhang |
ICASSP | 1 |
| 2019 | Robust Audio-visual Speech Recognition Using Bimodal Dfsmn with Multi-condition Training and Dropout RegularizationabstractAudio-visual speech recognition (AVSR) is thought to be one of the potential solutions for robust speech recognition, especially in noisy environments. Compared to audio only speech recognition, the major issues of AVSR include the lack of publicly available audio-visual corpora and the need of robust knowledge fusion of both speech and vision. In this work, based on the recently released NTCD-TIMIT audio-visual corpus, we address the challenges of AVSR through three aspects: 1) optimal integration of acoustic and visual information; 2) robust performance with multi-condition training; 3) robust modeling against missing visual information during decoding. We propose a bimodal-DFSMN to jointly learn feature fusion and acoustic modeling, and utilize a per-frame dropout approach to enhance the robustness of AVSR system against the missing of visual modality. In the experiments, we construct two setups based on the NTCD-TIMIT corpus that consists of 5 hours clean training data and 150 hours multi-condition training data, respectively. As a result, we achieve a phone error rate of 12.6% on clean test set and an average phone error rate of 26.2% on all test sets (clean, various SNRs, various noise types), which both dramatically improve the baseline performance in NTCD-TIMIT task. Shiliang Zhang, Bin Ma 0001, Lei Xie 0001 |
ICASSP | 1 |
| 2019 | Global-Local Temporal Representations for Video Person Re-IdentificationabstractThis paper proposes the Global-Local Temporal Representation (GLTR) to exploit the multi-scale temporal cues in video sequences for video person Re-Identification (ReID). GLTR is constructed by first modeling the short-term temporal cues among adjacent frames, then capturing the long-term relations among inconsecutive frames. Specifically, the short-term temporal cues are modeled by parallel dilated convolutions with different temporal dilation rates to represent the motion and appearance of pedestrian. The long-term relations are captured by a temporal self-attention model to alleviate the occlusions and noises in video sequences. The short and long-term temporal cues are aggregated as the final GLTR by a simple single-stream CNN. GLTR shows substantial superiority to existing features learned with body part cues or metric learning on four widely-used video ReID datasets. For instance, it achieves Rank-1 Accuracy of 87.02% on MARS dataset without re-ranking, better than current state-of-the art. Jianing Li 0001, Shiliang Zhang, Jingdong Wang 0001, Wen Gao 0001, Qi Tian 0001 |
ICCV | 2 |
| 2019 | Resolution-invariant Person Re-IdentificationabstractExploiting resolution invariant representation is critical for person Re-Identification (ReID) in real applications, where the resolutions of captured person images may vary dramatically. This paper learns person representations robust to resolution variance through jointly training a Foreground-Focus Super-Resolution (FFSR) module and a Resolution-Invariant Feature Extractor (RIFE) by end-to-end CNN learning. FFSR upscales the person foreground using a fully convolutional auto-encoder with skip connections learned with a foreground focus training loss. RIFE adopts two feature extraction streams weighted by a dual-attention block to learn features for low and high resolution images, respectively. These two complementary modules are jointly trained, leading to a strong resolution invariant representation. We evaluate our methods on five datasets containing person images at a large range of resolutions, where our methods show substantial superiority to existing solutions. For instance, we achieve Rank-1 accuracy of 36.4% and 73.3% on CAVIAR and MLR-CUHK03, outperforming the state-of-the art by 2.9% and 2.6%, respectively. Shunan Mao, Shiliang Zhang, Ming Yang 0007 |
IJCAI | 2 |
| 2019 | Audio Tagging with Compact Feedforward Sequential Memory Network and Audio-to-Audio Ratio Based Data Augmentation
Zhiying Huang, Shiliang Zhang |
INTERSPEECH | 2 |
| 2019 | Towards Language-Universal Mandarin-English Speech Recognition
Shiliang Zhang, Bin Ma 0001, Lei Xie 0001 |
INTERSPEECH | 1 |
| 2019 | Investigation of Transformer Based Spelling Correction Model for CTC-Based End-to-End Mandarin Speech Recognition
Shiliang Zhang, Zhijie Yan |
INTERSPEECH | 1 |
| 2019 | EAGER: Edge-Aided imaGe undERstanding SystemabstractImage understanding is a fundamental task for many multimedia and computer vision applications, such as self-driving, multimedia retrieval, and augmented reality, etc. In this paper, we demonstrate that edge detection could aid image understanding tasks such as semantic segmentation, optical flow estimation, and object proposal generation. Based on our recent research efforts on edge detection, we develop a robust and efficient Edge-Aided imaGe undERstanding system named as EAGER. EAGER is built on a compact and efficient edge detection module, which is constructed with a bi-directional cascade network, multi-scale feature enhancement, and layer-specific training supervision, respectively. Based on detected edges, EAGER achieves accurate semantic segment, optical flow estimation, as well as object bounding-box proposal generation for user-uploaded images and videos. Shiliang Zhang |
ICMR | 3 |
| 2019 | DR2-Net: Deep Residual Reconstruction Network for image compressive sensing
Hantao Yao, Shiliang Zhang, Yongdong Zhang 0001, Qi Tian 0001, Changsheng Xu |
Neurocomputing | 3 |
| 2019 | Deep Representation Learning With Part Loss for Person Re-IdentificationabstractLearning discriminative representations for unseen person images is critical for person Re-Identification (ReID). Most of current approaches learn deep representations in classification tasks, which essentially minimizes the empirical classification risk on the training set. As shown in our experiments, such representations easily get over-fitted on a discriminative human body part on the training set. To gain the discriminative power on unseen person images, we propose a deep representation learning procedure named Part Loss Network (PL-Net), to minimize both the empirical classification risk on training person images and the representation learning risk on unseen person images. The representation learning risk is evaluated by the proposed part loss, which automatically detects human body parts, and computes the person classification loss on each part separately. Compared with traditional global classification loss, simultaneously considering part loss enforces the deep network to learn representations for different body parts and gain the discriminative power on unseen persons. Experimental results on three person ReID datasets, i.e., Market1501, CUHK03, VIPeR, show that our representation outperforms existing deep representations. Hantao Yao, Shiliang Zhang, Richang Hong, Yongdong Zhang 0001, Changsheng Xu, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | GLAD: Global-Local-Alignment Descriptor for Scalable Person Re-IdentificationabstractThe huge variance of human pose and the misalign-ment of detected human images significantly increase the difficulty of pedestrian image matching in person Re-Identification (Re-ID). Moreover, the massive visual data being produced by surveillance video cameras requires highly efficient person Re-ID systems. Targeting to solve the first problem, this work proposes a robust and discriminative pedestrian image descriptor, namely, the Global-Local-Alignment Descriptor (GLAD). For the second problem, this work treats person Re-ID as image retrieval and proposes an efficient indexing and retrieval framework. GLAD explicitly leverages the local and global cues in the human body to generate a discriminative and robust representation. It consists of part extraction and descriptor learning modules, where several part regions are first detected and then deep neural networks are designed for representation learning on both the local and global regions. A hierarchical indexing and retrieval framework is designed to perform offline relevance mining to eliminate the huge person ID redundancy in the gallery set, and accelerate the online Re-ID procedure. Extensive experimental results on widely used public benchmark datasets show GLAD achieves competitive accuracy compared to the state-of-the-art methods. On a large-scale person, with the Re-ID dataset containing more than 520 K images, our retrieval framework significantly accelerates the online Re-ID procedure while also improving Re-ID accuracy. Therefore, this work has the potential to work better on person Re-ID tasks in real scenarios. Longhui Wei, Shiliang Zhang, Hantao Yao, Wen Gao 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Person Transfer GAN to Bridge Domain Gap for Person Re-IdentificationabstractAlthough the performance of person Re-Identification (ReID) has been significantly boosted, many challenging issues in real scenarios have not been fully investigated, e.g., the complex scenes and lighting variations, viewpoint and pose changes, and the large number of identities in a camera network. To facilitate the research towards conquering those issues, this paper contributes a new dataset called MSMT171 with many important features, e.g., 1) the raw videos are taken by an 15-camera network deployed in both indoor and outdoor scenes, 2) the videos cover a long period of time and present complex lighting variations, and 3) it contains currently the largest number of annotated identities, i.e., 4,101 identities and 126,441 bounding boxes. We also observe that, domain gap commonly exists between datasets, which essentially causes severe performance drop when training and testing on different datasets. This results in that available training data cannot be effectively leveraged for new testing domains. To relieve the expensive costs of annotating new training samples, we propose a Person Transfer Generative Adversarial Network (PTGAN) to bridge the domain gap. Comprehensive experiments show that the domain gap could be substantially narrowed-down by the PTGAN. Longhui Wei, Shiliang Zhang, Wen Gao 0001, Qi Tian 0001 |
CVPR | 2 |
| 2018 | Deep Feed-Forward Sequential Memory Networks for Speech SynthesisabstractThe Bidirectional LSTM (BLSTM) RNN based speech synthesis system is among the best parametric Text-to-Speech (TTS) systems in terms of the naturalness of generated speech, especially the naturalness in prosody. However, the model complexity and inference cost of BLSTM prevents its usage in many runtime applications. Meanwhile, Deep Feed-forward Sequential Memory Networks (DFSMN) has shown its consistent out-performance over BLSTM in both word error rate (WER) and the runtime computation cost in speech recognition tasks. Since speech synthesis also requires to model long-term dependencies compared to speech recognition, in this paper, we investigate the Deep-FSMN (DFSMN) in speech synthesis. Both objective and subjective experiments show that, compared with BLSTM TTS method, the DFSMN system can generate synthesized speech with comparable speech quality while drastically reduce model complexity and speech generation time. Mengxiao Bi, Shiliang Zhang, Zhijie Yan |
ICASSP | 3 |
| 2018 | Deep-FSMN for Large Vocabulary Continuous Speech RecognitionabstractIn this paper, we present an improved feedforward sequential memory networks (FSMN) architecture, namely Deep-FSMN (DFSMN), by introducing skip connections between memory blocks in adjacent layers. These skip connections enable the information flow across different layers and thus alleviate the gradient vanishing problem when building very deep structure. As a result, DFSMN significantly benefits from these skip connections and deep structure. We have compared the performance of DFSMN to BLSTM both with and without lower frame rate (LFR) on several large speech recognition tasks, including English and Mandarin. Experimental results shown that DFSMN can consistently outperform BLSTM with dramatic gain, especially trained with LFR using CD-Phone as modeling units. In the 20000 hours Fisher (FSH) task, the proposed DFSMN can achieve a word error rate of 9.4% by purely using the cross-entropy criterion and decoding with a 3-gram language model, which achieves a 1.5% absolute improvement compared to the BLSTM. In a 20000 hours Mandarin recognition task, the LFR trained DFSMN can achieve more than 20% relative improvement compared to the LFR trained BLSTM. Moreover, we can easily design the lookahead filter order of the memory blocks in DFSMN to control the latency for real-time applications. Shiliang Zhang, Zhijie Yan, Li-Rong Dai 0001 |
ICASSP | 1 |
| 2018 | RAM: A Region-Aware Deep Model for Vehicle Re-IdentificationabstractPrevious works on vehicle Re-ID mainly focus on extracting global features and learning distance metrics. Because some vehicles commonly share same model and maker, it is hard to distinguish them based on their global appearances. Compared with the global appearance, local regions such as decorations and inspection stickers attached to the windshield, may be more distinctive for vehicle Re-ID. To embed the detailed visual cues in those local regions, we propose a Region-Aware deep Model (RAM). Specifically, in addition to extracting global features, RAM also extracts features from a series of local regions. As each local region conveys more distinctive visual cues, RAM encourages the deep model to learn discriminative features. We also introduce a novel learning algorithm to jointly use vehicle IDs, types/models, and colors to train the RAM. This strategy fuses more cues for training and results in more discriminative global and regional features. We evaluate our methods on two large-scale vehicle Re-ID datasets, i.e., VeRi and VehicleID. Experimental results show our methods achieve promising performance in comparison with recent works. Shiliang Zhang, Qingming Huang, Wen Gao 0001 |
ICME | 2 |
| 2018 | Compact Feedforward Sequential Memory Networks for Small-footprint Keyword Spotting
Mengzhe Chen, Shiliang Zhang, Haitao Yao |
INTERSPEECH | 2 |
| 2018 | Acoustic Modeling with DFSMN-CTC and Joint CTC-CE Learning
Shiliang Zhang |
INTERSPEECH | 1 |
| 2018 | VP-ReID: Vehicle and Person Re-Identification SystemabstractWith the capability of locating and tracking specific suspects or vehicles in a large camera network, person Re-Identification (ReID) and vehicle ReID show potential to be a key technology in smart surveillance system. They have been drawing lots of attentions from both academia and industry. To demonstrate our recent research progresses on those two tasks, we develop a robust and efficient person and video ReID system named as VP-ReID. This system is build based on our recent works including Deep Convolutional Neural Network design for discriminative feature extraction, efficient off-line indexing, as well as distance metric optimization for deep feature learning. Constructed upon those algorithms, VP-ReID identifies query vehicle and person efficiently and accurately from a large gallery set. Longhui Wei, Jianing Li 0001, Shiliang Zhang |
ICMR | 4 |
| 2018 | Multi-Task Learning with Low Rank Attribute Embedding for Multi-Camera Person Re-IdentificationabstractWe propose Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) to address the problem of person re-identification on multi-cameras. Re-identifications on different cameras are considered as related tasks, which allows the shared information among different tasks to be explored to improve the re-identification accuracy. The MTL-LORAE framework integrates low-level features with mid-level attributes as the descriptions for persons. To improve the accuracy of such description, we introduce the low-rank attribute embedding, which maps original binary attributes into a continuous space utilizing the correlative relationship between each pair of attributes. In this way, inaccurate attributes are rectified and missing attributes are recovered. The resulting objective function is constructed with an attribute embedding error and a quadratic loss concerning class labels. It is solved by an alternating optimization strategy. The proposed MTL-LORAE is tested on four datasets and is validated to outperform the existing methods with significant margins. Chi Su, Fan Yang 0016, Shiliang Zhang, Qi Tian 0001, Larry Davis 0001, Wen Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Multi-type attributes driven multi-camera person re-identification
Chi Su, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001 |
Pattern Recognit. | 2 |
| 2018 | Learning Affective Features With a Hybrid Deep Model for Audio-Visual Emotion RecognitionabstractEmotion recognition is challenging due to the emotional gap between emotions and audio-visual features. Motivated by the powerful feature learning ability of deep neural networks, this paper proposes to bridge the emotional gap by using a hybrid deep model, which first produces audio-visual segment features with Convolutional Neural Networks (CNNs) and 3D-CNN, then fuses audio-visual segment features in a Deep Belief Networks (DBNs). The proposed method is trained in two stages. First, CNN and 3D-CNN models pre-trained on corresponding large-scale image and video classification tasks are fine-tuned on emotion recognition tasks to learn audio and visual segment features, respectively. Second, the outputs of CNN and 3D-CNN models are combined into a fusion network built with a DBN model. The fusion network is trained to jointly learn a discriminative audio-visual segment feature representation. After average-pooling segment features learned by DBN to form a fixed-length global video feature, a linear Support Vector Machine is used for video emotion classification. Experimental results on three public audio-visual emotional databases, including the acted RML database, the acted eNTERFACE05 database, and the spontaneous BAUM-1s database, demonstrate the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues with CNN, 3D-CNN, and DBN for audio-visual emotion recognition. Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Interacting Tracklets for Multi-Object TrackingabstractIn this paper, we propose to exploit the interactions between non-associable tracklets to facilitate multi-object tracking. We introduce two types of tracklet interactions, close interaction and distant interaction. The close interaction imposes physical constraints between two temporally overlapping tracklets and more importantly, allows us to learn local classifiers to distinguish targets that are close to each other in the spatiotemporal domain. The distant interaction, on the other hand, accounts for the higher-order motion and appearance consistency between two temporally isolated tracklets. Our approach is modeled as a binary labeling problem and solved using the efficient Quadratic Pseudo-Boolean Optimization (QPBO). It yields promising tracking performance on the challenging PETS09 and MOT16 dataset. Our code will be made publicly available upon the acceptance of the manuscript. Long Lan, Xinchao Wang, Shiliang Zhang, Dacheng Tao, Wen Gao 0001, Thomas S. Huang |
IEEE Trans. Image Process. | 3 |
| 2018 | AutoBD: Automated Bi-Level Description for Scalable Fine-Grained Visual CategorizationabstractCompared with traditional image classification, fine-grained visual categorization is a more challenging task, because it targets to classify objects belonging to the same species, e.g., classify hundreds of birds or cars. In the past several years, researchers have made many achievements on this topic. However, most of them are heavily dependent on the artificial annotations, e.g., bounding boxes, part annotations, and so on. The requirement of artificial annotations largely hinders the scalability and application. Motivated to release such dependence, this paper proposes a robust and discriminative visual description named Automated Bi-level Description (AutoBD). “Bi-level” denotes two complementary part-level and object-level visual descriptions, respectively. AutoBD is “automated,” because it only requires the image-level labels of training images and does not need any annotations for testing images. Compared with the part annotations labeled by the human, the image-level labels can be easily acquired, which thus makes AutoBD suitable for large-scale visual categorization. Specifically, the part-level description is extracted by identifying the local region saliently representing the visual distinctiveness. The object-level description is extracted from object bounding boxes generated with a co-localization algorithm. Although only using the image-level labels, AutoBD outperforms the recent studies on two public benchmark, i.e., classification accuracy achieves 81.6% on CUB-200-2011 and 88.9% on Car-196, respectively. On the large-scale Birdsnap data set, AutoBD achieves the accuracy of 68%, which is currently the best performance to the best of our knowledge. Hantao Yao, Shiliang Zhang, Chenggang Yan 0001, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Speech Emotion Recognition Using Deep Convolutional Neural Network and Discriminant Temporal Pyramid MatchingabstractSpeech emotion recognition is challenging because of the affective gap between the subjective emotions and low-level features. Integrating multilevel feature learning and model training, deep convolutional neural networks (DCNN) has exhibited remarkable success in bridging the semantic gap in visual tasks like image classification, object detection. This paper explores how to utilize a DCNN to bridge the affective gap in speech signals. To this end, we first extract three channels of log Mel-spectrograms (static, delta, and delta delta) similar to the red, green, blue (RGB) image representation as the DCNN input. Then, the AlexNet DCNN model pretrained on the large ImageNet dataset is employed to learn high-level feature representations on each segment divided from an utterance. The learned segment-level features are aggregated by a discriminant temporal pyramid matching (DTPM) strategy. DTPM combines temporal pyramid matching and optimal Lp-norm pooling to form a global utterance-level feature representation, followed by the linear support vector machines for emotion classification. Experimental results on four public datasets, that is, EMO-DB, RML, eNTERFACE05, and BAUM-1s, show the promising performance of our DCNN model and the DTPM strategy. Another interesting finding is that the DCNN model pretrained for image applications performs reasonably good in affective speech feature extraction. Further fine tuning on the target emotional speech datasets substantially promotes recognition performance. Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Sequential Outlier Criterion for Sparsification of Online Adaptive FilteringabstractIn this paper, we deal with the learning problem when using an adaptive filtering method. For the learning system in filtering, the knowledge is obtained and updated based on the newly acquired information that is extracted and learned from the sequential samples over time. Effective measurement on the informativeness of a sample and reasonable subsequent treatment on the sample will improve the learning performance. This paper proposes a sequential outlier criterion for sparsification of online adaptive filtering. The method is proposed to achieve effective informativeness measurement of online filtering to obtain a more accurate and more compact network in the learning process. In the proposed method, the measurement on the samples' informativeness is established based on the historical sequentially adjacent samples, and then the informative-measured samples are treated individually by the learning system based on whether the sample is informative, redundant, or abnormal. With our method, a more sensible learning process can be achieved with valid knowledge extracted, and the optimal network in the learning system can be obtained. Simulations based on static function estimation, Mackey-Glass time series prediction, and Lorenz chaotic time series prediction demonstrate that the proposed method can provide more effective classification on samples and more accurate networks in online adaptive filtering. Shiliang Zhang, Hui Cao 0003, Xiali Hei 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | Pose-Driven Deep Convolutional Model for Person Re-identificationabstractFeature extraction and matching are two crucial components in person Re-Identification (ReID). The large pose deformations and the complex view variations exhibited by the captured person images significantly increase the difficulty of learning and matching of the features from person images. To overcome these difficulties, in this work we propose a Pose-driven Deep Convolutional (PDC) model to learn improved feature extraction and matching models from end to end. Our deep architecture explicitly leverages the human part cues to alleviate the pose variations and learn robust feature representations from both the global image and different local parts. To match the features from global human body and local body parts, a pose driven feature weighting sub-network is further designed to learn adaptive feature fusions. Extensive experimental analyses and results on three popular datasets demonstrate significant performance improvements of our model over all published state-of-the-art methods. Chi Su, Jianing Li 0001, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001 |
ICCV | 3 |
| 2017 | Large-scale person re-identification as retrievalabstractThis paper targets to bring together the research efforts on two fields that are growing actively in the past few years: multicamera person Re-Identification (ReID) and large-scale image retrieval. We demonstrate that the essentials of image retrieval and person ReID are the same, i.e., measuring the similarity between images. However, person ReID requires more discriminative and robust features to identify the subtle differences of different persons and overcome the large variance among images of the same person. Specifically, we propose a coarse-to-fine (C2F) framework and a Convolutional Neural Network structure named as Conv-Net to tackle the large-scale person ReID as an image retrieval task. Given a query person image, the C2F firstly employ Conv-Net to extract a compact descriptor and perform the coarse-level search. A robust descriptor conveying more spatial cues is hence extracted to perform the fine-level search. Extensive experimental results show that the proposed method outperforms existing methods on two public datasets. Further, the evaluation on a large-scale Person-520K dataset demonstrates that our work is significantly more efficient than existing works, e.g., only needs 180ms to identify a query person from 520K images. Hantao Yao, Shiliang Zhang, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Yu Wang 0089, Qi Tian 0001 |
ICME | 2 |
| 2017 | Gaussian Prediction Based Attention for Online End-to-End Speech Recognition
Junfeng Hou, Shiliang Zhang, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2017 | GLAD: Global-Local-Alignment Descriptor for Pedestrian RetrievalabstractThe huge variance of human pose and the misalignment of detected human images significantly increase the difficulty of person Re-Identification (Re-ID). Moreover, efficient Re-ID systems are required to cope with the massive visual data being produced by video surveillance systems. Targeting to solve these problems, this work proposes a Global-Local-Alignment Descriptor (GLAD) and an efficient indexing and retrieval framework, respectively. GLAD explicitly leverages the local and global cues in human body to generate a discriminative and robust representation. It consists of part extraction and descriptor learning modules, where several part regions are first detected and then deep neural networks are designed for representation learning on both the local and global regions. A hierarchical indexing and retrieval framework is designed to eliminate the huge redundancy in the gallery set, and accelerate the online Re-ID procedure. Extensive experimental results show GLAD achieves competitive accuracy compared to the state-of-the-art methods. Our retrieval framework significantly accelerates the online Re-ID procedure without loss of accuracy. Therefore, this work has potential to work better on person Re-ID tasks in real scenarios. Longhui Wei, Shiliang Zhang, Hantao Yao, Wen Gao 0001, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2017 | One-Shot Fine-Grained Instance RetrievalabstractFine-Grained Visual Categorization (FGVC) has achieved significant progress recently. However, the number of fine-grained species could be huge and dynamically increasing in real scenarios, making it difficult to recognize unseen objects under the current FGVC framework. This raises an open issue to perform large-scale fine-grained identification without a complete training set. Aiming to conquer this issue, we propose a retrieval task named One-Shot Fine-Grained Instance Retrieval (OSFGIR). "One-Shot" denotes the ability of identifying unseen objects through a fine-grained retrieval task assisted with an incomplete auxiliary training set. This paper first presents the detailed description to OSFGIR task and our collected OSFGIR-378K dataset. Next, we propose the Convolutional and Normalization Networks (CN-Nets) learned on the auxiliary dataset to generate a concise and discriminative representation. Finally, we present a coarse-to-fine retrieval framework consisting of three components, i.e., coarse retrieval, fine-grained retrieval, and query expansion, respectively. The framework progressively retrieves images with similar semantics, and performs fine-grained identification. Experiments show our OSFGIR framework achieves significantly better accuracy and efficiency than existing FGVC and image retrieval methods, thus could be a better solution for large-scale fine-grained object identification. Hantao Yao, Shiliang Zhang, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2017 | DSP: Discriminative Spatial Part modeling for Fine-Grained Visual Categorization
Hantao Yao, Dongming Zhang 0004, Jintao Li 0001, Jianshe Zhou, Shiliang Zhang, Yongdong Zhang 0001 |
Image Vis. Comput. | 5 |
| 2017 | Attributes driven tracklet-to-tracklet person re-identification using latent prototypes space mapping
Chi Su, Shiliang Zhang, Fan Yang 0016, Guangxiao Zhang, Qi Tian 0001, Wen Gao 0001, Larry Davis 0001 |
Pattern Recognit. | 2 |
| 2017 | Watch, attend and parse: An end-to-end neural network based approach to handwritten mathematical expression recognition
Jianshu Zhang 0001, Jun Du 0002, Shiliang Zhang, Dan Liu 0008, Yulong Hu, Jin-Shui Hu, Si Wei, Li-Rong Dai 0001 |
Pattern Recognit. | 3 |
| 2017 | Nonrecurrent Neural Structure for Long-Term DependenceabstractIn this paper, we propose a novel neural network structure, namely feedforward sequential memory networks (FSMN), to model long-term dependence in time series without using recurrent feedback. The proposed FSMN is a standard fully connected feedforward neural network equipped with some learnable memory blocks in its hidden layers. The memory blocks use a tapped-delay line structure to encode the long context information into a fixed-size representation as short-term memory mechanism which are somehow similar to the time-delay neural networks layers. We have evaluated the FSMNs in several standard benchmark tasks, including speech recognition and language modeling. Experimental results have shown that FSMNs outperform the conventional recurrent neural networks (RNN) while can be learned much more reliably and faster in modeling sequential signals like speech or language. Moreover, we also propose a compact feedforward sequential memory networks (cFSMN) by combining FSMN with low-rank matrix factorization and make a slight modification to the encoding method used in FSMNs in order to further simplify the network architecture. On the speech recognition Switchboard task, the proposed cFSMN structures can reduce the model size by 60% and speed up the learning by more than seven times while the model can still significantly outperform the popular bidirectional LSTMs for both frame-level cross-entropy criterion-based training and MMI-based sequence training. Shiliang Zhang, Cong Liu 0006, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001, Yu Hu 0003 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2016 | Deep Attributes Driven Multi-camera Person Re-identification
Chi Su, Shiliang Zhang, Junliang Xing, Wen Gao 0001, Qi Tian 0001 |
ECCV (2) | 2 |
| 2016 | Future Context Attention for Unidirectional LSTM Based Acoustic Model
Shiliang Zhang, Si Wei, Li-Rong Dai 0001 |
INTERSPEECH | 2 |
| 2016 | Compact Feedforward Sequential Memory Networks for Large Vocabulary Continuous Speech Recognition
Shiliang Zhang, Hui Jiang 0001, Shifu Xiong, Si Wei, Li-Rong Dai 0001 |
INTERSPEECH | 1 |
| 2016 | Multimodal Deep Convolutional Neural Network for Audio-Visual Emotion RecognitionabstractEmotion recognition is a challenging task because of the emotional gap between subjective emotion and the low-level audio-visual features. Inspired by the recent success of deep learning in bridging the semantic gap, this paper proposes to bridge the emotional gap based on a multimodal Deep Convolution Neural Network (DCNN), which fuses the audio and visual cues in a deep model. This multimodal DCNN is trained with two stages. First, two DCNN models pre-trained on large-scale image data are fine-tuned to perform audio and visual emotion recognition tasks respectively on the corresponding labeled speech and face data. Second, the outputs of these two DCNNs are integrated in a fusion network constructed by a number of fully-connected layers. The fusion network is trained to obtain a joint audio-visual feature representation for emotion recognition. Experimental results on the RML audio-visual database demonstrates the promising performance of the proposed method. To the best of our knowledge, this is an early work fusing audio and visual cues in DCNN for emotion recognition. Its success guarantees further research in this direction. Shiqing Zhang, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001 |
ICMR | 2 |
| 2016 | Hybrid Orthogonal Projection and Estimation (HOPE): A New Framework to Learn Neural NetworksabstractIn this paper, we propose a novel model for high-dimensional data, called the Hybrid Orthogonal Projection and Estimation (HOPE) model, which combines a linear orthogonal projection and a finite mixture model under a unified generative modeling framework. The HOPE model itself can be learned unsupervised from unlabelled data based on the maximum likelihood estimation as well as discriminatively from labelled data. More interestingly, we have shown the proposed HOPE models are closely related to neural networks (NNs) in a sense that each hidden layer can be reformulated as a HOPE model. As a result, the HOPE framework can be used as a novel tool to probe why and how NNs work, more importantly, to learn NNs in either supervised or unsupervised ways. In this work, we have investigated the HOPE framework to learn NNs for several standard tasks, including image recognition on MNIST and speech recognition on TIMIT. Experimental results have shown that the HOPE framework yields significant performance gains over the current state-of-the-art methods in various types of NN learning problems, including unsupervised feature learning, supervised or semi-supervised learning. Shiliang Zhang, Hui Jiang 0001, Li-Rong Dai 0001 |
J. Mach. Learn. Res. | 1 |
| 2016 | Coarse-to-Fine Description for Fine-Grained Visual CategorizationabstractRecent years have witnessed the significant advance in fine-grained visual categorization, which targets to classify the objects belonging to the same species. To capture enough subtle visual differences and build discriminative visual description, most of the existing methods heavily rely on the artificial part annotations, which are expensive to collect in real applications. Motivated to conquer this issue, this paper proposes a multi-level coarse-to-fine object description. This novel description only requires the original image as input, but could automatically generate visual descriptions discriminative enough for fine-grained visual categorization. This description is extracted from five sources representing coarse-to-fine visual clues: 1) original image is used as the source of global visual clue; 2) object bounding boxes are generated using convolutional neural network (CNN); 3) with the generated bounding box, foreground is segmented using the proposed k nearest neighbour-based co-segmentation algorithm; and 4) two types of part segmentations are generated by dividing the foreground with an unsupervised part learning strategy. The final description is generated by feeding these sources into CNN models and concatenating their outputs. Experiments on two public benchmark data sets show the impressive performance of this coarse-to-fine description, i.e., classification accuracy achieves 82.5% on CUB-200-2011, and 86.9% on fine-grained visual categorization-Aircraft, respectively, which outperform many recent works. Hantao Yao, Shiliang Zhang, Yongdong Zhang 0001, Jintao Li 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Multi-Task Learning with Low Rank Attribute Embedding for Person Re-IdentificationabstractWe propose a novel Multi-Task Learning with Low Rank Attribute Embedding (MTL-LORAE) framework for person re-identification. Re-identifications from multiple cameras are regarded as related tasks to exploit shared information to improve re-identification accuracy. Both low level features and semantic/data-driven attributes are utilized. Since attributes are generally correlated, we introduce a low rank attribute embedding into the MTL formulation to embed original binary attributes to a continuous attribute space, where incorrect and incomplete attributes are rectified and recovered to better describe people. The learning objective function consists of a quadratic loss regarding class labels and an attribute embedding error, which is solved by an alternating optimization procedure. Experiments on three person re-identification datasets have demonstrated that MTL-LORAE outperforms existing approaches by a large margin and produces state-of-the-art results. Chi Su, Fan Yang 0016, Shiliang Zhang, Qi Tian 0001, Larry Davis 0001, Wen Gao 0001 |
ICCV | 3 |
| 2015 | Rectified linear neural networks with tied-scalar regularization for LVCSR
Shiliang Zhang, Hui Jiang 0001, Si Wei, Li-Rong Dai 0001 |
INTERSPEECH | 1 |
| 2015 | Augmented Feature Fusion for Image Retrieval SystemabstractThe performance of current image retrieval system is largely determined by the quality and discriminative capability of features. Therefore, using what features and how to effectively combine the power of appropriate features are important in the system. We adopt the reciprocal neighbor based graph fusion approach for feature fusion. More importantly, we explicitly augment the original approach with the following two strategies: 1) we investigate the most suitable feature combinations on various datasets, including the deep learning feature, which has been popular for image retrieval recently; 2) we further improve the robustness of original graph fusion approach by the SVM prediction strategy. Yang Zhou 0017, Dan Zeng 0001, Shiliang Zhang, Qi Tian 0001 |
ICMR | 3 |
| 2015 | Multi-order visual phrase for scalable partial-duplicate visual search
Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Yong Rui |
Multim. Syst. | 1 |
| 2015 | Semantic-Aware Co-Indexing for Image RetrievalabstractIn content-based image retrieval, inverted indexes allow fast access to database images and summarize all knowledge about the database. Indexing multiple clues of image contents allows retrieval algorithms search for relevant images from different perspectives, which is appealing to deliver satisfactory user experiences. However, when incorporating diverse image features during online retrieval, it is challenging to ensure retrieval efficiency and scalability. In this paper, for large-scale image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. Specifically, for an initial set of inverted indexes of local features, we utilize semantic attributes to filter out isolated images and insert semantically similar images to this initial set. Encoding these two distinct and complementary cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online retrieval, because only local features but no semantic attributes are employed for the query. Hence, this co-indexing is different from existing image retrieval methods fusing multiple features or retrieval results. Extensive experiments and comparisons with recent retrieval methods manifest the competitive performance of our method. Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | An Attribute-Assisted Reranking Model for Web Image SearchabstractImage search reranking is an effective approach to refine the text-based image search result. Most existing reranking approaches are based on low-level visual features. In this paper, we propose to exploit semantic attributes for image search reranking. Based on the classifiers for all the predefined attributes, each image is represented by an attribute feature consisting of the responses from these classifiers. A hypergraph is then used to model the relationship between images by integrating low-level visual features and attribute features. Hypergraph ranking is then performed to order the images. Its basic principle is that visually similar images should have similar ranking scores. In this paper, we propose a visual-attribute joint hypergraph learning approach to simultaneously explore two information sources. A hypergraph is constructed to model the relationship of all images. We conduct experiments on more than 1,000 queries in MSRA-MMV2.0 data set. The experimental results demonstrate the effectiveness of our approach. Zhengjun Zha, Meng Wang 0001, Shiliang Zhang, Qi Tian 0001 |
IEEE Trans. Image Process. | 4 |
| 2015 | Cross Indexing With GroupletsabstractMost of the current image indexing systems for retrieval view a database as a set of individual images. It limits the flexibility of the retrieval framework to conduct sophisticated cross-image analysis, resulting in higher memory consumption and sub-optimal retrieval accuracy. To conquer this issue, we propose cross indexing with grouplets, where the core idea is to view the database images as a set of grouplets, each of which is defined as a group of highly relevant images. Because a grouplet groups similar images together, the number of grouplets is smaller than the number of images, thus naturally leading to less memory cost. Moreover, the definition of a grouplet could be based on customized relations, allowing for seamless integration of advanced image features and data mining techniques like the deep convolutional neural network (DCNN) in off-line indexing . To validate the proposed framework, we construct three different types of grouplets , which are respectively based on local similarity , regional relation, and global semantic modeling. Extensive experiments on public benchmark datasets demonstrate the efficiency and superior performance of our approach. Shiliang Zhang, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Hybrid-Indexing Multi-type Features for Large-Scale Image Search
Qingjun Luo, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001 |
ACCV (1) | 2 |
| 2014 | Improving deep neural networks for LVCSR using dropout and shrinking structureabstractRecently, the hybrid deep neural networks and hidden Markov models (DNN/HMMs) have achieved dramatic gains over the conventional GMM/HMMs method on various large vocabulary continuous speech recognition (LVCSR) tasks. In this paper, we propose two new methods to further improve the hybrid DNN/HMMs model: i) use dropout as pre-conditioner (DAP) to initialize DNN prior to back-propagation (BP) for better recognition accuracy; ii) employ a shrinking DNN structure (sDNN) with hidden layers decreasing in size from bottom to top for the purpose of reducing model size and expediting computation time. The proposed DAP method is evaluated in a 70-hour Mandarin transcription (PSC) task and the 309-hour Switchboard (SWB) task. Compared with the traditional greedy layer-wise pre-trained DNN, it can achieve about 10% and 6.8% relative recognition error reduction for PSC and SWB tasks respectively. In addition, we also evaluate sDNN as well as its combination with DAP on the SWB task. Experimental results show that these methods can reduce model size to 45% of original size and accelerate training and test time by 55%, without losing recognition accuracy. Shiliang Zhang, Yebo Bao, Hui Jiang 0001, Li-Rong Dai 0001 |
ICASSP | 1 |
| 2014 | Superimage: Packing Semantic-Relevant Images for Indexing and RetrievalabstractAs an important procedure in image retrieval, off-line indexing focuses on organizing relevant images together and making them easy to access. However, most of existing indexing strategies view database images individually and only consider partial relevance, i.e., either visual or semantic relevance among them. To overcome these issues and design better indexing strategy, we propose to package semantically relevant images into superimages, and then index superimages instead of single images. Superimage effectively packages multiple images into one new unit, hence significantly decreases the number of images to be indexed. This naturally saves the memory cost and retrieval time. To make the final index file discriminative to both visual and semantic relevances, we extract local descriptors from superimages and index them with inverted file. During online retrieval, we only need to extract local descriptors from queries, but could get semantic-aware retrieval results. This is because during our off-line indexing stage, both the semantically and visually relevant images are organized together. Therefore, our approach is superior to many online retrieval fusion algorithms. Experimental results on UKbench, Holidays, and one large-scale dataset all manifest the promising performance of our approach, i.e., competitive precision, better efficiency, and only about 1/2 memory consumption compared with state-of-the-arts. Qingjun Luo, Shiliang Zhang, Tiejun Huang 0001, Wen Gao 0001, Qi Tian 0001 |
ICMR | 2 |
| 2014 | Personalized Visual Vocabulary Adaption for Social Image RetrievalabstractWith the popularity of mobile devices and social networks, users can easily build their personalized image sets. Thus, personalized image analysis, indexing, and retrieval have become important topics in social media analysis. Because of users' diverse preferences, their personalized image sets are usually related to specific topics and show large feature distribution bias from general Internet images. Therefore, the visual vocabulary trained on general Internet images may could not fit across users' personalized image sets very well. To improve the image retrieval performance on personalized image sets, we propose the personalized visual vocabulary adaption which removes non-discriminative visual words and replaces them with more exact and discriminative ones, i.e., adapt a general vocabulary toward a specific user's image set. The proposed algorithm updates the visual vocabulary during off-line feature quantization, and operates on a limited number of visual words, hence shows satisfying efficiency. Extensive experiments of image search on public datasets demonstrate the efficiency and superior performance of our approach. Zhenxing Niu, Shiliang Zhang, Xinbo Gao 0001, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2014 | ObjectPatchNet: Towards scalable and semantic image annotation and retrieval
Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Wen Gao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2014 | Cascade Category-Aware Visual SearchabstractIncorporating image classification into image retrieval system brings many attractive advantages. For instance, the search space can be narrowed down by rejecting images in irrelevant categories of the query. The retrieved images can be more consistent in semantics by indexing and returning images in the relevant categories together. However, due to their different goals on recognition accuracy and retrieval scalability, it is hard to efficiently incorporate most image classification works into large-scale image search. To study this problem, we propose cascade category-aware visual search, which utilizes weak category clue to achieve better retrieval accuracy, efficiency, and memory consumption. To capture the category and visual clues of an image, we first learn category-visual words, which are discriminative and repeatable local features labeled with categories. By identifying category-visual words in database images, we are able to discard noisy local features and extract image visual and category clues, which are hence recorded in a hierarchical index structure. Our retrieval system narrows down the search space by: 1) filtering the noisy local features in query; 2) rejecting irrelevant categories in database; and 3) preforming discriminative visual search in relevant categories. The proposed algorithm is tested on object search, landmark search, and large-scale similar image search on the large-scale LSVRC10 data set. Although the category clue introduced is weak, our algorithm still shows substantial advantages in retrieval accuracy, efficiency, and memory consumption than the state-of-the-art. Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Yong Rui |
IEEE Trans. Image Process. | 1 |
| 2014 | USB: Ultrashort Binary Descriptor for Fast Visual Matching and RetrievalabstractCurrently, many local descriptors have been proposed to tackle a basic issue in computer vision: duplicate visual content matching. These descriptors either are represented as high-dimensional vectors relatively expensive to extract and compare or are binary codes limited in robustness. Bag-of-visual words (BoWs) model compresses local features into a compact representation that allows for fast matching and scalable indexing. However, the codebook training, high-dimensional feature extraction, and quantization significantly degrade the flexibility and efficiency of BoWs model. In this paper, we study an alternative to current local descriptors and BoWs model by extracting the ultrashort binary descriptor (USB) and a compact auxiliary spatial feature from each keypoint detected in images. A typical USB is a 24-bit binary descriptor, hence it directly quantizes visual clues of image keypoints to about 16 million unique IDs. USB allows fast image matching and indexing and avoids the expensive codebook training and feature quantization in BoWs model. The spatial feature complementarily captures the spatial configuration in neighbor region of each keypoint, hence is used to filter mismatched USBs in a cascade verification. In image matching task, USB shows promising accuracy and nearly one-order faster speed than SIFT. We also test USB in retrieval tasks on UKbench, Oxford5K, and 1.2 million distractor images. Comparisons with recent retrieval methods manifest the competitive accuracy, memory consumption, and significantly better efficiency of our approach. Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Yong Rui |
IEEE Trans. Image Process. | 1 |
| 2013 | Semantic-Aware Co-indexing for Image RetrievalabstractInverted indexes in image retrieval not only allow fast access to database images but also summarize all knowledge about the database, so that their discriminative capacity largely determines the retrieval performance. In this paper, for vocabulary tree based image retrieval, we propose a semantic-aware co-indexing algorithm to jointly embed two strong cues into the inverted indexes: 1) local invariant features that are robust to delineate low-level image contents, and 2) semantic attributes from large-scale object recognition that may reveal image semantic meanings. For an initial set of inverted indexes of local features, we utilize 1000 semantic attributes to filter out isolated images and insert semantically similar images to the initial set. Encoding these two distinct cues together effectively enhances the discriminative capability of inverted indexes. Such co-indexing operations are totally off-line and introduce small computation overhead to online query cause only local features but no semantic attributes are used for query. Experiments and comparisons with recent retrieval methods on 3 datasets, i.e., UKbench, Holidays, Oxford5K, and 1.3 million images from Flickr as distractors, manifest the competitive performance of our method. Shiliang Zhang, Ming Yang 0007, Xiaoyu Wang 0002, Yuanqing Lin, Qi Tian 0001 |
ICCV | 1 |
| 2013 | Learning attribute-aware dictionary for image classification and searchabstractBag-of-visual words (BoW) model has recently been well advocated for image classification and search. However, one critical limitation of existing BoW model is the lack of semantic information. To alleviate the impact of this issue, it is imperative to construct semantic-aware visual dictionary. In this paper, we propose a novel approach for learning visual word dictionary embedding intermediate-level semantics. Specifically, we first introduce an Attribute aware Dictionary Learning(AttrDL) scheme to learn multiple sub-dictionaries with specific semantic meanings. We divide training images into different sets and each represents a specific attribute. For each image set, an attribute-aware sub-vocabulary is learned. Hence, these resulting sub-vocabularies are more discriminative for semantics than the traditional vocabularies. Second, to get semantic-aware and discriminative BoW representation with the learned sub-vocabularies, we adopt the idea of L21-norm regularized sparse coding and recode the resulting sparse representation of each image. Experimental results show that the proposed scheme outperforms the state-of-the-art algorithms in both image classification and search tasks. Zhengjun Zha, Huan-Bo Luan, Shiliang Zhang, Qi Tian 0001 |
ICMR | 4 |
| 2013 | Edge-SIFT: Discriminative Binary Descriptor for Scalable Partial-Duplicate Mobile SearchabstractAs the basis of large-scale partial duplicate visual search on mobile devices, image local descriptor is expected to be discriminative, efficient, and compact. Our study shows that the popularly used histogram-based descriptors, such as scale invariant feature transform (SIFT) are not optimal for this task. This is mainly because histogram representation is relatively expensive to compute on mobile platforms and loses significant spatial clues, which are important for improving discriminative power and matching near-duplicate image patches. To address these issues, we propose to extract a novel binary local descriptor named Edge-SIFT from the binary edge maps of scale- and orientation-normalized image patches. By preserving both locations and orientations of edges and compressing the sparse binary edge maps with a boosting strategy, the final Edge-SIFT shows strong discriminative power with compact representation. Furthermore, we propose a fast similarity measurement and an indexing framework with flexible online verification. Hence, the Edge-SIFT allows an accurate and efficient image search and is ideal for computation sensitive scenarios such as a mobile image search. Experiments on a large-scale dataset manifest that the Edge-SIFT shows superior retrieval accuracy to Oriented BRIEF (ORB) and is superior to SIFT in the aspects of retrieval precision, efficiency, compactness, and transmission cost. Shiliang Zhang, Qi Tian 0001, Ke Lu 0002, Qingming Huang, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | ObjectBook construction for large-scale semantic-aware image retrievalabstractAutomatic image annotation assigns semantic labels to images thus presents great potential to achieve semantic-aware image retrieval. However, existing annotation algorithms are not scalable to this emerging need, both in terms of computational efficiency and the number of tags they can deal with. Facilitated by recent development of the large-scale image category recognition data such as ImageNet, we extrapolate from it a model for scalable image annotation and semantic-aware image retrieval, namely ObjectBook. The element in the ObjectBook, which is called an ObjectWord, is defined as a collection of discriminative image patches annotated with the corresponding objects. We take ObjectBook as a high-level semantic preserving visual vocabulary, and hence are able to easily develop efficient image annotation and inverted file indexing strategies for large-scale image collections. The proposed retrieval strategy is compared with state-of-the-art algorithms. Experimental results manifest that the ObjectBook is both discriminative and scalable for large-scale semantic-aware image retrieval. Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001 |
MMSP | 1 |
| 2011 | Modeling spatial and semantic cues for large-scale near-duplicated image retrieval
Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Wengang Zhou 0001, Qingming Huang, Houqiang Li, Wen Gao 0001 |
Comput. Vis. Image Underst. | 1 |
| 2011 | Building descriptive and discriminative visual codebook for large-scale image applications
Qi Tian 0001, Shiliang Zhang, Wengang Zhou 0001, Rongrong Ji, Bingbing Ni, Nicu Sebe |
Multim. Tools Appl. | 2 |
| 2011 | Generating Descriptive Visual Words and Visual Phrases for Large-Scale Image ApplicationsabstractBag-of-visual Words (BoWs) representation has been applied for various problems in the fields of multimedia and computer vision. The basic idea is to represent images as visual documents composed of repeatable and distinctive visual elements, which are comparable to the text words. Notwithstanding its great success and wide adoption, visual vocabulary created from single-image local descriptors is often shown to be not as effective as desired. In this paper, descriptive visual words (DVWs) and descriptive visual phrases (DVPs) are proposed as the visual correspondences to text words and phrases, where visual phrases refer to the frequently co-occurring visual word pairs. Since images are the carriers of visual objects and scenes, a descriptive visual element set can be composed by the visual words and their combinations which are effective in representing certain visual objects or scenes. Based on this idea, a general framework is proposed for generating DVWs and DVPs for image applications. In a large-scale image database containing 1506 object and scene categories, the visual words and visual word pairs descriptive to certain objects or scenes are identified and collected as the DVWs and DVPs. Experiments show that the DVWs and DVPs are informative and descriptive and, thus, are more comparable with the text words than the classic visual words. We apply the identified DVWs and DVPs in several applications including large-scale near-duplicated image retrieval, image search re-ranking, and object recognition. The combination of DVW and DVP performs better than the state of the art in large-scale near-duplicated image retrieval in terms of accuracy, efficiency and memory consumption. The proposed image search re-ranking algorithm: DWPRank outperforms the state-of-the-art algorithm by 12.4% in mean average precision and about 11 times faster in efficiency. Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2010 | Building pair-wise visual word tree for efficent image re-rankingabstractBag-of-visual Words (BoW) image representation is getting popular in computer vision and multimedia communities. However, experiments show that the traditional BoW representation is not as effective as it is desired. One of the most important reasons for its ineffectiveness is that, the traditional BoW representation lost the spatial information in images. To overcome this problem, we propose the pair-wise visual word tree, within which each visual word keeps both the appearance and spatial information between two interest points in image. Thus, the corresponding novel BoW representation preserves the spatial structure in image. Based on the pair-wise visual word tree, we propose an efficient topic word selection algorithm, which utilizes the Latent Semantic Analysis to discover the most expressive visual words for different image categories. An efficient strategy is then utilized to combine the selected topic words for image re-ranking. Massive experiments show that the novel BoW representation shows promising performance. Meanwhile, the proposed image re-ranking strategy shows the state-of-the-art precision and promising efficiency. Shiliang Zhang, Qingming Huang, Yijuan Lu, Wen Gao 0001, Qi Tian 0001 |
ICASSP | 1 |
| 2010 | Building contextual visual vocabulary for large-scale image applicationsabstractNot withstanding its great success and wide adoption in Bag-of-visual Words representation, visual vocabulary created from single image local features is often shown to be ineffective largely due to three reasons. First, many detected local features are not stable enough, resulting in many noisy and non-descriptive visual words in images. Second, single visual word discards the rich spatial contextual information among the local features, which has been proven to be valuable for visual matching. Third, the distance metric commonly used for generating visual vocabulary does not take the semantic context into consideration, which renders them to be prone to noise. To address these three confrontations, we propose an effective visual vocabulary generation framework containing three novel contributions: 1) we propose an effective unsupervised local feature refinement strategy; 2) we consider local features in groups to model their spatial contexts; 3) we further learn a discriminant distance metric between local feature groups, which we call discriminant group distance. This group distance is further leveraged to induce visual vocabulary from groups of local features. We name it contextual visual vocabulary, which captures both the spatial and semantic contexts. We evaluate the proposed local feature refinement strategy and the contextual visual vocabulary in two large-scale image applications: large-scale near-duplicate image retrieval on a dataset containing 1.5 million images and image search re-ranking tasks. Our experimental results show that the contextual visual vocabulary shows significant improvement over the classic visual vocabulary. Moreover, it outperforms the state-of-the-art Bundled Feature in the terms of retrieval precision, memory consumption and efficiency. Shiliang Zhang, Qingming Huang, Gang Hua 0001, Shuqiang Jiang, Wen Gao 0001, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2010 | Affective Visualization and Retrieval for Music VideoabstractIn modern times, music video (MV) has become an important favorite pastime to people because of its conciseness, convenience, and the ability to bring both audio and visual experiences to audiences. As the amount of MVs is explosively increasing, it has become an important task to develop new techniques for effective MV analysis, retrieval, and management. By stimulating the human affective response mechanism, affective video content analysis extracts the affective information contained in videos, and, with the affective information, natural, user-friendly, and effective MV access strategies could be developed. In this paper, a novel integrated system (i.MV) is proposed for personalized MV affective analysis, visualization, and retrieval. In i.MV, we not only perform the personalized MV affective analysis, which is a challenging and insufficiently covered problem in current affective content analysis field, but also propose novel affective visualization to convert the abstract affective states intuitive and friendly to users. Based on the affective analysis and visualization, affective information based MV retrieval is achieved. Both comprehensive experiments and subjective user studies on a large MV dataset demonstrate that our personalized affective analysis is more effective than the previous algorithms. In addition, affective visualization is proved to be more suitable for affective information-based MV retrieval than the commonly used affective state representation strategies. Shiliang Zhang, Qingming Huang, Shuqiang Jiang, Wen Gao 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 1 |
| 2009 | Utilizing affective analysis for efficient movie browsingabstractBecause of the fast increasing number of movies and long time span each movie lasts, novel methods should be developed to help users browse movies and find their desired clips effectively. Affective information in movies is closely related with users' experiences and preferences. Therefore, in this paper, we analyze the affective states of movies and propose affective information based movie browsing. Affective movie content analysis is challenging due to the great variety of movie contents and styles. To address this challenge, we first extract rich audio-visual features. Then, feature selection and affective modeling are carried out to select and map effective features into corresponding affective states. Finally, we propose novel Affective Visualization techniques which intuitively visualize affective states to achieve efficient and user-friendly movie browsing. Experiments on representative movie dataset demonstrate the effectiveness of our proposed methods. Shiliang Zhang, Qi Tian 0001, Qingming Huang, Wen Gao 0001, Shipeng Li 0001 |
ICIP | 1 |
| 2009 | Descriptive visual words and visual phrases for image applicationsabstractThe Bag-of-visual Words (BoW) image representation has been applied for various problems in the fields of multimedia and computer vision. The basic idea is to represent images as visual documents composed of repeatable and distinctive visual elements, which are comparable to the words in texts. However, massive experiments show that the commonly used visual words are not as expressive as the text words, which is not desirable because it hinders their effectiveness in various applications. In this paper, Descriptive Visual Words (DVWs) and Descriptive Visual Phrases (DVPs) are proposed as the visual correspondences to text words and phrases, where visual phrases refer to the frequently co-occurring visual word pairs. Since images are the carriers of visual objects and scenes, novel descriptive visual element set can be composed by the visual words and their combinations which are effective in representing certain visual objects or scenes. Based on this idea, a general framework is proposed for generating DVWs and DVPs from classic visual words for various applications. In a large-scale image database containing 1506 object and scene categories, the visual words and visual word pairs descriptive to certain scenes or objects are identified as the DVWs and DVPs. Experiments show that the DVWs and DVPs are compact and descriptive, thus are more comparable with the text words than the classic visual words. We apply the identified DVWs and DVPs in several applications including image retrieval, image re-ranking, and object recognition. The DVW and DVP combination outperforms the classic visual words by 19.5% and 80% in image retrieval and object recognition tasks, respectively. The DVW and DVP based image re-ranking algorithm: DWPRank outperforms the state-of-the-art VisualRank by 12.4% in accuracy and about 11 times faster in efficiency. Shiliang Zhang, Qi Tian 0001, Gang Hua 0001, Qingming Huang, Shipeng Li 0001 |
ACM Multimedia | 1 |
| 2008 | Affective MTV analysis based on arousal and valence featuresabstractNowadays, MTV has become an important favorite pastime to modern people because of its conciseness, convenience to play and the characteristic that can bring both audio and visual experiences to audiences. In this paper, we propose an affective MTV analysis framework, which realizes MTV affective state extraction, representation and clustering. Firstly, affective features are extracted from both audio and visual signals. Then, the affective state of each MTV is modeled with 2D dimensional affective model and visualized in the Arousal-Valence space. Finally the MTVs having similar affective states are clustered into same categories. The validity of proposed framework is proved by subjective user study. The comparisons between our selected features and those in related work prove that our features improve the performance by a significant margin. Shiliang Zhang, Qi Tian 0001, Shuqiang Jiang, Qingming Huang, Wen Gao 0001 |
ICME | 1 |
| 2008 | i.MTV: an integrated system for mtv affective analysisabstractIn modern time, MTV has become an important favorite pastime to people because of its conciseness, convenience and the ability to bring both audio and visual experiences to audiences. It has become an significant task to develop new techniques for natural, user-friendly, and effective MTV access. In this demo, an integrated system (i.MTV) is constructed for MTV Affective Analysis, Visualization, Retrieval, and User Profile Analysis. We not only perform the effective MTV affective analysis, but also propose novel Affective Visualization techniques to make the abstract affective states intuitive and friendly to users. Based on the affective analysis and visualization, MTV affective retrieval and management are achieved. Furthermore, novel methods are proposed for user affective preferences analysis and MTV recommendation. Shiliang Zhang, Qingming Huang, Qi Tian 0001, Shuqiang Jiang, Wen Gao 0001 |
ACM Multimedia | 1 |