VLDB 2026 Research / reviewers in the wild / expert
Sicheng Zhao
dblp:65/10574
· DBLP profile ↗
176ranked-venue papers
54as first author
83since 2021 · last 2026
0000-0001-5843-6411ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 105 · 29 first-author · 42 since 2021Artificial intelligence and machine learning · 90 · 28 first-author · 51 since 2021Computer networks · 16 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object DetectionabstractSource-Free Object Detection (SFOD) aims to adapt a source-pretrained object detector to a target domain without access to source data. However, existing SFOD methods predominantly rely on internal knowledge from the source model, which limits their capacity to generalize across domains and often results in biased pseudo-labels, thereby hindering both transferability and discriminability. In contrast, Vision Foundation Models (VFMs), pretrained on massive and diverse data, exhibit strong perception capabilities and broad generalization, yet their potential remains largely untapped in the SFOD setting. In this paper, we propose a novel SFOD framework that leverages VFMs as external knowledge sources to jointly enhance feature alignment and label quality. Specifically, we design three VFM-based modules: (1) Patch-weighted Global Feature Alignment (PGFA) distills global features from VFMs using patch-similarity–based weighting to enhance global feature transferability; (2) Prototype-based Instance Feature Alignment (PIFA) performs instance-level contrastive learning guided by momentum-updated VFM prototypes; and (3) Dual-source Enhanced Pseudo-label Fusion (DEPF) fuses predictions from detection VFMs and teacher models via an entropy-aware strategy to yield more reliable supervision. Extensive experiments on six benchmarks demonstrate that our method achieves state-of-the-art SFOD performance, validating the effectiveness of integrating VFMs to simultaneously improve transferability and discriminability. Huizai Yao, Sicheng Zhao, Pengteng Li, Shuo Lu, Weiyu Guo, Yunfan Lu, Yijie Xu, Hui Xiong 0001 |
AAAI | 2 |
| 2026 | Content-aware Information Compression and Selection for Whole Slide Image AnalysisabstractRecent advances in multi-instance learning (MIL) have demonstrated impressive performance in whole slide image (WSI) analysis. However, current methods search for cues and draw conclusions from all instances or regions, resulting in excessive redundant computation and suboptimal representation quality due to irrelevant and uninformative feature interference. To address these issues, we propose CICS, an efficient and general framework that performs compact information compression and selection for high-efficiency WSI analysis. In particular, CICS features two key components: (1) context-aware compression (CAC), which partitions the instance space into sub-regions and applies learnable compression to discard irrelevant components, reduce computational complexity while facilitating information selection, and (2) global-proximity selective attention (GPSA), which cherry-picks the most informative representation with a proximity-assisted global dynamic selection strategy. Building upon these innovations, CICS forms a plug-and-play module that reduces computational complexity through compact instance representations while improving feature quality by preserving the most informative cues. Extensive experiments on six WSI classification and survival prediction datasets show that CICS consistently improves the performance of multiple representative MIL methods. It achieves 2.5%, 7.7%, and 3.9% accuracy gain over the state-of-the-art Transformer-based TransMIL, Mamba-based MambaMIL, and graph-based WIKG methods on the ESCA dataset. Hongxun Yao, Sicheng Zhao, Yi Xiao 0003 |
AAAI | 3 |
| 2026 | Augmenting and contrasting distortion for open panoramic segmentation
Sicheng Zhao, Jiankun Zhu, Xi Chen 0110, Hongxun Yao |
Sci. China Inf. Sci. | 2 |
| 2026 | SfMamba: Efficient source-free domain adaptation via selective scan modeling
Xi Chen 0110, Hongxun Yao, Sicheng Zhao, Jiankun Zhu, Kui Jiang |
Expert Syst. Appl. | 3 |
| 2026 | CMPF: Harmonizing Cross-Model Prior Fusion for Open-Vocabulary Segmentation
Sicheng Zhao, Xi Chen 0110, Hongxun Yao, Haosen Yang 0003, Yanhao Zhang 0001, Sheng Jin 0002, Xiatian Zhu, Haonan Lu, Kui Jiang, Guiguang Ding |
Int. J. Comput. Vis. | 1 |
| 2026 | RepAttn3D: Re-parameterizing 3D attention with spatiotemporal augmentation for video understanding
Xiusheng Lu, Lechao Cheng, Sicheng Zhao, Ying Zheng 0009, Yongheng Wang, Guiguang Ding, Mingli Song |
Neural Networks | 3 |
| 2026 | CAIT: Triple-Win Compression Toward High Accuracy, Fast Inference, and Favorable Transferability for ViTsabstractVision Transformers (ViTs) have emerged as state-of-the-art models for various vision tasks recently. However, their heavy computation costs remain daunting for resource-limited devices. To address this, researchers have dedicated themselves to compressing redundant information in ViTs for acceleration. However, existing approaches generally sparsely drop redundant image tokens by token pruning or brutally remove channels by channel pruning, leading to a sub-optimal balance between model performance and inference speed. Moreover, they struggle when transferring compressed models to downstream vision tasks that require the spatial structure of images, such as semantic segmentation. To tackle these issues, we propose CAIT, a joint compression method for ViTs that achieves a harmonious blend of high accuracy, fast inference speed, and favorable transferability to downstream tasks. Specifically, we introduce an asymmetric token merging (ATME) strategy to effectively integrate neighboring tokens. It can successfully compress redundant token information while preserving the spatial structure of images. On top of it, we further design a consistent dynamic channel pruning (CDCP) strategy to dynamically prune unimportant channels in ViTs. Thanks to CDCP, insignificant channels in multi-head self-attention modules of ViTs can be pruned uniformly, significantly enhancing the model compression. Extensive experiments on multiple benchmark datasets show that our proposed method can achieve state-of-the-art performance across various ViTs. Hui Chen 0013, Zijia Lin, Sicheng Zhao, Jungong Han, Guiguang Ding |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | DINO-PCB: Two-stage vision foundation model pretraining and distillation for real-time circuit-board defect detection
Junjie Ke, Lihuo He, Jing Zhang 0037, Yuqi Ji, Hui Chen 0013, Jie Li 0001, Sicheng Zhao, Guiguang Ding, Xinbo Gao 0001 |
Pattern Recognit. | 8 |
| 2026 | CLNS: Camera-aware label noise suppression for unsupervised visible-infrared person re-identification
Sicheng Zhao, Wei Lu 0032, Sibao Chen 0001, Chris Ding, Futian Wang, Jin Tang 0001, Bin Luo 0001 |
Pattern Recognit. | 1 |
| 2026 | HEART: Emotionally Grounded Video Captioning via Hierarchical Emotion-Aligned RepresentationabstractEmotional Video Captioning (EVC) seeks to generate video descriptions that are both factually accurate and emotionally expressive. However, existing approaches often lack structured semantic grounding and fine-grained temporal modeling, leading to incomplete or emotionally inconsistent captions. To address these issues, we proposeHEART(HierarchicalEmotion-AlignedRepresentation withTemporal structure), a unified framework that jointly models hierarchical visual semantics and multi-scale temporal context. Specifically, HEART introduces a Hierarchical Semantic Extraction Module that decomposes visual content into entity-, action-, and event-level representations, providing a rich foundation for multi-level emotional alignment. A Temporal Pyramid Module captures short- and long-range temporal dependencies through multi-scale convolution, enabling temporally coherent captioning. Together, these components enable HEART to generate captions that are both emotionally grounded and temporally complete. To support this framework, we construct EmoStruct, a new benchmark dataset with fine-grained emotional annotations at the subject and predicate levels. Experiments on EmoStruct and public datasets demonstrate that HEART significantly outperforms prior methods in both semantic and emotional dimensions. Tingting Han 0003, Yuxuan Gong, Sicheng Zhao, Min Tan 0005, Zhou Yu 0001, Hongxun Yao |
IEEE Trans. Affect. Comput. | 3 |
| 2026 | LLMI3D: MLLM-Based 3D Perception From a Single 2D ImageabstractRecent advancements in autonomous driving, augmented reality, robotics, and embodied intelligence have necessitated 3D perception algorithms. However, current 3D perception methods, especially specialized small models, exhibit poor generalization in open scenarios. On the other hand, multimodal large language models (MLLMs) excel in general capacity but underperform in 3D tasks, due to weak 3D local spatial object perception, poor text-based geometric numerical output, and inability to handle camera focal variations. To address these challenges, we develop LLMI3D, and propose the following solutions: Spatial-Enhanced Local Feature Mining for better 3D spatial feature extraction, 3D Query Token-Derived Info Decoding for precise geometric regression, and Geometry Projection-Based 3D Reasoning for handling camera focal length variations. We are the first to adapt an MLLM for image-based 3D perception. Additionally, we have constructed the IG3D dataset, which provides fine-grained descriptions and question-answer annotations. Extensive experiments demonstrate that our LLMI3D achieves state-of-the-art performance, outperforming other methods by a large margin. We will publicly release our code, models, and dataset. Fan Yang 0083, Sicheng Zhao, Yanhao Zhang 0001, Hui Chen 0013, Haonan Lu, Jungong Han, Guiguang Ding |
IEEE Trans. Multim. | 2 |
| 2025 | Feature Denoising Diffusion Model for Blind Image Quality AssessmentabstractBlind Image Quality Assessment (BIQA) aims to evaluate image quality in line with human perception, without reference benchmarks. Currently, deep learning BIQA methods typically depend on using features from high-level tasks for transfer learning. However, the inherent differences between BIQA and these high-level tasks inevitably introduce noise into the quality-aware features. In this paper, we take an initial step toward exploring the diffusion model for feature denoising in BIQA, namely Perceptual Feature Diffusion for IQA (PFD-IQA), which aims to remove noise from quality-aware features. Specifically, 1) we propose a Perceptual Prior Discovery and Aggregation module to establish two auxiliary tasks to discover potential low-level features in images that are used to aggregate perceptual textual prompt conditions for the diffusion model. 2) we propose a Perceptual Conditional Feature Refinement strategy, which matches noisy features to predefined denoising trajectories and then performs exact feature denoising based on textual prompt conditions. By incorporating a lightweight denoiser and requiring only a few feature denoising steps (e.g., just five iterations), our PFD-IQA framework achieves superior performance across eight standard BIQA datasets, validating its effectiveness. Yan Zhang 0109, Yunhang Shen, Ke Li 0015, Runze Hu, Xiawu Zheng, Sicheng Zhao |
AAAI | 7 |
| 2025 | Bridge Then Begin Anew: Generating Target-Relevant Intermediate Model for Source-Free Visual Emotion AdaptationabstractVisual emotion recognition (VER), which aims at understanding humans' emotional reactions toward different visual stimuli, has attracted increasing attention. Given the subjective and ambiguous characteristics of emotion, annotating a reliable large-scale dataset is hard. For reducing reliance on data labeling, domain adaptation offers an alternative solution by adapting models trained on labeled source data to unlabeled target data. Conventional domain adaptation methods require access to source data. However, due to privacy concerns, source emotional data may be inaccessible. To address this issue, we propose an unexplored task: source-free domain adaptation (SFDA) for VER, which does not have access to source data during the adaptation process. To achieve this, we propose a novel framework termed Bridge then Begin Anew (BBA), which consists of two steps: domain-bridged model generation (DMG) and target-related model adaptation (TMA). First, the DMG bridges cross-domain gaps by generating an intermediate model, avoiding direct alignment between two VER datasets with significant differences. Then, the TMA begins training the target model anew to fit the target structure, avoiding the influence of source-specific knowledge. Extensive experiments are conducted on six SFDA settings for VER. The results demonstrate the effectiveness of BBA, which achieves remarkable performance gains compared with state-of-the-art SFDA methods and outperforms representative unsupervised domain adaptation approaches. Jiankun Zhu, Sicheng Zhao, Wenbo Tang, Zhaopan Xu, Tingting Han 0003, Pengfei Xu 0001, Hongxun Yao |
AAAI | 2 |
| 2025 | DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text RetrievalabstractThe parameter-efficient adaptation of the image-text pre-training model CLIP for video-text retrieval is a prominent area of research. While CLIP is focused on image-level vision-language matching, video-text retrieval demands comprehensive understanding at the video level. Three key discrepancies emerge in the transfer from image-level to video-level: vision, language, and alignment. However, existing methods mainly focus on vision while neglecting language and alignment. In this paper, we propose Discrepancy Reduction in Vision, Language, and Alignment (DiscoVLA), which simultaneously mitigates all three discrepancies. Specifically, we introduce Image-Video Features Fusion to integrate image-level and video-level features, effectively tackling both vision and language discrepancies. Additionally, we generate pseudo image captions to learn fine-grained image-level alignment. To mitigate alignment discrepancies, we propose Image-To-Video Alignment Distillation, which leverages image-level alignment knowledge to enhance video-level alignment. Extensive experiments demonstrate the superiority of our DiscoVLA. In particular, on MSRVTT with CLIP (ViT-B/16), DiscoVLA outperforms previous methods by 2.2% R@1 and 7.5% R@sum. The code is available at https://github.com/LunarShen/DsicoVLA. Leqi Shen, Guoqiang Gong, Tianxiang Hao 0001, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding |
CVPR | 7 |
| 2025 | Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionabstractMultimodal sentiment analysis has attracted extensive research attention as increasing numbers of users share images and texts to express their emotions and opinions on social media. Collecting large amounts of labeled sentiment data is an expensive and challenging task due to the high cost of labeling and unavoidable label ambiguity. Semi-supervised learning (SSL) is explored to utilize the extensive unlabeled data to alleviate the demand for annotation. However, unlike typical multimodal tasks, sentiment inconsistency between image and text degrades the performance of SSL algorithms. To address the issue, we propose SCRD, the first semi-supervised framework for image-text sentiment recognition. To better utilize the discriminative features of each modality, we decouple features into common and private parts. We then use the private features to train unimodal classifiers for enhanced modality-specific sentiment representation. Considering the complex relationships between modalities, we devise a modal selection-based attention module that adaptively identifies the dominant sentiment modality at the sample level to guide multimodal fusion. Furthermore, to prevent model predictions from over-relying on common features under the guidance of multimodal labels, we design a pseudo-label filtering strategy based on the matching degree of prediction and dominant modality. Extensive experiments and comparisons on five publicly available datasets demonstrate that SCRD outperforms state-of-the-art methods. Our code is released on https://github.com/wuyou-xia/Seek-Common-Ground-While-Reserving-Differences. Wuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng Yang |
CVPR | 3 |
| 2025 | HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility EvaluatorabstractAIGC images are prevalent across various fields, yet they frequently suffer from quality issues like artifacts and unnatural textures. Specialized models aim to predict defect region heatmaps but face two primary challenges: (1) lack of explainability, failing to provide reasons and analyses for subtle defects, and (2) inability to leverage common sense and logical reasoning, leading to poor generalization. Multimodal large language models (MLLMs) promise better comprehension and reasoning but face their own challenges: (1) difficulty in fine-grained defect localization due to the limitations in capturing tiny details, and (2) constraints in providing pixel-wise outputs necessary for precise heatmap generation. To address these challenges, we propose HEIE: a novel MLLM-Based Hierarchical Explainable Image Implausibility Evaluator. We introduce the CoT-Driven Explainable Trinity Evaluator, which integrates heatmaps, scores, and explanation outputs, using CoT to decompose complex tasks into subtasks of increasing difficulty and enhance interpretability. Our Adaptive Hierarchical Implausibility Mapper synergizes low-level image features with high-level mapper tokens from LLMs, enabling precise local-to-global hierarchical heatmap predictions through an uncertainty-based adaptive token approach. Moreover, we propose a new dataset: Expl-AIGI-Eval, designed to facilitate interpretable implausibility evaluation of AIGC images. Our method demonstrates state-of-the-art performance through extensive experiments. Our project is at https://yfthu.github.io/HEIE/. Fan Yang 0083, Ru Zhen, Yanhao Zhang 0001, Haoxiang Chen 0007, Haonan Lu, Sicheng Zhao, Guiguang Ding |
CVPR | 7 |
| 2025 | M3amba: Memory Mamba is All You Need for Whole Slide Image ClassificationabstractMulti-instance learning (MIL) has demonstrated impressive performance in whole slide image (WSI) analysis. However, existing approaches struggle with undesirable results and unbearable computational overhead due to the quadratic complexity of Transformers. Recently, Mamba has offered a feasible solution for modeling long-range dependencies with linear complexity. However, vanilla Mamba inherently suffers from contextual forgetting issues, making it ill-suited for capturing global dependencies across instances in large-scale WSIs. To address this, we propose a memory-driven Mamba network, dubbed M3amba, to fully explore the global latent relations among instances. Specifically, M3amba retains and iteratively updates historical information with a dynamic memory bank (DMB), thus overcoming the catastrophic forgetting defects of Mamba for long-term context representation. For better feature representation, M3amba involves an intra-group bidirectional Mamba (BiMamba) block to refine local interactions within groups. Meanwhile, we additionally perform cross-attention fusion to incorporate relevant historical information across groups, facilitating richer inter-group connections. The joint learning of inter- and intra-group representations with memory merits enables M3amba with a more powerful capability for achieving accurate and comprehensive WSI representation. Extensive experiments on four datasets demonstrate that M3amba outperforms the state-of-the-art by 6.2% and 7.0% in accuracy on the TCGA BRCA and TCGA Lung datasets while maintaining low computational costs. Kui Jiang, Yi Xiao 0003, Sicheng Zhao, Hongxun Yao |
CVPR | 4 |
| 2025 | Gaussian Constrained Diffeomorphic Deformation Network for Panoramic Semantic SegmentationabstractPanoramic semantic segmentation has garnered increasing attention due to its ability to provide comprehensive environmental perception. However, it requires a large number of annotated panoramic images to achieve satisfactory performance, which is costly. Recently, Domain Adaptation for Panoramic Semantic Segmentation (DA4PASS) has been proposed to reduce the reliance on annotated data by transferring segmentation models trained on annotated pinhole images to unlabelled panoramic images. Previous DA4PASS methods mainly focus on aligning features between pinhole and panoramic images, overlooking the unique appearance characteristics of panoramic images, particularly object distortion. To address the appearance discrepancies between pinhole and panoramic images, we propose Gaussian Constrained Diffeomorphic Deformation Network (GCDDN), which applies a panoramic deformation transformation obtained by Gaussian kernels to the annotated pinhole images. Specifically, GCDDN predicts multiple Gaussian kernels and performs first-order horizontal/vertical differences to obtain a naturally smooth and reversible panoramic deformation field, which is diffeomorphic. Due to its universality, GCDDN can be integrated into any domain adaptation (DA) method. Extensive experimental results demonstrate that integrating GCDDN leads to substantial improvements in both DA methods for pinhole images and those specifically designed for panoramic images, with a maximum gain of 1.80% in outdoor scenarios. Code is available at https://github.com/jingjiang02/GCDDN. Jiankun Zhu, Zhaopan Xu, Xi Chen 0110, Sicheng Zhao, Hongxun Yao |
ICASSP | 5 |
| 2025 | Open-Vocabulary Visual Emotion Adaptation via Prompt LearningabstractVisual Emotion Recognition (VER) aims to identify emotions from visual content and has garnered significant attention in recent years due to its wide-ranging applications. Although deep learning-based methods have shown success in VER, they require extensive labeled data, which is costly. Unsupervised Domain Adaptation (UDA) methods can reduce reliance on annotated data by transferring models trained on labeled datasets to unlabeled data. However, these methods assume that the source and target domains share the same label space. In practice, this assumption is often violated due to the inherent ambiguity and subjectivity in emotion labeling. To address this limitation, we propose a novel prompt learning paradigm for open-vocabulary visual emotion UDA, termed Domain-specific Ensemble Prompting (DSEP). DSEP leverages psychological emotion models to unify emotion labels into a common space in an ensemble manner, enhancing the open-vocabulary capabilities of UDA. It then combines ensemble label prompts with domain-specific content prompts to achieve open-vocabulary UDA. To our knowledge, we are the first to explore open-vocabulary adaptation for VER. Extensive experiments demonstrate that DSEP consistently outperforms state-of-the-art methods across four public benchmarks. Zhaopan Xu, Sicheng Zhao, Xiaojiang Peng, Hongxun Yao |
ICASSP | 2 |
| 2025 | Learning Class Prototypes for Visual Emotion RecognitionabstractVisual emotion recognition (VER), which aims at understanding humans’ emotional reactions toward different visual stimuli, has attracted increasing attention. However, because of the subjectivity and complex nature of emotion, existing VER methods suffer from one or more of the following problems: 1) semantic gap: the large affective gap between visual clues and the emotional expressions; 2) overfitting: the lack of model robustness due to unclear features in the emotional category samples; 3) label ambiguity: the overlap between categories caused by diverse emotional responses. To address these limitations, we present a novel VER method named ProtoEmotion (PoE), exploring discriminative emotional representations by jointly learning prototypes of textual emotional expressions and visual features. Specifically, text prototypes build explicit textual features for each emotion category by extracting prototypes of learnable prompts from multiple aspects, reducing semantic differences. The visual prototypes capture the most defining image features of each category, providing a more robust and discriminative feature representation, while bringing samples closer together to reduce overfitting. In addition, to alleviate the label ambiguity, we propose a label smoothing algorithm based on the prototype distance. Extensive experiments demonstrate the effectiveness of PoE, which outperforms the state-of-the-art by 1.37% on FI and 1.52% on EmotionROI datasets. Jiankun Zhu, Sicheng Zhao, Zhaopan Xu, Wenbo Tang, Hongxun Yao |
ICASSP | 2 |
| 2025 | From Easy to Hard: Progressive Active Learning Framework for Infrared Small Target Detection with Single Point SupervisionabstractRecently, single-frame infrared small target (SIRST) detection with single point supervision has drawn wide-spread attention. However, the latest label evolution with single point supervision (LESPS) framework suffers from instability, excessive label evolution, and difficulty in exerting embedded network performance. Inspired by organisms gradually adapting to their environment and continuously accumulating knowledge, we construct an innovative Progressive Active Learning (PAL) framework, which drives the existing SIRST detection networks progressively and actively recognizes and learns harder samples. Specifically, to avoid the early low-performance model leading to the wrong selection of hard samples, we propose a model pre-start concept, which focuses on automatically selecting a portion of easy samples and helping the model have basic task-specific learning capabilities. Meanwhile, we propose a refined dual-update strategy, which can promote reasonable learning of harder samples and continuous refinement of pseudo-labels. In addition, to alleviate the risk of excessive label evolution, a decay factor is reasonably introduced, which helps to achieve a dynamic balance between the expansion and contraction of target annotations. Extensive experiments show that existing SIRST detection networks equipped with our PAL framework have achieved state-of-the-art (SOTA) results on multiple public datasets. Furthermore, our PAL framework can build an efficient and stable bridge between full supervision and single point supervision tasks. Our code is available at https://github.com/YuChuang1205/PAL Chuang Yu 0003, Jinmiao Zhao, Yunpeng Liu 0001, Sicheng Zhao, Yimian Dai, Xiangyu Yue 0001 |
ICCV | 4 |
| 2025 | GMMamba: Group Masking Mamba for Whole Slide Image Classification
Hongxun Yao, Kui Jiang, Yi Xiao 0003, Sicheng Zhao |
ICCV | 5 |
| 2025 | TempMe: Video Temporal Token Merging for Efficient Text-Video RetrievalabstractMost text-video retrieval methods utilize the text-image pre-trained models like CLIP as a backbone. These methods process each sampled frame independently by the image encoder, resulting in high computational overhead and limiting practical deployment. Addressing this, we focus on efficient text-video retrieval by tackling two key challenges: 1. From the perspective of trainable parameters, current parameter-efficient fine-tuning methods incur high inference costs; 2. From the perspective of model complexity, current token compression methods are mainly designed for images to reduce spatial redundancy but overlook temporal redundancy in consecutive frames of a video. To tackle these challenges, we propose Temporal Token Merging (TempMe), a parameter-efficient and training-inference efficient text-video retrieval architecture that minimizes trainable parameters and model complexity. Specifically, we introduce a progressive multi-granularity framework. By gradually combining neighboring clips, we reduce spatio-temporal redundancy and enhance temporal modeling across different frames, leading to improved efficiency and performance. Extensive experiments validate the superiority of our TempMe. Compared to previous parameter-efficient text-video retrieval methods, TempMe achieves superior performance with just 0.50M trainable parameters. It significantly reduces output tokens by 95% and GFLOPs by 51%, while achieving a 1.8X speedup and a 4.4% R-Sum improvement. With full fine-tuning, TempMe achieves a significant 7.9% R-Sum improvement, trains 1.57X faster, and utilizes 75.2% GPU memory usage. The code is available at https://github.com/LunarShen/TempMe. Leqi Shen, Tianxiang Hao 0001, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding |
ICLR | 4 |
| 2025 | An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception CapabilityabstractThe advancements in Multimodal Large Language Models (MLLMs) have enabled various multimodal tasks to be addressed under a zero-shot paradigm. This paradigm sidesteps the cost of model fine-tuning, emerging as a dominant trend in practical application. Nevertheless, Multimodal Sentiment Analysis (MSA), a pivotal challenge in the quest for general artificial intelligence, fails to accommodate this convenience. The zero-shot paradigm exhibits undesirable performance on MSA, casting doubt on whether MLLMs can perceive sentiments as competent as supervised models. By extending the zero-shot paradigm to In-Context Learning (ICL) and conducting an in-depth study on configuring demonstrations, we validate that MLLMs indeed possess such capability. Specifically, three key factors that cover demonstrations' retrieval, presentation, and distribution are comprehensively investigated and optimized. A sentimental predictive bias inherent in MLLMs is also discovered and later effectively counteracted. By complementing each other, the devised strategies for three factors result in average accuracy improvements of 15.9% on six MSA datasets against the zero-shot paradigm and 11.2% against the random ICL baseline. Daiqing Wu, Dongbao Yang, Sicheng Zhao, Can Ma, Yu Zhou 0015 |
ICML | 3 |
| 2025 | Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual VariationsabstractVision-language models (VLMs) exhibit remarkable zero-shot capabilities but struggle with distribution shifts in downstream tasks when labeled data is unavailable, which has motivated the development of Test-Time Adaptation (TTA) to improve VLMs' performance during inference without annotations. Among various TTA approaches, cache-based methods show promise by preserving historical knowledge from low-entropy samples in a dynamic cache and fostering efficient adaptation. However, these methods face two critical reliability challenges: (1) entropy often becomes unreliable under distribution shifts, causing error accumulation in the cache and degradation in adaptation performance; (2) the final predictions may be unreliable due to inflexible decision boundaries that fail to accommodate large downstream shifts. To address these challenges, we propose a Reliable Test-time Adaptation (ReTA) method that integrates two complementary strategies to enhance reliability from two perspectives. First, to mitigate the unreliability of entropy as a sample selection criterion for cache construction, we introduce Consistency-aware Entropy Reweighting (CER), which incorporates consistency constraints to weight entropy during cache updating. While conventional approaches rely solely on low entropy for cache prioritization and risk introducing noise, our method leverages predictive consistency to maintain a high-quality cache and facilitate more robust adaptation. Second, we present Diversity-driven Distribution Calibration (DDC), which models class-wise text embeddings as multivariate Gaussian distributions, enabling adaptive decision boundaries for more accurate predictions across visually diverse content. Extensive experiments demonstrate that ReTA consistently outperforms state-of-the-art methods, particularly under real-world distribution shifts. Hui Chen 0013, Yizhe Xiong, Mengyao Lyu, Zijia Lin, Shuaicheng Niu, Sicheng Zhao, Jungong Han, Guiguang Ding |
ACM Multimedia | 8 |
| 2025 | Emotion in a Bottle: Information Bottleneck Guided Disentanglement for Emotion Domain AdaptationabstractVisual emotion recognition (VER), which aims to understand human emotional reactions to different visual stimuli, has garnered increasing attention. However, the inherent ambiguity of emotional features presents significant challenges for data annotation in supervised learning paradigms. To address this limitation, emotion domain adaptation (EDA) facilitates knowledge transfer from labeled source domains to unlabeled target domains. Recently, large visual-language models such as CLIP have demonstrated impressive transfer performance on traditional UDA tasks. However, when generalizing to more abstract concepts such as emotion, the misalignment between CLIP and emotion spaces greatly affects the model performance. To address these challenges, we propose a CLIP-based emotion disentanglement (EmoD) framework designed for EDA. Leveraging perspectives from information bottleneck theory, EmoD implements a disentangler network that extracts emotion-specific features while removing redundant emotion-agnostic information. It also incorporates cross-domain feature alignment to reduce the affective gap between domains. Experimental evaluations in six EDA settings demonstrate that EmoD achieves state-of-the-art performance, surpassing traditional CLIP-based UDA methods by an average of 2.53%. Jiankun Zhu, Sicheng Zhao, Lulu Tian, Xi Chen 0110, Hongxun Yao |
ACM Multimedia | 2 |
| 2025 | Mitigating Hallucinations in Multi-modal Large Language Models via Image Token Attention-Guided DecodingabstractXinhao Xu, Hui Chen, Mengyao Lyu, Sicheng Zhao, Yizhe Xiong, Zijia Lin, Jungong Han, Guiguang Ding. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hui Chen 0013, Mengyao Lyu, Sicheng Zhao, Yizhe Xiong, Zijia Lin, Jungong Han, Guiguang Ding |
NAACL (Long Papers) | 4 |
| 2025 | FastVID: Dynamic Density Pruning for Fast Video Large Language ModelsabstractVideo Large Language Models have demonstrated strong video understanding capabilities, yet their practical deployment is hindered by substantial inference costs caused by redundant video tokens.
Existing pruning techniques fail to effectively exploit the spatiotemporal redundancy present in video data.
To bridge this gap, we perform a systematic analysis of video redundancy from two perspectives: temporal context and visual context.
Leveraging these insights, we propose Dynamic Density Pruning for Fast Video LLMs termed FastVID.
Specifically, FastVID dynamically partitions videos into temporally ordered segments to preserve temporal structure and applies a density-based token pruning strategy to maintain essential spatial and temporal information.
Our method significantly reduces computational overhead while maintaining temporal and visual integrity.
Extensive evaluations show that FastVID achieves state-of-the-art performance across various short- and long-video benchmarks on leading Video LLMs, including LLaVA-OneVision, LLaVA-Video, Qwen2-VL, and Qwen2.5-VL.
Notably, on LLaVA-OneVision-7B, FastVID effectively prunes $\textbf{90.3\%}$ of video tokens, reduces FLOPs to $\textbf{8.3\%}$, and accelerates the LLM prefill stage by $\textbf{7.1}\times$, while maintaining $\textbf{98.0\%}$ of the original accuracy.
The code is available at https://github.com/LunarShen/FastVID. Leqi Shen, Guoqiang Gong, Pengzhang Liu, Sicheng Zhao, Guiguang Ding |
NeurIPS | 6 |
| 2025 | AD2: Anomaly Detection During Training an Distillation-Based Anomaly Detection Model
Kai Chen 0044, Xiaowang Wang, Huiyue Yang, Hui Chen 0013, Yuwang Wang, Sicheng Zhao, Guiguang Ding |
PRCV (12) | 6 |
| 2025 | Correction: Multi-source-free Domain Adaptive Object Detection
Sicheng Zhao, Huizai Yao, Chuang Lin 0003, Yue Gao 0002, Guiguang Ding |
Int. J. Comput. Vis. | 1 |
| 2025 | Universal Federated Domain Adaptation Through One-vs-All Self-Supervision for Internet of ThingsabstractIn practical Internet of Things (IoT) applications, deep neural networks (DNNs) often encounter challenges arising from covariate shifts (differences in feature distributions) and category shifts (discrepancies in label spaces), which significantly degrade their generalization performance. To mitigate these issues, universal federated domain adaptation (UFDA) techniques have been proposed to train a global model that can classify known and unknown categories while keeping data private. Nevertheless, most existing methods still struggle to precisely identify samples belonging to unknown classes in the target domain due to the unavailability of data from the source domain clients. To address these challenges, we propose a novel method, termed one-vs-all self-supervision (OSS) for IoT scenario. Specifically, OSS mainly consists of following three components. First, one-vs-all pseudo-label generation is proposed to generate high-quality pseudo-labels by leveraging source client models. Subsequently, we design a category-diverse strategy to aggregate the source models by assigning appropriate weights to each source domain client. Finally, we implement a target self-supervised learning strategy to refine feature alignment with respect to cluster centers. Comprehensive experiments are performed on four benchmark datasets: Office-31, Office-Home, VisDA-2017+ImageCLEF-DA, and Digits. The results show that our proposed OSS method achieves state-of-the-art performance in UFDA, significantly enhancing the recognition accuracy. Haojin Liao, Qiang Wang 0051, Sicheng Zhao, Tengfei Xing, Runbo Hu |
IEEE Internet Things J. | 3 |
| 2025 | Large Language Model Enhanced Logic Tensor Network for Stance Detection
Genan Dai, Jiayu Liao, Sicheng Zhao, Xianghua Fu, Xiaojiang Peng, Hu Huang 0009, Bowen Zhang 0005 |
Neural Networks | 3 |
| 2025 | Adversarial temporal sentence grounding by learning from external data
Tingting Han 0003, Kai Wang 0036, Jun Yu 0002, Sicheng Zhao, Jianping Fan 0001 |
Pattern Recognit. | 4 |
| 2025 | GraphMamba: Whole slide image classification meets graph-driven selective state space model
Hongxun Yao, Sicheng Zhao, Kui Jiang, Yi Xiao 0003 |
Pattern Recognit. | 3 |
| 2025 | Dynamic Causal Disentanglement Model for Dialogue Emotion DetectionabstractEmotion detection is a critical technology extensively employed in diverse fields. While the incorporation of commonsense knowledge has proven beneficial for existing emotion detection methods, dialogue-based emotion detection encounters numerous difficulties and challenges due to human agency and the variability of dialogue content. In dialogues, human emotions tend to accumulate in bursts. However, they are often implicitly expressed. This implies that many genuine emotions remain concealed within a plethora of unrelated words and dialogues. In this paper, we propose a Dynamic Causal Disentanglement Model founded on the separation of hidden variables, which effectively decomposes the content of dialogues and investigates the temporal accumulation of emotions, thereby enabling more precise emotion recognition. First, we introduce a novel Causal Directed Acyclic Graph (DAG) to establish the correlation between hidden emotional information and other observed elements. Subsequently, our approach utilizes pre-extracted personal states and utterance topics as guiding factors for the distribution of hidden variables, aiming to separate irrelevant ones. Specifically, we propose a Dynamic Causal Disentanglement Model to infer the propagation of utterances and hidden variables, enabling the accumulation of emotion-related information throughout the conversation. To guide this disentanglement process, we leverage the GPT4.0 and LSTM networks to extract utterance topics and personal states as observed information. Finally, we test our approach on popular datasets in dialogue emotion detection and relevant experimental results verified the model's superiority. Yuting Su 0001, Weizhi Nie, Sicheng Zhao, Anan Liu |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | SDRS: Sentiment-Aware Disentangled Representation Shifting for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) aims to leverage the complementary information from multiple modalities for affective understanding of user-generated videos. Existing methods mainly focused on designing sophisticated feature fusion strategies to integrate the separately extracted multimodal representations, ignoring the interference of the information irrelevant to sentiment. In this paper, we propose to disentangle the unimodal representations into sentiment-specific and sentiment-independent features, the former of which are fused for the MSA task. Specifically, we design a novel Sentiment-aware Disentangled Representation Shifting framework, termed SDRS, with two components.Interactive sentiment-aware representation disentanglementaims to extract sentiment-specific feature representations for each nonverbal modality by considering the contextual influence of other modalities with the newly developed cross-attention autoencoder.Attentive cross-modal representation shiftingtries to shift the textual representation in a latent token space using the nonverbal sentiment-specific representations after projection. The shifted representation is finally employed to fine-tune a pre-trained language model for multimodal sentiment analysis. Extensive experiments are conducted on three public benchmark datasets, i.e., CMU-MOSI, CMU-MOSEI, and CH-SIMS. The results demonstrate that the proposed SDRS framework not only obtains state-of-the-art results based solely on multimodal labels but also outperforms the methods that additionally require the labels of each modality. Sicheng Zhao, Zhenhua Yang, Henglin Shi, Lingpengkun Meng, Bing Qin 0001, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Action-Driven Semantic Representation and Aggregation for Video CaptioningabstractVideo captioning, a challenging task that entails generating natural language descriptions of visual content, often fails to effectively grasp the essence of action semantics. To harness the power of action detection to facilitate a deeper understanding of the video content, we propose an action-driven method, named Hierarchical Semantic Representation and Aggregation (HSRA) network. This method explicitly exploits action clues with a hierarchical semantic representation module, which models visual semantics in a three-level structure: “object-action-event”. By employing learnable action queries, our approach injects extensive action semantics into the model, thereby enabling more accurate and context-rich captions. To further enhance semantic alignment and understanding, we introduce a semantic aggregation composed of a semantic interaction module and a semantic refinement module. This component facilitates the alignment of semantics across different levels and emphasizes key information, ultimately leading to significant improvements in semantic consistency between the video and generated captions. We performed extensive evaluations on two well-established public datasets, MSVD and MSR-VTT, and the findings consistently demonstrate that our proposed HSRA network outperforms contemporary state-of-the-art methods. Tingting Han 0003, Yaochen Xu, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Guest Editorial: Special Issue on Fuzzy Affective Computing Systems
Sicheng Zhao, Hongxun Yao, Xinde Li, James Z. Wang 0001, Björn W. Schuller |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | DAR-Prompt: Dynamic Regulation in Prompt Tuning for Multi-Label Zero-Shot LearningabstractPrompt tuning achieves superior performance across a wide range of tasks, including multi-label zero-shot classification. Existing approaches employ multiple prompts to acquire comprehensive knowledge from categories, demonstrating state-of-the-art performance and significant computational efficiency. However, two main challenges still exist in these methods that impede the full potential of generalization. First, the class imbalance is not carefully addressed. Despite some efforts to adopt re-weighted loss functions to alleviate the positive-negative imbalance, such strategies tend to exacerbate the class imbalance by over-suppression of labels with fewer samples and overfitting to dominant classes. Second, the multi-prompt methods neglect the interactions between prompts during parameter optimization, underestimating the potential of prompts and leading to suboptimal performance. To address these issues, we present a novel framework named Dynamic Regulation in Prompt Tuning (DAR-Prompt). DAR-Prompt introduces three dynamic components: semantic regulator and debiased regulator to address the class imbalance, along with contrastive gradient regularization to enhance feature separation through prompt interactions during the backward pass. Specifically, the semantic regulator generates class-adaptive thresholds to compensate for tail classes and mitigate over-suppression, while the debiased regulator focuses on learning biased classes by rectifying overconfident predictions. Moreover, we apply dynamic regularization to the gradient update directions of prompts to promote orthogonality, thereby enhancing feature distinctiveness. Extensive experiments on several benchmarks show that our method can achieve state-of-the-art performance, well demonstrating its effectiveness and superiority. Code is available at https://github.com/Evelyn1ywliang/DAR-Prompt. Hui Chen 0013, Zijia Lin, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding |
IEEE Trans. Image Process. | 8 |
| 2025 | Source-Free Object Detection With Detection TransformerabstractSource-Free Object Detection (SFOD) enables knowledge transfer from a source domain to an unsupervised target domain for object detection without access to source data. Most existing SFOD approaches are either confined to conventional object detection (OD) models like Faster R-CNN or designed as general solutions without tailored adaptations for novel OD architectures, especially Detection Transformer (DETR). In this paper, we introduce Feature Reweighting ANd Contrastive Learning NetworK (FRANCK), a novel SFOD framework specifically designed to perform query-centric feature enhancement for DETRs. FRANCK comprises four key components: 1) an Objectness Score-based Sample Reweighting (OSSR) module that computes attention-based objectness scores on multi-scale encoder feature maps, reweighting the detection loss to emphasize less-recognized regions; 2) a Contrastive Learning with Matching-based Memory Bank (CMMB) module that integrates multi-level features into memory banks, enhancing class-wise contrastive learning; 3) an Uncertainty-weighted Query-fused Feature Distillation (UQFD) module that improves feature distillation through prediction quality reweighting and query feature fusion; and 4) an improved self-training pipeline with a Dynamic Teacher Updating Interval (DTUI) that optimizes pseudo-label quality. By leveraging these components, FRANCK effectively adapts a source-pre-trained DETR model to a target domain with enhanced robustness and generalization. Extensive experiments on several widely used benchmarks demonstrate that our method achieves state-of-the-art performance, highlighting its effectiveness and compatibility with DETR-based SFOD models. Huizai Yao, Sicheng Zhao, Shuo Lu, Hui Chen 0013, Tengfei Xing, Chenggang Yan 0001, Jianhua Tao 0001, Guiguang Ding |
IEEE Trans. Image Process. | 2 |
| 2025 | Cross-Modality Prompts: Few-Shot Multi-Label Recognition With Single-Label TrainingabstractFew-shot multi-label recognition (FS-MLR) presents a significant challenge due to the need to assign multiple labels to images with limited examples. Existing methods often struggle to balance the learning of novel classes and the retention of knowledge from base classes. To address this issue, we propose a novel Cross-Modality Prompts (CMP) approach. Unlike conventional methods that rely on additional semantic information to mitigate the impact of limited samples, our approach leverages multimodal prompts to adaptively tune the feature extraction network. A new FS-MLR benchmark is also proposed, which includes single-label training and multi-label testing, accompanied by benchmark datasets constructed from MS-COCO and NUS-WIDE. Extensive experiments on these datasets demonstrate the superior performance of our CMP approach, highlighting its effectiveness and adaptability. Our results show that CMP outperforms CoOp on the MS-COCO dataset with a maximal improvement of 19.47% and 23.94% in mAPharmonicfor 5-way 1-shot and 5-way 5-shot settings, respectively. Zixuan Ding, Hui Chen 0013, Tianxiang Hao 0001, Yizhe Xiong, Sicheng Zhao, Qiang Zhang 0020, Jungong Han |
IEEE Trans. Multim. | 6 |
| 2025 | Temporal Modeling With Frozen Vision-Language Foundation Models for Parameter-Efficient Text-Video RetrievalabstractTemporal modeling plays an important role in the effective adaption of the powerful pretrained text-image foundation model into text-video retrieval. However, existing methods often rely on additional heavy trainable modules, such as transformer or BiLSTM, which are inefficient. In contrast, we avoid introducing such heavy components by leveraging frozen foundation models. To this end, we propose temporal modeling with frozen vision-language foundation models (TFVL) to model the temporal dynamics with fixed encoders. Specifically, text encoder temporal modeling (TextTemp) and image encoder temporal modeling (ImageTemp) apply frozen text and image encoders within the video head and video backbone, respectively. TextTemp uses a frozen text encoder to interpret frame representations as "visual words" within a temporal "sentence," capturing temporal dependencies. On the other hand, ImageTemp uses a frozen image encoder to treat all frame tokens as a unified visual entity, learning spatiotemporal information. The total trainable parameters of our method, comprising a lightweight projection and several prompt tokens, are significantly fewer than those in other existing methods. We evaluate the effectiveness of our method on MSR-VTT, DiDeMo, ActivityNet, and LSMDC. Compared with full fine-tuning on MSR-VTT, our TFVL achieves an average 3.25% gain in R@1 with merely 0.35% of the parameters. Extensive experiments demonstrate that the proposed TFVL outperforms state-of-the-art methods with significantly fewer parameters. Leqi Shen, Tianxiang Hao 0001, Pengzhang Liu, Sicheng Zhao, Jungong Han, Guiguang Ding |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Mixed Attention and Channel Shift Transformer for Efficient Action RecognitionabstractThe practical use of the Transformer-based methods for processing videos is constrained by the high computing complexity. Although previous approaches adopt the spatiotemporal decomposition of 3D attention to mitigate the issue, they suffer from the drawback of neglecting the majority of visual tokens. This article presents a novel mixed attention operation that subtly fuses the random, spatial, and temporal attention mechanisms. The proposed random attention stochastically samples video tokens in a simple yet effective way, complementing other attention methods. Furthermore, since the attention operation concentrates on learning long-distance relationships, we employ the channel shift operation to encode short-term temporal characteristics. Our model can provide more comprehensive motion representations thanks to the amalgamation of these techniques. Experimental results show that the proposed method produces competitive action recognition results with low computational overhead on both large-scale and small-scale public video datasets. Xiusheng Lu, Yanbin Hao, Lechao Cheng, Sicheng Zhao, Yutao Liu 0002, Mingli Song |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Spatio-Temporal Attention for Text-Video RetrievalabstractText-video retrieval, a fundamental task for associating textual descriptions with video content, has become increasingly important in the video domain. Most existing methods focus on the single-modality features only considering the knowledge within individual video or text modalities, often neglecting cross-modal interactions. However, a text description corresponds to a specific spatio-temporal content within a video, involving a certain segment of a frame sequence and distinct sub-regions within these frames. Therefore, we focus on the text-conditioned video features to bridge the modality gap. In this article, we propose Spatio-Temporal Attention for video-text retrieval, termed STAttn, which utilizes textual information to focus on the spatio-temporal video content. Our final text-conditioned video features are generated from the text-related video frames and the text-related regions within these frames. First, we propose the Spatial Text-Attention Module (STAM) to learn the spatial information within video frames. STAM introduces the text-related salient patches to capture more fine-grained details. Second, we propose the Temporal Text-Attention Module (TTAM) to learn the temporal relationships between video frames. Temporal Triplet loss is proposed in TTAM to enhance the attention toward the text-related frames. Thus, the two modules learn the text-related spatio-temporal content from both intra-frame and inter-frame aspects. Extensive experiments on three benchmark datasets, MSRVTT, ActivityNet, and DiDeMo, demonstrate that our STAttn outperforms the state-of-the-art methods. Leqi Shen, Sicheng Zhao, Pengzhang Liu, Yongjun Bao, Guiguang Ding |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Geometry-Guided Domain Generalization for Monocular 3D Object DetectionabstractMonocular 3D object detection (M3OD) is important for autonomous driving. However, existing deep learning-based methods easily suffer from performance degradation in real-world scenarios due to the substantial domain gap between training and testing. M3OD's domain gaps are complex, including camera intrinsic parameters, extrinsic parameters, image appearance, etc. Existing works primarily focus on the domain gaps of camera intrinsic parameters, ignoring other key factors. Moreover, at the feature level, conventional domain invariant learning methods generally cause the negative transfer issue, due to the ignorance of dependency between geometry tasks and domains. To tackle these issues, in this paper, we propose MonoGDG, a geometry-guided domain generalization framework for M3OD, which effectively addresses the domain gap at both camera and feature levels. Specifically, MonoGDG consists of two major components. One is geometry-based image reprojection, which mitigates the impact of camera discrepancy by unifying intrinsic parameters, randomizing camera orientations, and unifying the field of view range. The other is geometry-dependent feature disentanglement, which overcomes the negative transfer problems by incorporating domain-shared and domain-specific features. Additionally, we leverage a depth-disentangled domain discriminator and a domain-aware geometry regression attention mechanism to account for the geometry-domain dependency. Extensive experiments on multiple autonomous driving benchmarks demonstrate that our method achieves state-of-the-art performance in domain generalization for M3OD. Fan Yang 0083, Hui Chen 0013, Sicheng Zhao, Chenghao Zhang 0006, Guiguang Ding |
AAAI | 4 |
| 2024 | X-ReID: Cross-Instance Transformer for Identity-Level Person Re-IdentificationabstractCurrently, most existing person re-identification methods use instance-level features, which are extracted only from a single image. However, these instance-level features can easily ignore the discriminative information because the appearance of each identity varies greatly in different images. Thus, it is necessary to exploit identity-level features, which can be shared across different images of each identity. In this paper, we propose a novel training framework, named X-ReID, to promote instance-level features to identity-level features by employing cross-attention to incorporate information from one image to another of the same identity, thus more unified and discriminative pedestrian information can be obtained. Extensive experiments on benchmark datasets show the superiority of our method over existing works. Particularly, on the challenging MSMT17, our proposed method gains 1.1% mAP improvements when compared to the second place. Leqi Shen, Sicheng Zhao, Zhelun Shen, Tianshi Xu, Guiguang Ding |
ICME | 3 |
| 2024 | More is Better: Deep Domain Adaptation with Multiple Sources
Sicheng Zhao, Hui Chen 0013, Hu Huang 0009, Pengfei Xu 0013, Guiguang Ding |
IJCAI | 1 |
| 2024 | Multi-Label Learning with Block Diagonal LabelsabstractCollecting large-scale multi-label data with full labels is difficult for real-world scenarios. Many existing studies have tried to address the issue of missing labels caused by annotation but ignored the difficulties encountered during the annotation process. We find that the high annotation workload can be attributed to two reasons: (1) Annotators are required to identify labels on widely varying visual concepts. (2) Exhaustively annotating the entire dataset with all the labels becomes notably difficult and time-consuming. In this paper, we propose a new setting, i.e. block diagonal labels, to reduce the workload on both sides. The numerous categories can be divided into different subsets based on semantics and relevance. Each annotator can only focus on its own subset of labels so that only a small set of highly relevant labels are required to be annotated per image. To deal with the issue of such missing labels, we introduce a simple yet effective method that does not require any prior knowledge of the dataset. In practice, we propose an Adaptive Pseudo-Labeling method to predict the unknown labels with less noise. Formal analysis is conducted to evaluate the superiority of our setting. Extensive experiments are conducted to verify the effectiveness of our method on multiple widely used benchmarks. Leqi Shen, Sicheng Zhao, Hui Chen 0013, Jundong Zhou, Pengzhang Liu, Yongjun Bao, Guiguang Ding |
ACM Multimedia | 2 |
| 2024 | Label-Efficient Emotion and Sentiment AnalysisabstractEmotion and sentiment analysis (ESA) assists machines to serve humans more intelligently. However, collecting large-scale high-quality datasets for training ESA models in a supervised manner is expensive, time-consuming, and difficult in practice. This tutorial focuses on the label-efficient ESA (LeESA) learning methods. Specifically, we first introduce the stimuli and characteristics of emotion and then illustrate seven typical training paradigms, followed by applications and future directions of LeESA. Sicheng Zhao, Guoli Jia, Xiaopeng Hong, Jianhua Tao 0001 |
ACM Multimedia | 1 |
| 2024 | MACP: Efficient Model Adaptation for Cooperative PerceptionabstractVehicle-to-vehicle (V2V) communications have greatly enhanced the perception capabilities of connected and automated vehicles (CAVs) by enabling information sharing to "see through the occlusions", resulting in significant performance improvements. However, developing and training complex multi-agent perception models from scratch can be expensive and unnecessary when existing single-agent models show remarkable generalization capabilities. In this paper, we propose a new framework termed MACP, which equips a single-agent pre-trained model with cooperation capabilities. We approach this objective by identifying the key challenges of shifting from single-agent to cooperative settings, adapting the model by freezing most of its parameters and adding a few lightweight modules. We demonstrate in our experiments that the proposed framework can effectively utilize cooperative observations and outperform other state-of-the-art approaches in both simulated and real-world cooperative perception benchmarks while requiring substantially fewer tunable parameters with reduced communication costs. Our ource code is available at https://github.com/PurdueDigitalTwin/MACP. Yunsheng Ma, Juanwu Lu, Can Cui 0009, Sicheng Zhao, Wenqian Ye, Ziran Wang |
WACV | 4 |
| 2024 | Big data-driven TBM tunnel intelligent construction system with automated-compliance-checking (ACC) optimization
Sicheng Zhao, Yadong Xue, Hehua Zhu |
Expert Syst. Appl. | 2 |
| 2024 | Multi-source-free Domain Adaptive Object Detection
Sicheng Zhao, Huizai Yao, Chuang Lin 0003, Yue Gao 0002, Guiguang Ding |
Int. J. Comput. Vis. | 1 |
| 2024 | Mixed Resolution Network with hierarchical motion modeling for efficient action recognition
Xiusheng Lu, Sicheng Zhao, Lechao Cheng, Ying Zheng 0009, Xueqiao Fan, Mingli Song |
Knowl. Based Syst. | 2 |
| 2024 | Video summarization via knowledge-aware multimodal deep networks
Jiehang Xie, Xuanbai Chen, Sicheng Zhao, Shao-Ping Lu |
Knowl. Based Syst. | 3 |
| 2024 | Effective Video Summarization by Extracting Parameter-Free Motion AttentionabstractVideo summarization remains a challenging task despite increasing research efforts. Traditional methods focus solely on long-range temporal modeling of video frames, overlooking important local motion information that cannot be captured by frame-level video representations. In this article, we propose the Parameter-free Motion Attention Module (PMAM) to exploit the crucial motion clues potentially contained in adjacent video frames, using a multi-head attention architecture. The PMAM requires no additional training for model parameters, leading to an efficient and effective understanding of video dynamics. Moreover, we introduce the Multi-feature Motion Attention Network (MMAN), integrating the PMAM with local and global multi-head attention based on object-centric and scene-centric video representations. The synergistic combination of local motion information, extracted by the proposed PMAM, with long-range interactions modeled by the local and global multi-head attention mechanism, can significantly enhance the performance of video summarization. Extensive experimental results on the benchmark datasets, SumMe and TVSum, demonstrate that the proposed MMAN outperforms other state-of-the-art methods, resulting in remarkable performance gains. Tingting Han 0003, Jun Yu 0002, Zhou Yu 0001, Sicheng Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Learning Deep Hierarchical Features with Spatial Regularization for One-Class Facial Expression RecognitionabstractExisting methods on facial expression recognition (FER) are mainly trained in the setting when multi-class data is available. However, to detect the alien expressions that are absent during training, this type of methods cannot work. To address this problem, we develop a Hierarchical Spatial One Class Facial Expression Recognition Network (HS-OCFER) which can construct the decision boundary of a given expression class (called normal class) by training on only one-class data. Specifically, HS-OCFER consists of three novel components. First, hierarchical bottleneck modules are proposed to enrich the representation power of the model and extract detailed feature hierarchy from different levels. Second, multi-scale spatial regularization with facial geometric information is employed to guide the feature extraction towards emotional facial representations and prevent the model from overfitting extraneous disturbing factors. Third, compact intra-class variation is adopted to separate the normal class from alien classes in the decision space. Extensive evaluations on 4 typical FER datasets from both laboratory and wild scenarios show that our method consistently outperforms state-of-the-art One-Class Classification (OCC) approaches. Bingjun Luo, Sicheng Zhao, Xibin Zhao, Yue Gao 0002 |
AAAI | 4 |
| 2023 | Confidence-based Visual Dispersal for Few-shot Unsupervised Domain AdaptationabstractUnsupervised domain adaptation aims to transfer knowledge from a fully-labeled source domain to an unlabeled target domain. However, in real-world scenarios, providing abundant labeled data even in the source domain can be infeasible due to the difficulty and high expense of annotation. To address this issue, recent works consider the Few-shot Unsupervised Domain Adaptation (FUDA) where only a few source samples are labeled, and conduct knowledge transfer via self-supervised learning methods. Yet existing methods generally overlook that the sparse label setting hinders learning reliable source knowledge for transfer. Additionally, the learning difficulty difference in target samples is different but ignored, leaving hard target samples poorly classified. To tackle both deficiencies, in this paper, we propose a novel Confidence-based Visual Dispersal Transfer learning method (C-VisDiT) for FUDA. Specifically, C-VisDiT consists of a cross-domain visual dispersal strategy that transfers only high-confidence source knowledge for model adaptation and an intra-domain visual dispersal strategy that guides the learning of hard target samples with easy ones. We conduct extensive experiments on Office-31, Office-Home, VisDA-C, and Domain- Net benchmark datasets and the results demonstrate that the proposed C-VisDiT significantly outperforms state-of- the-art FUDA methods. Our code is available at https://github.com/Bostoncake/C-VisDiT. Yizhe Xiong, Hui Chen 0013, Zijia Lin, Sicheng Zhao, Guiguang Ding |
ICCV | 4 |
| 2023 | Domain consensual contrastive learning for few-shot universal domain adaptation
Haojin Liao, Qiang Wang 0051, Sicheng Zhao, Tengfei Xing, Runbo Hu |
Appl. Intell. | 3 |
| 2023 | Deep learning-based covert brain infarct detection from multiple MRI sequences
Sicheng Zhao, Hamid F. Bagce, Vadim Spektor, Yen Chou, Clarissa D. Morales, Hao Yang 0019, Jingchen Ma, Lawrence H. Schwartz, Jennifer J. Manly, Richard P. Mayeux, Adam M. Brickman, Jose D. Gutierrez, Binsheng Zhao |
Neurocomputing | 1 |
| 2023 | Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding EmotionabstractThe emergence of artificial emotional intelligence technology is revolutionizing the fields of computers and robotics, allowing for a new level of communication and understanding of human behavior that was once thought impossible. While recent advancements in deep learning have transformed the field of computer vision, automated understanding of evoked or expressed emotions in visual media remains in its infancy. This foundering stems from the absence of a universally accepted definition of "emotion," coupled with the inherently subjective nature of emotions and their intricate nuances. In this article, we provide a comprehensive, multidisciplinary overview of the field of emotion analysis in visual media, drawing on insights from psychology, engineering, and the arts. We begin by exploring the psychological foundations of emotion and the computational principles that underpin the understanding of emotions from images and videos. We then review the latest research and systems within the field, accentuating the most promising approaches. We also discuss the current technological challenges and limitations of emotion analysis, underscoring the necessity for continued investigation and innovation. We contend that this represents a "Holy Grail" research problem in computing and delineate pivotal directions for future inquiry. Finally, we examine the ethical ramifications of emotion-understanding technologies and contemplate their potential societal impacts. Overall, this article endeavors to equip readers with a deeper understanding of the domain of emotion analysis in visual media and to inspire further research and development in this captivating and rapidly evolving field. James Z. Wang 0001, Sicheng Zhao, Chenyan Wu, Reginald B. Adams Jr., Michelle G. Newman, Tal Shafir, Rachelle Tsachor |
Proc. IEEE | 2 |
| 2023 | Toward Label-Efficient Emotion and Sentiment AnalysisabstractEmotion and sentiment play a central role in various human activities, such as perception, decision-making, social interaction, and logical reasoning. Developing artificial emotional intelligence (AEI) for machines is becoming a bottleneck in human–computer interaction. The first step of AEI is to recognize the emotion and sentiment that are conveyed in different affective signals. Traditional supervised emotion and sentiment analysis (ESA) methods, especially deep learning-based ones, usually require large-scale labeled training data. However, due to the essential subjectivity, complexity, uncertainty and ambiguity, and subtlety, collecting such annotations is expensive, time-consuming, and difficult in practice. In this article, we introduce label-efficient ESA from the computational perspective. First, we present a hierarchical taxonomy for label-efficient learning based on the availability of sample labels, emotion categories, and data domains during training. Second, for each of the seven paradigms, i.e., unsupervised, semisupervised, weakly supervised, low-shot, incremental, domain-adaptive, and domain-generalizable ESA, we give the definition, summarize existing methods, and present our views on the quantitative and qualitative comparison. Finally, we provide several promising real-world applications, followed by unsolved challenges and potential future directions. Sicheng Zhao, Xiaopeng Hong, Jufeng Yang, Guiguang Ding |
Proc. IEEE | 1 |
| 2023 | Multimodal Sentiment Analysis With Image-Text Interaction NetworkabstractMore and more users are getting used to posting images and text on social networks to share their emotions or opinions. Accordingly, multimodal sentiment analysis has become a research topic of increasing interest in recent years. Typically, there exist affective regions that evoke human sentiment in an image, which are usually manifested by corresponding words in people's comments. Similarly, people also tend to portray the affective regions of an image when composing image descriptions. As a result, the relationship between image affective regions and the associated text is of great significance for multimodal sentiment analysis. However, most of the existing multimodal sentiment analysis approaches simply concatenate features from image and text, which could not fully explore the interaction between them, leading to suboptimal results. Motivated by this observation, we propose a new image-text interaction network (ITIN) to investigate the relationship between affective image regions and text for multimodal sentiment analysis. Specifically, we introduce a cross-modal alignment module to capture region-word correspondence, based on which multimodal features are fused through an adaptive cross-modal gating module. Moreover, considering the complementary role of context information on sentiment analysis, we integrate the individual-modal contextual feature representations for achieving more reliable prediction. Extensive experimental results and comparisons on public datasets demonstrate that the proposed model is superior to the state-of-the-art methods. Tong Zhu 0003, Leida Li, Jufeng Yang, Sicheng Zhao, Hantao Liu, Jiansheng Qian |
IEEE Trans. Multim. | 4 |
| 2023 | Multimodal Emotion Classification With Multi-Level Semantic Reasoning NetworkabstractNowadays, people are accustomed to posting images and associated text for expressing their emotions on social networks. Accordingly, multimodal sentiment analysis has drawn increasingly more attention. Most of the existing image-text multimodal sentiment analysis methods simply predict the sentiment polarity. However, the same sentiment polarity may correspond to quite different emotions, such as happiness vs. excitement and disgust vs. sadness. Therefore, sentiment polarity is ambiguous and may not convey the accurate emotions that people want to express. Psychological research has shown that objects and words are emotional stimuli and that semantic concepts can affect the role of stimuli. Inspired by this observation, this paper presents a new MUlti-Level SEmantic Reasoning network (MULSER) for fine-grained image-text multimodal emotion classification, which not only investigates the semantic relationship among objects and words respectively, but also explores the semantic relationship between regional objects and global concepts. For image modality, we first build graphs to extract objects and global representation, and employ a graph attention module to perform bilevel semantic reasoning. Then, a joint visual graph is built to learn the regional-global semantic relations. For text modality, we build a word graph and further apply graph attention to reinforce the interdependencies among words in a sentence. Finally, a cross-modal attention fusion module is proposed to fuse semantic-enhanced visual and textual features, based on which informative multimodal representations are obtained for fine-grained emotion classification. The experimental results on public datasets demonstrate the superiority of the proposed model over the state-of-the-art methods. Tong Zhu 0003, Leida Li, Jufeng Yang, Sicheng Zhao, Xiao Xiao 0007 |
IEEE Trans. Multim. | 4 |
| 2023 | How to Use In-Band Network Telemetry Wisely: Network-Wise Orchestration of Sel-INTabstractAs a promising network monitoring technique, in-band network telemetry (INT) helps to visualize networks in a fine-grained and real-time manner. Meanwhile, to address the overheads of INT, people have proposed a few selective INT (Sel-INT) approaches that only select a portion of packets in each flow to insert INT fields and distribute different types of INT fields over the selected packets. In this paper, we study how to use Sel-INT wisely in a network such that the tradeoff between monitoring accuracy/coverage and INT overheads can be balanced well. Specifically, we try to orchestrate the Sel-INT schemes of flows in both network- and flow-levels. For the network-level optimization, we model it as an INT planning problem in which the Sel-INT schemes of flows should be determined to maximize the information gain of INT as well as minimize the bandwidth overheads of INT. We formulate an integer linear programming (ILP) model to tackle the problem, prove its$\mathcal {N} \mathcal {P}$-hardness, and leverage Lagrangian relaxation to design a polynomial-time approximation algorithm for it. The flow-level optimization considers a dynamic network environment, and we propose to combine deep learning (DL) based traffic prediction with Sel-INT, such that the Sel-INT scheme of each individual flow can be updated timely and adaptively. We implement the proposal in a small but real network testbed and experimentally demonstrate self-adaptive orchestration of Sel-INT with it. Shaofei Tang, Sicheng Zhao, Xiaoqin Pan, Zuqing Zhu |
IEEE/ACM Trans. Netw. | 2 |
| 2022 | Modeling long-term video semantic distribution for temporal action proposal generation
Tingting Han 0003, Sicheng Zhao, Xiaoshuai Sun, Jun Yu 0002 |
Neurocomputing | 2 |
| 2022 | Affective Image Content Analysis: Two Decades Review and New PerspectivesabstractImages can convey rich semantics and induce various emotions in viewers. Recently, with the rapid advancement of emotional intelligence and the explosive growth of visual data, extensive research efforts have been dedicated to affective image content analysis (AICA). In this survey, we will comprehensively review the development of AICA in the recent two decades, especially focusing on the state-of-the-art methods with respect to three main challenges - the affective gap, perception subjectivity, and label noise and absence. We begin with an introduction to the key emotion representation models that have been widely employed in AICA and description of available datasets for performing evaluation with quantitative comparison of label noise and dataset bias. We then summarize and compare the representative approaches on (1) emotion feature extraction, including both handcrafted and deep features, (2) learning methods on dominant emotion recognition, personalized emotion prediction, emotion distribution learning, and learning from noisy data or few labels, and (3) AICA based applications. Finally, we discuss some challenges and promising research directions in the future, such as image content and context understanding, group emotion clustering, and viewer-image interaction. Sicheng Zhao, Xingxu Yao, Jufeng Yang, Guoli Jia, Guiguang Ding, Tat-Seng Chua, Björn W. Schuller, Kurt Keutzer |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | CLN: Cross-Domain Learning Network for 2D Image-Based 3D Shape RetrievalabstractRetrieving 3D shapes based on 2D images is a challenging research topic, due to the significant gap between different domains. Recently, various approaches have been proposed to handle this problem. However, the majority of methods target the cross-domain retrieval task as a pure domain adaptation problem, which focuses on the alignment but ignores the visual relevance between the 2D images and their corresponding 3D shapes. To fundamentally decrease the divergence between different domains, we propose a novel cross-domain learning network (CLN) for 2D image-based 3D shape retrieval task. First, we estimate the pose information from the 2D image to guide the view rendering of 3D shapes, which increases the visual correlations of the cross-domain data to eliminate the divergence between them. Second, we introduce a novel joint learning network, considering both the domain-specific characteristics and the cross-domain interactions for data alignment, which further compensates for the gap between different domains by controlling the distance of intra- and inter-classes. After the metric learning process, discriminative descriptors of images and shapes are generated for the cross-domain retrieval task. To prove the effectiveness and robustness of the proposed method, we conduct extensive experiments on the MI3DOR, SHREC’13, and SHREC’14 datasets. The experimental results demonstrate the superiority of our proposed method, and significant improvements have been achieved compared with state-of-the-art methods. Weizhi Nie, Yue Zhao 0042, Jie Nie, Anan Liu, Sicheng Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Deep Correlated Joint Network for 2-D Image-Based 3-D Model RetrievalabstractIn this article, we propose a novel deep correlated joint network (DCJN) approach for 2-D image-based 3-D model retrieval. First, the proposed method can jointly learn two distinct deep neural networks, which are trained for individual modalities to learn two deep nonlinear transformations for visual feature extraction from the co-embedding feature space. Second, we propose the global loss function for the DCJN, consisting of a discriminative loss and a correlation loss. The discriminative loss aims to minimize the intraclass distance of the extracted features and maximize the interclass distance of such features to a large margin within each modality, while the correlation loss focuses on mitigating the distribution discrepancy across different modalities. Consequently, the proposed method can realize cross-modality feature extraction guided by the defined global loss function to benefit the similarity measure between 2-D images and 3-D models. For a comparison experiment, we contribute the current largest 2-D image-based 3-D model retrieval dataset. Moreover, the proposed method was further evaluated on three popular benchmarks, including the 3-D Shape Retrieval Contest 2014, 2016, and 2018 benchmarks. The extensive comparison experimental results demonstrate the superiority of this method over the state-of-the-art methods. Weizhi Nie, Anan Liu, Sicheng Zhao, Yue Gao 0002 |
IEEE Trans. Cybern. | 3 |
| 2022 | Emotional Semantics-Preserved and Feature-Aligned CycleGAN for Visual Emotion AdaptationabstractThanks to large-scale labeled training data, deep neural networks (DNNs) have obtained remarkable success in many vision and multimedia tasks. However, because of the presence of domain shift, the learned knowledge of the well-trained DNNs cannot be well generalized to new domains or datasets that have few labels. Unsupervised domain adaptation (UDA) studies the problem of transferring models trained on one labeled source domain to another unlabeled target domain. In this article, we focus on UDA in visual emotion analysis for both emotion distribution learning and dominant emotion classification. Specifically, we design a novel end-to-end cycle-consistent adversarial model, called CycleEmotionGAN++. First, we generate an adapted domain to align the source and target domains on the pixel level by improving CycleGAN with a multiscale structured cycle-consistency loss. During the image translation, we propose a dynamic emotional semantic consistency loss to preserve the emotion labels of the source images. Second, we train a transferable task classifier on the adapted domain with feature-level alignment between the adapted and target domains. We conduct extensive UDA experiments on the Flickr-LDL and Twitter-LDL datasets for distribution learning and ArtPhoto and Flickr and Instagram datasets for emotion classification. The results demonstrate the significant improvements yielded by the proposed CycleEmotionGAN++ compared to state-of-the-art UDA approaches. Sicheng Zhao, Xuanbai Chen, Xiangyu Yue 0001, Chuang Lin 0003, Pengfei Xu 0013, Ravi Krishna, Jufeng Yang, Guiguang Ding, Alberto L. Sangiovanni-Vincentelli, Kurt Keutzer |
IEEE Trans. Cybern. | 1 |
| 2022 | Personalized Image Aesthetics Assessment via Meta-Learning With Bilevel Gradient OptimizationabstractTypical image aesthetics assessment (IAA) is modeled for the generic aesthetics perceived by an "average" user. However, such generic aesthetics models neglect the fact that users' aesthetic preferences vary significantly depending on their unique preferences. Therefore, it is essential to tackle the issue for personalized IAA (PIAA). Since PIAA is a typical small sample learning (SSL) problem, existing PIAA models are usually built by fine-tuning the well-established generic IAA (GIAA) models, which are regarded as prior knowledge. Nevertheless, this kind of prior knowledge based on "average aesthetics" fails to incarnate the aesthetic diversity of different people. In order to learn the shared prior knowledge when different people judge aesthetics, that is, learn how people judge image aesthetics, we propose a PIAA method based on meta-learning with bilevel gradient optimization (BLG-PIAA), which is trained using individual aesthetic data directly and generalizes to unknown users quickly. The proposed approach consists of two phases: 1) meta-training and 2) meta-testing. In meta-training, the aesthetics assessment of each user is regarded as a task, and the training set of each task is divided into two sets: 1) support set and 2) query set. Unlike traditional methods that train a GIAA model based on average aesthetics, we train an aesthetic meta-learner model by bilevel gradient updating from the support set to the query set using many users' PIAA tasks. In meta-testing, the aesthetic meta-learner model is fine-tuned using a small amount of aesthetic data of a target user to obtain the PIAA model. The experimental results show that the proposed method outperforms the state-of-the-art PIAA metrics, and the learned prior model of BLG-PIAA can be quickly adapted to unseen PIAA tasks. Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, Guangming Shi |
IEEE Trans. Cybern. | 4 |
| 2022 | A Review of Single-Source Deep Unsupervised Visual Domain AdaptationabstractLarge-scale labeled training datasets have enabled deep neural networks to excel across a wide range of benchmark vision tasks. However, in many applications, it is prohibitively expensive and time-consuming to obtain large quantities of labeled data. To cope with limited labeled training data, many have attempted to directly apply models trained on a large-scale labeled source domain to another sparsely labeled or unlabeled target domain. Unfortunately, direct transfer across domains often performs poorly due to the presence of domain shift or dataset bias. Domain adaptation (DA) is a machine learning paradigm that aims to learn a model from a source domain that can perform well on a different (but related) target domain. In this article, we review the latest single-source deep unsupervised DA methods focused on visual tasks and discuss new perspectives for future research. We begin with the definitions of different DA strategies and the descriptions of existing benchmark datasets. We then summarize and compare different categories of single-source unsupervised DA methods, including discrepancy-based methods, adversarial discriminative methods, adversarial generative methods, and self-supervision-based methods. Finally, we discuss future research directions with challenges and possible solutions. Sicheng Zhao, Xiangyu Yue 0001, Shanghang Zhang, Bo Li 0080, Han Zhao 0002, Bichen Wu, Ravi Krishna, Joseph Gonzalez 0001, Alberto L. Sangiovanni-Vincentelli, Sanjit A. Seshia, Kurt Keutzer |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | On the Bilevel Optimization for Remapping Virtual Networks in an HOE-DCNabstractHybrid optical/electrical datacenter network (HOE-DCN) uses the inter-rack networks that consist of both electrical Ethernet switches and optical cross-connects (OXCs), for better cost-efficiency and scalability. Meanwhile, to provision dynamic network services well, the operator of an HOE-DCN needs to deploy virtual networks (VNTs) and remap them adaptively. Therefore, this work studies the problem of VNT remapping in an HOE-DCN from a novel perspective, i.e., the remapping schemes should be optimized for not only the network status after the remapping but also the transition to realize it. Specifically, we model this problem as a bilevel optimization, where the upper-level optimization aims at selecting proper virtual machines (VMs) to migrate such that the estimated latency of VM migration can be minimized, and the lower-level optimization determines the actual scheme of VNT remapping for minimizing the number of resource hot-spots. We first formulate a bilevel mixed integer linear programming (BMILP) model for the bilevel optimization, and then propose a polynomial time algorithm based on enumeration to solve it approximately. Extensive simulations verify the effectiveness of our proposal. Hao Yang 0019, Xiaoqin Pan, Sicheng Zhao, Binjie Ge, Zuqing Zhu |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2021 | ePointDA: An End-to-End Simulation-to-Real Domain Adaptation Framework for LiDAR Point Cloud SegmentationabstractDue to its robust and precise distance measurements, LiDAR plays an important role in scene understanding for autonomous driving. Training deep neural networks (DNNs) on LiDAR data requires large-scale point-wise annotations, which are time-consuming and expensive to obtain. Instead, simulation-to-real domain adaptation (SRDA) trains a DNN using unlimited synthetic data with automatically generated labels and transfers the learned model to real scenarios. Existing SRDA methods for LiDAR point cloud segmentation mainly employ a multi-stage pipeline and focus on feature-level alignment. They require prior knowledge of real-world statistics and ignore the pixel-level dropout noise gap and the spatial feature gap between different domains. In this paper, we propose a novel end-to-end framework, named ePointDA, to address the above issues. Specifically, ePointDA consists of three modules: self-supervised dropout noise rendering, statistics-invariant and spatially-adaptive feature alignment, and transferable segmentation learning. The joint optimization enables ePointDA to bridge the domain shift at the pixel-level by explicitly rendering dropout noise for synthetic LiDAR and at the feature-level by spatially aligning the features between different domains, without requiring the real-world statistics. Extensive experiments adapting from synthetic GTA-LiDAR to real KITTI and SemanticKITTI demonstrate the superiority of ePointDA for LiDAR point cloud segmentation. Sicheng Zhao, Yezhen Wang, Bo Li 0080, Bichen Wu, Yang Gao 0029, Pengfei Xu 0013, Trevor Darrell, Kurt Keutzer |
AAAI | 1 |
| 2021 | Spatio-temporal Contrastive Domain Adaptation for Action RecognitionabstractCompared with image-based UDA, video-based UDA is comprehensive to bridge the domain shift on both spatial representation and temporal dynamics. Most previous works focus on short-term modeling and alignment with frame-level or clip-level features, which is not discriminative sufficiently for video-based UDA tasks. To address these problems, in this paper we propose to establish the cross-modal domain alignment via self-supervised contrastive framework, i.e., spatio-temporal contrastive domain adaptation (STCDA), to learn the joint clip-level and video-level representation alignment. Since the effective representation is modeled from unlabeled data by self-supervised learning (SSL), spatio-temporal contrastive learning (STCL) is proposed to explore the useful long-term feature representation for classification, using self-supervision setting trained from the contrastive clip/video pairs with positive or negative properties. Besides, we involve a novel domain metric scheme, i.e., video-based contrastive alignment (VCA), to optimize the category-aware video-level alignment and generalization between source and target. The proposed STCDA achieves stat-of-the-art results on several UDA benchmarks for action recognition. Sicheng Zhao, Jing-Yu Yang 0002, Huanjing Yue, Pengfei Xu 0013, Runbo Hu |
CVPR | 2 |
| 2021 | Domain-Invariant Disentangled Network for Generalizable Object DetectionabstractWe address the problem of domain generalizable object detection, which aims to learn a domain-invariant detector from multiple "seen" domains so that it can generalize well to other "unseen" domains. The generalization ability is crucial in practical scenarios especially when it is difficult to collect data. Compared to image classification, domain generalization in object detection has seldom been explored with more challenges brought by domain gaps on both image and instance levels. In this paper, we propose a novel generalizable object detection model, termed Domain-Invariant Disentangled Network (DIDN). In contrast to directly aligning multiple sources, we integrate a disentangled network into Faster R-CNN. By disentangling representations on both image and instance levels, DIDN is able to learn domain-invariant representations that are suitable for generalized object detection. Furthermore, we design a cross-level representation reconstruction to complement this two-level disentanglement so that informative object representations could be preserved. Extensive experiments are conducted on five benchmark datasets and the results demonstrate that our model achieves state-of-the-art performances on domain generalization for object detection. Chuang Lin 0003, Zehuan Yuan, Sicheng Zhao, Peize Sun, Changhu Wang, Jianfei Cai 0001 |
ICCV | 3 |
| 2021 | Multi-Source Domain Adaptation for Object DetectionabstractTo reduce annotation labor associated with object detection, an increasing number of studies focus on transferring the learned knowledge from a labeled source domain to another unlabeled target domain. However, existing methods assume that the labeled data are sampled from a single source domain, which ignores a more generalized scenario, where labeled data are from multiple source domains. For the more challenging task, we propose a unified Faster R-CNN based framework, termed Divide-and-Merge Spindle Network (DMSN), which can simultaneously enhance domain invariance and preserve discriminative power. Specifically, the framework contains multiple source subnets and a pseudo target subnet. First, we propose a hierarchical feature alignment strategy to conduct strong and weak alignments for low- and high-level features, respectively, considering their different effects for object detection. Second, we develop a novel pseudo subnet learning algorithm to approximate optimal parameters of pseudo target subset by weighted combination of parameters in different source subnets. Finally, a consistency regularization for region proposal network is proposed to facilitate each subnet to learn more abstract invariances. Extensive experiments on different adaptation scenarios demonstrate the effectiveness of the proposed model. Xingxu Yao, Sicheng Zhao, Pengfei Xu 0013, Jufeng Yang |
ICCV | 2 |
| 2021 | Predicting Cutterhead Torque for TBM based on Different Characteristics and AGA-Optimized LSTM-MLPabstractAdaptive adjustment of excavation parameters makes a significant role in the process of tunneling by tunnel boring machine (TBM), which ensures the tunneling carried out safely and efficiently. Though substantial effort has been devoted to this area, there is still a lack of a comprehensive method for TBM data analysis. In this paper, we analyzed the TBM data from different perspectives. The data source is from the Songhua River Water Conveyance Project. In order to facilitate the processing and analysis of the data, we proposed the concepts of rising characteristic interval (RCI) and stable characteristic interval (SCI), which are the first 30 seconds of the rising stage and one sixth of the center part of the stable stage respectively. As a key parameter, the cutterhead torque (T), which reflects the interaction between the cutter and the soil, is selected as our prediction target. In order to forecast the value of T in the SCIs, the time series characteristic and the non time series (mean and variance) characteristic of the important excavation parameters in the RCIs are analyzed. A sequential combination of long short-term memory (LSTM) and multi-layer perceptrons (MLP), LSTM-MLP for short, is used to make a comprehensive analysis of the two characteristics. Notably, adaptive genetic algorithm (AGA) was employed to optimize the topology structure and the hyper parameters of our neural network, which ensures the convergence of the basic genetic algorithm and maintains the diversity of the population at the same time. The experimental results indicate that, LSTM-MLP performs better in comparison with LSTM network and backpropagation neural network (BPNN, a kind of MLP). Our work provides a reference for the control and optimization of TBM’s excavation parameters. To make our results fully reproducible, all the relevant source codes and the preprocessed dataset are publicly available at https://github.com/Dandelionslove/LSTM MLP for TBM. Shuangli Zhang, Qingfeng Du, Sicheng Zhao |
SMC | 3 |
| 2021 | Curriculum CycleGAN for Textual Sentiment Domain Adaptation with Multiple SourcesabstractSentiment analysis of user-generated reviews or comments on products and services in social networks can help enterprises to analyze the feedback from customers and take corresponding actions for improvement. To mitigate large-scale annotations on the target domain, domain adaptation (DA) provides an alternate solution by learning a transferable model from other labeled source domains. Existing multi-source domain adaptation (MDA) methods either fail to extract some discriminative features in the target domain that are related to sentiment, neglect the correlations of different sources and the distribution difference among different sub-domains even in the same source, or cannot reflect the varying optimal weighting during different training stages. In this paper, we propose a novel instance-level MDA framework, named curriculum cycle-consistent generative adversarial network (C-CycleGAN), to address the above issues. Specifically, C-CycleGAN consists of three components: (1) pre-trained text encoder which encodes textual input from different domains into a continuous representation space, (2) intermediate domain generator with curriculum instance-level adaptation which bridges the gap across source and target domains, and (3) task classifier trained on the intermediate domain for final sentiment classification. C-CycleGAN transfers source samples at instance-level to an intermediate domain that is closer to the target domain with sentiment semantics preserved and without losing discriminative features. Further, our dynamic instance-level weighting mechanisms can assign the optimal weights to different source samples in each training stage. We conduct extensive experiments on three benchmark datasets and achieve substantial gains over state-of-the-art DA approaches. Our source code is released at: https://github.com/WArushrush/Curriculum-CycleGAN. Sicheng Zhao, Xiangyu Yue 0001, Jufeng Yang, Ravi Krishna, Pengfei Xu 0013, Kurt Keutzer |
WWW | 1 |
| 2021 | MADAN: Multi-source Adversarial Domain Aggregation Network for Domain Adaptation
Sicheng Zhao, Bo Li 0080, Pengfei Xu 0013, Xiangyu Yue 0001, Guiguang Ding, Kurt Keutzer |
Int. J. Comput. Vis. | 1 |
| 2021 | Sketch-specific data augmentation for freehand sketch recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Shengping Zhang, Sicheng Zhao, Fatih Porikli |
Neurocomputing | 5 |
| 2021 | 3D Pose Estimation Based on Reinforce Learning for 2D Image-Based 3D Model RetrievalabstractIn this paper, we propose a novel characteristic view selection model (CVSM) to address the 2D image-based 3D object retrieval problem. This work includes two key contributions: 1) we propose a novel reinforcement learning model to estimate the 3D pose based on a 2D image; and 2) we render the pose-specific model to generate a representative angle view for retrieval applications. First, we define state, policy, action and reward functions to train an agent with the reinforcement learning framework, by which the agent can effectively reduce the computational cost of the characteristic view selection and directly obtain the 3D model pose. Second, to resolve the problem of computing similarity in the cross-domain between the virtual 3D model view and the real query image, we project them into the skeleton domain, and the skeleton information can effectively bridge the gap between the image and 3D model view for cross-media retrieval. To demonstrate the performance of our approach, we compare with some classic 3D pose estimation methods using the popular Pascal3D dataset. To demonstrate the performance of our approach in model retrieval, we collect a new dataset that includes pairs of 2D images and 3D objects, where 3D objects are based on the ModelNet40 dataset and 2D images are based on the ImageNet dataset, and we experiment with our method using the SHREC 2018 and SHREC 2019 databases. The experimental results demonstrate the superiority of our method. Weizhi Nie, Wen-Wu Jia, Wenhui Li 0001, Anan Liu, Sicheng Zhao |
IEEE Trans. Multim. | 5 |
| 2021 | C-GCN: Correlation Based Graph Convolutional Network for Audio-Video Emotion RecognitionabstractWith the development of both hardware and deep neural network technologies, tremendous improvements have been achieved in the performance of automatic emotion recognition (AER) based on the video data. However, AER is still a challenging task due to subtle expression, abstract concept of emotion and the representation of multi-modal information. Most proposed approaches focus on the multi-modal feature learning and fusion strategy, which pay more attention to the characteristic of a single video and ignore the correlation among the videos. To explore this correlation, in this paper, we propose a novel correlation-based graph convolutional network (C-GCN) for AER, which can comprehensively consider the correlation of the intra-class and inter-class videos for feature learning and information fusion. More specifically, we introduce the graph model to represent the correlation among the videos. This correlated information can help to improve the discrimination of node features in the progress of graph convolutional network. Meanwhile, the multi-head attention mechanism is applied to predict the hidden relationship among the videos, which can strengthen the inter-class correlation to improve the performance of classifiers. The C-GCN is evaluated on the AFEW datasets and eNTERFACE 05 dataset. The final experimental results demonstrate the superiority of our proposed method over the state-of-the-art methods. Weizhi Nie, Minjie Ren, Jie Nie, Sicheng Zhao |
IEEE Trans. Multim. | 4 |
| 2021 | APSE: Attention-Aware Polarity-Sensitive Embedding for Emotion-Based Image RetrievalabstractWith the popularity of social media, an increasing number of people are accustomed to expressing their feelings and emotions online using images and videos. An emotion-based image retrieval (EBIR) system is useful for obtaining visual contents with desired emotions from a massive repository. Existing EBIR methods mainly focus on modeling the global characteristics of visual content without considering the crucial role of informative regions of interest in conveying emotions. Further, they ignore the hierarchical relationships between coarse polarities and fine categories of emotions. In this paper, we design an attention-aware polarity-sensitive embedding (APSE) network to address these issues. First, we develop a hierarchical attention mechanism to automatically discover and model the informative regions of interest. Specifically, both polarity- and emotion-specific attended representations are aggregated for discriminative feature embedding. Second, we propose a generated emotion-pair (GEP) loss to simultaneously consider the inter- and intra-polarity relationships of the emotion labels. Moreover, we adaptively generate negative examples of different hard levels in the feature space guided by the attention module to further improve the performance of feature embedding. Extensive experiments on four popular benchmark datasets demonstrate that the proposed APSE method outperforms the state-of-the-art EBIR approaches by a large margin. Xingxu Yao, Sicheng Zhao, Yukun Lai, Dongyu She, Jie Liang 0007, Jufeng Yang |
IEEE Trans. Multim. | 2 |
| 2020 | Heterogeneous Transfer Learning with Weighted Instance-Correspondence DataabstractInstance-correspondence (IC) data are potent resources for heterogeneous transfer learning (HeTL) due to the capability of bridging the source and the target domains at the instance-level. To this end, people tend to use machine-generated IC data, because manually establishing IC data is expensive and primitive. However, existing IC data machine generators are not perfect and always produce the data that are not of high quality, thus hampering the performance of domain adaption. In this paper, instead of improving the IC data generator, which might not be an optimal way, we accept the fact that data quality variation does exist but find a better way to use the data. Specifically, we propose a novel heterogeneous transfer learning method named Transfer Learning with Weighted Correspondence (TLWC), which utilizes IC data to adapt the source domain to the target domain. Rather than treating IC data equally, TLWC can assign solid weights to each IC data pair depending on the quality of the data. We conduct extensive experiments on HeTL datasets and the state-of-the-art results verify the effectiveness of TLWC. Xiaoming Jin, Guiguang Ding, Jungong Han, Jiyong Zhang 0001, Sicheng Zhao |
AAAI | 7 |
| 2020 | Multi-Source Domain Adaptation for Visual Sentiment ClassificationabstractExisting domain adaptation methods on visual sentiment classification typically are investigated under the single-source scenario, where the knowledge learned from a source domain of sufficient labeled data is transferred to the target domain of loosely labeled or unlabeled data. However, in practice, data from a single source domain usually have a limited volume and can hardly cover the characteristics of the target domain. In this paper, we propose a novel multi-source domain adaptation (MDA) method, termed Multi-source Sentiment Generative Adversarial Network (MSGAN), for visual sentiment classification. To handle data from multiple source domains, it learns to find a unified sentiment latent space where data from both the source and target domains share a similar distribution. This is achieved via cycle consistent adversarial learning in an end-to-end manner. Extensive experiments conducted on four benchmark datasets demonstrate that MSGAN significantly outperforms the state-of-the-art MDA approaches for visual sentiment classification. Chuang Lin 0003, Sicheng Zhao, Lei Meng 0001, Tat-Seng Chua |
AAAI | 2 |
| 2020 | An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated VideosabstractEmotion recognition in user-generated videos plays an important role in human-centered computing. Existing methods mainly employ traditional two-stage shallow pipeline, i.e. extracting visual and/or audio features and training classifiers. In this paper, we propose to recognize video emotions in an end-to-end manner based on convolutional neural networks (CNNs). Specifically, we develop a deep Visual-Audio Attention Network (VAANet), a novel architecture that integrates spatial, channel-wise, and temporal attentions into a visual 3D CNN and temporal attentions into an audio 2D CNN. Further, we design a special classification loss, i.e. polarity-consistent cross-entropy loss, based on the polarity-emotion hierarchy constraint to guide the attention generation. Extensive experiments conducted on the challenging VideoEmotion-8 and Ekman-6 datasets demonstrate that the proposed VAANet outperforms the state-of-the-art approaches for video emotion recognition. Our source code is released at: https://github.com/maysonma/VAANet. Sicheng Zhao, Yunsheng Ma, Jufeng Yang, Tengfei Xing, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
AAAI | 1 |
| 2020 | Multi-Source Distilling Domain AdaptationabstractDeep neural networks suffer from performance decay when there is domain shift between the labeled source domain and unlabeled target domain, which motivates the research on domain adaptation (DA). Conventional DA methods usually assume that the labeled data is sampled from a single source distribution. However, in practice, labeled data may be collected from multiple sources, while naive application of the single-source DA algorithms may lead to suboptimal solutions. In this paper, we propose a novel multi-source distilling domain adaptation (MDDA) network, which not only considers the different distances among multiple sources and the target, but also investigates the different similarities of the source samples to the target ones. Specifically, the proposed MDDA includes four stages: (1) pre-train the source classifiers separately using the training data from each source; (2) adversarially map the target into the feature space of each source respectively by minimizing the empirical Wasserstein distance between source and target; (3) select the source training samples that are closer to the target to fine-tune the source classifiers; and (4) classify each encoded target feature by corresponding source classifier, and aggregate different predictions using respective domain weight, which corresponds to the discrepancy between each source and target. Extensive experiments are conducted on public DA benchmarks, and the results demonstrate that the proposed MDDA significantly outperforms the state-of-the-art approaches. Our source code is released at: https://github.com/daoyuan98/MDDA. Sicheng Zhao, Guangzhi Wang, Shanghang Zhang, Yaxian Li, Zhichao Song, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
AAAI | 1 |
| 2020 | On the Parallel Reconfiguration of Virtual Networks in Hybrid Optical/Electrical Datacenter NetworksabstractRecently, hybrid optical/electrical datacenter networks (HOE-DCNs) have been considered as a promising DCN architecture, because they merge the merits of electrical packet switching (EPS) and optical circuit switching (OCS). This paper considers the reconfiguration of virtual networks (VNTs) in an HOE-DCN to address the dynamic nature of emerging network services. Specifically, we study the problem that given the original and new virtual network embedding (VNE) schemes of several VNTs, how to schedule parallel virtual machine (VM) migrations in batches to realize the VNT reconfiguration within the shortest time. We design two algorithms to reconFigure the inter-rack network in an HOE-DCN in steps, schedule VMs to migrate accordingly, and allocate bandwidth to the VM migrations. The first algorithm uses the one-shot approach, where all the VM migrations are conducted in parallel within the shortest possible time. We formulate a linear programming (LP) to solve the bandwidth allocations in it exactly. Next, to relieve the bandwidth competition introduced by the one-shot approach, we propose the second algorithm by leveraging the multi-shot approach, i.e., invoking multiple batches of parallel VM migrations such that the reconfiguration time can be further reduced. Extensive simulations verify the effectiveness of our proposals. Sicheng Zhao, Xiaoqin Pan, Zuqing Zhu |
GLOBECOM | 1 |
| 2020 | Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal SpaceabstractBoth images and music can convey rich semantics and are widely used to induce specific emotions. Matching images and music with similar emotions might help to make emotion perceptions more vivid and stronger. Existing emotion-based image and music matching methods either employ limited categorical emotion states which cannot well reflect the complexity and subtlety of emotions, or train the matching model using an impractical multi-stage pipeline. In this paper, we study end-to-end matching between image and music based on emotions in the continuous valence-arousal (VA) space. First, we construct a large-scale dataset, termed Image-Music-Emotion-Matching-Net (IMEMNet), with over 140K image-music pairs. Second, we propose cross-modal deep continuous metric learning (CDCML) to learn a shared latent embedding space which preserves the cross-modal similarity relationship in the continuous matching space. Finally, we refine the embedding space by further preserving the single-modal emotion relationship in the VA spaces of both images and music. The metric learning in the embedding space and task regression in the label space are jointly optimized for both cross-modal matching and single-modal VA prediction. The extensive experiments conducted on IMEMNet demonstrate the superiority of CDCML for emotion-based image and music matching as compared to the state-of-the-art approaches. Sicheng Zhao, Yaxian Li, Xingxu Yao, Weizhi Nie, Pengfei Xu 0013, Jufeng Yang, Kurt Keutzer |
ACM Multimedia | 1 |
| 2020 | IExpressNet: Facial Expression Recognition with Incremental ClassesabstractExisting methods on facial expression recognition (FER) are mainly trained in the setting when all expression classes are fixed in advance. However, in real applications, expression classes are becoming increasingly fine-grained and incremental. To deal with sequential expression classes, we can fine-tune or re-train these models, but this often results in poor performance or large computing resources consumption. To address these problems, we develop an Incremental Facial Expression Recognition Network (IExpressNet), which can learn a competitive multi-class classifier at any time with a lower requirement of computing resources. Specifically, IExpressNet consists of two novel components. First, we construct an exemplar set by dynamically selecting representative samples from old expression classes. Then, the exemplar set and new expression classes samples constitute the training set. Second, we design a novel center-expression-distilled loss. As for facial expression in the wild, center-expression-distilled loss enhances the discriminative power of the deeply learned features and prevents catastrophic forgetting. Extensive experiments are conducted on two large-scale FER datasets in the wild, RAF-DB and AffectNet. The results demonstrate the superiority of the proposed method as compared to state-of-the-art incremental learning approaches. Bingjun Luo, Sicheng Zhao, Shihui Ying, Xibin Zhao, Yue Gao 0002 |
ACM Multimedia | 3 |
| 2020 | Actionness-pooled Deep-convolutional Descriptor for fine-grained action recognition
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Sicheng Zhao, Wei Yu 0004 |
Neurocomputing | 5 |
| 2020 | TVENet: Temporal variance embedding network for fine-grained action representation
Tingting Han 0003, Hongxun Yao, Wenlong Xie, Xiaoshuai Sun, Sicheng Zhao, Jun Yu 0002 |
Pattern Recognit. | 5 |
| 2020 | Discrete Probability Distribution Prediction of Image Emotions with Shared Sparse LearningabstractComputationally modelling the affective content of images has been extensively studied recently because of its wide applications in entertainment, advertisement, and education. Significant progress has been made on designing discriminative features to bridge the affective gap. Assuming that viewers can reach a consensus on the emotion of images, most existing works focused on assigning the dominant emotion category or the average dimension values to an image. However, the image emotions perceived by viewers are subjective by nature with the influence of personal and situational factors. In this paper, we propose a novel machine learning approach that characterizes the categorical image emotions as a discrete probability distribution (DPD). To associate emotion with the visual features extracted from images, we present shared sparse learning to learn the combination coefficients, with which the DPD of an unseen image is predicted by linearly combining the DPDs of the training images. Furthermore, we extend our method to the setup where multi-features are available and learn the optimal weights for each feature to reflect the importance of different features. Extensive experiments are carried out on Abstract, Emotion6 and IESN datasets and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches. Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Xin Zhao 0020, Youbao Tang, Jungong Han, Hongxun Yao, Qingming Huang |
IEEE Trans. Affect. Comput. | 1 |
| 2020 | Personality-Assisted Multi-Task Learning for Generic and Personalized Image Aesthetics AssessmentabstractTraditional image aesthetics assessment (IAA) approaches mainly predict the average aesthetic score of an image. However, people tend to have different tastes on image aesthetics, which is mainly determined by their subjective preferences. As an important subjective trait, personality is believed to be a key factor in modeling individual's subjective preference. In this paper, we present a personality-assisted multi-task deep learning framework for both generic and personalized image aesthetics assessment. The proposed framework comprises two stages. In the first stage, a multi-task learning network with shared weights is proposed to predict the aesthetics distribution of an image and Big-Five (BF) personality traits of people who like the image. The generic aesthetics score of the image can be generated based on the predicted aesthetics distribution. In order to capture the common representation of generic image aesthetics and people's personality traits, a Siamese network is trained using aesthetics data and personality data jointly. In the second stage, based on the predicted personality traits and generic aesthetics of an image, an inter-task fusion is introduced to generate individual's personalized aesthetic scores on the image. The performance of the proposed method is evaluated using two public image aesthetics databases. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts in both generic and personalized IAA tasks. Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Weisi Lin |
IEEE Trans. Image Process. | 3 |
| 2020 | ACMNet: Adaptive Confidence Matching Network for Human Behavior Analysis via Cross-modal RetrievalabstractCross-modality human behavior analysis has attracted much attention from both academia and industry. In this article, we focus on the cross-modality image-text retrieval problem for human behavior analysis, which can learn a common latent space for cross-modality data and thus benefit the understanding of human behavior with data from different modalities. Existing state-of-the-art cross-modality image-text retrieval models tend to be fine-grained region-word matching approaches, where they begin with measuring similarities for each image region or text word followed by aggregating them to estimate the global image-text similarity. However, it is observed that such fine-grained approaches often encounter the similarity bias problem, because they only consider matched text words for an image region or matched image regions for a text word for similarity calculation, but they totally ignore unmatched words/regions, which might still be salient enough to affect the global image-text similarity. In this article, we propose an Adaptive Confidence Matching Network (ACMNet), which is also a fine-grained matching approach, to effectively deal with such a similarity bias. Apart from calculating the local similarity for each region(/word) with its matched words(/regions), ACMNet also introduces a confidence score for the local similarity by leveraging the global text(/image) information, which is expected to help measure the semantic relatedness of the region(/word) to the whole text(/image). Moreover, ACMNet also incorporates the confidence scores together with the local similarities in estimating the global image-text similarity. To verify the effectiveness of ACMNet, we conduct extensive experiments and make comparisons with state-of-the-art methods on two benchmark datasets, i.e., Flickr30k and MS COCO. Experimental results show that the proposed ACMNet can outperform the state-of-the-art methods by a clear margin, which well demonstrates the effectiveness of the proposed ACMNet in human behavior analysis and the reasonableness of tackling the mentioned similarity bias issue. Hui Chen 0013, Guiguang Ding, Zijia Lin, Sicheng Zhao, Xiaopeng Gu, Wenyuan Xu 0001, Jungong Han |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | On Parallel and Hitless vSDN ReconfigurationabstractThe symbiosis of network virtualization and software-defined networking (SDN) enables an infrastructure provider (InP) to build various virtual software defined networks (vSDNs) over a shared substrate network (SNT). To handle a dynamic network environment, the InP may need to reconfigure the mapping schemes of vSDNs for a variety of reasons. Although previous studies have addressed how to calculate the new virtual network embedding (VNE) schemes for vSDN reconfiguration under different objectives, the transition to migrate vSDNs from their original VNE schemes to new ones is still under-explored. Hence, this article studies how to realize parallel and hitless vSDN reconfiguration, by leveraging the “makebefore-break” scenario. We come up with a generic solution to optimize the transition to remap vSDNs to new VNE schemes, such that the remappings can be done in the parallel, hitless and resource-efficient manner, as long as the new VNE schemes are feasible. More specifically, our proposal is the multi-stage parallel vSDN reconfiguration based on maximal connected reconfigurable subgraph (MCRSG). To ensure the efficiency of our proposal, we formulate the optimization for selecting MCRSGs to reconfigure in each stage, and prove the NP-hardness of the problem. Then, we design an approximation algorithm based on Lagrangian relaxation to solve it time-efficiently. Extensive simulations verify that the proposed algorithm can obtain nearoptimal solutions quickly. In addition to the algorithmic study, we also realize our multi-stage parallel vSDN reconfiguration in a practical NVH system, and demonstrate its performance in a real network testbed. Our experimental study identifies in what condition losing of packets during remapping would be inevitable, studies the tradeoff between reconfiguration latency and packet loss rate, and reveal an empirical method to adjust key parameters of our NVH system, for adapting to various network environments. Sicheng Zhao, Zuqing Zhu |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Dual-View Ranking with Hardness Assessment for Zero-Shot LearningabstractZero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL. Guiguang Ding, Jungong Han, Xiaohan Ding, Sicheng Zhao, Zheng Wang 0001, Chenggang Yan 0001, Qionghai Dai |
AAAI | 5 |
| 2019 | A Neural Multi-Task Learning Framework to Jointly Model Medical Named Entity Recognition and NormalizationabstractState-of-the-art studies have demonstrated the superiority of joint modeling over pipeline implementation for medical named entity recognition and normalization due to the mutual benefits between the two processes. To exploit these benefits in a more sophisticated way, we propose a novel deep neural multi-task learning framework with explicit feedback strategies to jointly model recognition and normalization. On one hand, our method benefits from the general representations of both tasks provided by multi-task learning. On the other hand, our method successfully converts hierarchical tasks into a parallel multi-task setting while maintaining the mutual supports between tasks. Both of these aspects improve the model performance. Experimental results demonstrate that our method performs significantly better than state-of-theart approaches on two publicly available medical literature datasets. Sendong Zhao, Ting Liu 0001, Sicheng Zhao, Fei Wang 0001 |
AAAI | 3 |
| 2019 | CycleEmotionGAN: Emotional Semantic Consistency Preserved CycleGAN for Adapting Image EmotionsabstractDeep neural networks excel at learning from large-scale labeled training data, but cannot well generalize the learned knowledge to new domains or datasets. Domain adaptation studies how to transfer models trained on one labeled source domain to another sparsely labeled or unlabeled target domain. In this paper, we investigate the unsupervised domain adaptation (UDA) problem in image emotion classification. Specifically, we develop a novel cycle-consistent adversarial model, termed CycleEmotionGAN, by enforcing emotional semantic consistency while adapting images cycleconsistently. By alternately optimizing the CycleGAN loss, the emotional semantic consistency loss, and the target classification loss, CycleEmotionGAN can adapt source domain images to have similar distributions to the target domain without using aligned image pairs. Simultaneously, the annotation information of the source images is preserved. Extensive experiments are conducted on the ArtPhoto and FI datasets, and the results demonstrate that CycleEmotionGAN significantly outperforms the state-of-the-art UDA approaches. Sicheng Zhao, Chuang Lin 0003, Pengfei Xu 0013, Sendong Zhao, Ravi Krishna, Guiguang Ding, Kurt Keutzer |
AAAI | 1 |
| 2019 | On Application-Aware and On-Demand Service Composition in Heterogenous NFV EnvironmentsabstractIn this work, we try to further enhance the flexibility and cost-effectiveness of network function virtualization (NFV) by considering virtual network function service chaining (vNF-SC) in a heterogeneous NFV environment that can instantiate vNFs on virtual machines (VMs), docker containers, and SmartNICs. Specifically, we first lay out the network model, build a real network testbed that supports vNF deployment on kernel-based VMs, docker containers, and commercial SmartNICs, and then conduct experiments to measure the throughput and latency of traffic processing and memory usage of four types of vNFs implemented on them. Next, based on the measurement results, we formulate an integer linear programming (ILP) model to optimize the application-aware vNF-SC provisioning in the heterogenous NFV environment. Finally, we design and perform two experiments to demonstrate that our heterogeneous NFV environment can combine the advantages of VM/docker container/SmartNIC to provide enhanced flexibility for realizing on-demand and application-aware vNF-SC composition. Kai Han 0003, Lipei Liang, Sicheng Zhao, Zuqing Zhu |
GLOBECOM | 5 |
| 2019 | Attention-Aware Polarity Sensitive Embedding for Affective Image RetrievalabstractImages play a crucial role for people to express their opinions online due to the increasing popularity of social networks. While an affective image retrieval system is useful for obtaining visual contents with desired emotions from a massive repository, the abstract and subjective characteristics make the task challenging. To address the problem, this paper introduces an Attention-aware Polarity Sensitive Embedding (APSE) network to learn affective representations in an end-to-end manner. First, to automatically discover and model the informative regions of interest, we develop a hierarchical attention mechanism, in which both polarity- and emotion-specific attended representations are aggregated for discriminative feature embedding. Second, we present a weighted emotion-pair loss to take the inter- and intra-polarity relationships of the emotional labels into consideration. Guided by attention module, we weight the sample pairs adaptively which further improves the performance of feature embedding. Extensive experiments on four popular benchmark datasets show that the proposed method performs favorably against the state-of-the-art approaches. Xingxu Yao, Dongyu She, Sicheng Zhao, Jie Liang 0007, Yukun Lai, Jufeng Yang |
ICCV | 3 |
| 2019 | Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization Without Accessing Target Domain DataabstractWe propose to harness the potential of simulation for semantic segmentation of real-world self-driving scenes in a domain generalization fashion. The segmentation network is trained without any information about target domains and tested on the unseen target domains. To this end, we propose a new approach of domain randomization and pyramid consistency to learn a model with high generalizability. First, we propose to randomize the synthetic images with styles of real images in terms of visual appearances using auxiliary datasets, in order to effectively learn domain-invariant representations. Second, we further enforce pyramid consistency across different "stylized" images and within an image, in order to learn domain-invariant and scale-invariant features, respectively. Extensive experiments are conducted on generalization from GTA and SYNTHIA to Cityscapes, BDDS, and Mapillary; and our method achieves superior results over the state-of-the-art techniques. Remarkably, our generalization results are on par with or even better than those obtained by state-of-the-art simulation-to-real domain adaptation methods, which access the target domain data at training time. Xiangyu Yue 0001, Yang Zhang 0035, Sicheng Zhao, Alberto L. Sangiovanni-Vincentelli, Kurt Keutzer, Boqing Gong |
ICCV | 3 |
| 2019 | Zero-Shot Emotion Recognition via Affective Structural EmbeddingabstractImage emotion recognition attracts much attention in recent years due to its wide applications. It aims to classify the emotional response of humans, where candidate emotion categories are generally defined by specific psychological theories, such as Ekman's six basic emotions. However, with the development of psychological theories, emotion categories become increasingly diverse, fine-grained, and difficult to collect samples. In this paper, we investigate zero-shot learning (ZSL) problem in the emotion recognition task, which tries to recognize the new unseen emotions. Specifically, we propose a novel affective-structural embedding framework, utilizing mid-level semantic representation, i.e., adjective-noun pairs (ANP) features, to construct an affective embedding space. By doing this, the learned intermediate space can narrow the semantic gap between low-level visual and high-level semantic features. In addition, we introduce an affective adversarial constraint to retain the discriminative capacity of visual features and the affective structural information of semantic features during training process. Our method is evaluated on five widely used affective datasets and the perimental results show the proposed algorithm outperforms the state-of-the-art approaches. Chi Zhan, Dongyu She, Sicheng Zhao, Ming-Ming Cheng, Jufeng Yang |
ICCV | 3 |
| 2019 | Personality Driven Multi-task Learning for Image Aesthetic AssessmentabstractWith the prevalence of convolutional neural networks (CNNs), assessing the aesthetics of an image has gained great advances recently. Individual users often have different aesthetic preferences on images, which we believe are mainly affected by their personality traits. However, most of the current aesthetics models predict a generic aesthetic score based on handcrafted and/or learned feature representations, which are unified and thus cannot reflect the individual differences during image aesthetic rating. In this paper, we propose an end-to-end personality driven multi-task deep learning model to address this problem. Firstly, both image aesthetics and personality traits are learned from the proposed multi-task model. Then the personality features are employed to modulate the aesthetics features, producing the optimal generic image aesthetics scores. The experimental results on two public databases show that the proposed method is superior to the state-of-the-art approaches. Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Allen Tan |
ICME | 3 |
| 2019 | SqueezeSegV2: Improved Model Structure and Unsupervised Domain Adaptation for Road-Object Segmentation from a LiDAR Point CloudabstractEarlier work demonstrates the promise of deep-learning-based approaches for point cloud segmentation; however, these approaches need to be improved to be practically useful. To this end, we introduce a new model SqueezeSegV2. With an improved model structure, SqueezeSetV2 is more robust against dropout noises in LiDAR point cloud and therefore achieves significant accuracy improvement. Training models for point cloud segmentation requires large amounts of labeled data, which is expensive to obtain. To sidestep the cost of data collection and annotation, simulators such as GTA-V can be used to create unlimited amounts of labeled, synthetic data. However, due to domain shift, models trained on synthetic data often do not generalize well to the real world. Existing domain-adaptation methods mainly focus on images and most of them cannot be directly applied to point clouds. We address this problem with a domain-adaptation training pipeline consisting of three major components: 1) learned intensity rendering, 2) geodesic correlation alignment, and 3) progressive domain calibration. When trained on real data, our new model exhibits segmentation accuracy improvements of 6.0-8.6% over the original SqueezeSeg. When training our new model on synthetic data using the proposed domain adaptation pipeline, we nearly double test accuracy on real-world data, from 29.0% to 57.4%. Our source code and synthetic dataset are open sourced. https://github.com/xuanyuzhou98/SqueezeSegV2. Bichen Wu, Xuanyu Zhou, Sicheng Zhao, Xiangyu Yue 0001, Kurt Keutzer |
ICRA | 3 |
| 2019 | Cross-Modal Image-Text Retrieval with Semantic ConsistencyabstractCross-modal image-text retrieval has been a long-standing challenge in the multimedia community. Existing methods explore various complicated embedding spaces to assess the semantic similarity between a given image-text pair, but consider no/little about the consistency across them. To remedy this situation, we introduce the idea of semantic consistency for learning various embedding spaces jointly. Specifically, similar to the previous works, we start by constructing two different embedding spaces, namely the image-grounded embedding space and the text-grounded embedding space. However, instead of learning these two embedding spaces separately, we incorporate a semantic consistency constraint in the common ranking objective function such that both embedding spaces can be learned simultaneously and benefit from each other to gain performance improvement. We conduct extensive experiments on three benchmark datasets, \ie Flickr8k, Flickr30k and MS COCO. Results show that our model outperforms the state-of-the-art models on all three datasets, which can well demonstrate the effectiveness and superiority of the introduction of semantic consistency. Our source code is released at: \urlhttps://github.com/HuiChen24/SemanticConsistency. Hui Chen 0013, Guiguang Ding, Zijia Lin, Sicheng Zhao, Jungong Han |
ACM Multimedia | 4 |
| 2019 | PDANet: Polarity-consistent Deep Attention Network for Fine-grained Visual Emotion RegressionabstractExisting methods on visual emotion analysis mainly focus on coarse-grained emotion classification, i.e. assigning an image with a dominant discrete emotion category. However, these methods cannot well reflect the complexity and subtlety of emotions. In this paper, we study the fine-grained regression problem of visual emotions based on convolutional neural networks (CNNs). Specifically, we develop a Polarity-consistent Deep Attention Network (PDANet), a novel network architecture that integrates attention into a CNN with an emotion polarity constraint. First, we propose to incorporate both spatial and channel-wise attentions into a CNN for visual emotion regression, which jointly considers the local spatial connectivity patterns along each channel and the interdependency between different channels. Second, we design a novel regression loss, i.e. polarity-consistent regression (PCR) loss, based on the weakly supervised emotion polarity to guide the attention generation. By optimizing the PCR loss, PDANet can generate a polarity preserved attention map and thus improve the emotion regression performance. Extensive experiments are conducted on the IAPS, NAPS, and EMOTIC datasets, and the results demonstrate that the proposed PDANet outperforms the state-of-the-art approaches by a large margin for fine-grained visual emotion regression. Our source code is released at: https://github.com/ZizhouJia/PDANet. Sicheng Zhao, Zizhou Jia, Hui Chen 0013, Leida Li, Guiguang Ding, Kurt Keutzer |
ACM Multimedia | 1 |
| 2019 | Multi-source Domain Adaptation for Semantic SegmentationabstractSimulation-to-real domain adaptation for semantic segmentation has been actively studied for various applications such as autonomous driving. Existing methods mainly focus on a single-source setting, which cannot easily handle a more practical scenario of multiple sources with different distributions. In this paper, we propose to investigate multi-source domain adaptation for semantic segmentation. Specifically, we design a novel framework, termed Multi-source Adversarial Domain Aggregation Network (MADAN), which can be trained in an end-to-end manner. First, we generate an adapted domain for each source with dynamic semantic consistency while aligning at the pixel-level cycle-consistently towards the target. Second, we propose sub-domain aggregation discriminator and cross-domain cycle discriminator to make different adapted domains more closely aggregated. Finally, feature-level alignment is performed between the aggregated domain and target domain while training the segmentation network. Extensive experiments from synthetic GTA and SYNTHIA to real Cityscapes and BDDS datasets demonstrate that the proposed MADAN model outperforms state-of-the-art approaches. Our source code is released at: https://github.com/Luodian/MADAN. Sicheng Zhao, Bo Li 0080, Xiangyu Yue 0001, Pengfei Xu 0013, Runbo Hu, Kurt Keutzer |
NeurIPS | 1 |
| 2019 | Action recognition with multi-scale trajectory-pooled 3D convolutional descriptors
Xiusheng Lu, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Shengping Zhang |
Multim. Tools Appl. | 3 |
| 2019 | An improved ridge regression algorithm and its application in predicting TV ratings
Sicheng Zhao, Xiuping Wu, Yun Zhai |
Multim. Tools Appl. | 2 |
| 2019 | Learning Descriptors With Cube Loss for View-Based 3-D Object Retrievalabstract3-D object retrieval has been a hot research topic in recent years. Within such a field, view-based approaches are attracting increasing attention because of the flexibility of data representation as well as the reported state-of-the-art performance. One of the most important issues related to view-based 3-D object retrieval is how to learn embedding features that are discriminative across classes while being compactly distributed within each class. In this paper, we analyze the difference between the two tasks of classification and retrieval, and propose a novel way to learn a view-pooling feature via a triplet network. In addition, we propose a new loss, named cube loss, which is able to sample a number of triplets equal to the cube of the samples in a batch. With the new loss, both hard-negative and hard-positive pairs can be effectively investigated. The experimental results on the ModelNet benchmark demonstrate that the proposed method achieves superior performance compared to state-of-the-art approaches. Dong Wang 0030, Hongxun Yao, Federico Tombari, Sicheng Zhao, Bin Wang 0032, Hong Liu 0002 |
IEEE Trans. Multim. | 4 |
| 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event RepresentationabstractRepresentation of videos is essential since it conveys an understanding of video content and enables many higher level tasks to be tackled efficiently. However, it is challenging to propose a rational representation for complex event videos, as most video information is either noisy or redundant. In this paper, we propose a compact event representation method that can concisely describe the inner modes of events. We deem that an optimal event representation scheme should reflect the long-term and high-level visual semantics (visual topics) of events, so different from previous frame-level video semantics representation methods and concept-based video representation methods, we investigate the problem from the perspective of segment-level video representations. We then present three appealing properties of segment-level visual semantics. Based on the observation, we propose different algorithms that rely on a novel deep-visual-word-based video encoding method to discover latent discriminative patterns of events. Finally, our multi-mode event representation is obtained by concatenating the discovered patterns as inner modes. We adopt our event representation for representative event parts mining, which can highlight the visual topics of events and remarkably prune the raw videos. We validate our event representation method based on complex event detection task. Experimental results on two standard benchmarking datasets, MED11 and CCV Dataset, show that the proposed method can significantly outperform the state-of-the-art approaches. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Tingting Han 0003, Sicheng Zhao, Tat-Seng Chua |
IEEE Trans. Multim. | 5 |
| 2019 | Proactive and Hitless vSDN Reconfiguration to Balance Substrate TCAM Utilization: From Algorithm Design to System PrototypeabstractThe combination of network virtualization and software-defined networking enables an infrastructure provider to create software-defined virtual networks (vSDNs) over a shared substrate network (SNT), for supporting new network services more timely and cost-effectively. Meanwhile, as both the services and traffic in the Internet are becoming more and more dynamic, how to properly maintain vSDNs in a dynamic network environment exhibits increasing importance but still has not been fully explored. In this paper, we conduct a study on how to realize proactive and hitless vSDN reconfiguration to balance the utilization of ternary content-addressable memory (TCAM) in a dynamic SNT. Specifically, we consider both algorithm design and system prototyping. From the algorithmic perspective, we try to solve the problems of “what to reconfigure” and “how to reconfigure”. A selection algorithm is designed to proactively choose the virtual switches (vSWs) that should be migrated to other substrate switches for balancing TCAM utilization, i.e., solving what to reconfigure. Then, for the problem of how to reconfigure, i.e., where to re-map the selected vSWs and the virtual links connecting to them, we formulate a mixed integer linear programming model to solve it exactly, and design two heuristics to improve time efficiency. Next, we move to the system part, implement the proposed algorithms in our protocol-oblivious forwarding enabled network virtualization hypervisor system, and conduct experiments to demonstrate proactive and hitless vSDN reconfiguration. The experimental results indicate that our proposal does make vSDN reconfiguration transparent to the vSDNs' virtual controllers and proactive, and when reconfiguring a vSDN with live traffic, it achieves hitless operations without traffic disruption. Sicheng Zhao, Deyun Li, Kai Han 0003, Zuqing Zhu |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2019 | Personalized Emotion Recognition by Personality-Aware High-Order Learning of Physiological SignalsabstractDue to the subjective responses of different subjects to physical stimuli, emotion recognition methodologies from physiological signals are increasingly becoming personalized. Existing works mainly focused on modeling the involved physiological corpus of each subject, without considering the psychological factors, such as interest and personality. The latent correlation among different subjects has also been rarely examined. In this article, we propose to investigate the influence of personality on emotional behavior in a hypergraph learning framework. Assuming that each vertex is a compound tuple (subject, stimuli), multi-modal hypergraphs can be constructed based on the personality correlation among different subjects and on the physiological correlation among corresponding stimuli. To reveal the different importance of vertices, hyperedges, and modalities, we learn the weights for each of them. As the hypergraphs connect different subjects on the compound vertices, the emotions of multiple subjects can be simultaneously recognized. In this way, the constructed hypergraphs are vertex-weighted multi-modal multi-task ones. The estimated factors, referred to as emotion relevance, are employed for emotion recognition. We carry out extensive experiments on the ASCERTAIN dataset and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art emotion recognition approaches. Sicheng Zhao, Amir Gholami, Guiguang Ding, Yue Gao 0002, Jungong Han, Kurt Keutzer |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2018 | Temporal-Difference Learning With Sampling Baseline for Image CaptioningabstractThe existing methods for image captioning usually train the language model under the cross entropy loss, which results in the exposure bias and inconsistency of evaluation metric. Recent research has shown these two issues can be well addressed by policy gradient method in reinforcement learning domain attributable to its unique capability of directly optimizing the discrete and non-differentiable evaluation metric. In this paper, we utilize reinforcement learning method to train the image captioning model. Specifically, we train our image captioning model to maximize the overall reward of the sentences by adopting the temporal-difference (TD) learning method, which takes the correlation between temporally successive actions into account. In this way, we assign different values to different words in one sampled sentence by a discounted coefficient when back-propagating the gradient with the REINFORCE algorithm, enabling the correlation between actions to be learned. Besides, instead of estimating a "baseline" to normalize the rewards with another network, we utilize the reward of another Monte-Carlo sample as the "baseline" to avoid high variance. We show that our proposed method can improve the quality of generated captions and outperforms the state-of-the-art methods on the benchmark dataset MS COCO in terms of seven evaluation metrics. Hui Chen 0013, Guiguang Ding, Sicheng Zhao, Jungong Han |
AAAI | 3 |
| 2018 | Shift: A Zero FLOP, Zero Parameter Alternative to Spatial ConvolutionsabstractNeural networks rely on convolutions to aggregate spatial information. However, spatial convolutions are expensive in terms of model size and computation, both of which grow quadratically with respect to kernel size. In this paper, we present a parameter-free, FLOP-free "shift" operation as an alternative to spatial convolutions. We fuse shifts and point-wise convolutions to construct end-to-end trainable shift-based modules, with a hyperparameter characterizing the tradeoff between accuracy and efficiency. To demonstrate the operation's efficacy, we replace ResNet's 3x3 convolutions with shift-based modules for improved CIFAR10 and CIFAR100 accuracy using 60% fewer parameters; we additionally demonstrate the operation's resilience to parameter reduction on ImageNet, outperforming ResNet family members. We finally show the shift operation's applicability across domains, achieving strong performance with fewer parameters on image classification, face verification and style transfer. Bichen Wu, Alvin Wan, Xiangyu Yue 0001, Peter H. Jin, Sicheng Zhao, Noah Golmant, Amir Gholami, Joseph Gonzalez 0001, Kurt Keutzer |
CVPR | 5 |
| 2018 | Virtualization of Table Resources in Programmable Data Plane with Global ConsiderationabstractIn this work, we try to address the problem of memory fragmentation in ternary content addressable memory (TCAM) in programmable data plane (PDP), by designing and implementing a novel network hypervisor for PDP, namely, TPVX. TPVX realizes the virtualization of table resources in PDP with global consideration, i.e., when mapping tenant flow tables to physical switches, TPVX considers their table sizes and the pre-formatted sub-tables in the physical network to improve TCAM utilization and avoid memory fragmentation. Our experimental results verify that with TPVX, the utilization of the table resources in PDP can be improved dramatically and the extra processing latency due to the newly-introduced overheads can be maintained well simultaneously. Yuhan Xue, Shengru Li, Kai Han 0003, Sicheng Zhao, Huibai Huang, Shui Yu 0001, Zuqing Zhu |
GLOBECOM | 4 |
| 2018 | Make Big Data Applications More Reliable: Hitless vSDN Migration to Avoid TCAM DepletionabstractWith network virtualization, an infrastructure provider can create virtual software-defined networks (vSDNs) over a shared substrate network and lease them to service providers (SPs). This enables the SPs to run their Big Data applications in a short time-to-market, flexible and cost-effective way. However, in a dynamic network, both the instances of vSDNs and the traffic in each vSDN can change over time, which would degrade the optimality of the embedding schemes of vSDNs and even cause serious reliability issues. Therefore, in this work, we extend our protocol-oblivious forwarding (POF) based network virtualization hypervisor (NVH) system (i.e., PVX) to realize hitless vSDN migration to avoid ternary content addressable memory (TCAM) depletion. Specifically, we design the PVX system to realize the vSDN migration that is transparent to the controllers of vSDNs and would cause zero or very few packet losses. The proposed PVX is then implemented in a real network testbed, and we conduct experiments to verify its effectiveness. Sicheng Zhao, Kai Han 0003, Zuqing Zhu |
ICC | 1 |
| 2018 | Show, Observe and Tell: Attribute-driven Attention Model for Image CaptioningabstractDespite the fact that attribute-based approaches and attention-based approaches have been proven to be effective in image captioning, most attribute-based approaches simply predict attributes independently without taking the co-occurrence dependencies among attributes into account. Besides, most attention-based captioning models directly leverage the feature map extracted from CNN, in which many features may be redundant in relation to the image content. In this paper, we focus on training a good attribute-inference model via the recurrent neural network (RNN) for image captioning, where the co-occurrence dependencies among attributes can be maintained. The uniqueness of our inference model lies in the usage of a RNN with the visual attention mechanism to \textit{observe} the image before generating captions. Additionally, it is noticed that compact and attribute-driven features will be more useful for the attention-based captioning model. To this end, we extract the context feature for each attribute, and guide the captioning model adaptively attend to these context features. We verify the effectiveness and superiority of the proposed approach over the other captioning approaches by conducting massive experiments and comparisons on MS COCO image captioning dataset. Hui Chen 0013, Guiguang Ding, Zijia Lin, Sicheng Zhao, Jungong Han |
IJCAI | 4 |
| 2018 | Implicit Non-linear Similarity Scoring for Recognizing Unseen ClassesabstractRecognizing unseen classes is an important task for real-world applications, due to: 1) it is common that some classes in reality have no labeled image exemplar for training; and 2) novel classes emerge rapidly. Recently, to address this task many zero-shot learning (ZSL) approaches have been proposed where explicit linear scores, like inner product score, are employed to measure the similarity between a class and an image. We argue that explicit linear scoring (ELS) seems too weak to capture complicated image-class correspondence. We propose a simple yet effective framework, called Implicit Non-linear Similarity Scoring (ICINESS). In particular, we train a scoring network which uses image and class features as input, fuses them by hidden layers, and outputs the similarity. Based on the universal approximation theorem, it can approximate the true similarity function between images and classes if a proper structure is used in an implicit non-linear way, which is more flexible and powerful. With ICINESS framework, we implement ZSL algorithms by shallow and deep networks, which yield consistently superior results. Guiguang Ding, Jungong Han, Sicheng Zhao, Bin Wang 0021 |
IJCAI | 4 |
| 2018 | Affective Image Content Analysis: A Comprehensive SurveyabstractImages can convey rich semantics and induce strong emotions in viewers. Recently, with the explosive growth of visual data, extensive research efforts have been dedicated to affective image content analysis (AICA). In this paper, we review the state-of-the-art methods comprehensively with respect to two main challenges -- affective gap and perception subjectivity. We begin with an introduction to the key emotion representation models that have been widely employed in AICA. Available existing datasets for performing evaluation are briefly described. We then summarize and compare the representative approaches on emotion feature extraction, personalized emotion prediction, and emotion distribution learning. Finally, we discuss some future research directions. Sicheng Zhao, Guiguang Ding, Qingming Huang, Tat-Seng Chua, Björn W. Schuller, Kurt Keutzer |
IJCAI | 1 |
| 2018 | Personality-Aware Personalized Emotion Recognition from Physiological SignalsabstractEmotion recognition methodologies from physiological signals are increasingly becoming personalized, due to the subjective responses of different subjects to physical stimuli. Existing works mainly focused on modelling the involved physiological corpus of each subject, without considering the psychological factors. The latent correlation among different subjects has also been rarely examined. We propose to investigate the influence of personality on emotional behavior in a hypergraph learning framework. Assuming that each vertex is a compound tuple (subject, stimuli), multi-modal hypergraphs can be constructed based on the personality correlation among different subjects and on the physiological correlation among corresponding stimuli. To reveal the different importance of vertices, hyperedges, and modalities, we assign each of them with weights. The emotion relevance learned on the vertex-weighted multi-modal multi-task hypergraphs is employed for emotion recognition. We carry out extensive experiments on the ASCERTAIN dataset and the results demonstrate the superiority of the proposed method. Sicheng Zhao, Guiguang Ding, Jungong Han, Yue Gao 0002 |
IJCAI | 1 |
| 2018 | ASMMC-MMAC 2018: The Joint Workshop of 4th the Workshop on Affective Social Multimedia Computing and first Multi-Modal Affective Computing of Large-Scale Multimedia Data WorkshopabstractAffective social multimedia computing is an emergent research topic for both affective computing and multimedia research communities. Social multimedia is fundamentally changing how we communicate, interact, and collaborate with other people in our daily lives. Social multimedia contains much affective information. Effective extraction of affective information from social multimedia can greatly help social multimedia computing (e.g., processing, index, retrieval, and understanding). Besides, with the rapid development of digital photography and social networks, people get used to sharing their lives and expressing their opinions online. As a result, user-generated social media data, including text, images, audios, and videos, grow rapidly, which urgently demands advanced techniques on the management, retrieval, and understanding of these data. Dong-Yan Huang, Sicheng Zhao, Björn W. Schuller, Hongxun Yao, Jianhua Tao 0001, Min Xu 0001, Lei Xie 0001, Qingming Huang |
ACM Multimedia | 2 |
| 2018 | EmotionGAN: Unsupervised Domain Adaptation for Learning Discrete Probability Distributions of Image EmotionsabstractDeep neural networks have performed well on various benchmark vision tasks with large-scale labeled training data; however, such training data is expensive and time-consuming to obtain. Due to domain shift or dataset bias, directly transferring models trained on a large-scale labeled source domain to another sparsely labeled or unlabeled target domain often results in poor performance. In this paper, we consider the domain adaptation problem in image emotion recognition. Specifically, we study how to adapt the discrete probability distributions of image emotions from a source domain to a target domain in an unsupervised manner. We develop a novel adversarial model for emotion distribution learning, termed EmotionGAN, which alternately optimizes the Generative Adversarial Network (GAN) loss, semantic consistency loss, and regression loss. The EmotionGAN model can adapt source domain images such that they appear as if they were drawn from the target domain, while preserving the annotation information. Extensive experiments are conducted on the FlickrLDL and TwitterLDL datasets, and the results demonstrate the superiority of the proposed method as compared to state-of-the-art approaches. Sicheng Zhao, Xin Zhao 0020, Guiguang Ding, Kurt Keutzer |
ACM Multimedia | 1 |
| 2018 | Transmission time analysis for hybrid V2V and V2I communications in multi-lane vehicular networksabstractAlong with the increasing data demands in vehicular networks (VNs), how to satisfy large-amount data transmission under the requirements of low delay and high reliability has gained a lot of attentions. In this paper, we investigate the key performance metrics when vehicles download the required data by hybrid Vehicle-to-Vehicle (V2V) and Vehicle-to-Infrastructure (V2I) patterns in multi-lane scenario. Analytical results are derived on the transmission time and the number of handovers. In addition, in order to realize the minimum number of handovers in V2V transmission process, we put forward a Maximum Single Download Time (MSDT) handover strategy. Finally, simulations and discussions are presented to validate the performance of the proposed strategy under different transmission models, and to prove its advantage by comparing with some other typical strategies. Sicheng Zhao, Xuefei Zhang 0003, Yan Han 0006, Xiaofeng Tao 0001 |
WCNC | 1 |
| 2018 | Rediscover flowers structurally
Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004 |
Multim. Tools Appl. | 4 |
| 2018 | Exploring part-aware segmentation for fine-grained visual categorization
Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Yanhao Zhang 0001 |
Multim. Tools Appl. | 4 |
| 2018 | Off-the-shelf CNN features for 3D object retrieval
Dong Wang 0030, Bin Wang 0032, Sicheng Zhao, Hongxun Yao, Hong Liu 0002 |
Multim. Tools Appl. | 3 |
| 2018 | Guest Editorial: Large-scale 3D Multimedia Analysis and Applications
Sicheng Zhao, Jun Zhang 0018, Yi Zhen |
Multim. Tools Appl. | 1 |
| 2018 | Evaluating attributed personality traits from scene perception probability
Hancheng Zhu, Leida Li, Sicheng Zhao |
Pattern Recognit. Lett. | 3 |
| 2018 | Event patches: Mining effective parts for event detection and understanding
Wenlong Xie, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Tingting Han 0003 |
Signal Process. | 3 |
| 2018 | Distinctive action sketch for human action recognition
Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Fatih Porikli |
Signal Process. | 4 |
| 2018 | Predicting Personalized Image Emotion Perceptions in Social NetworksabstractImages can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning (RMTHG) is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua |
IEEE Trans. Affect. Comput. | 1 |
| 2018 | Real-Time Multimedia Social Event Detection in MicroblogabstractDetecting events from massive social media data in social networks can facilitate browsing, search, and monitoring of real-time events by corporations, governments, and users. The short, conversational, heterogeneous, and real-time characteristics of social media data bring great challenges for event detection. The existing event detection approaches rely mainly on textual information, while the visual content of microblogs and the intrinsic correlation among the heterogeneous data are scarcely explored. To deal with the above challenges, we propose a novel real-time event detection method by generating an intermediate semantic level from social multimedia data, named microblog clique (MC), which is able to explore the high correlations among different microblogs. Specifically, the proposed method comprises three stages. First, the heterogeneous data in microblogs is formulated in a hypergraph structure. Hypergraph cut is conducted to group the highly correlated microblogs with the same topics as the MCs, which can address the information inadequateness and data sparseness issues. Second, a bipartite graph is constructed based on the generated MCs and the transfer cut partition is performed to detect the events. Finally, for new incoming microblogs, incremental hypergraph is constructed based on the latest MCs to generate new MCs, which are classified by bipartite graph partition into existing events or new ones. Extensive experiments are conducted on the events in the Brand-Social-Net dataset and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches. Sicheng Zhao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua |
IEEE Trans. Cybern. | 1 |
| 2018 | Real-Time Scalable Visual Tracking via Quadrangle Kernelized Correlation FiltersabstractCorrelation filter (CF) has been widely used in tracking tasks due to its simplicity and high efficiency. However, conventional CF-based trackers fail to handle the scale variation that occurs when the targeted object is moving, which is one of the most notable unsolved problems of visual object tracking. In this paper, we propose a scalable visual tracking algorithm based on kernelized correlation filters, referred to as quadrangle kernelized correlation filters (QKCF). Unlike existing complicated scalable trackers that either perform the correlation filtering operation multiple times or extract many candidate windows at various scales, our tracker intends to estimate the scale of the object based on the positions of its four corners, which can be detected using a new Gaussian training output matrix within one filtering process. After obtaining four peak values corresponding to the four corners, we measure the detection confidence of each part response by evaluating its spatial and temporal smoothness. On top of it, a weighted Bayesian inference framework is employed to estimate the final location and size of the bounding box from the response matrix, where the weights are synchronized with the calculated detection likelihoods. Experiments are performed on the OTB-100 data set and 16 benchmark sequences with significant scale variations. The results demonstrate the superiority of the proposed method in terms of both effectiveness and robustness, compared with the state-of-the-art methods. Guiguang Ding, Wenshuo Chen, Sicheng Zhao, Jungong Han, Qiaoyan Liu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Reference Based LSTM for Image CaptioningabstractImage captioning is an important problem in artificial intelligence, related to both computer vision and natural language processing. There are two main problems in existing methods: in the training phase, it is difficult to find which parts of the captions are more essential to the image; in the caption generation phase, the objects or the scenes are sometimes misrecognized. In this paper, we consider the training images as the references and propose a Reference based Long Short Term Memory (R-LSTM) model, aiming to solve these two problems in one goal. When training the model, we assign different weights to different words, which enables the network to better learn the key information of the captions. When generating a caption, the consensus score is utilized to exploit the reference information of neighbor images, which might fix the misrecognition and make the descriptions more natural-sounding. The proposed R-LSTM model outperforms the state-of-the-art approaches on the benchmark dataset MS COCO and obtains top 2 position on 11 of the 14 metrics on the online test server. Minghai Chen, Guiguang Ding, Sicheng Zhao, Hui Chen 0013, Qiang Liu 0016, Jungong Han |
AAAI | 3 |
| 2017 | Leveraging Protocol-Oblivious Forwarding (POF) to Realize NFV-Assisted Mobility ManagementabstractIn the Internet, to enable emerging network services such as e-Health to be delivered to numerous people anytime and anywhere, mobility management plays an important role. In this work, we leverage the protocol-oblivious forwarding (POF) to design a novel network system that can realize highly-efficient mobility management and overcome the drawbacks of existing approaches. Moreover, we explore the forwarding plane programmability provided by POF to realize dynamic virtual network function (vNF) deployment for traffic adaption. We implement our system and conduct experiments to verify that it functions well for NFV-assisted mobility management and achieves reduced handover latency and enhanced QoE. Kai Han 0003, Shengru Li, Shaofei Tang, Huibai Huang, Sicheng Zhao, Zuqing Zhu |
GLOBECOM | 6 |
| 2017 | Approximating Discrete Probability Distribution of Image Emotions by Multi-Modal Features FusionabstractExisting works on image emotion recognition mainly assigned the dominant emotion category or average dimension values to an image based on the assumption that viewers can reach a consensus on the emotion of images. However, the image emotions perceived by viewers are subjective by nature and highly related to the personal and situational factors. On the other hand, image emotions can be conveyed by different features, such as semantics and aesthetics. In this paper, we propose a novel machine learning approach that formulates the categorical image emotions as a discrete probability distribution (DPD). To associate emotions with the extracted visual features, we present a weighted multi-modal shared sparse leaning to learn the combination coefficients, with which the DPD of an unseen image can be predicted by linearly integrating the DPDs of the training images. The representation abilities of different modalities are jointly explored and the optimal weight of each modality is automatically learned. Extensive experiments on three datasets verify the superiority of the proposed method, as compared to the state-of-the-art. Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Jungong Han |
IJCAI | 1 |
| 2017 | Learning Visual Emotion Distributions via Multi-Modal Features FusionabstractCurrent image emotion recognition works mainly classified the images into one dominant emotion category, or regressed the images with average dimension values by assuming that the emotions perceived among different viewers highly accord with each other. However, due to the influence of various personal and situational factors, such as culture background and social interactions, different viewers may react totally different from the emotional perspective to the same image. In this paper, we propose to formulate the image emotion recognition task as a probability distribution learning problem. Motivated by the fact that image emotions can be conveyed through different visual features, such as aesthetics and semantics, we present a novel framework by fusing multi-modal features to tackle this problem. In detail, weighted multi-modal conditional probability neural network (WMMCPNN) is designed as the learning model to associate the visual features with emotion probabilities. By jointly exploring the complementarity and learning the optimal combination coefficients of different modality features, WMMCPNN could effectively utilize the representation ability of each uni-modal feature. We conduct extensive experiments on three publicly available benchmarks and the results demonstrate that the proposed method significantly outperforms the state-of-the-art approaches for emotion distribution prediction. Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Jungong Han |
ACM Multimedia | 1 |
| 2017 | Large-scale image retrieval with Sparse Embedded Hashing
Guiguang Ding, Jile Zhou, Zijia Lin, Sicheng Zhao, Jungong Han |
Neurocomputing | 5 |
| 2017 | Text image deblurring via two-tone prior
Xiaolei Jiang, Hongxun Yao, Sicheng Zhao |
Neurocomputing | 3 |
| 2017 | View-based 3D object retrieval with discriminative views
Dong Wang 0030, Bin Wang 0032, Sicheng Zhao, Hongxun Yao, Hong Liu 0002 |
Neurocomputing | 3 |
| 2017 | Actor identification via mining representative actions
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004, Shengping Zhang |
Neurocomputing | 4 |
| 2017 | Discovering discriminative patches for free-hand sketch analysis
Ying Zheng 0009, Hongxun Yao, Sicheng Zhao, Yasi Wang |
Multim. Syst. | 3 |
| 2017 | Towards more efficient and flexible face image deblurring using robust salient face landmark detection
Yinghao Huang, Hongxun Yao, Sicheng Zhao, Yanhao Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Exploring Coherent Motion Patterns via Structured Trajectory Learning for Crowd Mood ModelingabstractCrowd behavior analysis has recently attracted extensive attention in research. However, the existing research mainly focuses on investigating motion patterns in crowds, while the emotional aspects of crowd behaviors are left unexplored. Analyzing the emotion of crowd behaviors is indeed extremely important, as it uncovers the social moods that are beneficial for video surveillance. In this paper, we propose a novel crowd representation termed crowd mood. Crowd mood is established based upon the discovery that the social emotional hypothesis of crowd behaviors can be revealed by investigating the spacing interactions and the structural levels of motion patterns in crowds. To this end, we first learn the structured trajectories of crowds by particle advection using low-rank approximation with group sparsity constraint, which implicitly characterizes the coherent motion patterns. Second, rich emotional motion features are explicitly extracted and fused by support vector regression to reflect social characteristics. In particular, we construct weighted features in a boosted manner by considering the features' significance. Finally, crowd mood is intuitively presented as affective curves to track the emotion states of the crowd dynamics, which is robust to noise, sensitive to semantic shift, and compact for pattern expressions. Extensive evaluations on crowd video data sets demonstrate that our approach effectively models crowd mood and achieves significantly better results with comparisons to several alternative and state-of-the-art approaches for various tasks, i.e., crowd mood classification, global abnormal mood detection, and crowd emotion matching. Yanhao Zhang 0001, Rongrong Ji, Sicheng Zhao, Qingming Huang, Jiebo Luo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Continuous Probability Distribution Prediction of Image Emotions via Multitask Shared Sparse RegressionabstractPrevious works on image emotion analysis mainly focused on predicting the dominant emotion category or the average dimension values of an image for affective image classification and regression. However, this is often insufficient in various real-world applications, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the continuous probability distribution of image emotions which are represented in dimensional valence-arousal space. We carried out large-scale statistical analysis on the constructed Image-Emotion-Social-Net dataset, on which we observed that the emotion distribution can be well-modeled by a Gaussian mixture model. This model is estimated by an expectation-maximization algorithm with specified initializations. Then, we extract commonly used emotion features at different levels for each image. Finally, we formalize the emotion distribution prediction task as a shared sparse regression (SSR) problem and extend it to multitask settings, named multitask shared sparse regression (MTSSR), to explore the latent information between different prediction tasks. SSR and MTSSR are optimized by iteratively reweighted least squares. Experiments are conducted on the Image-Emotion-Social-Net dataset with comparisons to three alternative baselines. The quantitative results demonstrate the superiority of the proposed method. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Guiguang Ding |
IEEE Trans. Multim. | 1 |
| 2016 | Affective Computing and Applications of Image Emotion PerceptionsabstractImages can convey rich semantics and evoke strong emotions in viewers. The research of my PhD thesis focuses on image emotion computing (IEC), which aims to predict the emotion perceptions of given images. The development of IEC is greatly constrained by two main challenges: affective gap and subjective evaluation. Previous works mainly focused on finding features that can express emotions better to bridge the affective gap, such as elements-of-art based features and shape features. According to the emotion representation models, including categorical emotion states (CES) and dimensional emotion space (DES), three different tasks are traditionally performed on IEC: affective image classification, regression and retrieval. The state-of-the-art methods on the three above tasks are image-centric, focusing on the dominant emotions for the majority of viewers. For my PhD thesis, I plan to answer the following questions: (1) Compared to the low-level elements-of-art based features, can we find some higher level features that are more interpretable and have stronger link to emotions? (2) Are the emotions that are evoked in viewers by an image subjective and different? If they are, how can we tackle the user-centric emotion prediction? (3) For image-centric emotion computing, can we predict the emotion distribution instead of the dominant emotion category? Sicheng Zhao, Hongxun Yao |
AAAI | 1 |
| 2016 | User-Centric Affective Computing of Image Emotion PerceptionsabstractWe propose to predict the personalized emotion perceptions of images for each viewer. Different factors that may influence emotion perceptions, including visual content, social context, temporal evolution, and location influence are jointly investigated via the presented rolling multi-task hypergraph learning. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate the superiority of the proposed method, as compared to state-of-the-art. Sicheng Zhao, Hongxun Yao, Wenlong Xie, Xiaolei Jiang |
AAAI | 1 |
| 2016 | Mining representative actions for actor identificationabstractPrevious works on actor identification mainly focused on static features based on face identification and costume detection, without considering the abundant dynamic information contained in videos. In this paper, we propose a novel method to mine representative actions of each actor, and show the remarkable power of such actions for actor identification task. Videos are firstly divided into shots and represented by BoW based on spatial-temporal features. Then we integrate the prototype theory with SVM to rank the shots and obtain the representative actions. Our method for actor identification combines representative actions with actors' appearance. We validate the method on episodes of the TV series "The Big Bang Theory". The experimental results show that the representative actions are consistent with human judgements and can greatly improve the matching performance as complementary to existing handcrafted static features for actor identification. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Tingting Han 0003 |
ICASSP | 4 |
| 2016 | Crowd video retrieval via deep attribute-embedding graph rankingabstractSince the number of surveillance cameras in public areas increases very fast, massive crowd videos are captured and shared, which brings an urgent need to retrieve these videos efficiently and effectively. However, most recent research on crowd video mainly focused on crowd behavior understanding and abnormal detection. In this study, as the very first attempt, we propose a crowd video retrieval method via deep attribute-embedding graph ranking. Group profiling attributes are capable of reflecting rich crowd patterns in videos. To deeply embed the specific relationship and manifold structure of crowd patterns, we integrate graph ranking, optimized weights learning and deep metric transforming in a unified regularization framework for crowd video retrieval. To sufficiently explore the effects of multiple attributes and their complementation in crowds, we devise several scene-independent visual descriptors specifically for each crowd video. Interpretable descriptors are categorized into different levels and structures as group profiling attributes, according to semantic properties of crowd patterns. Extensive experiments conducted on CUHK crowd dataset demonstrate the effectiveness and superiority of the proposed approach. Yanhao Zhang 0001, Sicheng Zhao, Rongrong Ji, Xiusheng Lu, Hongxun Yao, Qingming Huang |
ICME | 3 |
| 2016 | A VHR scene classification method integrating sparse PCA and saliency computingabstractUnderstanding a scene provided by very high resolution (VHR) satellite imagery has become a more and more challenging problem. In this paper, we propose a new method for scene classification based on saliency computing of patches sampling from the VHR images. Sparse principal component analysis (sPCA) is then adopted to select the corresponding informative salient patches for image scene representation. The proposed technique for selecting informative salient patches is efficient and robust for scene understanding. We conduct experiments on the public UC Merced benchmark dataset, which contains 21 different areal categories with sub-meter resolution. Experimental results demonstrate the effectiveness of the proposed method, as compared with several state-of-the-art methods. Souleyman Chaib, Yanfeng Gu, Hongxun Yao, Sicheng Zhao |
IGARSS | 4 |
| 2016 | Image Emotion ComputingabstractImages can convey rich semantics and induce strong emotions in viewers. My research aims to predict image emotions from different aspects with respect to two main challenges: affective gap and subjective evaluation. To bridge the affective gap, we extract emotion features based on principles-of-art to recognize image-centric dominant emotions. As the emotions that are induced in viewers by an image are highly subjective and different, we propose to predict user-centric personalized emotion perceptions for each viewer and image-centric emotion probability distribution for each image. To tackle the subjective evaluation issue, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Different types of factors may influence personalized image emotion perceptions, including visual content, social context, temporal evolution and location influence. We make an initial attempt to jointly combine them by the proposed rolling multi-task hypergraph learning. Both discrete and continuous emotion distributions are modelled via shared sparse learning. Further, several potential applications based on image emotions are designed and implemented. Sicheng Zhao |
ACM Multimedia | 1 |
| 2016 | Predicting Personalized Emotion Perceptions of Social ImagesabstractImages can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Wenlong Xie, Xiaolei Jiang, Tat-Seng Chua |
ACM Multimedia | 1 |
| 2016 | Exploring Discriminative Views for 3D Object Retrieval
Dong Wang 0030, Bin Wang 0032, Sicheng Zhao, Hongxun Yao, Hong Liu 0002 |
MMM (1) | 3 |
| 2016 | Affective Computing of Image Emotion PerceptionsabstractNo abstract available. Sicheng Zhao |
WSDM | 1 |
| 2016 | Unsupervised discovery of crowd activities by saliency-based clustering
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Yanhao Zhang 0001 |
Neurocomputing | 4 |
| 2016 | Auto-encoder based dimensionality reduction
Yasi Wang, Hongxun Yao, Sicheng Zhao |
Neurocomputing | 3 |
| 2016 | Event causality extraction based on connectives analysis
Sendong Zhao, Ting Liu 0001, Sicheng Zhao, Yiheng Chen, Jian-Yun Nie |
Neurocomputing | 3 |
| 2016 | Logo information recognition in large-scale social media data
Fanglin Wang, Shuhan Qi, Sicheng Zhao |
Multim. Syst. | 4 |
| 2016 | Multi-modal microblog classification via multi-task learning
Sicheng Zhao, Hongxun Yao, Sendong Zhao, Xuesong Jiang, Xiaolei Jiang |
Multim. Tools Appl. | 1 |
| 2015 | Predicting discrete probability distribution of image emotionsabstractMost existing works on affective image classification tried to assign a dominant emotion category to an image. However, this is often insufficient, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the probability distribution of categorical image emotions. Firstly we extract commonly used features of different levels for each image. Then we formulize the emotion distribution prediction as a shared sparse leaning problem, which is optimized by iteratively reweighted least squares. Besides, we introduce three baseline algorithms. Experiments are carried out on a dataset of peer rated abstract paintings and the results demonstrate the superiority of our proposed method, as compared to some state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Xiaolei Jiang, Xiaoshuai Sun |
ICIP | 1 |
| 2015 | Distinctive action sketchabstractRecent developments in the field of image and video processing have led to a renewed interest in sketch correlated research. With that, there have emerged considerable solid evidence which revealed the significance of sketch to us. However, there have been few profound discussions on sketch based action analysis so far. In this paper, we present a framework of converting human actions to commendable sketches and discovering distinctive action sketches. The action sketches should satisfy three characteristics: sketchability, objective-ness and consistency. Primitive sketches are prepared according to the structured forests based fast edge detection. Meanwhile, we take advantage of Struck to accomplish adaptive object tracking in parallel. After that, we propose a method to ensure the spatio-temporal consistency between every sequential action sketches. On completion of previous stages, the process of distinctive mining of action sketches is carried out. The experimental results show that our approach has a promising potential. Ying Zheng 0009, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao |
ICIP | 4 |
| 2015 | "Clustering of Dancelets": Towards Video Recommendation Based on Dance StylesabstractDance is a special and important type of action, composed of abundant and various action elements. However, the recommendation of dance videos on the web are still not well studied. It is hard to realize it in the way of traditional methods using associated texts or static features of video content. In this paper, we study the problem focusing on extraction and representation of action information in dances. We propose to recommend dance videos based on the automatically discovered ``Dance Styles'', which play a significant role in characterizing different types of dances. To bridge the semantic gap of video content and mid-level concept, style, we take advantage of a mid-level action representation method, and extract representative patches as ``Dancelets'', a sort of intermediation between videos and the concepts. Furthermore, we propose to employ Motion Boundaries as saliency priors and sparsely extract patches containing more representative information to generate a set of dancelet candidates. Dancelets are then discovered by Normalized-cut method, which is superior in grouping visually similar patterns into the same clusters. For the fast and effective recommendation, a random forest-based index is built, and the ranking results are derived according to the matching results in all the leaf notes. Extensive experiments validated on the web dance videos demonstrate the effectiveness of the proposed methods for dance style discovery and video recommendation based on styles. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001, Sicheng Zhao, Xiusheng Lu, Yinghao Huang, Wenlong Xie |
ACM Multimedia | 5 |
| 2015 | Predicting Continuous Probability Distribution of Image Emotions in Valence-Arousal SpaceabstractPrevious works on image emotion analysis mainly focused on assigning a dominated emotion category or the average dimension values to an image for affective image classification and regression. However, this is often insufficient in many applications, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the continuous probability distribution of dimensional image emotions represented in valence-arousal space. By the statistical analysis on the constructed Image-Emotion-Social-Net dataset, we represent the emotion distribution as a Gaussian mixture model (GMM), which is estimated by the EM algorithm. Then we extract commonly used features of different levels for each image. Finally, we formulize the emotion distribution prediction as a multi-task shared sparse regression (MTSSR) problem, which is optimized by iteratively reweighted least squares. Besides, we introduce three baseline algorithms. Experiments conducted on the Image-Emotion-Social-Net dataset demonstrate the superiority of the proposed method, as compared to some state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Xiaolei Jiang |
ACM Multimedia | 1 |
| 2015 | Multimedia Social Event Detection in Microblog
Yue Gao 0002, Sicheng Zhao, Yang Yang 0002, Tat-Seng Chua |
MMM (1) | 2 |
| 2015 | Strategy for aesthetic photography recommendation via collaborative composition modelabstractIn this study, the authors propose a collaborative composition model for automatically recommending suitable positions and poses in the scene of photography taken by amateurs. By analysing aesthetic‐aware features, the authors' strategy jointly takes attention and geometry composition into account to learn the aesthetic manifestation knowledge of professional photographers. Firstly, aesthetic composition representation exploits the strength of visual saliency to explicitly encode the spatial correlation of the professional photos. Secondly, ℓ2 regularised least square is adopted to constrain the representation coefficients, which provides a fast solution in selecting aesthetic candidates collaboratively. In addition, a novel confidence measure scheme is further designed based on reconstruction errors and the reference photos are updated adaptively according to the composition rules. Both qualitative and quantitative evaluations show that the model performs well for the portrait photographing recommendation. Yanhao Zhang 0001, Qingming Huang, Sicheng Zhao, Xiusheng Lu, Xiaoshuai Sun, Hongxun Yao |
IET Comput. Vis. | 4 |
| 2015 | Strategy for dynamic 3D depth data matching towards robust action retrieval
Sicheng Zhao, Lujun Chen, Hongxun Yao, Yanhao Zhang 0001, Xiaoshuai Sun |
Neurocomputing | 1 |
| 2015 | View-based 3D object retrieval via multi-modal graph learning
Sicheng Zhao, Hongxun Yao, Yanhao Zhang 0001, Yasi Wang, Shaohui Liu |
Signal Process. | 1 |
| 2014 | Exploring Principles-of-Art Features For Image Emotion RecognitionabstractEmotions can be evoked in humans by images. Most previous works on image emotion analysis mainly used the elements-of-art-based low-level visual features. However, these features are vulnerable and not invariant to the different arrangements of elements. In this paper, we investigate the concept of principles-of-art and its influence on image emotions. Principles-of-art-based emotion features (PAEF) are extracted to classify and score image emotions for understanding the relationship between artistic principles and emotions. PAEF are the unified combination of representation features derived from different principles, including balance, emphasis, harmony, variety, gradation, and movement. Experiments on the International Affective Picture System (IAPS), a set of artistic photography and a set of peer rated abstract paintings, demonstrate the superiority of PAEF for affective image classification and regression (with about 5% improvement on classification accuracy and 0.2 decrease in mean squared error), as compared to the state-of-the-art approaches. We then utilize PAEF to analyze the emotions of master paintings, with promising results. Sicheng Zhao, Yue Gao 0002, Xiaolei Jiang, Hongxun Yao, Tat-Seng Chua, Xiaoshuai Sun |
ACM Multimedia | 1 |
| 2014 | Affective Image Retrieval via Multi-Graph LearningabstractImages can convey rich emotions to viewers. Recent research on image emotion analysis mainly focused on affective image classification, trying to find features that can classify emotions better. We concentrate on affective image retrieval and investigate the performance of different features on different kinds of images in a multi-graph learning framework. Firstly, we extract commonly used features of different levels for each image. Generic features and features derived from elements-of-art are extracted as low-level features. Attributes and interpretable principles-of-art based features are viewed as mid-level features, while semantic concepts described by adjective noun pairs and facial expressions are extracted as high-level features. Secondly, we construct single graph for each kind of feature to test the retrieval performance. Finally, we combine the multiple graphs together in a regularization framework to learn the optimized weights of each graph to efficiently explore the complementation of different features. Extensive experiments are conducted on five datasets and the results demonstrate the effectiveness of the proposed method. Sicheng Zhao, Hongxun Yao, Yanhao Zhang 0001 |
ACM Multimedia | 1 |
| 2013 | Flexible Presentation of Videos Based on Affective Content Analysis
Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun, Xiaolei Jiang, Pengfei Xu 0001 |
MMM (1) | 1 |
| 2013 | Video classification and recommendation based on affective analysis of viewers
Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun |
Neurocomputing | 1 |
| 2012 | Cookie-Proxy: A Scheme to Prevent SSLStrip Attack
Sendong Zhao, Ding Wang 0002, Sicheng Zhao, Chunguang Ma |
ICICS | 3 |
| 2011 | Affective Video Classification Based on Spatio-temporal Feature FusionabstractIn this paper, we propose a novel affective video classification method based on facial expression recognition by learning the spatio-temporal feature fusion of actors' and viewers' facial expressions. For spatial features, we integrate Haar-like features into compositional ones according to the features' correlation, and train a mid classifier during the period. Then this process is embedded into improved AdaBoost learning algorithm to obtain spatial features. And for temporal feature fusion, we adopt hidden dynamic conditional random fields (HDCRFs) based on HCRFs by introducing time dimension variable. Finally spatial features are embedded into HDCRFs to recognize facial expressions. Experiments on the well-known Cohn-Kanada database show that the proposed method has a promising recognition performance. And affective classification experimental results on our own videos show that most subjects are satisfied with the classification results. Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun |
ICIG | 1 |
| 2011 | Video indexing and recommendation based on affective analysis of viewersabstractMost previous works on video indexing and recommendation were only based on the content of video itself, without considering the affective analysis of viewers, which is an efficient and important way to reflect viewers' attitudes, feelings and evaluations of videos. In this paper, we propose a novel method to index and recommend videos based on affective analysis, mainly on facial expression recognition of viewers. We first build a facial expression recognition classifier by embedding the process of building compositional Haar-like features into hidden conditional random fields (HCRFs). Then we extract viewers' facial expressions frame by frame through the videos, collected from the camera when viewers are watching videos, to obtain the affections of viewers. Finally, we draw the affective curve which tells the process of affection changes. Through the curve, we segment each video into affective sections, give the indexing result of the videos, and list recommendation points from views' aspect. Experiments on our collected database from the web show that the proposed method has a promising performance. Sicheng Zhao, Hongxun Yao, Xiaoshuai Sun, Pengfei Xu 0001, Xianming Liu 0005, Rongrong Ji |
ACM Multimedia | 1 |