Joey Tianyi Zhou

dblp:123/5110 · also Joey Zhou, Tianyi Zhou 0007 · DBLP profile ↗
← Back
198ranked-venue papers
17as first author
146since 2021 · last 2026
0000-0002-4675-7055ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 135 · 14 first-author · 93 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 6 first-author · 46 since 2021Computer networks · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Systems, architecture and hardware · 7 · 6 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Robust Semi-paired Multimodal Learning for Cross-modal Retrieval
abstract
Cross-modal retrieval is a fundamental application of multi-modal learning that has achieved remarkable success with large-scale well-paired data. However, in practice, it is costly to collect large-scale well-paired data. To alleviate the dependence on the amount of paired data, in this paper, we study a practical learning paradigm: semi-paired cross-modal learning (SPL), which utilizes both a small amount of paired data and a large amount of unpaired data to enhance cross-modal learning directly and is more accessible in practice. To achieve this, we take image-text retrieval as an example and propose a novel Robust Cross-modal Semi-paired Learning method (RCSL) by addressing two challenges. To be specific, i) to overcome the under-optimization issue caused by too little paired data, we present Semi-paired Discriminative Learning (SDL) to fully learn visual-semantic associations from a small amount of image-text pairs by preserving the alignment and uniformity of modality representations. ii) To mine visual-semantic correspondences from unpaired data, RCSL first constructs pseudo-paired correlations across different modalities by nearest neighbor association. However, this may introduce noisy correspondences (NCs) due to inaccurate pseudo signals, which could degrade the model's performance. To tackle NCs, we devise Robust Cross-correlation Mining (RCM) based on the risk minimization criterion to robustly and explicitly learn visual-semantic associations from pseudo-paired data, thus boosting cross-modal learning. Finally, we conduct extensive experiments on four datasets, i.e., three widely used benchmark datasets of Flickr30K, MS-COCO, CC152K, and a newly constructed real-world dataset Drone-SP, to demonstrate the effectiveness of RCSL under semi-paired and noisy settings.
Yuan Sun 0016, Xi Peng 0001, Dezhong Peng, Joey Tianyi Zhou, Xiaomin Song, Peng Hu 0002
AAAI5
2026 Poisoned Distillation: Injecting Backdoors into Distilled Datasets Without Raw Data Access
abstract
Dataset distillation (DD) condenses large datasets into smaller synthetic ones to enhance training efficiency and reducing bandwidth. DD enables models to achieve comparable performance to those trained on the raw full dataset, making it popular for data sharing. Existing work shows that injecting backdoors during the distillation process can threaten downstream models. However, these studies assume attackers can have access to the raw dataset and interfere with the entire distillation process, which is unrealistic. In contrast, this work is the first to address a more realistic and concerning threat: attackers may intercept the dataset distribution process, inject backdoors into the distilled datasets, and redistribute them to users. While distilled datasets were previously considered resistant to backdoor attacks, we demonstrate that they remain vulnerable to such attacks. Furthermore, we show that attackers do not even require access to any raw data to inject the backdoors successfully within one minute. Specifically, our approach reconstructs conceptual archetypes for each class from the model trained on the distilled dataset. Backdoors are then injected into these archetypes to update the distilled dataset. Moreover, we ensure the updated dataset not only retains the backdoor but also preserves the original optimization trajectory, thus maintaining the knowledge of the raw dataset. To achieve this, a hybrid loss is designed to integrate backdoor information along the benign optimization trajectory, ensuring that previously learned information is not forgotten. Extensive experiments demonstrate that distilled datasets are highly vulnerable to our attack, with risks pervasive across various raw datasets, distillation methods, and downstream training strategies.
Ming Yan 0007, Joey Tianyi Zhou
AAAI4
2026 Agentic Spatio-Temporal Grounding via Collaborative Reasoning
abstract
Spatio-Temporal Video Grounding (STVG) aims to retrieve the spatio-temporal tube of a target object or person in a video given a text query. Most existing approaches perform frame-wise spatial localization within a predicted temporal span, resulting in redundant computation, heavy supervision requirements, and limited generalization. Weakly-supervised variants mitigate annotation costs but remain constrained by the dataset-level train-and-fit paradigm with an inferior performance. To address these challenges, we propose the Agentic Spatio-Temporal Grounder (ASTG) framework for the task of STVG in an open-world and zero-shot setting. Specifically, two specialized agents SRA (Spatial Reasoning Agent) and TRA (Temporal Reasoning Agent) constructed leveraging modern Multi-modal Large Language Models (MLLMs) work collaboratively to retrieve the target tube in an autonomous and self-guided manner. Following a propose-and-evaluation paradigm, ASTG duly decouples spatio-temporal reasoning and automates the tube extraction, verification and temporal localization processes. With a dedicated visual memory and dialogue context, ASTG achieves architectural efficiency by minimizing the number of reasoning calls compared to exhaustive per-frame reasoning, eliminating the logical redundancy inherent in joint-reasoning systems. Experiments on popular benchmarks demonstrate the superiority of the proposed approach where it outperforms existing weakly-supervised and zero-shot approaches by a margin and is comparable to some of the fully-supervised methods.
Heng Zhao 0004, Yew-Soon Ong, Joey Tianyi Zhou
SIGIR3
2026 LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Multilingual Text-Centric VQA
abstract
Multilingual Text-Centric Visual Question Answering (TEC-VQA) has become crucial for real-world applications, as it requires fine-grained understanding and reasoning over multilingual scene text. Recent advances in vision-language models (VLMs) have demonstrated strong potential in tackling multimodal tasks. However, most existing approaches rely primarily on textual Chain-of-Thought (CoT) and provide limited support for multilingual multimodal reasoning. To address this gap, we introduce LaV-CoT, the first Language-aware Visual CoT framework with Multi-Aspect Reward Optimization. LaV-CoT incorporates an interpretable multi-stage reasoning pipeline consisting of text summary with bounding box, language identification, spatial object-level captioning, and step-by-step logical reasoning. To improve reasoning accuracy and cross-lingual generalization, we propose a novel verifiable Multi-Aspect Reward Optimization in addition to supervised fine-tuning that incorporates rewards for linguistic consistency, structural fidelity, and response accuracy. Extensive evaluations on public datasets, including MMMB, Multilingual MMBench, and MTVQA, show that LaV-CoT outperforms open-source models of similar size by up to ~9.5% accuracy, even surpassing open-source models more than twice its size, and further exceeding several state-of-the-art proprietary models. Moreover, LaV-CoT has been integrated into our online Intelligent Document Processing platform. A further online A/B test demonstrates an \(\sim\)8.7% improvement in acceptance rate, validating its effectiveness in industrial deployment and commercial applications. Our code is available at this https://github.com/HJNVR/LaV-CoT repository.
Zhiya Tan, Shutao Gong, Fanwei Zeng, Joey Tianyi Zhou, Changtao Miao, Huazhe Tan, Weibin Yao, Jianshu Li
WWW5
2026 Mapping text to multiplex graph: Prompt compression as Lévy walk-guided graph pruning
Yaxin Gao, Yao Lu 0041, Jinhong Deng, Jiaqi Nie, Jian Zhang 0023, Zhaowei Zhu, Shanqing Yu, Qi Xuan 0001, Joey Tianyi Zhou
Knowl. Based Syst.10
2026 Auto-Clustering with Continuous Distribution Estimation on Centroids
Yuangang Pan, Yinghua Yao, Atsushi Nitanda, Joey Tianyi Zhou, Ivor W. Tsang
Mach. Learn.4
2026 Dynamic-Aware video distillation: Adaptive temporal partitioning based on video semantics for edge device
Yinjie Zhao, Heng Zhao 0004, Yew-Soon Ong, Bihan Wen, Joey Tianyi Zhou
Neural Networks5
2026 3D Hand Pose Estimation via Articulated Anchor-to-Joint 3D Local Regressors
abstract
In this paper, we propose to address monocular 3D hand pose estimation from a single RGB or depth image via articulated anchor-to-joint 3D local regressors, in form of A2J-Transformer+. The key idea is to make the local regressors (i.e., anchor points) in 3D space be aware of hand's local fine details and global articulated context jointly, to facilitate predicting their 3D offsets toward hand joints with linear weighted aggregation for joint localization. Our intuition is that, local fine details help to estimate accurate offset but may suffer from the issues including serious occlusion, confusing similar patterns, and overfitting risk. On the other hand, hand's global articulated context can essentially provide additional descriptive clues and constraints to alleviate these issues. To set anchor points adaptively in 3D space, A2J-Transformer+ runs in a 2-stage manner. At the first stage, since the input modality property anchor points distribute more densely on X-Y plane, it leads to lower prediction accuracy along Z direction compared with those in the X and Y directions. To alleviate this, at the second stage anchor points are set near the joints yielded by the first stage evenly along X, Y, and Z directions. This treatment brings two main advantages: (1) balancing the prediction accuracy along X, Y, and Z directions, and (2) ensuring the anchor-joint offsets are of small values relatively easy to estimate. Wide-range experiments on three RGB hand datasets (InterHand2.6 M, HO-3D V2 and RHP) and three depth hand datasets (NYU, ICVL and HANDS 2017) verify A2J-Transformer+'s superiority and generalization ability for different modalities (i.e., RGB and depth) and hand cases (i.e., single hand, interacting hands, and hand-object interaction), even outperforming model-based manners. The test on ITOP dataset reveals that, A2J-Transformer+ can also be applied to 3D human pose estimation task.
Changlong Jiang, Yang Xiao 0007, Jinghong Zheng 0002, Haohong Kuang, Cunlin Wu, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Pattern Anal. Mach. Intell.9
2026 Safe Image Generation via Lightweight Concept Erasure in Diffusion Models
abstract
Text-to-image diffusion models have achieved remarkable progress in image synthesis, but their potential misuse for generating unauthorized or harmful content has raised growing safety concerns. This has created an urgent need for safe diffusion-based image generation methods that can selectively suppress sensitive concepts while preserving the model's general generative capability. Existing concept erasure approaches typically rely on either model fine-tuning or closed-form editing. However, they often suffer from two major limitations: (1 insufficient or excessive erasure, where the former fails to suppress target concepts and the latter disrupts benign semantics; and (2 degradation of non-target concepts, where removing target concepts undermines the generation of unrelated concepts, especially in multi-concept scenarios. To address these issues, we propose the Singular Value Eraser (SVEraser), a lightweight concept erasure module that removes specific concepts by optimizing singular-value offsets of weight matrices. Operating in a compact yet expressive singular-value space, SVEraser enables precise concept removal while reducing side effects on unrelated content. Moreover, once trained for different concepts, multiple SVErasers can be flexibly combined for multi-concept erasure. To further reduce interference, we introduce an eraser activation mechanism that adaptively selects the appropriate SVErasers during inference based on the input prompt. Extensive experiments on copyrighted objects, artistic styles, and explicit content demonstrate that our method achieves accurate target concept removal while preserving non-target semantics, providing a practical and reliable solution for safe diffusion-based image generation.
Xiaoyu Geng, Shuaixiong Hui, Joey Tianyi Zhou, Zheng Wang 0007
IEEE Trans. Image Process.4
2026 PhyTrace: Tracing Physical Inconsistency in AI-Generated Images via ISP Emulation
abstract
The high realism of AI-generated images has emerged as a significant cybersecurity threat. While existing detection methods have achieved some success, most rely on fixed models that are incapable of adapting to new data or generating model updates. This paper overcomes these limitations by shifting the focus to the fundamental imaging process of real images: Image Signal Processing (ISP). Unlike real images, AI-generated images do not undergo this process, making them more susceptible to physical variations within ISP modules. By analyzing how ISP sub-modules influence the physical characteristics of imaging, we simulate the ISP mapping process to amplify the differences in physical responses between real and AI-generated images during ISP transformations. Tracing these physical differences, we propose PhyTrace, a novel training-free method for detecting AI-generated images. PhyTrace enforces physical consistency constraints within ISP, operates independently of specific datasets and generative models, and effectively detects a wide range of AI-generated images. PhyTrace reveals distinct distribution patterns of real and AI-generated images. Extensive experiments on 18 test sets demonstrate that our method outperforms prior approaches in average precision and generalization, offering a robust solution for AI-generated image detection in open-world scenarios.
Wenxuan Liu 0008, Danni Xu, Joey Tianyi Zhou, Zheng Wang 0007
IEEE Trans. Image Process.4
2026 Single-Domain Generalization via Path Flatness-Aware Optimization of Loss Landscapes
abstract
Domain generalization (DG) methods traditionally rely on multiple source domains to achieve the robust performance across unseen target domains. However, single-DG (SDG) presents a more practical paradigm by learning from a single source domain, addressing scenarios where access to multiple domains is limited. While existing SDG approaches primarily focus on data augmentation and style transfer techniques to enhance the model robustness, these methods often incur substantial computational overhead and may inadequately capture the complexity of real-world domain shifts. In this article, we propose path flatness-aware optimization (PFO), an optimization framework that addresses the fundamental challenges of SDG. Unlike conventional approaches that rely on the synthetic data generation, PFO identifies and exploits regions of flat minima within the optimization landscape of deep neural networks. The framework employs an iterative optimization strategy to construct a path through the parameter space along which an ensemble of candidate models achieves the minimal empirical risk. The initialization of this optimization path is achieved through the strategic interconnection of model instances, each originating from carefully selected anchor points that are computationally determined through the systematic analysis of classification decision manifolds. This optimization path serves as a mechanism for implicit distribution alignment between source and target domains within the loss landscape, consequently enhancing the model's capacity for cross-DG. Empirical evaluation on multiple benchmark datasets demonstrates significant performance improvements in cross-DG, validating the efficacy of our approach.
Zizhou Wang, Yan Wang 0015, Yangqin Feng, Jiawei Du 0002, Joey Tianyi Zhou, Rick Siow Mong Goh, Yong Liu 0026, Liangli Zhen
IEEE Trans. Neural Networks Learn. Syst.5
2025 KPL: Training-Free Medical Knowledge Mining of Vision-Language Models
abstract
Visual Language Models such as CLIP excel in image recognition due to extensive image-text pre-training. However, applying the CLIP inference in zero-shot classification, particularly for medical image diagnosis, faces challenges due to: 1) the inadequacy of representing image classes solely with single category names; 2) the modal gap between the visual and text spaces generated by CLIP encoders. Despite attempts to enrich disease descriptions with large language models, the lack of class-specific knowledge often leads to poor performance. In addition, empirical evidence suggests that existing proxy learning methods for zero-shot image classification on natural image datasets exhibit instability when applied to medical datasets. To tackle these challenges, we introduce the Knowledge Proxy Learning (KPL) to mine knowledge from CLIP. KPL is designed to leverage CLIP's multimodal understandings for medical image classification through Text Proxy Optimization and Multimodal Proxy Learning. Specifically, KPL retrieves image-relevant knowledge descriptions from the constructed knowledge-enhanced base to enrich semantic text proxies. It then harnesses input images and these descriptions, encoded via CLIP, to stably generate multimodal proxies that boost the zero-shot classification performance. Extensive experiments conducted on both medical and natural image datasets demonstrate that KPL enables effective zero-shot image classification, outperforming all baselines. These findings highlight the great potential in this paradigm of mining knowledge from CLIP for medical image classification and broader areas.
Tianxiang Hu, Jiawei Du 0002, Ruiyuan Zhang, Joey Tianyi Zhou, Zuozhu Liu
AAAI5
2025 Understanding Large Language Model Vulnerabilities to Social Bias Attacks
abstract
Large Language Models (LLMs) have become foundational in human-computer interaction, demonstrating remarkable linguistic capabilities across various tasks. However, there is a growing concern about their potential to perpetuate social biases present in their training data. In this paper, we comprehensively investigate the vulnerabilities of contemporary LLMs to various social bias attacks, including prefix injection, refusal suppression, and learned attack prompts. We evaluate popular models such as LLaMA-2, GPT-3.5, and GPT-4 across gender, racial, and religious bias types. Our findings reveal that models are generally more susceptible to gender bias attacks compared to racial or religious biases. We also explore novel aspects such as cross-bias and multiple-bias attacks, finding varying degrees of transferability across bias types. Additionally, our results show that larger models and pretrained base models often exhibit higher susceptibility to bias attacks. These insights contribute to the development of more inclusive and ethically responsible LLMs, emphasizing the importance of understanding and mitigating potential bias vulnerabilities. We offer recommendations for model developers and users to enhance the robustness of LLMs against social bias attacks.
Jiaxu Zhao 0002, Fanghua Ye 0001, Joey Tianyi Zhou, Mykola Pechenizkiy
ACL (1)6
2025 DiffPO: Diffusion-styled Preference Optimization for Inference Time Alignment of Large Language Models
abstract
Ruizhe Chen, Wenhao Chai, Zhifei Yang, Xiaotian Zhang, Ziyang Wang, Tony Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ruizhe Chen, Wenhao Chai, Zhifei Yang 0004, Tony Q. S. Quek, Joey Tianyi Zhou, Soujanya Poria, Zuozhu Liu
ACL (1)7
2025 Learning with Noisy Triplet Correspondence for Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) enables editable image search by integrating a query pair—a reference image ref and a textual modification mod—to retrieve a target image tar that reflects the intended change. While existing CIR methods have shown promising performance using well-annotated triplets ⟨ref,mod,tar⟩, almost all of them implicitly assume these triplets are accurately associated with each other. In practice, however, this assumption is often violated due to the limited knowledge of annotators, inevitably leading to incorrect textual modifications and resulting in a practical yet less-touched problem: noisy triplet correspondence (NTC). To tackle this challenge, we propose a Task-oriented Modification Enhancement framework (TME) to learn robustly from noisy triplets, which comprises three key modules: Robust Fusion Query (RFQ), Pseudo Text Enhancement (PTE), and Task-Oriented Prompt (TOP). Specifically, to mitigate the adverse impact of noise, RFQ employs a sample selection strategy to divide the training triplets into clean and noisy sets, thus enhancing the reliability of the training data for robust learning. To further leverage the noisy data instead of discarding it, PTE unifies the triplet noise as an adapter mismatch problem, thereby adjusting mod to align with ref and tar in the mismatched triplet. Finally, TOP replaces ref in the clean set with a trainable prompt, which is then concatenated with mod to form a query independent of the visual reference, aiming to mitigate visually irrelevant noise. Extensive experiments on two domain-specific datasets demonstrate the robustness and superiority of TME, particularly in noisy scenarios.
Changhao He, Xiting Liu, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002
CVPR4
2025 ProjAttacker: A Configurable Physical Adversarial Attack for Face Recognition via Projector
abstract
Previous physical adversarial attacks have shown that carefully crafted perturbations can deceive face recognition systems, revealing critical security vulnerabilities. However, these attacks often struggle to impersonate multiple targets and frequently fail to bypass liveness detection. For example, attacks using human-skin masks [28] are challenging to fabricate, inconvenient to swap between users, and often fail liveness detection due to facial occlusions. A projector, however, can generate content-rich light without obstructing the face, making it ideal for non-intrusive attacks. Thus, we propose a novel physical adversarial attack using a projector and explore the superposition of projected and natural light to create adversarial facial images. This approach eliminates the need for physical artifacts on the face, effectively overcoming these limitations. Specifically, our proposed ProjAttacker generates adversarial 3D textures that are projected onto human faces. To ensure physical realizability, we introduce a light reflection function that models complex optical interactions between projected light and human skin, accounting for reflection and diffraction effects. Furthermore, we incorporate camera Image Signal Processing (ISP) simulation to maintain the robustness of adversarial perturbations across real-world diverse imaging conditions. Comprehensive evaluations conducted in both digital and physical scenarios validate the effectiveness of our method.
Yuanwei Liu, Hui Wei 0004, Ruqi Xiao, Weijian Ruan, Xingxing Wei 0001, Joey Tianyi Zhou, Zheng Wang 0007
CVPR7
2025 SSpMV: A Sparsity-aware SpMV Framework Empowered by Multimodal Machine Learning
abstract
Sparse Matrix-Vector Multiplication (SpMV) is an essential sparse operation in scientific computing and artificial intelligence. Efficiently adapting SpMV algorithms to diverse matrices and architectures requires a framework capable of accurately recognizing sparse patterns and selecting the optimal implementation. In this work, we introduce Sparsity-aware SpMV (SSpMV), a framework that integrates expert-designed features with multimodal representations to adaptively predict the best-performing algorithm and parameters. For this purpose, we design a multimodal neural network called MM-Adapter, to capture diverse modalities to represent the computational features of SpMV. Experimental results demonstrate that MMAdapter achieves the highest accuracy of $81.05 \%$, outperforming existing SpMV prediction models. Furthermore, SSpMV consistently delivers substantial performance improvements over state-of-the-art sparse libraries across various multi-core platforms.
Shengle Lin, Chubo Liu, Yan Ding 0004, Joey Tianyi Zhou, Kenli Li 0001, Wangdong Yang
DAC4
2025 CondenseLM: LLMs-driven Text Dataset Condensation via Reward Matching
abstract
Dataset condensation has emerged as a promising technique to improve data efficiency under limited data budgets.However, when applied to the text level, existing methods struggle to compress more information into samples through optimization.Thus, these methods provide no obvious advantage over simpler coreset selection despite their high computational cost.In this paper, we introduce CondenseLM, a novel paradigm for both effective and efficient textlevel dataset condensation.Our framework employs an LLMs-driven approach to sidestep the inherent limitations of existing methods, successfully generating more informative and less biased samples.In addition, it incorporates reward matching to align the LLMs-condensed dataset with the original dataset, maximizing representability and coverage.We conducted extensive experiments on SST-2, MNLI, AG News, and IMDB.Our approach outperforms both coreset selection and existing dataset condensation methods by large margins while also substantially reducing the computational cost.
Yew-Soon Ong, Joey Tianyi Zhou
EMNLP3
2025 RESCUE: Crowd Evacuation Simulation via Controlling SDM-United Characters
Joey Tianyi Zhou, Hongbo Kang, Wenguo Weng, Yukun Lai, Kun Li 0001
ICCV2
2025 Breaking Class Barriers: Efficient Dataset Distillation via Inter-Class Feature Compensator
abstract
Dataset distillation has emerged as a technique aiming to condense informative features from large, natural datasets into a compact and synthetic form. While recent advancements have refined this technique, its performance is bottlenecked by the prevailing class-specific synthesis paradigm. Under this paradigm, synthetic data is optimized exclusively for a pre-assigned one-hot label, creating an implicit class barrier in feature condensation. This leads to inefficient utilization of the distillation budget and oversight of inter-class feature distributions, which ultimately limits the effectiveness and efficiency, as demonstrated in our analysis. To overcome these constraints, this paper presents the Inter-class Feature Compensator (INFER), an innovative distillation approach that transcends the class-specific data-label framework widely utilized in current dataset distillation methods. Specifically, INFER leverages a Universal Feature Compensator (UFC) to enhance feature integration across classes, enabling the generation of multiple additional synthetic instances from a single UFC input. This significantly improves the efficiency of the distillation budget. Moreover, INFER enriches inter-class interactions during the distillation, thereby enhancing the effectiveness and generalizability of the distilled data. By allowing for the linear interpolation of labels similar to those in the original dataset, INFER meticulously optimizes the synthetic data and dramatically reduces the size of soft labels in the synthetic dataset to almost zero, establishing a new benchmark for efficiency and effectiveness in dataset distillation. In practice, INFER demonstrates state-of-the-art performance across benchmark datasets. For instance, in the $\texttt{ipc} = 50$ setting on ImageNet-1k with the same compression level, it outperforms SRe2L by 34.5\% using ResNet18. Codes are available at https://github.com/zhangxin-xd/UFC.
Xin Zhang 0092, Jiawei Du 0002, Ping Liu 0004, Joey Tianyi Zhou
ICLR4
2025 Deep Unsupervised Hashing via External Guidance
abstract
Recently, deep unsupervised hashing has gained considerable attention in image retrieval due to its advantages in cost-free data labeling, computational efficiency, and storage savings. Although existing methods achieve promising performance by leveraging inherent visual structures within the data, they primarily focus on learning discriminative features from unlabeled images through limited internal knowledge, resulting in an intrinsic upper bound on their performance. To break through this intrinsic limitation, we propose a novel method, called Deep Unsupervised Hashing with External Guidance (DUH-EG), which incorporates external textual knowledge as semantic guidance to enhance discrete representation learning. Specifically, our DUH-EG: i) selects representative semantic nouns from an external textual database by minimizing their redundancy, then matches images with them to extract more discriminative external features; and ii) presents a novel bidirectional contrastive learning mechanism to maximize agreement between hash codes in internal and external spaces, thereby capturing discrimination from both external and intrinsic structures in Hamming space. Extensive experiments on four benchmark datasets demonstrate that our DUH-EG remarkably outperforms existing state-of-the-art hashing methods.
Qihong Song, XitingLiu, Hongyuan Zhu 0002, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002
ICML4
2025 AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
abstract
Recent advances in AI-generated content have fueled the rise of highly realistic synthetic videos, posing severe risks to societal trust and digital integrity. Existing benchmarks for video authenticity detection typically suffer from limited realism, insufficient scale, and inadequate complexity, failing to effectively evaluate modern vision-language models against sophisticated forgeries. To address this critical gap, we introduce AEGIS, a novel large-scale benchmark explicitly targeting the detection of hyper-realistic and semantically nuanced AI-generated videos. AEGIS comprises over 10,000 rigorously curated real and synthetic videos generated by diverse, state-of-the-art generative models, including Stable Video Diffusion, CogVideoX-5B, KLing, and Sora, encompassing open-source and proprietary architectures. In particular, AEGIS features specially constructed challenging subsets enhanced with GPT-4o-refined prompts, creating unprecedentedly realistic scenarios for rigorous robustness evaluation. Furthermore, we provide multimodal annotations spanning Semantic-Authenticity Descriptions, Motion Features, and Low-level Visual Features, facilitating authenticity detection and supporting downstream tasks such as multimodal fusion and forgery localization. Extensive experiments using advanced vision-language models demonstrate limited detection capabilities on the most challenging subsets of AEGIS, highlighting the dataset's unique complexity and realism beyond the current generalization capabilities of existing models. In essence, AEGIS establishes an indispensable evaluation benchmark, fundamentally advancing research toward developing genuinely robust, reliable, and broadly generalizable video authenticity detection methodologies capable of addressing real-world forgery threats. Our dataset is avaliable on https://huggingface.co/datasets/Clarifiedfish/AEGIS.
Xin Zhang 0092, Joey Tianyi Zhou
ACM Multimedia3
2025 MFFI: Multi-Dimensional Face Forgery Image Dataset for Real-World Scenarios
abstract
Rapid advances in Artificial Intelligence Generated Content (AIGC) have enabled increasingly sophisticated face forgeries, posing a significant threat to social security. However, current Deepfake detection methods are limited by constraints in existing datasets, which lack the diversity necessary in real-world scenarios. Specifically, these data sets fall short in four key areas: unknown of advanced forgery techniques, variability of facial scenes, richness of real data, and degradation of real-world propagation. To address these challenges, we propose the Multi-dimensional Face Forgery Image (MFFI ) dataset, tailored for real-world scenarios. MFFI enhances realism based on four strategic dimensions: 1) Wider Forgery Methods; 2) Varied Facial Scenes; 3) Diversified Authentic Data; 4) Multi-level Degradation Operations. MFFI integrates 50 different forgery methods and contains 1024K image samples. Benchmark evaluations show that MFFI outperforms existing public datasets in terms of scene complexity, cross-domain generalization capability, and detection difficulty gradients. These results validate the technical advance and practical utility of MFFI in simulating real-world conditions. The dataset and additional details are publicly available at https://github.com/inclusionConf/MFFI.
Changtao Miao, Weiwei Feng, Qi Chu 0001, Jianshu Li, Yunfeng Diao, Wei Zhou 0021, Joey Tianyi Zhou, Xiaoshuai Hao
ACM Multimedia11
2025 Learning to Be a Doctor: Searching for Effective Medical Agent Architectures
abstract
Large Language Model (LLM)-based agents have demonstrated strong capabilities across a wide range of tasks, and their application in the medical domain holds particular promise due to the demand for high generalizability and reliance on interdisciplinary knowledge. However, existing medical agent systems often rely on static, manually crafted workflows that lack the flexibility to accommodate diverse diagnostic requirements and adapt to emerging clinical scenarios. Motivated by the success of automated machine learning (AutoML), this paper introduces a novel framework for the automated design of medical agent architectures. Specifically, we define a hierarchical and expressive agent search space that enables dynamic workflow adaptation through structured modifications at the node, structural, and framework levels. Our framework conceptualizes medical agents as graph-based architectures composed of diverse, functional node types and supports iterative self-improvement guided by diagnostic feedback. Experimental results on skin disease diagnosis tasks demonstrate that the proposed method effectively evolves workflow structures and significantly enhances diagnostic accuracy over time. This work represents the first fully automated framework for medical agent architecture design and offers a scalable, adaptable foundation for deploying intelligent agents in real-world clinical environments.
Yangyang Zhuang, Wenjia Jiang, Ze Yang 0002, Joey Tianyi Zhou, Chi Zhang 0007
ACM Multimedia5
2025 SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
abstract
Multimodal Large Language Models (MLLMs) typically process a large number of visual tokens, leading to considerable computational overhead, even though many of these tokens are redundant. Existing visual token pruning methods primarily focus on selecting the most salient tokens based on attention scores, resulting in the semantic incompleteness of the selected tokens. In this paper, we propose a novel visual token pruning strategy, called **S**aliency-**C**overage **O**riented token **P**runing for **E**fficient MLLMs (SCOPE), to jointly model both the saliency and coverage of the selected visual tokens to better preserve semantic completeness. Specifically, we introduce a set-coverage for a given set of selected tokens, computed based on the token relationships. We then define a token-coverage gain for each unselected token, quantifying how much additional coverage would be obtained by including it. By integrating the saliency score into the token-coverage gain, we propose our SCOPE score and iteratively select the token with the highest SCOPE score. We conduct extensive experiments on multiple vision-language understanding benchmarks using the LLaVA-1.5 and LLaVA-Next models. Experimental results demonstrate that our method consistently outperforms prior approaches.
Jinhong Deng, Wen Li 0001, Joey Tianyi Zhou, Yang He 0002
NeurIPS3
2025 Beyond Modality Collapse: Representation Blending for Multimodal Dataset Distillation
abstract
Multimodal Dataset Distillation (MDD) seeks to condense large-scale image-text datasets into compact surrogates while retaining their effectiveness for cross-modal learning. Despite recent progress, existing MDD approaches often suffer from ***Modality Collapse***, characterized by over-concentrated intra-modal representations and enlarged distributional gap across modalities. In this paper, at the first time, we identify this issue as stemming from a fundamental conflict between the over-compression behavior inherent in dataset distillation and the cross-modal supervision imposed by contrastive objectives. To alleviate modality collapse, we introduce **RepBlend**, a novel MDD framework that weakens overdominant cross-modal supervision via representation blending, thereby significantly enhancing intra-modal diversity. Additionally, we observe that current MDD methods impose asymmetric supervision across modalities, resulting in biased optimization. To address this, we propose symmetric projection trajectory matching, which synchronizes the optimization dynamics using modality-specific projection heads, thereby promoting balanced supervision and enhancing cross-modal alignment. Experiments on Flickr-30K and MS-COCO show that RepBlend consistently outperforms prior state-of-the-art MDD methods, achieving significant gains in retrieval performance (e.g., +9.4 IR@10, +6.3 TR@10 under the 100-pair setting) and offering up to 6.7$\times$ distillation speedup.
Xin Zhang 0092, Ziruo Zhang, Jiawei Du 0002, Zuozhu Liu, Joey Tianyi Zhou
NeurIPS5
2025 PandaPose: 3D Human Pose Lifting from a Single Image via Propagating 2D Pose Prior to 3D Anchor Space
abstract
3D human pose lifting from a single RGB image is a challenging task in 3D vision. Existing methods typically establish a direct joint-to-joint mapping from 2D to 3D poses based on 2D features. This formulation suffers from two fundamental limitations: inevitable error propagation from input predicted 2D pose to 3D predictions and inherent difficulties in handling self-occlusion cases. In this paper, we propose PandaPose, a 3D human pose lifting approach via propagating 2D pose prior to 3D anchor space as the unified intermediate representation. Specifically, our 3D anchor space comprises: (1) Joint-wise 3D anchors in the canonical coordinate system, providing accurate and robust priors to mitigate 2D pose estimation inaccuracies. (2) Depth-aware joint-wise feature lifting that hierarchically integrates depth information to resolve self-occlusion ambiguities. (3) The anchor-feature interaction decoder that incorporates 3D anchors with lifted features to generate unified anchor queries encapsulating joint-wise 3D anchor set, visual cues and geometric depth information. The anchor queries are further employed to facilitate anchor-to-joint ensemble prediction. Experiments on three well-established benchmarks (i.e., Human3.6M, MPI-INF-3DHP and 3DPW) demonstrate the superiority of our proposition. The substantial reduction in error by 14.7% compared to SOTA methods on the challenging conditions of Human3.6M and qualitative comparisons further showcase the effectiveness and robustness of our approach.
Jinghong Zheng 0002, Changlong Jiang, Yang Xiao 0007, Jiaqi Li 0007, Haohong Kuang, Ran Wang 0005, Zhiguo Cao 0001, Joey Tianyi Zhou
NeurIPS10
2025 PVP: Polar Representation Boost for 3D Semantic Occupancy Prediction
abstract
Recently, representations based on polar coordinates have exhibited promising characteristics for 3D perceptual tasks. In addition to Cartesian-based methods, representing surrounding spaces through polar grids offers a compelling alternative in these tasks. This approach is advantageous for its ability to represent larger areas while preserving greater detail of nearby spaces. However, polar-based methods are inherently challenged by the issue of feature distortion due to the non-uniform division inherent to polar representation. To harness the advantages of po-lar representation while addressing its challenges, we propose Polar Voxel Occupancy Predictor (PVP), a novel 3D multimodal occupancy predictor operating in polar coordinates. PVP mitigates the issues of feature distortion and misalignment across different modalities with the following two design elements: 1) Global Represent Propagation (GRP) module, which incorporates global spatial information into the intermediate 3D volume, taking into account the prior spatial structure. It then employs Global Decomposed Attention to accurately propagate features to their correct locations. 2) Plane Decomposed Convolution (PD-Conv), which simplifies 3D distortions in polar coordinates by replacing 3D convolution with a series of 2D convolutions. With these straightforward yet impactful modifications, our PVP surpasses state-of-the-art works by significant margins-improving by 1.9% mIoU and 2.9% IoU over LiDAR-only methods, and by 7.9% mIoU and 6.8% IoU over multimodal methods on the OpenOccupancy dataset.
Yujing Xue, Jiawei Du 0002, Joey Tianyi Zhou
WACV4
2025 Correction to: Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu
Mach. Learn.6
2025 STADe: Sensory Temporal Action Detection via Temporal-Spectral Representation Learning
abstract
Temporal action detection (TAD) is a vital challenge in computer vision and the Internet of Things, aiming to detect and identify actions within temporal sequences. While TAD has primarily been associated with video data, its applications can also be extended to sensor data, opening up opportunities for various real-world applications. However, applying existing TAD models to sensory signals presents distinct challenges such as varying sampling rates, intricate pattern structures, and subtle, noise-prone patterns. In response to these challenges, we propose a Sensory Temporal Action Detection (STADe) model. STADe leverages Fourier kernels and adaptive frequency filtering to adaptively capture the nuanced interplay of temporal and frequency features underlying complex patterns. Moreover, STADe embraces adaptability by employing deep fusion at varying resolutions and scales, making it versatile enough to accommodate diverse data characteristics, such as the wide spectrum of sampling rates and action durations encountered in sensory signals. Unlike conventional models with unidirectional category-to-proposal dependencies, STADe adopts a cross-cascade predictor to introduce bidirectional and temporal dependencies within categories. To extensively evaluate STADe and promote future research in sensory TAD, we establish three diverse datasets using various sensors, featuring diverse sensor types, action categories, and sampling rates. Experiments across one public and our three new datasets demonstrate STADe's superior performance over state-of-the-art TAD models in sensory TAD tasks.
Bing Li 0002, Haotian Duan, Yun Liu 0011, Le Zhang 0001, Wei Cui 0002, Joey Tianyi Zhou
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Multilingual-Prompt-Guided Directional Feature Learning for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection has gained attention for its effective performance and cost-efficient annotation, using video-level labels to distinguish between normal and abnormal patterns. However, challenges arise from the diversity and incompleteness of anomalous events, complicating feature learning. Vision-language models offer promising approaches, but designing precise prompts remains difficult. This is because accommodating the diverse range of normal and anomalous scenarios in real-world settings is challenging, and the workload is significant. To tackle these issues, we propose integrating multilingualism and multiple prompts to improve feature learning. By utilizing prompts in various languages to define "anomaly" and "normalcy," we tackle these concepts across different linguistic domains. In each domain, multiple prompts are employed for adaptive top-K prompt selection of snippets. To enhance visual feature learning, a multi-granularity attention module combining Transformer and Mamba is designed. Mamba's long-range adaptation selection builds fine-grained temporal correlations among coarse-grained snippets, while Transformer enhances fine-grained information guided by coarse-grained information. Alongside a multilingual prompt guidance loss, we introduce a gradual directional loss to jointly optimize visual feature distribution and the top-K prompt selection. Our method demonstrates effectiveness on four video datasets and provides generalizability analyses on two medical datasets, including EMG and ECG temporal data.
Chizhuo Xiao, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 L2A: Learning Affinity From Attention for Weakly Supervised Continual Semantic Segmentation
abstract
Despite significant advances in continual semantic segmentation (CSS), they still rely on the pixel-level annotation to train models, which is time-consuming and labor-intensive. Continual learning from image-level labels is an emerging scheme in continual semantic segmentation to reduce the annotation cost. However, the incomplete and coarse pseudo-labels are insufficient to train a model to maintain a balance between stability and plasticity. To solve these issues, we propose a novel end-to-end framework based on Transformer, called L2A, for Weakly Supervised Continual Semantic Segmentation (WSCSS). In particular, to generate reliable annotations from the image-level supervision, we introduce a semantic affinity from multi-head self-attention (SA-MHSA) module to capture the semantic relationships among adjacent image coordinates. Subsequently, this acquired semantic affinity is employed to refine the initial pseudo labels of new classes trained with the image-level annotations. Furthermore, to minimize catastrophic forgetting, we propose a semantic drift compensation (SDC) strategy to optimize the pseudo-label generation process, which can effectively improve the alignment of object boundaries across both new and old categories. Comprehensive experiments conducted on the PASCAL VOC 2012 and COCO datasets demonstrate the superiority of our framework in existing WSCSS scenarios and a newly proposed challenge protocol, as well as remains competitive compared to the pixel-level supervised CSS methods.
Hao Liu 0065, Yong Zhou 0003, Bing Liu 0016, Ming Yan 0007, Joey Tianyi Zhou
IEEE Trans. Circuits Syst. Video Technol.5
2025 Rhythmer: Ranking-Based Skill Assessment With Rhythm-Aware Transformer
abstract
Ranking-based skill assessment is an essential component of video understanding. In this task lacking precise procedure annotations, existing methods place greater emphasis on evaluating the procedure quality via manually normalizing the execution duration. However, the inherent duration-related procedural patterns will undergo alteration. Experimentally, we discover that distinct duration biases are prevalent in duration-sensitive skills, such as those in medical and everyday life. Hence, duration information is crucial for ranking-based skill assessment when dealing with varying durations. Additionally, similar execution processes tend to have closer execution durations. Thus, another critical factor lies in extracting duration-related procedural information alongside similar durations. It is defined as mining rhythm patterns, which are inspired by music rhythms including various duration and duration-related procedures. In our work, a rhythm-aware transformer is proposed to mine the rhythm patterns adaptively. Given pairwise inputs, a co-attention module is designed to mutually highlight duration-related procedure information when comparing pairwise input videos with similar durations, and adaptively attenuate the efficacy when confronted with pairwise inputs featuring significantly different durations. A rhythm-encoding module further embeds duration information into the concatenation of raw features and co-attention features. Following these features, the transformer decoder is designed to learn duration-related queries supervised by a novel duration grouping loss among various duration groups. The experimental results demonstrate that the rhythm-aware transformer is effective for ranking-based skill assessment.
Zhuang Luo, Yang Xiao 0007, Feng Yang 0012, Joey Tianyi Zhou, Zhiwen Fang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Trustworthy Visual-Textual Retrieval
abstract
Visual-textual retrieval, as a link between computer vision and natural language processing, aims at jointly learning visual-semantic relevance to bridge the heterogeneity gap across visual and textual spaces. Existing methods conduct retrieval only relying on the ranking of pairwise similarities, but they cannot self-evaluate the uncertainty of retrieved results, resulting in unreliable retrieval and hindering interpretability. To address this problem, we propose a novel Trust-Consistent Learning framework (TCL) to endow visual-textual retrieval with uncertainty evaluation for trustworthy retrieval. More specifically, TCL first models the matching evidence according to cross-modal similarity to estimate the uncertainty for cross-modal uncertainty-aware learning. Second, a simple yet effective consistency module is presented to enforce the subjective opinions of bidirectional learning to be consistent for high reliability and accuracy. Finally, extensive experiments are conducted to demonstrate the superiority and generalizability of TCL on six widely-used benchmark datasets, i.e., Flickr30K, MS-COCO, MSVD, MSR-VTT, ActivityNet, and DiDeMo. Furthermore, some qualitative experiments are carried out to provide comprehensive and insightful analyses for trustworthy visual-textual retrieval, verifying the reliability and interoperability of TCL. The code is available in https://github.com/QinYang79/TCL.
Lifu Huang, Dezhong Peng, Bohan Jiang, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002
IEEE Trans. Image Process.5
2025 Language-Guided 3-D Action Feature Learning Without Ground-Truth Sample Class Label
abstract
This work pays the first research effort to leverage point cloud sequence-based Self-supervised 3-D Action Feature Learning (S3AFL), under text's cross-modality weak supervision. We intend to fill the huge performance gap between point cloud sequence and 3-D skeleton-based manners. The key intuition derives from the observation that skeleton-based manners actually hold the human pose's high-level knowledge that leads to attention on the body's joint-aware local parts. Inspired by this, we propose to introduce the text's weak supervision of high-level semantics into a point cloud sequence-based paradigm. With RGB-point cloud pair sequence acquired via RGB-D camera, text sequence is first generated from RGB component using pretrained image captioning model, as auxiliary weak supervision. Then, S3AFL runs in a cross and intra-modality contrastive learning (CL) way. To resist text's missing and redundant semantics, feature learning is conducted in a multistage way with semantic refinement. Essentially, text is only required for training. To facilitate the feature's representation power on fine-grained actions, a multirank max-pooling (MR-MP) way is also proposed for the point set network to better maintain discriminative clues. Experiments verify that the text's weak supervision can facilitate performance by 10.8%, 10.4%, and 8.0% on NTU RGB+D 60, 120, and N-UCLA at most. The performance gap between point cloud sequence and skeleton-based manners has been remarkably narrowed down. The idea of transferring text's weak supervision to S3AFL can also be applied to a skeleton manner, with strong generality. The source code is available at https://github.com/tangent-T/W3AMT.
Yang Xiao 0007, Xingyu Tong 0002, Tingbing Yan, Zhiguo Cao 0001, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.7
2024 PointCVaR: Risk-Optimized Outlier Removal for Robust 3D Point Cloud Classification
abstract
With the growth of 3D sensing technology, the deep learning system for 3D point clouds has become increasingly important, especially in applications such as autonomous vehicles where safety is a primary concern. However, there are growing concerns about the reliability of these systems when they encounter noisy point clouds, either occurring naturally or introduced with malicious intent. This paper highlights the challenges of point cloud classification posed by various forms of noise, from simple background noise to malicious adversarial/backdoor attacks that can intentionally skew model predictions. While there's an urgent need for optimized point cloud denoising, current point outlier removal approaches, an essential step for denoising, rely heavily on handcrafted strategies and are not adapted for higher-level tasks, such as classification. To address this issue, we introduce an innovative point outlier cleansing method that harnesses the power of downstream classification models. Using gradient-based attribution analysis, we define a novel concept: point risk. Drawing inspiration from tail risk minimization in finance, we recast the outlier removal process as an optimization problem, named PointCVaR. Extensive experiments show that our proposed technique not only robustly filters diverse point cloud outliers but also consistently and significantly enhances existing robust methods for point cloud classification. A notable feature of our approach is its effectiveness in defending against the latest threat of backdoor attacks in point clouds.
Junchi Lu, Henghui Ding, Changsheng Sun, Joey Tianyi Zhou, Yeow Meng Chee
AAAI5
2024 Noisy-Correspondence Learning for Text-to-Image Person Re-Identification
abstract
Text-to-image person re-identification (TIReID) is a compelling topic in the cross-modal community, which aims to retrieve the target person based on a textual query. Although numerous TIReID methods have been proposed and achieved promising performance, they implicitly assume the training image-text pairs are correctly aligned, which is not always the case in real-world scenarios. In practice, the image-text pairs inevitably exist under-correlated or even false-correlated, a.k.a noisy correspondence (NC), due to the low quality of the images and annotation errors. To address this problem, we propose a novel Robust Dual Embedding method (RDE) that can learn robust visual-semantic associations even with NC. Specifically, RDE consists of two main components: 1) A Confident Consensus Division (CCD) module that leverages the dual-grained decisions of dual embedding modules to obtain a consensus set of clean training data, which enables the model to learn correct and reliable visual-semantic associations. 2) A Triplet Alignment Loss (TAL) relaxes the conventional Triplet Ranking loss with the hardest negative samples to a log-exponential upper bound over all negative ones, thus preventing the model collapse under NC and can also focus on hard-negative samples for promising performance. We conduct extensive experiments on three public benchmarks, namely CUHK-PEDES, ICFG-PEDES, and RSTPReID, to evaluate the performance and robustness of our RDE. Our method achieves state-of-the-art results both with and without synthetic noisy correspondences on all three datasets. Code is available at https://github.com/QinYang79/RDE.
Yingke Chen, Dezhong Peng, Xi Peng 0001, Joey Tianyi Zhou, Peng Hu 0002
CVPR5
2024 Spanning Training Progress: Temporal Dual-Depth Scoring (TDDS) for Enhanced Dataset Pruning
abstract
Dataset pruning aims to construct a coreset capable of achieving performance comparable to the original, full dataset. Most existing dataset pruning methods rely on snapshot-based criteria to identify representative samples, often resulting in poor generalization across various pruning and cross-architecture scenarios. Recent studies have addressed this issue by expanding the scope of training dynamics considered, including factors such as forgetting event and probability change, typically using an averaging approach. However, these works struggle to integrate a broader range of training dynamics without overlooking well-generalized samples, which may not be sufficiently highlighted in an averaging manner. In this study, we propose a novel dataset pruning method termed as Temporal Dual-Depth Scoring (TDDS), to tackle this problem. TDDS utilizes a dual-depth strategy to achieve a balance between incorporating extensive training dynamics and identifying representative samples for dataset pruning. In the first depth, we estimate the series of each sample's individual contributions spanning the training progress, ensuring comprehensive integration of training dynamics. In the second depth, we focus on the variability of the sample-wise contributions identified in the first depth to highlight well- generalized samples. Extensive experiments conducted on CIFAR and ImageNet datasets verify the superiority of TDDS over previous SOTA methods. Specifically on CIFAR-100, our method achieves 54.51% accuracy with only 10% training data, surpassing baselines methods by more than 12.69%. Our codes are available at https://github.com/zhangxin-xd/Dataset-Pruning-TDDS.
Xin Zhang 0092, Jiawei Du 0002, Yunsong Li 0001, Weiying Xie, Joey Tianyi Zhou
CVPR5
2024 STSP: Spatial-Temporal Subspace Projection for Video Class-Incremental Learning
Hao Cheng 0016, Siyuan Yang 0001, Chong Wang 0011, Joey Tianyi Zhou, Alex Chichung Kot, Bihan Wen
ECCV (28)4
2024 Direct Distillation Between Different Domains
Jialiang Tang, Shuo Chen 0003, Gang Niu 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Chen Gong 0002, Masashi Sugiyama
ECCV (80)5
2024 CrossGLG: LLM Guides One-Shot Skeleton-Based 3D Action Recognition in a Cross-Level Manner
Tingbing Yan, Wenzheng Zeng, Yang Xiao 0007, Xingyu Tong 0002, Zhiwen Fang, Zhiguo Cao 0001, Joey Tianyi Zhou
ECCV (20)8
2024 MedCoT: Medical Chain of Thought via Hierarchical Expert
abstract
Artificial intelligence has advanced in Medical Visual Question Answering (Med-VQA), but prevalent research tends to focus on the accuracy of the answers, often overlooking the reasoning paths and interpretability, which are crucial in clinical settings.Besides, current Med-VQA algorithms, typically reliant on singular models, lack the robustness needed for real-world medical diagnostics which usually require collaborative expert evaluation.To address these shortcomings, this paper presents MedCoT, a novel hierarchical expert verification reasoning chain method designed to enhance interpretability and accuracy in biomedical imaging inquiries.MedCoT is predicated on two principles: The necessity for explicit reasoning paths in Med-VQA and the requirement for multi-expert review to formulate accurate conclusions.The methodology involves an Initial Specialist proposing diagnostic rationales, followed by a Follow-up Specialist who validates these rationales, and finally, a consensus is reached through a vote among a sparse Mixture of Experts within the locally deployed Diagnostic Specialist, which then provides the definitive diagnosis.Experimental evaluations on four standard Med-VQA datasets demonstrate that MedCoT surpasses existing state-of-the-art approaches, providing significant improvements in performance and interpretability.Code is released at https: //github.com/JXLiu-AI/MedCoT.
Jiawei Du 0002, Joey Tianyi Zhou, Zuozhu Liu
EMNLP4
2024 Shortcuts Arising from Contrast: Towards Effective and Lightweight Clean-Label Attacks in Prompt-Based Learning
abstract
Prompt-based learning paradigm has been shown to be vulnerable to backdoor attacks.Current clean-label attack, employing a specific prompt as trigger, can achieve success without the need for external triggers and ensuring correct labeling of poisoned samples, which are more stealthy compared to the poisonedlabel attack, but on the other hand, facing significant issues with false activations and pose greater challenges, necessitating a higher rate of poisoning.Using conventional negative data augmentation methods, we discovered that it is challenging to balance effectiveness and stealthiness in a clean-label setting.In addressing this issue, we are inspired by the notion that a backdoor acts as a shortcut, and posit that this shortcut stems from the contrast between the trigger and the data utilized for poisoning.In this study, we propose a method named Contrastive Shortcut Injection (CSI), by leveraging activation values, integrates trigger design and data selection strategies to craft stronger shortcut features.With extensive experiments on fullshot and few-shot text classification tasks, we empirically validate CSI's high effectiveness and high stealthiness at low poisoning rates.
Xiaopeng Xie, Ming Yan 0007, Xiwen Zhou, Chenlong Zhao, Suli Wang, Joey Tianyi Zhou
EMNLP7
2024 Video-Text Prompting for Weakly Supervised Spatio-Temporal Video Grounding
abstract
Weakly-supervised Spatio-Temporal Video Grounding(STVG) aims to localize target object tube given a text query, without densely annotated training data.Existing methods extract each candidate tube feature independently by cropping objects from video frame feature, discarding all contextual information such as position change and inter-entity relationship.In this paper, we propose Video-Text Prompting(VTP) to construct candidate feature.Instead of cropping tube region from feature map, we draw visual markers(e.g.red circle) over objects tubes as video prompts; corresponding text prompt(e.g. in red circle) is also inserted after the subject word of query text to highlight its presence.Nevertheless, each candidate feature may look similar without cropping.To address this, we further propose Contrastive VTP(CVTP) by introducing negative contrastive samples whose candidate object is erased instead of being highlighted; by comparing the difference between VTP candidate and the contrastive sample, the gap of matching score between correct candidate and the rest is enlarged.Extensive experiments and ablations are conducted on several STVG datasets and our results surpass existing weakly-supervised methods by a great margin, demonstrating the effectiveness of our proposed methods.
Heng Zhao 0004, Yinjie Zhao, Bihan Wen, Yew-Soon Ong, Joey Tianyi Zhou
EMNLP5
2024 Multisize Dataset Condensation
abstract
While dataset condensation effectively enhances training efficiency, its application in on-device scenarios brings unique challenges. 1) Due to the fluctuating computational resources of these devices, there's a demand for a flexible dataset size that diverges from a predefined size. 2) The limited computational power on devices often prevents additional condensation operations. These two challenges connect to the "subset degradation problem" in traditional dataset condensation: a subset from a larger condensed dataset is often unrepresentative compared to directly condensing the whole dataset to that smaller size. In this paper, we propose Multisize Dataset Condensation (MDC) by **compressing $N$ condensation processes into a single condensation process to obtain datasets with multiple sizes.** Specifically, we introduce an "adaptive subset loss" on top of the basic condensation loss to mitigate the "subset degradation problem". Our MDC method offers several benefits: 1) No additional condensation process is required; 2) reduced storage requirement by reusing condensed images. Experiments validate our findings on networks including ConvNet, ResNet and DenseNet, and datasets including SVHN, CIFAR-10, CIFAR-100 and ImageNet. For example, we achieved 5.22%-6.40% average accuracy gains on condensing CIFAR-10 to ten images per class. Code is available at: [https://github.com/he-y/Multisize-Dataset-Condensation](https://github.com/he-y/Multisize-Dataset-Condensation).
Yang He 0002, Lingao Xiao, Joey Tianyi Zhou, Ivor W. Tsang
ICLR3
2024 Data-independent Module-aware Pruning for Hierarchical Vision Transformers
abstract
Hierarchical vision transformers (ViTs) have two advantages over conventional ViTs. First, hierarchical ViTs achieve linear computational complexity with respect to image size by local self-attention. Second, hierarchical ViTs create hierarchical feature maps by merging image patches in deeper layers for dense prediction. However, existing pruning methods ignore the unique properties of hierarchical ViTs and use the magnitude value as the weight importance. This approach leads to two main drawbacks. First, the "local" attention weights are compared at a "global" level, which may cause some "locally" important weights to be pruned due to their relatively small magnitude "globally". The second issue with magnitude pruning is that it fails to consider the distinct weight distributions of the network, which are essential for extracting coarse to fine-grained features at various hierarchical levels. To solve the aforementioned issues, we have developed a Data-independent Module-Aware Pruning method (DIMAP) to compress hierarchical ViTs. To ensure that "local" attention weights at different hierarchical levels are compared fairly in terms of their contribution, we treat them as a **module** and examine their contribution by analyzing their information distortion. Furthermore, we introduce a novel weight metric that is solely based on weights and does not require input images, thereby eliminating the **dependence** on the patch merging process. Our method validates its usefulness and strengths on Swin Transformers of different sizes on ImageNet-1k classification. Notably, the top-5 accuracy drop is only 0.07% when we remove 52.5% FLOPs and 52.7% parameters of Swin-B. When we reduce 33.2% FLOPs and 33.2% parameters of Swin-S, we can even achieve a 0.8% higher relative top-5 accuracy than the original model. Code is available at: [https://github.com/he-y/Data-independent-Module-Aware-Pruning](https://github.com/he-y/Data-independent-Module-Aware-Pruning).
Yang He 0002, Joey Tianyi Zhou
ICLR2
2024 FedLoGe: Joint Local and Generic Federated Learning under Long-tailed Data
abstract
Federated Long-Tailed Learning (Fed-LT), a paradigm wherein data collected from decentralized local clients manifests a globally prevalent long-tailed distribution, has garnered considerable attention in recent times. In the context of Fed-LT, existing works have predominantly centered on addressing the data imbalance issue to enhance the efficacy of the generic global model while neglecting the performance at the local level. In contrast, conventional Personalized Federated Learning (pFL) techniques are primarily devised to optimize personalized local models under the presumption of a balanced global data distribution. This paper introduces an approach termed Federated Local and Generic Model Training in Fed-LT (FedLoGe), which enhances both local and generic model performance through the integration of representation learning and classifier alignment within a neural collapse framework. Our investigation reveals the feasibility of employing a shared backbone as a foundational framework for capturing overarching global trends, while concurrently employing individualized classifiers to encapsulate distinct refinements stemming from each client’s local features. Building upon this discovery, we establish the Static Sparse Equiangular Tight Frame Classifier (SSE-C), inspired by neural collapse principles that naturally prune extraneous noisy features and foster the acquisition of potent data representations. Furthermore, leveraging insights from imbalance neural collapse's classifier norm patterns, we develop Global and Local Adaptive Feature Realignment (GLA-FR) via an auxiliary global classifier and personalized Euclidean norm transfer to align global features with client preferences. Extensive experimental results on CIFAR-10/100-LT, ImageNet, and iNaturalist demonstrate the advantage of our method over state-of-the-art pFL and Fed-LT approaches.
Zikai Xiao, Zihan Chen 0001, Liyinglan Liu, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Wanlu Liu, Howard H. Yang, Zuozhu Liu
ICLR5
2024 Sentiment Confidence Separation: A Trust-Optimized Framework for Multimodal Sentiment Classification
abstract
The Multimodal Sentiment Classification (MSC) task aims to discern sentiments from diverse data sources. Existing efforts focus on integrating multimodal features and enhancing representation learning for improved recognition. The widespread use of MSC, particularly in risk-associated domains, highlights the need for heightened trustworthiness in predictions. However, most current MSC models often provide elevated confidence regardless of whether the prediction is correct or not, with less emphasis on whether this confidence reasonably reflects the model’s certainty in predictions. This paper proposes a novel confidence optimization framework, Sentiment Confidence Separation (SCS), which helps address unreliability in MSC models by making the correct and incorrect predictions output discriminative confidences. SCS comprises Confidence Separation Loss (CSL) and Flatness-Based Separation Optimization (FBSO), facilitating reliable and precise predictions. Comprehensive experimentation validates the efficacy of the proposed approach across multiple mainstream datasets.
Zemin Tang, Zhibang Yang, Xu Zhou 0001, Cen Chen 0001, Joey Tianyi Zhou
ICME6
2024 Ladder-of-Thought: Using Knowledge as Steps to Elevate Stance Detection
abstract
Stance detection aims to determine the attitude or viewpoint expressed in a document regarding a specific target. Recent advancements in Large Language Models (LLMs), such as Chain-of-Thought (CoT) prompting, have improved the reasoning capabilities of these models by integrating intermediate rationales. However, the efficacy of CoT can be limited by the model’s internal knowledge, resulting in inaccurate rationales that compromise the subsequent stance prediction. This limitation could further lead to hallucinations, where LLMs produce unfaithful responses and erroneous reasoning, affecting the output’s reliability and precision. Moreover, CoT can be challenging to implement on smaller language models with constrained knowledge and reasoning depth, which raises concerns about efficiency. In response to these issues, we propose the Ladder-of-Thought (LoT), a novel framework using knowledge as steps to elevate stance detection. LoT implements a triple-phase Progressive Optimization Framework: 1) External Knowledge Injection, which aims to enrich the model’s intrinsic knowledge base; 2) Intermediate Knowledge Generation, allowing the model to generate more accurate and dependable intermediate knowledge to enhance the downstream prediction; and 3) Downstream Fine-tuning & Prediction, which aims to improve the model’s prediction accuracy. This sequential approach symbolizes ascending a ladder, with each phase representing a progressive step towards achieving optimal reasoning and prediction performance. Our empirical results have demonstrated that LoT achieves state-of-the-art results in zero-shot/few-shot and in-target stance detection, marking a 16% improvement over ChatGPT and a 10% enhancement compared to ChatGPT with CoT on stance detection task.
Kairui Hu, Ming Yan 0007, Wen Haw Chong, Yong Keong Yap, Cuntai Guan, Joey Tianyi Zhou, Ivor W. Tsang
IJCNN6
2024 Evolution-aware VAriance (EVA) Coreset Selection for Medical Image Classification
abstract
In the medical field, managing high-dimensional massive medical imaging data and performing reliable medical analysis from it is a critical challenge, especially in resource-limited environments such as remote medical facilities and mobile devices. This necessitates effective dataset compression techniques to reduce storage, transmission, and computational cost. However, existing coreset selection methods are primarily designed for natural image datasets, and exhibit doubtful effectiveness when applied to medical image datasets due to challenges such as intra-class variation and inter-class similarity. In this paper, we propose a novel coreset selection strategy termed as Evolution-aware VAriance (EVA), which captures the evolutionary process of model training through a dual-window approach and reflects the fluctuation of sample importance more precisely through variance measurement. Extensive experiments on medical image datasets demonstrate the effectiveness of our strategy over previous SOTA methods, especially at high compression rates. EVA achieves 98.27% accuracy with only 10% training data, compared to 97.20% for the full training set. None of the compared baseline methods can exceed Random at 5% selection rate, while EVA outperforms Random by 5.61%, showcasing its potential for efficient medical image analysis.
Yuxin Hong, Xiao Zhang 0006, Xin Zhang 0092, Joey Tianyi Zhou
ACM Multimedia4
2024 Diversity-Driven Synthesis: Enhancing Dataset Distillation through Directed Weight Adjustment
abstract
The sharp increase in data-related expenses has motivated research into condensing datasets while retaining the most informative features. Dataset distillation has thus recently come to the fore. This paradigm generates synthetic datasets that are representative enough to replace the original dataset in training a neural network. To avoid redundancy in these synthetic datasets, it is crucial that each element contains unique features and remains diverse from others during the synthesis stage. In this paper, we provide a thorough theoretical and empirical analysis of diversity within synthesized datasets. We argue that enhancing diversity can improve the parallelizable yet isolated synthesizing approach. Specifically, we introduce a novel method that employs dynamic and directed weight adjustment techniques to modulate the synthesis process, thereby maximizing the representativeness and diversity of each synthetic instance. Our method ensures that each batch of synthetic data mirrors the characteristics of a large, varying subset of the original dataset. Extensive experiments across multiple datasets, including CIFAR, Tiny-ImageNet, and ImageNet-1K, demonstrate the superior performance of our method, highlighting its effectiveness in producing diverse and representative synthetic datasets with minimal computational expense. Our code is available at https://github.com/AngusDujw/Diversity-Driven-Synthesis.
Jiawei Du 0002, Xin Zhang 0092, Wenxin Huang, Joey Tianyi Zhou
NeurIPS5
2024 The Best of Both Worlds: On the Dilemma of Out-of-distribution Detection
abstract
Out-of-distribution (OOD) detection is essential for model trustworthiness which aims to sensitively identity semantic OOD samples and robustly generalize for covariate-shifted OOD samples. However, we discover that the superior OOD detection performance of state-of-the-art methods is achieved by secretly sacrificing the OOD generalization ability. The classification accuracy frequently collapses catastrophically when even slight noise is encountered. Such a phenomenon violates the motivation of trustworthiness and significantly limits the model's deployment in the real world. What is the hidden reason behind such a limitation? In this work, we theoretically demystify the "\textit{sensitive-robust}" dilemma that lies in previous OOD detection methods. Consequently, a theory-inspired algorithm is induced to overcome such a dilemma. By decoupling the uncertainty learning objective from a Bayesian perspective, the conflict between OOD detection and OOD generalization is naturally harmonized and a dual-optimized performance could be expected. Empirical studies show that our method achieves superior performance on commonly used benchmarks. To our best knowledge, this work is the first principled OOD detection method that achieves state-of-the-art OOD detection performance without sacrificing OOD generalization ability. Our code is available at https://github.com/QingyangZhang/DUL.
Qiuxuan Feng, Joey Tianyi Zhou, Yatao Bian, Qinghua Hu, Changqing Zhang 0002
NeurIPS3
2024 Rethinking the Reliability of Post-hoc Calibration Methods Under Subpopulation Shift
Huan Ma 0006, Changqing Zhang 0002, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu
PRICAI (2)6
2024 C2Net: content-dependent and -independent cross-attention network for anomaly detection in videos
Jiafei Liang, Yang Xiao 0007, Joey Tianyi Zhou, Feng Yang 0012, Zhiwen Fang
Appl. Intell.3
2024 A principled framework for explainable multimodal disentanglement
Zongbo Han, Tao Luo 0014, Huazhu Fu, Qinghua Hu, Joey Tianyi Zhou, Changqing Zhang 0002
Inf. Sci.5
2024 SAR: Sharpness-Aware minimization for enhancing DNNs' Robustness against bit-flip errors
Changbao Zhou, Jiawei Du 0002, Ming Yan 0007, Hengshan Yue, Xiaohui Wei 0002, Joey Tianyi Zhou
J. Syst. Archit.6
2024 A concise but high-performing network for image guided depth completion in autonomous driving
Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Joey Tianyi Zhou
Knowl. Based Syst.7
2024 Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu
Mach. Learn.6
2024 Blessing few-shot segmentation via semi-supervised learning with noisy support images
Runtong Zhang, Hongyuan Zhu 0002, Hanwang Zhang, Chen Gong 0002, Joey Tianyi Zhou, Fanman Meng
Pattern Recognit.5
2024 You Will Never Walk Alone: One-Shot 3D Action Recognition With Point Cloud Sequence
abstract
In this work, we pay the first effort to address one-shot 3D action recognition in point cloud sequence, without skeleton information. The main contribution lies in two folders. First, a novel one-shot classification approach that considers the feature distribution of 3D action is proposed. We find that, for different 3D actions their dimensional-wise feature distributions are generally in Gaussian form and similar action categories hold approximate feature distributions. Accordingly, K-nearest base classes’ mean value and covariance matrix information help to form one-shot novel class’s pseudo feature distribution. To alleviate the potential ambiguous problem within nearest neighbor search, we divide the base classes into subsets via C-means clustering to facilitate the similarity measure to novel class. Meanwhile, the feature distribution of base class’s whole set and subsets will be jointly considered for generating novel class’s pseudo feature distribution. Multi-dimensional Gaussian sampling is conducted on the acquired pseudo feature distribution for feature-level data augmentation, to make one-shot novel class “never walk alone” for leveraging classifier training. Secondly to better characterize fine-grained 3D action, a temporal attention method is proposed, via introducing vision Transformer (ViT) to capture action’s discriminative short-term motion pattern with densely sampled short-term 3DV (3D dynamic voxel) features along temporal dimension. Experiments on NTU RGB+D 120 and 60 verify superiority of our approach. It outperforms state-of-the-art skeleton-based methods by 13.9% at most. The source code is available athttps://github.com/Tong-XY/YNWA.
Xingyu Tong 0002, Yang Xiao 0007, Jianyu Yang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 Learning Student Network Under Universal Label Noise
abstract
Data-free knowledge distillation aims to learn a small student network from a large pre-trained teacher network without the aid of original training data. Recent works propose to gather alternative data from the Internet for training student network. In a more realistic scenario, the data on the Internet contains two types of label noise, namely: 1) closed-set label noise, where some examples belong to the known categories but are mislabeled; and 2) open-set label noise, where the true labels of some mislabeled examples are outside the known categories. However, the latter is largely ignored by existing works, leading to limited student network performance. Therefore, this paper proposes a novel data-free knowledge distillation paradigm by utilizing a webly-collected dataset under universal label noise, which means both closed-set and open-set label noise should be tackled. Specifically, we first split the collected noisy dataset into clean set, closed noisy set, and open noisy set based on the prediction uncertainty of various data types. For the closed-set noisy examples, their labels are refined by teacher network. Meanwhile, a noise-robust hybrid contrastive learning is performed on the clean set and refined closed noisy set to encourage student network to learn the categorical and instance knowledge inherited by teacher network. For the open-set noisy examples unexplored by previous work, we regard them as unlabeled and conduct self-supervised learning on them to enrich the supervision signal for student network. Intensive experimental results on image classification tasks demonstrate that our approach can achieve superior performance to state-of-the-art data-free knowledge distillation methods.
Jialiang Tang, Ning Jiang 0002, Hongyuan Zhu 0002, Joey Tianyi Zhou, Chen Gong 0002
IEEE Trans. Image Process.4
2024 TaiChiNet: Negative-Positive Cross-Attention Network for Breast Lesion Segmentation in Ultrasound Images
abstract
Breast lesion segmentation in ultrasound images is essential for computer-aided breast-cancer diagnosis. To improve the segmentation performance, most approaches design sophisticated deep-learning models by mining the patterns of foreground lesions and normal backgrounds simultaneously or by unilaterally enhancing foreground lesions via various focal losses. However, the potential of normal backgrounds is underutilized, which could reduce false positives by compacting the feature representation of all normal backgrounds. From a novel viewpoint of bilateral enhancement, we propose a negative-positive cross-attention network to concentrate on normal backgrounds and foreground lesions, respectively. Derived from the complementing opposites of bipolarity in TaiChi, the network is denoted as TaiChiNet, which consists of the negative normal-background and positive foreground-lesion paths. To transmit the information across the two paths, a cross-attention module, a complementary MLP-head, and a complementary loss are built for deep-layer features, shallow-layer features, and mutual-learning supervision, separately. To the best of our knowledge, this is the first work to formulate breast lesion segmentation as a mutual supervision task from the foreground-lesion and normal-background views. Experimental results have demonstrated the effectiveness of TaiChiNet on two breast lesion segmentation datasets with a lightweight architecture. Furthermore, extensive experiments on the thyroid nodule segmentation and retinal optic cup/disc segmentation datasets indicate the application potential of TaiChiNet.
Jinting Wang, Jiafei Liang, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang, Feng Yang 0012
IEEE J. Biomed. Health Informatics4
2024 MENet: Multi-Modal Mapping Enhancement Network for 3D Object Detection in Autonomous Driving
abstract
To achieve more accurate perception performance, LiDAR and camera are gradually chosen to improve 3D object detection simultaneously. However, it is still a non-trivial task to build an effective fusion mechanism, and this is hindering the development of multi-modal based method. Especially, the mapping relationship construction between two modalities is far from fully explored. Canonical cross-modal mapping suffers from failure when the calibration matrix is incorrect, and it also greatly wastes the amount and density of RGB image information. This paper aims to extend the traditional one-to-one alignment relationship between LiDAR and camera. For all projected point clouds, we enhance their cross-modal mapping relationship through aggregating color-texture related feature and shape-contour related feature. Further, a mapping pyramid is proposed to leverage the semantic representation of the image feature at different stages. Based on the above mapping enhancement strategies, our method increases the engagement rate of image. Finally, we design a fusion module based on an attention mechanism to improve the point cloud feature with the auxiliary image feature. Extensive experiments on the KITTI dataset and SUN-RGBD dataset show that our model achieves satisfactory 3D object detection, especially for categories with sparse point clouds compared with other multi-modal fusion networks.
Moyun Liu, Youping Chen, Jingming Xie, Yang Zhang 0053, Zhenshan Bing, Genghang Zhuang, Kai Huang 0001, Joey Tianyi Zhou
IEEE Trans. Intell. Transp. Syst.10
2024 Democratizing Federated WiFi-Based Human Activity Recognition Using Hypothesis Transfer
abstract
Human activity recognition (HAR) is a crucial task in IoT systems with applications ranging from surveillance and intruder detection to home automation and more. Recently, non-invasive HAR utilizing WiFi signals has gained considerable attention due to advancements in ubiquitous WiFi technologies. However, recent studies have revealed significant privacy risks associated with WiFi signals, raising concerns about bio-information leakage. To address these concerns, the decentralized paradigm, particularly federated learning (FL), has emerged as a promising approach for training HAR models while preserving data privacy. Nevertheless, FL models may struggle in end-user environments due to substantial domain discrepancies between the source training data and the target end-user environment. This discrepancy arises from the sensitivity of WiFi signals to environmental changes, resulting in notable domain shifts. As a consequence, FL-based HAR approaches often face challenges when deployed in real-world WiFi environments. Albeit there are pioneer attempts on federated domain adaptation, they typically require non-trivial communication and computation cost, which is prohibitively expensive especially considering edge-based hardware equipment of end-user environment. In this paper, we propose a model to democratize the WiFi-based HAR system by enhancing recognition accuracy in unannotated end-user environments while prioritizing data privacy. Our model leverages the hypothesis transfer and a lightweight hypothesis ensemble to mitigate negative transfer. We prove a tighter theoretical upper bound compared to existing multi-source federated domain adaptation models. Extensive experiments shows our model improves the average accuracy by approximately 10 absolute percentage points in both cross-person and cross-environment settings comparing several state-of-the-art baselines.
Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Min Wu 0008, Joey Tianyi Zhou
IEEE Trans. Mob. Comput.6
2024 RCT: Resource Constrained Training for Edge AI
abstract
Efficient neural network training is essential for in situ training of edge artificial intelligence (AI) and carbon footprint reduction in general. Train neural network on the edge is challenging because there is a large gap between limited resources on edge and the resource requirement of current training methods. Existing training methods are based on the assumption that the underlying computing infrastructure has sufficient memory and energy supplies. These methods involve two copies of the model parameters, which is usually beyond the capacity of on-chip memory in processors. The data movement between off-chip and on-chip memory incurs large amounts of energy. We propose resource constrained training (RCT) to realize resource-efficient training for edge devices and servers. RCT only keeps a quantized model throughout the training so that the memory requirement for model parameters in training is reduced. It adjusts per-layer bitwidth dynamically to save energy when a model can learn effectively with lower precision. We carry out experiments with representative models and tasks in image classification, natural language processing, and crowd counting applications. Experiments show that on average, 8-15-bit weight update is sufficient for achieving SOTA performance in these applications. RCT saves 63.5%-80% memory for model parameters and saves more energy for communications. Through experiments, we observe that the common practice on the first/last layer in model compression does not apply to efficient training. Also, interestingly, the more challenging a dataset is, the lower bitwidth is required for efficient training.
Tian Huang, Tao Luo 0014, Ming Yan 0007, Joey Tianyi Zhou, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.4
2024 Beyond Pattern Variance: Unsupervised 3-D Action Representation Learning With Point Cloud Sequence
abstract
This work pays the first research effort to address unsupervised 3-D action representation learning with point cloud sequence, which is different from existing unsupervised methods that rely on 3-D skeleton information. Our proposition is built on the state-of-the-art 3-D action descriptor 3-D dynamic voxel (3DV) with contrastive learning (CL). The 3DV can compress the point cloud sequence into a compact point cloud of 3-D motion information. Spatiotemporal data augmentations are conducted on it to drive CL. However, we find that existing CL methods (e.g., SimCLR or MoCo v2) often suffer from high pattern variance toward the augmented 3DV samples from the same action instance, that is, the augmented 3DV samples are still of high feature complementarity after CL, while the complementary discriminative clues within them have not been well exploited yet. To address this, a feature augmentation adapted CL (FACL) approach is proposed, which facilitates 3-D action representation via concerning the features from all augmented 3DV samples jointly, in spirit of feature augmentation. FACL runs in a global-local way: one branch learns global feature that involves the discriminative clues from the raw and augmented 3DV samples, and the other focuses on enhancing the discriminative power of local feature learned from each augmented 3DV sample. The global and local features are fused to characterize 3-D action jointly via concatenation. To fit FACL, a series of spatiotemporal data augmentation approaches is also studied on 3DV. Wide-range experiments verify the superiority of our unsupervised learning method for 3-D action feature learning. It outperforms the state-of-the-art skeleton-based counterparts by 6.4% and 3.6% with the cross-setup and cross-subject test settings on NTU RGB+D 120, respectively. The source code is available at https://github.com/tangent-T/FACL.
Yang Xiao 0007, Yancheng Wang 0002, Jianyu Yang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Neural Networks Learn. Syst.7
2024 Efficient Spiking Neural Networks With Radix Encoding
abstract
Spiking neural networks (SNNs) have advantages in latency and energy efficiency over traditional artificial neural networks (ANNs) due to their event-driven computation mechanism and the replacement of energy-consuming weight multiplication with addition. However, to achieve high accuracy, it usually requires long spike trains to ensure accuracy, usually more than 1000 time steps. This offsets the computation efficiency brought by SNNs because a longer spike train means a larger number of operations and larger latency. In this article, we propose a radix-encoded SNN, which has ultrashort spike trains. Specifically, it is able to use less than six time steps to achieve even higher accuracy than its traditional counterpart. We also develop a method to fit our radix encoding technique into the ANN-to-SNN conversion approach so that we can train radix-encoded SNNs more efficiently on mature platforms and hardware. Experiments show that our radix encoding can achieve 25× improvement in latency and 1.7% improvement in accuracy compared to the state-of-the-art method using the VGG-16 network on the CIFAR-10 dataset.
Zhehui Wang, Xiaozhe Gu, Rick Siow Mong Goh, Joey Tianyi Zhou, Tao Luo 0014
IEEE Trans. Neural Networks Learn. Syst.4
2024 EDCompress: Energy-Aware Model Compression for Dataflows
abstract
Edge devices demand low energy consumption, cost, and small form factor. To efficiently deploy convolutional neural network (CNN) models on the edge device, energy-aware model compression becomes extremely important. However, existing work did not study this problem well because of the lack of considering the diversity of dataflow types in hardware architectures. In this article, we propose EDCompress (EDC), an energy-aware model compression method for various dataflows. It can effectively reduce the energy consumption of various edge devices, with different dataflow types. Considering the very nature of model compression procedures, we recast the optimization process to a multistep problem and solve it by reinforcement learning algorithms. We also propose a multidimensional multistep (MDMS) optimization method, which shows higher compressing capability than the traditional multistep method. Experiments show that EDC could improve 20x, 17x, and 26x energy efficiency in VGG-16, MobileNet, and LeNet-5 networks, respectively, with negligible loss of accuracy. EDC could also indicate the optimal dataflow type for specific neural networks in terms of energy consumption, which can guide the deployment of CNN on hardware.
Zhehui Wang, Tao Luo 0014, Rick Siow Mong Goh, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.4
2024 GREnet: Gradually REcurrent Network With Curriculum Learning for 2-D Medical Image Segmentation
abstract
Medical image segmentation is a vital stage in medical image analysis. Numerous deep-learning methods are booming to improve the performance of 2-D medical image segmentation, owing to the fast growth of the convolutional neural network. Generally, the manually defined ground truth is utilized directly to supervise models in the training phase. However, direct supervision of the ground truth often results in ambiguity and distractors as complex challenges appear simultaneously. To alleviate this issue, we propose a gradually recurrent network with curriculum learning, which is supervised by gradual information of the ground truth. The whole model is composed of two independent networks. One is the segmentation network denoted as GREnet, which formulates 2-D medical image segmentation as a temporal task supervised by pixel-level gradual curricula in the training phase. The other is a curriculum-mining network. To a certain degree, the curriculum-mining network provides curricula with an increasing difficulty in the ground truth of the training set by progressively uncovering hard-to-segmentation pixels via a data-driven manner. Given that segmentation is a pixel-level dense-prediction challenge, to the best of our knowledge, this is the first work to function 2-D medical image segmentation as a temporal task with pixel-level curriculum learning. In GREnet, the naive UNet is adopted as the backbone, while ConvLSTM is used to establish the temporal link between gradual curricula. In the curriculum-mining network, UNet++ supplemented by transformer is designed to deliver curricula through the outputs of the modified UNet++ at different layers. Experimental results have demonstrated the effectiveness of GREnet on seven datasets, i.e., three lesion segmentation datasets in dermoscopic images, an optic disc and cup segmentation dataset and a blood vessel segmentation dataset in retinal images, a breast lesion segmentation dataset in ultrasound images, and a lung segmentation dataset in computed tomography (CT).
Jinting Wang, Yujiao Tang, Yang Xiao 0007, Joey Tianyi Zhou, Zhiwen Fang, Feng Yang 0012
IEEE Trans. Neural Networks Learn. Syst.4
2024 Word2Pix: Word to Pixel Cross-Attention Transformer in Visual Grounding
abstract
Current one-stage methods for visual grounding encode the language query as one holistic sentence embedding before fusion with visual features for target localization. Such a formulation provides insufficient ability to model query at the word level, and therefore is prone to neglect words that may not be the most important ones for a sentence but are critical for the referred object. In this article, we propose Word2Pix: a one-stage visual grounding network based on the encoder-decoder transformer architecture that enables learning for textual to visual feature correspondence via word to pixel attention. Each word from the query sentence is given an equal opportunity when attending to visual pixels through multiple stacks of transformer decoder layers. In this way, the decoder can learn to model the language query and fuse language with the visual features for target prediction simultaneously. We conduct the experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets, and the proposed Word2Pix outperforms the existing one-stage methods by a notable margin. The results obtained also show that Word2Pix surpasses the two-stage visual grounding models, while at the same time keeping the merits of the one-stage paradigm, namely, end-to-end training and fast inference speed. Code is available at https://github.com/azurerain7/Word2Pix.
Heng Zhao 0004, Joey Tianyi Zhou, Yew-Soon Ong
IEEE Trans. Neural Networks Learn. Syst.2
2024 GrapHAR: A Lightweight Human Activity Recognition Model by Exploring the Sub-Carrier Correlations
abstract
Human activity recognition (HAR) is an important task due to its far-reaching applications, such as surveillance, healthcare systems, and human-computer interaction. Recently, Channel State Information (CSI)-based HAR has attracted increasing attention in the research community due to its ubiquitous availability, good user privacy, and fewer constraints on working conditions. Most of the existing methods for CSI-based HAR use various deep learning models, such as Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM), and Transformers, to distinguish activities based on their temporal patterns. Despite their remarkable effectiveness, these methods solely focus on temporal patterns while ignoring the correlations among sub-carriers. This limitation prevents them from achieving further performance improvement. Moreover, recent works often involve advanced yet massive and inefficient neural architectures, like Transformers, to obtain satisfactory recognition accuracy. The performance gain is traded off with a steep increase in model complexity, which leads to low efficacy and high training/inference costs outsides the small time window. To address these issues, we propose a lightweight CSI-based HAR model. Our model makes the first effort to explore the graphical correlations of CSI sub-carriers, working in conjunction with a temporal causal convolution module. The high efficacy design enables our model to be highly effective without requiring excessive model complexity. Extensive experiments conducted on four real-world datasets demonstrate that our model outperforms state-of-the-art methods, including a strong Transformer-based baseline. It achieves an average improvement of 8 percentage points in recognition accuracy, with only 10% of the parameters compared to the Transformer-based method (4.95M vs. 49.24M). Additionally, our model is significantly faster, with empirical training and execution times at least 2.07 times faster than the baseline.
Wei Meng 0002, Zhicong Liu, Bing Li 0002, Wei Cui 0002, Joey Tianyi Zhou, Le Zhang 0001
IEEE Trans. Wirel. Commun.5
2023 Minimizing the Accumulated Trajectory Error to Improve Dataset Distillation
abstract
Model-based deep learning has achieved astounding successes due in part to the availability of large-scale real-world data. However, processing such massive amounts of data comes at a considerable cost in terms of computations, storage, training and the search for good neural architectures. Dataset distillation has thus recently come to the fore. This paradigm involves distilling information from large real-world datasets into tiny and compact synthetic datasets such that processing the latter ideally yields similar performances as the former. State-of-the-art methods primarily rely on learning the synthetic dataset by matching the gradients obtained during training between the real and synthetic data. However, these gradient-matching methods suffer from the so-called accumulated trajectory error caused by the discrepancy between the distillation and subsequent evaluation. To mitigate the adverse impact of this accumulated trajectory error, we propose a novel approach that encourages the optimization algorithm to seek a flat trajectory. We show that the weights trained on synthetic data are robust against the accumulated errors perturbations with the regularization towards the flat trajectory. Our method, called Flat Trajectory Distillation (FTD), is shown to boost the performance of gradient-matching methods by up to 4.7% on a subset of images of the ImageNet dataset with higher resolution images. We also validate the effectiveness and generalizability of our method with datasets of different resolutions and demonstrate its applicability to neural architecture search. Code is available at. https://github.com/AngusDujw/FTD-distillation.
Jiawei Du 0002, Yidi Jiang, Vincent Y. F. Tan, Joey Tianyi Zhou, Haizhou Li 0001
CVPR4
2023 A2J-Transformer: Anchor-to-Joint Transformer Network for 3D Interacting Hand Pose Estimation from a Single RGB Image
abstract
3D interacting hand pose estimation from a single RGB image is a challenging task, due to serious self-occlusion and inter-occlusion towards hands, confusing similar appearance patterns between 2 hands, ill-posed joint position mapping from 2D to 3D, etc.. To address these, we propose to extend A2J-the state-of-the-art depth-based 3D single hand pose estimation method-to RGB domain under interacting hand condition. Our key idea is to equip A2J with strong local-global aware ability to well capture interacting hands' local fine details and global articulated clues among joints jointly. To this end, A2J is evolved under Transformer's non-local encoding-decoding framework to build A2J- Transformer. It holds 3 main advantages over A2J. First, self-attention across local anchor points is built to make them global spatial context aware to better capture joints' articulation clues for resisting occlusion. Secondly, each anchor point is regarded as learnable query with adaptive feature learning for facilitating pattern fitting capacity, instead of having the same local representation with the others. Last but not least, anchor point locates in 3D space instead of 2D as in A2J, to leverage 3D pose prediction. Experiments on challenging InterHand 2.6M demonstrate that, A2J-Transformer can achieve state-of-the-art model-free performance (3.38mm MPJPE advancement in 2-hand case) and can also be applied to depth domain with strong generalization. The code is avaliable at https://github.com/ChanglongJiangGit/A2J-Transformer.
Changlong Jiang, Yang Xiao 0007, Cunlin Wu, Jinghong Zheng 0002, Zhiguo Cao 0001, Joey Tianyi Zhou
CVPR7
2023 Real-time Multi-person Eyeblink Detection in the Wild for Untrimmed Video
abstract
Real-time eyeblink detection in the wild can widely serve for fatigue detection, face anti-spoofing, emotion analysis, etc. The existing research efforts generally focus on single-person cases towards trimmed video. However, multi-person scenario within untrimmed videos is also important for practical applications, which has not been well concerned yet. To address this, we shed light on this research field for the first time with essential contributions on dataset, theory, and practices. In particular, a large-scale dataset termed MPEblink that involves 686 untrimmed videos with 8748 eyeblink events is proposed under multi-person conditions. The samples are captured from uncon-strainedfilms to reveal “in the wild“ characteristics. Meanwhile, a real-time multi-person eyeblink detection method is also proposed. Being different from the existing counter-parts, our proposition runs in a one-stage spatio-temporal way with end-to-end learning capacity. Specifically, it simultaneously addresses the sub-tasks of face detection, face tracking, and human instance-level eyeblink detection. This paradigm holds 2 main advantages: (1) eyeblink features can be facilitated via the face's global context (e.g., head pose and illumination condition) with joint optimization and interaction, and (2) addressing these sub-tasks in parallel instead of sequential manner can save time remarkably to meet the real-time running requirement. Experiments on MPEblink verify the essential challenges of real-time multi-person eyeblink detection in the wild for untrimmed video. Our method also outperforms existing approaches by large margins and with a high inference speed.
Wenzheng Zeng, Yang Xiao 0007, Sicheng Wei, Jinfang Gan, Xintao Zhang, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou
CVPR8
2023 Frequency Guidance Matters in Few-Shot Learning
abstract
Few-shot classification aims to learn a discriminative feature representation to recognize unseen classes with few labeled support samples. While most few-shot learning methods focus on exploiting the spatial information of image samples, frequency representation has also been proven essential in classification tasks. In this paper, we investigate the effect of different frequency components on the few-shot learning tasks. To enhance the performance and generalizability of few-shot methods, we propose a novel Frequency-Guided Few-shot Learning framework (dubbed FGFL), which leverages the task-specific frequency components to adaptively mask the corresponding image information, with a novel multi-level metric learning strategy including a triplet loss among original, masked and unmasked image as well as a contrastive loss between masked and original support and query sets to exploit more discriminative information. Extensive experiments on four benchmarks under several few-shot scenarios, i.e., standard, cross-dataset, cross-domain, and coarse-to-fine annotated classification, are conducted. Both qualitative and quantitative results show that our proposed FGFL scheme can attend to the class-discriminative frequency components, thus integrating those information towards more effective and generalizable few-shot learning.
Hao Cheng 0016, Siyuan Yang 0001, Joey Tianyi Zhou, Lanqing Guo, Bihan Wen
ICCV3
2023 Robust Geometry-Preserving Depth Estimation Using Differentiable Rendering
abstract
In this study, we address the challenge of 3D scene structure recovery from monocular depth estimation. While traditional depth estimation methods leverage labeled datasets to directly predict absolute depth, recent advancements advocate for mix-dataset training, enhancing generalization across diverse scenes. However, such mixed dataset training yields depth predictions only up to an unknown scale and shift, hindering accurate 3D reconstructions. Existing solutions necessitate extra 3D datasets or geometry-complete depth annotations, constraints that limit their versatility. In this paper, we propose a learning framework that trains models to predict geometry-preserving depth without requiring extra data or annotations. To produce realistic 3D structures, we render novel views of the reconstructed scenes and design loss functions to promote depth estimation consistency across different views. Comprehensive experiments underscore our framework’s superior generalization capabilities, surpassing existing state-of-the-art methods on several benchmark datasets without leveraging extra training information. Moreover, our innovative loss functions empower the model to autonomously recover domain-specific scale-and-shift coefficients using solely unlabeled images.
Chi Zhang 0007, Wei Yin 0006, Gang Yu 0002, Zhibin Wang 0004, Tao Chen 0003, Joey Tianyi Zhou, Chunhua Shen
ICCV7
2023 Semi-Supervised Few-Shot Segmentation with Noisy Support Images
abstract
Motivated by the semi-supervised learning that uses the unlabeled data and pseudo annotations to improve the image classification, this paper proposes a new semi-supervised few-shot segmentation (FSS) framework of which the training process uses not only the annotated images, but also the unlabeled images, e.g. images from other available datasets, to enhance the training of the FSS model. Furthermore, in the test phase, more support images and pseudo-annotations can also be generated by the proposed framework to enrich the support set of novel classes and therefore benefit the inference. However, unlabeled images are not a free lunch. The noisy intra-class samples and inter-class samples existed in the unlabeled images as well as the interferences of the bad quality of pseudo annotations make it difficult to utilize the correct images and pseudo annotations for a certain class. To this end, we further propose a ranking algorithm consisting of an inter-class confidence term and an intra-class confidence term to efficiently utilize the pseudo annotations of the class with high quality. Extensive experiments on COCO-20idataset demonstrate that the proposed semi-supervised FSS framework is superior to many state-of-the-art methods.
Runtong Zhang, Hongyuan Zhu 0002, Hanwang Zhang, Chen Gong 0002, Joey Tianyi Zhou, Fanman Meng
ICIP5
2023 Meta Knowledge Condensation for Federated Learning
Ping Liu 0004, Xin Yu 0002, Joey Tianyi Zhou
ICLR3
2023 Calibrating Multimodal Learning
abstract
Multimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, we identify current multimodal classification methods suffer from unreliable predictive confidence that tend to rely on partial modalities when estimating confidence. Specifically, we find that the confidence estimated by current models could even increase when some modalities are corrupted. To address the issue, we introduce an intuitive principle for multimodal learning, i.e., the confidence should not increase when one modality is removed. Accordingly, we propose a novel regularization technique, i.e., Calibrating Multimodal Learning (CML) regularization, to calibrate the predictive confidence of previous methods. This technique could be flexibly equipped by existing models and improve the performance in terms of confidence calibration, classification accuracy, and model robustness.
Huan Ma 0006, Changqing Zhang 0002, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu
ICML6
2023 dugMatting: Decomposed-Uncertainty-Guided Matting
abstract
Cutting out an object and estimating its opacity mask, known as image matting, is a key task in image and video editing. Due to the highly ill-posed issue, additional inputs, typically user-defined trimaps or scribbles, are usually needed to reduce the uncertainty. Although effective, it is either time consuming or only suitable for experienced users who know where to place the strokes. In this work, we propose a decomposed-uncertainty-guided matting (dugMatting) algorithm, which explores the explicitly decomposed uncertainties to efficiently and effectively improve the results. Basing on the characteristic of these uncertainties, the epistemic uncertainty is reduced in the process of guiding interaction (which introduces prior knowledge), while the aleatoric uncertainty is reduced in modeling data distribution (which introduces statistics for both data and possible noise). The proposed matting framework relieves the requirement for users to determine the interaction areas by using simple and efficient labeling. Extensively quantitative and qualitative results validate that the proposed method significantly improves the original matting algorithms in terms of both efficiency and efficacy.
Jiawei Wu 0001, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou
ICML6
2023 Provable Dynamic Fusion for Low-Quality Multimodal Data
abstract
The inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learning paradigm. Despite its widespread use, theoretical justifications in this field are still notably lacking. Can we design a provably robust multimodal fusion method? This paper provides theoretical understandings to answer this question under a most popular multimodal fusion framework from the generalization perspective. We proceed to reveal that several uncertainty estimation solutions are naturally available to achieve robust multimodal fusion. Then a novel multimodal fusion framework termed Quality-aware Multimodal Fusion (QMF) is proposed, which can improve the performance in terms of classification accuracy and model robustness. Extensive experimental results on multiple benchmarks can support our findings.
Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng 0001
ICML6
2023 TSegFormer: 3D Tooth Segmentation in Intraoral Scans with Geometry Guided Transformer
Huimin Xiong, Kunle Li, Kaiyuan Tan, Yang Feng 0011, Joey Tianyi Zhou, Jin Hao, Haochao Ying, Jian Wu 0001, Zuozhu Liu
MICCAI (6)5
2023 Towards Distribution-Agnostic Generalized Category Discovery
abstract
Data imbalance and open-ended distribution are two intrinsic characteristics of the real visual world. Though encouraging progress has been made in tackling each challenge separately, few works dedicated to combining them towards real-world scenarios. While several previous works have focused on classifying close-set samples and detecting open-set samples during testing, it's still essential to be able to classify unknown subjects as human beings. In this paper, we formally define a more realistic task as distribution-agnostic generalized category discovery (DA-GCD): generating fine-grained predictions for both close- and open-set classes in a long-tailed open-world setting. To tackle the challenging problem, we propose a Self-**Ba**lanced **Co**-Advice co**n**trastive framework (BaCon), which consists of a contrastive-learning branch and a pseudo-labeling branch, working collaboratively to provide interactive supervision to resolve the DA-GCD task. In particular, the contrastive-learning branch provides reliable distribution estimation to regularize the predictions of the pseudo-labeling branch, which in turn guides contrastive learning through self-balanced knowledge transfer and a proposed novel contrastive loss. We compare BaCon with state-of-the-art methods from two closely related fields: imbalanced semi-supervised learning and generalized category discovery. The effectiveness of BaCon is demonstrated with superior performance over all baselines and comprehensive analysis across various datasets. Our code is publicly available.
Jianhong Bai, Zuozhu Liu, Hualiang Wang, Ruizhe Chen, Lianrui Mu, Xiaomeng Li 0001, Joey Tianyi Zhou, Yang Feng 0011, Jian Wu 0001, Haoji Hu
NeurIPS7
2023 Fast Model DeBias with Machine Unlearning
abstract
Recent discoveries have revealed that deep neural networks might behave in a biased manner in many real-world scenarios. For instance, deep networks trained on a large-scale face recognition dataset CelebA tend to predict blonde hair for females and black hair for males. Such biases not only jeopardize the robustness of models but also perpetuate and amplify social biases, which is especially concerning for automated decision-making processes in healthcare, recruitment, etc., as they could exacerbate unfair economic and social inequalities among different groups. Existing debiasing methods suffer from high costs in bias labeling or model re-training, while also exhibiting a deficiency in terms of elucidating the origins of biases within the model. To this respect, we propose a fast model debiasing method (FMD) which offers an efficient approach to identify, evaluate and remove biases inherent in trained models. The FMD identifies biased attributes through an explicit counterfactual concept and quantifies the influence of data samples with influence functions. Moreover, we design a machine unlearning-based strategy to efficiently and effectively remove the bias in a trained model with a small counterfactual dataset. Experiments on the Colored MNIST, CelebA, and Adult Income datasets demonstrate that our method achieves superior or competing classification accuracies compared with state-of-the-art retraining-based methods while attaining significantly fewer biases and requiring much less debiasing cost. Notably, our method requires only a small external dataset and updating a minimal amount of model parameters, without the requirement of access to training data that may be too large or unavailable in practice.
Ruizhe Chen, Huimin Xiong, Jianhong Bai, Tianxiang Hu, Jin Hao, Yang Feng 0011, Joey Tianyi Zhou, Jian Wu 0001, Zuozhu Liu
NeurIPS8
2023 Sequential Subset Matching for Dataset Distillation
abstract
Dataset distillation is a newly emerging task that synthesizes a small-size dataset used in training deep neural networks (DNNs) for reducing data storage and model training costs. The synthetic datasets are expected to capture the essence of the knowledge contained in real-world datasets such that the former yields a similar performance as the latter. Recent advancements in distillation methods have produced notable improvements in generating synthetic datasets. However, current state-of-the-art methods treat the entire synthetic dataset as a unified entity and optimize each synthetic instance equally . This static optimization approach may lead to performance degradation in dataset distillation. Specifically, we argue that static optimization can give rise to a coupling issue within the synthetic data, particularly when a larger amount of synthetic data is being optimized. This coupling issue, in turn, leads to the failure of the distilled dataset to extract the high-level features learned by the deep neural network (DNN) in the latter epochs. In this study, we propose a new dataset distillation strategy called Sequential Subset Matching (SeqMatch), which tackles this problem by adaptively optimizing the synthetic data to encourage sequential acquisition of knowledge during dataset distillation. Our analysis indicates that SeqMatch effectively addresses the coupling issue by sequentially generating the synthetic instances, thereby enhancing its performance significantly. Our proposed SeqMatch outperforms state-of-the-art methods in various datasets, including SVNH, CIFAR-10, CIFAR-100, and Tiny ImageNet.
Jiawei Du 0002, Qin Shi 0005, Joey Tianyi Zhou
NeurIPS3
2023 You Only Condense Once: Two Rules for Pruning Condensed Datasets
abstract
Dataset condensation is a crucial tool for enhancing training efficiency by reducing the size of the training dataset, particularly in on-device scenarios. However, these scenarios have two significant challenges: 1) the varying computational resources available on the devices require a dataset size different from the pre-defined condensed dataset, and 2) the limited computational resources often preclude the possibility of conducting additional condensation processes. We introduce You Only Condense Once (YOCO) to overcome these limitations. On top of one condensed dataset, YOCO produces smaller condensed datasets with two embarrassingly simple dataset pruning rules: Low LBPE Score and Balanced Construction. YOCO offers two key advantages: 1) it can flexibly resize the dataset to fit varying computational constraints, and 2) it eliminates the need for extra condensation processes, which can be computationally prohibitive. Experiments validate our findings on networks including ConvNet, ResNet and DenseNet, and datasets including CIFAR-10, CIFAR-100 and ImageNet. For example, our YOCO surpassed various dataset condensation and dataset pruning methods on CIFAR-10 with ten Images Per Class (IPC), achieving 6.98-8.89% and 6.31-23.92% accuracy gains, respectively. The code is available at: [https://github.com/he-y/you-only-condense-once](https://github.com/he-y/you-only-condense-once).
Yang He 0002, Lingao Xiao, Joey Tianyi Zhou
NeurIPS3
2023 Cross-modal Active Complementary Learning with Self-refining Correspondence
abstract
Recently, image-text matching has attracted more and more attention from academia and industry, which is fundamental to understanding the latent correspondence across visual and textual modalities. However, most existing methods implicitly assume the training pairs are well-aligned while ignoring the ubiquitous annotation noise, a.k.a noisy correspondence (NC), thereby inevitably leading to a performance drop. Although some methods attempt to address such noise, they still face two challenging problems: excessive memorizing/overfitting and unreliable correction for NC, especially under high noise. To address the two problems, we propose a generalized Cross-modal Robust Complementary Learning framework (CRCL), which benefits from a novel Active Complementary Loss (ACL) and an efficient Self-refining Correspondence Correction (SCC) to improve the robustness of existing methods. Specifically, ACL exploits active and complementary learning losses to reduce the risk of providing erroneous supervision, leading to theoretically and experimentally demonstrated robustness against NC. SCC utilizes multiple self-refining processes with momentum correction to enlarge the receptive field for correcting correspondences, thereby alleviating error accumulation and achieving accurate and stable corrections. We carry out extensive experiments on three image-text benchmarks, i.e., Flickr30K, MS-COCO, and CC152K, to verify the superior robustness of our CRCL against synthetic and real-world noisy correspondences.
Yuan Sun 0016, Dezhong Peng, Joey Tianyi Zhou, Xi Peng 0001, Peng Hu 0002
NeurIPS4
2023 Fed-GraB: Federated Long-tailed Learning with Self-Adjusting Gradient Balancer
abstract
Data privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset; if the datasets can be globally aggregated, they jointly exhibit a long-tailed distribution. Under such a setting, existing federated optimization and/or centralized long-tailed learning methods hardly apply due to challenges in (a) characterizing the global long-tailed distribution under privacy constraints and (b) adjusting the local learning strategy to cope with the head-tail imbalance. In response, we propose a method termed $\texttt{Fed-GraB}$, comprised of a Self-adjusting Gradient Balancer (SGB) module that re-weights clients' gradients in a closed-loop manner, based on the feedback of global long-tailed distribution evaluated by a Direct Prior Analyzer (DPA) module. Using $\texttt{Fed-GraB}$, clients can effectively alleviate the distribution drift caused by data heterogeneity during the model training process and obtain a global model with better performance on the minority classes while maintaining the performance of the majority classes. Extensive experiments demonstrate that $\texttt{Fed-GraB}$ achieves state-of-the-art performance on representative datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist.
Zikai Xiao, Zihan Chen 0001, Songshang Liu, Hualiang Wang, Yang Feng 0011, Jin Hao, Joey Tianyi Zhou, Jian Wu 0001, Howard H. Yang, Zuozhu Liu
NeurIPS7
2023 Multi-spectral template matching based object detection in a few-shot learning manner
Chen Feng 0002, Zhiguo Cao 0001, Yang Xiao 0007, Zhiwen Fang, Joey Tianyi Zhou
Inf. Sci.5
2023 TC-SEPM: Characterizing soft error resilience of CNNs on Tensor Cores from program and microarchitecture perspectives
Xiaohui Wei 0002, Changbao Zhou, Hengshan Yue, Joey Tianyi Zhou
J. Syst. Archit.4
2023 Toward Communication-Efficient Digital Twin via AI-Powered Transmission and Reconstruction
abstract
Digital twin technology has recently gathered pace in engineering communities as it allows for the convergence of the real structure and its digital counterpart. 3D point cloud data is a more effective way to describe the real world and to reconstruct the digital counterpart than the conventional 2D images or 360-degree images. Large-scale, e.g., city-scale digital twins, typically collect point cloud data via internet-of-things (IoT) devices and transmit it over wireless networks. However, the existing wireless transmission technology can not carry real-time point cloud transmission for digital twin reconstruction due to mass data volume, high processing overheads, and low delay-tolerance. We propose a novel artificial intelligence (AI) powered end-to-end framework, termed AIRec, for efficient digital twin communication from point cloud compression, wireless channel coding, and digital twin reconstruction. AIRec adopts the encoder-decoder architecture. In the encoder, a novel importance-aware pooling scheme is designed to adaptively select important points with learnable thresholds to reduce the transmission volume. We also design a novel noise-aware joint source and channel coding is proposed to adaptively adjust the transmission strategy based on SNR and map the features to error-resilient channel symbols for wireless transmission to achieve a good tradeoff between the transmission rate and reconstruction quality. The decoder can accurately reconstruct the digital twins from the received symbols. Extensive experiments of typical datasets and comparison with baselines show that we achieve a good reconstruction quality under$24\times $compression ratio.
Cen Chen 0002, Xulei Yang, Joey Tianyi Zhou, Tao Zhang 0019, Yangfan Li 0001
IEEE J. Sel. Areas Commun.4
2023 Contrastive domain adaptation with consistency match for automated pneumonia diagnosis
Yangqin Feng, Zizhou Wang, Xinxing Xu, Yan Wang 0015, Huazhu Fu, Shaohua Li 0003, Liangli Zhen, Xiaofeng Lei, Yingnan Cui, Jordan Zheng Ting Sim, Yonghan Ting, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Cher Heng Tan
Medical Image Anal.12
2023 Trusted Multi-View Classification With Dynamic Evidential Fusion
abstract
Existing multi-view classification algorithms focus on promoting accuracy by exploiting different views, typically integrating them into common representations for follow-up tasks. Although effective, it is also crucial to ensure the reliability of both the multi-view integration and the final decision, especially for noisy, corrupted and out-of-distribution data. Dynamically assessing the trustworthiness of each view for different samples could provide reliable integration. This can be achieved through uncertainty estimation. With this in mind, we propose a novel multi-view classification algorithm, termed trusted multi-view classification (TMC), providing a new paradigm for multi-view learning by dynamically integrating different views at an evidence level. The proposed TMC can promote classification reliability by considering evidence from each view. Specifically, we introduce the variational Dirichlet to characterize the distribution of the class probabilities, parameterized with evidence from different views and integrated with the Dempster-Shafer theory. The unified learning framework induces accurate uncertainty and accordingly endows the model with both reliability and robustness against possible noise or corruption. Both theoretical and experimental results validate the effectiveness of the proposed model in accuracy, robustness and trustworthiness.
Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 DifFormer: Multi-Resolutional Differencing Transformer With Dynamic Ranging for Time Series Analysis
abstract
Time series analysis is essential to many far-reaching applications of data science and statistics including economic and financial forecasting, surveillance, and automated business processing. Though being greatly successful of Transformer in computer vision and natural language processing, the potential of employing it as the general backbone in analyzing the ubiquitous times series data has not been fully released yet. Prior Transformer variants on time series highly rely on task-dependent designs and pre-assumed "pattern biases", revealing its insufficiency in representing nuanced seasonal, cyclic, and outlier patterns which are highly prevalent in time series. As a consequence, they can not generalize well to different time series analysis tasks. To tackle the challenges, we propose DifFormer, an effective and efficient Transformer architecture that can serve as a workhorse for a variety of time-series analysis tasks. DifFormer incorporates a novel multi-resolutional differencing mechanism, which is able to progressively and adaptively make nuanced yet meaningful changes prominent, meanwhile, the periodic or cyclic patterns can be dynamically captured with flexible lagging and dynamic ranging operations. Extensive experiments demonstrate DifFormer significantly outperforms state-of-the-art models on three essential time-series analysis tasks, including classification, regression, and forecasting. In addition to its superior performances, DifFormer also excels in efficiency - a linear time/memory complexity with empirically lower time consumption.
Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Ce Zhu, Wei Wang 0011, Ivor W. Tsang, Joey Tianyi Zhou
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 Temporal Sentence Grounding in Videos: A Survey and Future Directions
abstract
Temporal sentence grounding in videos (TSGV), a.k.a., natural language video localization (NLVL) or video moment retrieval (VMR), aims to retrieve a temporal moment that semantically corresponds to a language query from an untrimmed video. Connecting computer vision and natural language, TSGV has drawn significant attention from researchers in both communities. This survey attempts to provide a summary of fundamental concepts in TSGV and current research status, as well as future research directions. As the background, we present a common structure of functional components in TSGV, in a tutorial style: from feature extraction from raw video and language query, to answer prediction of the target moment. Then we review the techniques for multimodal understanding and interaction, which is the key focus of TSGV for effective alignment between the two modalities. We construct a taxonomy of TSGV techniques and elaborate the methods in different categories with their strengths and weaknesses. Lastly, we discuss issues with the current TSGV research and share our insights about promising research directions.
Hao Zhang 0048, Aixin Sun, Joey Tianyi Zhou
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Accelerating Attention Mechanism on FPGAs based on Efficient Reconfigurable Systolic Array
abstract
Transformer model architectures have recently received great interest in natural language, machine translation, and computer vision, where attention mechanisms are their building blocks. However, the attention mechanism is expensive because of its intensive matrix computations and complicated data flow. The existing hardware architecture has some disadvantages for the computing structure of attention, such as inflexibility and low efficiency. Most of the existing papers accelerate attention by reducing the amount of computation through various pruning algorithms, which will affect the results to a certain extent with different sparsity. This paper proposes the hardware accelerator for the multi-head attention (MHA) on field-programmable gate arrays (FPGAs) with reconfigurable architecture, efficient systolic array, and hardware-friendly radix-2 softmax. We propose a novel method called Four inputs Processing Element (FPE) to double the computation rate of the data-aware systolic array (SA) and make it efficient and load balance. Especially, the computation framework is well designed to ensure the utilization of SA efficiently. Our design is evaluated on a Xilinx Alveo U250 card, and the proposed architecture achieves 51.3×, 17.3× improvement in latency, and 54.4×, 17.9× energy savings compared to CPU and GPU.
Wenhua Ye, Xu Zhou 0001, Joey Tianyi Zhou, Cen Chen 0002, Kenli Li 0001
ACM Trans. Embed. Comput. Syst.3
2023 Eyelid's Intrinsic Motion-Aware Feature Learning for Real-Time Eyeblink Detection in the Wild
abstract
Real-time eyeblink detection in the wild is a recently emerged challenging task that suffers from dramatic variations in face attribute, pose, illumination, camera view and distance, etc. One key issue is to well characterize eyelid’s intrinsic motion (i.e., approaching and departure between upper and lower eyelid) robustly, under unconstrained conditions. Towards this, a novel eyelid’s intrinsic motion-aware feature learning approach is proposed. Our proposition lies in 3 folds. First, the feature extractor is led to focus on informative eye region adaptively via introducing visual attention in a coarse-to-fine way, to guarantee robustness and fine-grained descriptive ability jointly. Then, 2 constraints are proposed to make feature learning be aware of eyelid’s intrinsic motion. Particularly, one concerns the fact that the inter-frame feature divergence within eyeblink processes should be greater than non-eyeblink ones to better reveal eyelid’s intrinsic motion. The other constraint minimizes the inter-frame feature divergence of non-eyeblink samples, to suppress motion clues due to head or camera movement, illumination change, etc. Meanwhile, concerning the high ambiguity between eyeblink and non-eyeblink samples, soft sample labels are acquired via self-knowledge distillation to conduct feature learning with finer supervision than the hard ones. The experiments verify that, our proposition is significantly superior to the state-of-the-art ones (i.e., advantage on F1-score over 7%) and with real-time running efficiency. It is also of strong generalization capacity towards constrained conditions. The source code is available athttps://github.com/wenzhengzeng/blink_eyelid.
Wenzheng Zeng, Yang Xiao 0007, Guilei Hu, Zhiguo Cao 0001, Sicheng Wei, Zhiwen Fang, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Inf. Forensics Secur.7
2023 Graph Neural Networks With Triple Attention for Few-Shot Learning
abstract
Recent advances in Graph Neural Networks (GNNs) have achieved superior results in many challenging tasks, such as few-shot learning. Despite its capacity to learn and generalize a model from only a few annotated samples, GNN is limited in scalability, as deep GNN models usually suffer from severe over-fitting and over-smoothing. In this work, we propose a novel GNN framework with atriple-attention mechanism,i.e.node self-attention, neighbor attention, and layer memory attention, to tackle these challenges. We provide both theoretical analysis and illustrations to explain why the proposed attentive modules can improve GNN scalability for few-shot learning tasks. Our experiments show that the proposed Attentive GNN model outperforms the state-of-the-art few-shot learning methods using both GNN and non-GNN approaches. The improvement is consistent over the mini-ImageNet, tiered-ImageNet, CUB-200-2011, and Flowers-102 benchmarks, using both ConvNet-4 and ResNet-12 backbones, and under both the inductive and transductive settings. Furthermore, we demonstrate the superiority of our method for few-shot fine-grained and semi-supervised classification tasks with extensive experiments. The code for this work is publicly available athttps://github.com/chenghao-ch94/AGNN.
Hao Cheng 0016, Joey Tianyi Zhou, Wee-Peng Tay, Bihan Wen
IEEE Trans. Multim.2
2023 CoRec: An Efficient Internet Behavior-based Recommendation Framework with Edge-cloud Collaboration on Deep Convolution Neural Networks
abstract
Both accurate and fast mobile recommendation systems based on click behaviors analysis are crucial in e-business. Deep learning has achieved state-of-the-art accuracy and the traditional wisdom often hosts these computation-intensive models in powerful cloud centers. However, the cloud-only approaches put significant computational pressure on cloud servers and increase the latency in heavy-load scenarios. Moreover, existing work often adopts RNN structures to model behaviors that suffer from low processing speed for under-utilization of parallel devices such as GPUs. In this work, we propose an efficient internet behavior-based recommendation framework with edge-cloud collaboration on deep CNNs (CoRec) to improve both the accuracy and speed for mobile recommendation. A novel convolutional interest network (CIN) improves the accuracy by modeling the long- and short-term interests and accelerates the prediction through parallel-friendly convolutions. To further improve the serving throughput and latency, a novel device-cloud collaboration strategy reduces workloads by pre-computing and caching long-term interests in the cloud offline and real-time computation of short-term interests in devices. Extensive experiments on real-world datasets show that CoRec significantly outperforms the state-of-the-art methods in accuracy and has achieved at least an order of magnitude improvement in latency and throughput compared to cloud-only RNN-based approaches for long behaviors.
Yangfan Li 0001, Kenli Li 0001, Wei Wei 0006, Joey Tianyi Zhou, Cen Chen 0002
ACM Trans. Sens. Networks4
2022 Perceiving the World: Question-guided Reinforcement Learning for Text-based Games
abstract
Text-based games provide an interactive way to study natural language processing.While deep reinforcement learning has shown effectiveness in developing the game playing agent, the low sample efficiency and the large action space remain to be the two major challenges that hinder the DRL from being applied in the real world.In this paper, we address the challenges by introducing world-perceiving modules, which automatically decompose tasks and prune actions by answering questions about the environment.We then propose a two-phase training framework to decouple language learning from reinforcement learning, which further improves the sample efficiency.The experimental results show that the proposed method significantly improves the performance and sample efficiency.Besides, it shows robustness against compound error and limited pre-training data.
Yunqiu Xu, Ling Chen 0006, Yali Du 0001, Joey Tianyi Zhou, Chengqi Zhang
ACL (1)5
2022 C3P: Cross-Domain Pose Prior Propagation for Weakly Supervised 3D Human Pose Estimation
Cunlin Wu, Yang Xiao 0007, Boshen Zhang, Zhiguo Cao 0001, Joey Tianyi Zhou
ECCV (5)6
2022 Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
Jiawei Du 0002, Hanshu Yan, Jiashi Feng, Joey Tianyi Zhou, Liangli Zhen, Rick Siow Mong Goh, Vincent Y. F. Tan
ICLR4
2022 Deep Unfolding for Compressed Sensing with Denoiser
abstract
Recent years have witnessed increasingly more exercises and uses of deep unfolding network (DUN) in image compressed sensing (CS) due to its high performance and interpretability. However, the existing DUN does not make full use of more flexible regularization methods. Besides, the intermediate information generated during the iterations of the DUN, which is crucial for the quality improvement of image reconstruction, has been largely overlooked in the existing methods. To alleviate this problem, we propose a novel DUN for image CS with regularization by denoising which casts half quadratic splitting (HQS) algorithm into the neural network. Further, we design an information collection strategy to leverage the useful information generated during the iterations. The information is provided to the image denoiser of the proposed network, which could enhance the image processing ability of the denoiser. The extensive experiments demonstrate that the proposed method is more efficient and achieves state-of-the-art reconstruction quality.
Joey Tianyi Zhou, Xiao Zhang 0006, Yu Zhou 0027
ICME2
2022 Sharpness-Aware Training for Free
abstract
Modern deep neural networks (DNNs) have achieved state-of-the-art performances but are typically over-parameterized. The over-parameterization may result in undesirably large generalization error in the absence of other customized training strategies. Recently, a line of research under the name of Sharpness-Aware Minimization (SAM) has shown that minimizing a sharpness measure, which reflects the geometry of the loss landscape, can significantly reduce the generalization error. However, SAM-like methods incur a two-fold computational overhead of the given base optimizer (e.g. SGD) for approximating the sharpness measure. In this paper, we propose Sharpness-Aware Training for Free, or SAF, which mitigates the sharp landscape at almost zero additional computational cost over the base optimizer. Intuitively, SAF achieves this by avoiding sudden drops in the loss in the sharp local minima throughout the trajectory of the updates of the weights. Specifically, we suggest a novel trajectory loss, based on the KL-divergence between the outputs of DNNs with the current weights and past weights, as a replacement of the SAM's sharpness measure. This loss captures the rate of change of the training loss along the model's update trajectory. By minimizing it, SAF ensures the convergence to a flat minimum with improved generalization capabilities. Extensive empirical results show that SAF minimizes the sharpness in the same way that SAM does, yielding better results on the ImageNet dataset with essentially the same computational cost as the base optimizer.
Jiawei Du 0002, Daquan Zhou, Jiashi Feng, Vincent Y. F. Tan, Joey Tianyi Zhou
NeurIPS5
2022 Multi-Scale Adaptive Network for Single Image Denoising
abstract
Multi-scale architectures have shown effectiveness in a variety of tasks thanks to appealing cross-scale complementarity. However, existing architectures treat different scale features equally without considering the scale-specific characteristics, \textit{i.e.}, the within-scale characteristics are ignored in the architecture design. In this paper, we reveal this missing piece for multi-scale architecture design and accordingly propose a novel Multi-Scale Adaptive Network (MSANet) for single image denoising. Specifically, MSANet simultaneously embraces the within-scale characteristics and the cross-scale complementarity thanks to three novel neural blocks, \textit{i.e.}, adaptive feature block (AFeB), adaptive multi-scale block (AMB), and adaptive fusion block (AFuB). In brief, AFeB is designed to adaptively preserve image details and filter noises, which is highly expected for the features with mixed details and noises. AMB could enlarge the receptive field and aggregate the multi-scale information, which meets the need of contextually informative features. AFuB devotes to adaptively sampling and transferring the features from one scale to another scale, which fuses the multi-scale features with varying characteristics from coarse to fine. Extensive experiments on both three real and six synthetic noisy image datasets show the superiority of MSANet compared with 12 methods. The code could be accessed from https://github.com/XLearning-SCU/2022-NeurIPS-MSANet.
Yuanbiao Gou, Peng Hu 0002, Jiancheng Lv 0001, Joey Tianyi Zhou, Xi Peng 0001
NeurIPS4
2022 Adversarial Semantic Hallucination for Domain Generalized Semantic Segmentation
abstract
Convolutional neural networks typically perform poorly when the test (target domain) and training (source domain) data have significantly different distributions. While this problem can be mitigated by using the target domain data to align the source and target domain feature representations, the target domain data may be unavailable due to privacy concerns. Consequently, there is a need for methods that generalize well despite restricted access to target domain data during training. In this work, we propose an adversarial semantic hallucination approach (ASH), which combines a class-conditioned hallucination module and a semantic segmentation module. Since the segmentation performance varies across different classes, we design a semantic-conditioned style hallucination module to generate affine transformation parameters from semantic information in the segmentation probability maps of the source domain image. Unlike previous adaptation approaches, which treat all classes equally, ASH considers the class-wise differences. The segmentation module and the hallucination module compete adversarially, with the hallucination module generating increasingly "difficult" stylized images to challenge the segmentation module. In response, the segmentation module improves as it is trained with generated samples at an appropriate class-wise difficulty level. Our results on the Cityscapes and Mapillary benchmark datasets show that our method is competitive with state of the art work. Code is made available at https://github.com/gabriel-tjio/ASH.
Gabriel Tjio, Ping Liu 0004, Joey Tianyi Zhou, Rick Siow Mong Goh
WACV3
2022 Dynamic thresholding for video anomaly detection
abstract
Abstract Anomaly detection is one of the most important applications in video surveillance that involves the temporal localisation of anomaly events in unannotated video sequences. By learning the normal patterns to generate frames and calculating their reconstruction error relative to the ground truth, a frame can be recognised as being abnormal if the reconstruction error exceeds a threshold. Most existing works use a fixed threshold that computes over all the testing data to determine the anomalies. However, fixed threshold strategy cannot address the challenges brought by the dynamic environment, e.g. changes in illumination conditions. In this paper, a dynamic thresholding algorithm (DTA) is proposed, which is fully data‐driven and capable of automatically determining thresholds such that the developed anomaly detection system can flexibly adapt to different scenarios. The proposed DTA is independent of the backbone network and can be easily incorporated into most existing video anomaly detection models to help identify the appropriate thresholds. On both synthetic and real‐world datasets, the experimental results show that with the proposed DTA, the video anomaly detection methods achieve a better performance considering the changes in dynamic environment.
Diyang Jia, Xiao Zhang 0006, Joey Tianyi Zhou, Pan Lai, Yifei Wei
IET Image Process.3
2022 Enhanced gradient learning for deep neural networks
abstract
Abstract Deep neural networks have achieved great success in both computer vision and natural language processing tasks. How to improve the gradient flows is crucial in training very deep neural networks. To address this challenge, a gradient enhancement approach is proposed through constructing the short circuit neural connections. The proposed short circuit is a unidirectional neural connection that back propagates the sensitivities rather than gradients in neural networks from the deep layers to the shallow layers. Moreover, the short circuit is further formulated as a gradient truncation operation in its connecting layers, which can be plugged into the backbone models without introducing extra training parameters. Extensive experiments demonstrate that the deep neural networks, with the help of short circuit connection, gain a large margin of improvement over the baselines on both computer vision and natural language processing tasks. The work provides the promising solution to the low‐resource scenarios, such as, intelligence transport systems of computer vision, question answering of natural language processing.
Ming Yan 0007, Jianxi Yang, Cen Chen 0001, Joey Tianyi Zhou, Yi Pan 0001, Zeng Zeng
IET Image Process.4
2022 Introduction to the Special Issue on edge intelligence: Neurocomputing meets edge computing
Zeng Zeng, Cen Chen 0002, Bharadwaj Veeravalli, Keqin Li 0001, Joey Tianyi Zhou
Neurocomputing5
2022 XAI Beyond Classification: Interpretable Neural Clustering
abstract
In this paper, we study two challenging problems in explainable AI (XAI) and data clustering. The first is how to directly design a neural network with inherent interpretability, rather than giving post-hoc explanations of a black-box model. The second is implementing discrete $k$-means with a differentiable neural network that embraces the advantages of parallel computing, online clustering, and clustering-favorable representation learning. To address these two challenges, we design a novel neural network, which is a differentiable reformulation of the vanilla $k$-means, called inTerpretable nEuraL cLustering (TELL). Our contributions are threefold. First, to the best of our knowledge, most existing XAI works focus on supervised learning paradigms. This work is one of the few XAI studies on unsupervised learning, in particular, data clustering. Second, TELL is an interpretable, or the so-called intrinsically explainable and transparent model. In contrast, most existing XAI studies resort to various means for understanding a black-box model with post-hoc explanations. Third, from the view of data clustering, TELL possesses many properties highly desired by $k$-means, including but not limited to online clustering, plug-and-play module, parallel computing, and provable convergence. Extensive experiments show that our method achieves superior performance comparing with 14 clustering approaches on three challenging data sets. The source code could be accessed at www.pengxi.me.
Xi Peng 0001, Yunfan Li 0003, Ivor W. Tsang, Hongyuan Zhu 0002, Jiancheng Lv 0001, Joey Tianyi Zhou
J. Mach. Learn. Res.6
2022 APT: The master-copy-free training method for quantised neural network on edge devices
Tian Huang, Tao Luo 0014, Joey Tianyi Zhou
J. Parallel Distributed Comput.3
2022 Hardware-software co-exploration with racetrack memory based in-memory computing for CNN inference in embedded systems
Benjamin Chen Ming Choong, Tao Luo 0014, Cheng Liu 0008, Bingsheng He, Wei Zhang 0012, Joey Tianyi Zhou
J. Syst. Archit.6
2022 Adversarial multimodal fusion with attention mechanism for skin lesion classification using clinical and dermoscopic images
Yan Wang 0015, Yangqin Feng, Lei Zhang 0005, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Liangli Zhen
Medical Image Anal.4
2022 Deep Partial Multi-View Learning
abstract
Although multi-view learning has made significant progress over the past few decades, it is still challenging due to the difficulty in modeling complex correlations among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Networks (CPM-Nets), which aims to fully and flexibly take advantage of multiple partial views. We first provide a formal definition of completeness and versatility for multi-view representation and then theoretically prove the versatility of the learned latent representations. For completeness, the task of learning latent multi-view representation is specifically translated to a degradation process by mimicking data transmission, such that the optimal tradeoff between consistency and complementarity across different views can be implicitly achieved. Equipped with adversarial strategy, our model stably imputes missing views, encoding information from all views for each sample to be encoded into latent representation to further enhance the completeness. Furthermore, a nonparametric classification loss is introduced to produce structured representations and prevent overfitting, which endows the algorithm with promising generalization under view-missing cases. Extensive experimental results validate the effectiveness of our algorithm over existing state of the arts for classification, representation learning and data imputation.
Changqing Zhang 0002, Yajie Cui, Zongbo Han, Joey Tianyi Zhou, Huazhu Fu, Qinghua Hu
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Natural Language Video Localization: A Revisit in Span-Based Question Answering Framework
abstract
Natural Language Video Localization (NLVL) aims to locate a target moment from an untrimmed video that semantically corresponds to a text query. Existing approaches mainly solve the NLVL problem from the perspective of computer vision by formulating it as ranking, anchor, or regression tasks. These methods suffer from large performance degradation when localizing on long videos. In this work, we address the NLVL from a new perspective, i.e., span-based question answering (QA), by treating the input video as a text passage. We propose a video span localizing network (VSLNet), on top of the standard span-based QA framework (named VSLBase), to address NLVL. VSLNet tackles the differences between NLVL and span-based QA through a simple yet effective query-guided highlighting (QGH) strategy. QGH guides VSLNet to search for the matching video span within a highlighted region. To address the performance degradation on long videos, we further extend VSLNet to VSLNet-L by applying a multi-scale split-and-concatenation strategy. VSLNet-L first splits the untrimmed video into short clip segments; then, it predicts which clip segment contains the target moment and suppresses the importance of other segments. Finally, the clip segments are concatenated, with different confidences, to locate the target moment accurately. Extensive experiments on three benchmark datasets show that the proposed VSLNet and VSLNet-L outperform the state-of-the-art methods; VSLNet-L addresses the issue of performance degradation on long videos. Our study suggests that the span-based QA framework is an effective strategy to solve the NLVL problem.
Hao Zhang 0048, Aixin Sun, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Locality-Aware Crowd Counting
abstract
Imbalanced data distribution in crowd counting datasets leads to severe under-estimation and over-estimation problems, which has been less investigated in existing works. In this paper, we tackle this challenging problem by proposing a simple but effective locality-based learning paradigm to produce generalizable features by alleviating sample bias. Our proposed method is locality-aware in two aspects. First, we introduce a locality-aware data partition (LADP) approach to group the training data into different bins via locality-sensitive hashing. As a result, a more balanced data batch is then constructed by LADP. To further reduce the training bias and enhance the collaboration with LADP, a new data augmentation method called locality-aware data augmentation (LADA) is proposed where the image patches are adaptively augmented based on the loss. The proposed method is independent of the backbone network architectures, and thus could be smoothly integrated with most existing deep crowd counting approaches in an end-to-end paradigm to boost their performance. We also demonstrate the versatility of the proposed method by applying it for adversarial defense. Extensive experiments verify the superiority of the proposed method over the state of the arts.
Joey Tianyi Zhou, Le Zhang 0001, Jiawei Du 0002, Xi Peng 0001, Zhiwen Fang, Hongyuan Zhu 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Hierarchical Graph Neural Networks for Few-Shot Learning
abstract
Recent graph neural network (GNN) based methods for few-shot learning (FSL) represent the samples of interest as a fully-connected graph and conduct reasoning on the nodes flatly, which ignores the hierarchical correlations among nodes. However, real-world categories may have hierarchical structures, and for FSL, it is important to extract the distinguishing features of the categories from individual samples. To explore this, we propose a novel hierarchical graph neural network (HGNN) for FSL, which consists of three parts, i.e., bottom-up reasoning, top-down reasoning, and skip connections, to enable the efficient learning of multi-level relationships. For the bottom-up reasoning, we design intra-class k-nearest neighbor pooling (intra-class knnPool) and inter-class knnPool layers, to conduct hierarchical learning for both the intra- and inter-class nodes. For the top-down reasoning, we propose to utilize graph unpooling (gUnpool) layers to restore the down-sampled graph into its original size. Skip connections are proposed to fuse multi-level features for the final node classification. The parameters of HGNN are learned by episodic training with the signal of node losses, which aims to train a well-generalizable model for recognizing unseen classes with few labeled data. Experimental results on benchmark datasets have demonstrated that HGNN outperforms other state-of-the-art GNN based methods significantly, for both transductive and non-transductive FSL tasks. The dataset as well as the source code can be downloaded online1
Cen Chen 0002, Kenli Li 0001, Wei Wei 0006, Joey Tianyi Zhou, Zeng Zeng
IEEE Trans. Circuits Syst. Video Technol.4
2022 Feature Selection With Multi-Source Transfer
abstract
Feature selection aims at choosing a subset of features to represent the original feature space. In practice, however, it is hard to achieve desirable performance due to limited training data. To alleviate this issue, we propose a novel problem named feature selection with multi-source transfer where the privileged information from another data source or modality– only available during the training phase, is exploited to improve the performance of feature selection. To be exact, we propose a novel objective function that formulates the privileged information into feature selection. Moreover, an efficient optimization algorithm is introduced to solve the proposed problem of high dimension. Extensive experimental results demonstrate that the proposed algorithm significantly outperforms several popular algorithms, especially when the training data size and the selected feature size are small.
Joey Tianyi Zhou
IEEE Trans. Circuits Syst. Video Technol.1
2022 Deep-LIFT: Deep Label-Specific Feature Learning for Image Annotation
abstract
Image annotation aims to jointly predict multiple tags for an image. Although significant progress has been achieved, existing approaches usually overlook aligning specific labels and their corresponding regions due to the weak supervised information (i.e., "bag of labels" for regions), thus failing to explicitly exploit the discrimination from different classes. In this article, we propose the deep label-specific feature (Deep-LIFT) learning model to build the explicit and exact correspondence between the label and the local visual region, which improves the effectiveness of feature learning and enhances the interpretability of the model itself. Deep-LIFT extracts features for each label by aligning each label and its region. Specifically, Deep-LIFTs are achieved through learning multiple correlation maps between image convolutional features and label embeddings. Moreover, we construct two variant graph convolutional networks (GCNs) to further capture the interdependency among labels. Empirical studies on benchmark datasets validate that the proposed model achieves superior performance on multilabel classification over other existing state-of-the-art methods.
Junbing Li, Changqing Zhang 0002, Joey Tianyi Zhou, Huazhu Fu, Shuyin Xia, Qinghua Hu
IEEE Trans. Cybern.3
2022 Point Adversarial Self-Mining: A Simple Method for Facial Expression Recognition
abstract
In this article, we propose a simple yet effective approach, called point adversarial self mining (PASM), to improve the recognition accuracy in facial expression recognition (FER). Unlike previous works focusing on designing specific architectures or loss functions to solve this problem, PASM boosts the network capability by simulating human learning processes: providing updated learning materials and guidance from more capable teachers. Specifically, to generate new learning materials, PASM leverages a point adversarial attack method and a trained teacher network to locate the most informative position related to the target task, generating harder learning samples to refine the network. The searched position is highly adaptive since it considers both the statistical information of each sample and the teacher network capability. Other than being provided new learning materials, the student network also receives guidance from the teacher network. After the student network finishes training, the student network changes its role and acts as a teacher, generating new learning materials and providing stronger guidance to train a better student network. The adaptive learning materials generation and teacher/student update can be conducted more than one time, improving the network capability iteratively. Extensive experimental results validate the efficacy of our method over the existing state of the arts for FER.
Ping Liu 0004, Yuewei Lin, Zibo Meng, Weihong Deng, Joey Tianyi Zhou, Yi Yang 0001
IEEE Trans. Cybern.6
2022 ECML: An Ensemble Cascade Metric-Learning Mechanism Toward Face Verification
abstract
Face verification can be regarded as a two-class fine-grained visual-recognition problem. Enhancing the feature's discriminative power is one of the key problems to improve its performance. Metric-learning technology is often applied to address this need while achieving a good tradeoff between underfitting, and overfitting plays a vital role in metric learning. Hence, we propose a novel ensemble cascade metric-learning (ECML) mechanism. In particular, hierarchical metric learning is executed in a cascade way to alleviate underfitting. Meanwhile, at each learning level, the features are split into nonoverlapping groups. Then, metric learning is executed among the feature groups in the ensemble manner to resist overfitting. Considering the feature distribution characteristics of faces, a robust Mahalanobis metric-learning method (RMML) with a closed-form solution is additionally proposed. It can avoid the computation failure issue on an inverse matrix faced by some well-known metric-learning approaches (e.g., KISSME). Embedding RMML into the proposed ECML mechanism, our metric-learning paradigm (EC-RMML) can run in the one-pass learning manner. The experimental results demonstrate that EC-RMML is superior to state-of-the-art metric-learning methods for face verification. The proposed ECML mechanism is also applicable to other metric-learning approaches.
Fu Xiong, Yang Xiao 0007, Zhiguo Cao 0001, Yancheng Wang 0002, Joey Tianyi Zhou, Jianxin Wu 0001
IEEE Trans. Cybern.5
2022 Person Re-Identification With Hierarchical Discriminative Spatial Aggregation
abstract
Practically, person re-identification (re-ID) may suffer from the critical spatial misalignment problem due to inaccurate human detection, variation on human pose and camera viewpoint, etc. To address this, a hierarchical discriminative spatial aggregation method is proposed. The key idea is to conduct spatial aggregation on local human parts via global average-pooling to acquire the strong spatial misalignment tolerance, with VALD encoding on the local parts for facilitating discriminative power jointly. This proposition is built on NetVLAD to ensure end-to-end deep learning capacity. Due to the fine-grained property of person re-ID task that has not been well concerned by the original NetVLAD model for scene recognition, a feature refinement layer that consists of 1 fully-connected (FC) layer and 2 batch normalization (BN) layers is added on top of the raw NetVLAD layer to enhance the discriminative power and training convergence. And, a human body occlusion and background component dropout manner is also proposed to resist the effect of serious occlusion. Technically, a refined codeword initialization manner is proposed to alleviate the potential codeword imbalance problem caused by naive random initialization. The proposed discriminative spatial aggregation approach is then conducted on multi-resolution convolutional feature map layers hierarchically via early feature fusion, to involve richer semantic and fine-grained visual clues jointly. Wide-range experiments on 6 datasets (i.e., CUHK03, DukeMTMC-reID, Occluded-DukeMTMC, Market-1501, MSMT17 and Occluded-REID) verifies the effectiveness of our proposition. The source code and supporting material is available athttps://github.com/zmyme/HDSA-reID.
Yang Xiao 0007, Fu Xiong, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou
IEEE Trans. Inf. Forensics Secur.7
2022 Memory-Assistant Collaborative Language Understanding for Artificial Intelligence of Things
abstract
Artificial intelligence shows promising efforts in collaborating the language models with the artificial intelligence of things (AIoT), promoting the edging intelligence on natural language understanding. To adapt to the limited computational resources in AIoT, the large language models (e.g., transformer) are compressed into light-weight models, which always results in poor feature representation and unsatisfactory performance on downstream tasks, especially on those low-resource language understanding tasks. To address the above issues, we propose a method named memory-assistant multi-task learning (MAMT), where an auxiliary memory module is introduced to promote multitask learning (MT), which serves as a surrogate of target domain representation and performs instance-level weighted MT. More importantly, our MAMT module is in a plug-and-play fashion. Thus, researchers can plug in it to conduct collaborative training and plug it out for AIoT model inference without extra computation burdens. Experiments demonstrate that MAMT significantly improves the performance of light-weight transformer models and show its superiority over the state-of-the-arts on eight GLUE subtasks.
Ming Yan 0007, Cen Chen 0002, Jiawei Du 0002, Xi Peng 0001, Joey Tianyi Zhou, Zeng Zeng
IEEE Trans. Ind. Informatics5
2022 Deep Supervised Domain Adaptation for Pneumonia Diagnosis From Chest X-Ray Images
abstract
Pneumonia is one of the most common treatable causes of death, and early diagnosis allows for early intervention. Automated diagnosis of pneumonia can therefore improve outcomes. However, it is challenging to develop high-performance deep learning models due to the lack of well-annotated data for training. This paper proposes a novel method, called Deep Supervised Domain Adaptation (DSDA), to automatically diagnose pneumonia from chest X-ray images. Specifically, we propose to transfer the knowledge from a publicly available large-scale source dataset (ChestX-ray14) to a well-annotated but small-scale target dataset (the TTSH dataset). DSDA aligns the distributions of the source domain and the target domain according to the underlying semantics of the training samples. It includes two task-specific sub-networks for the source domain and the target domain, respectively. These two sub-networks share the feature extraction layers and are trained in an end-to-end manner. Unlike most existing domain adaptation approaches that perform the same tasks in the source domain and the target domain, we attempt to transfer the knowledge from a multi-label classification task in the source domain to a binary classification task in the target domain. To evaluate the effectiveness of our method, we compare it with several existing peer methods. The experimental results show that our method can achieve promising performance for automated pneumonia diagnosis.
Yangqin Feng, Xinxing Xu, Yan Wang 0015, Xiaofeng Lei, Soo Kng Teo, Jordan Zheng Ting Sim, Yonghan Ting, Liangli Zhen, Joey Tianyi Zhou, Yong Liu 0026, Cher Heng Tan
IEEE J. Biomed. Health Informatics9
2022 Anomaly Detection With Bidirectional Consistency in Videos
abstract
The core component of most anomaly detectors is a self-supervised model, tasked with modeling patterns included in training samples and detecting unexpected patterns as the anomalies in testing samples. To cope with normal patterns, this model is typically trained with reconstruction constraints. However, the model has the risk of overfitting to training samples and being sensitive to hard normal patterns in the inference phase, which results in irregular responses at normal frames. To address this problem, we formulate anomaly detection as a mutual supervision problem. Due to collaborative training, the complementary information of mutual learning can alleviate the aforementioned problem. Based on this motivation, a SIamese generative network (SIGnet), including two subnetworks with the same architecture, is proposed to simultaneously model the patterns of the forward and backward frames. During training, in addition to traditional constraints on improving the reconstruction performance, a bidirectional consistency loss based on the forward and backward views is designed as the regularization term to improve the generalization ability of the model. Moreover, we introduce a consistency-based evaluation criterion to achieve stable scores at the normal frames, which will benefit detecting anomalies with fluctuant scores in the inference phase. The results on several challenging benchmark data sets demonstrate the effectiveness of our proposed method.
Zhiwen Fang, Jiafei Liang, Joey Tianyi Zhou, Yang Xiao 0007, Feng Yang 0012
IEEE Trans. Neural Networks Learn. Syst.3
2022 Discriminative Multi-View Dynamic Image Fusion for Cross-View 3-D Action Recognition
abstract
Dramatic imaging viewpoint variation is the critical challenge toward action recognition for depth video. To address this, one feasible way is to enhance view-tolerance of visual feature, while still maintaining strong discriminative capacity. Multi-view dynamic image (MVDI) is the most recently proposed 3-D action representation manner that is able to compactly encode human motion information and 3-D visual clue well. However, it is still view-sensitive. To leverage its performance, a discriminative MVDI fusion method is proposed by us via multi-instance learning (MIL). Specifically, the dynamic images (DIs) from different observation viewpoints are regarded as the instances for 3-D action characterization. After being encoded using Fisher vector (FV), they are then aggregated by sum-pooling to yield the representative 3-D action signature. Our insight is that viewpoint aggregation helps to enhance view-tolerance. And, FV can map the raw DI feature to the higher dimensional feature space to promote the discriminative power. Meanwhile, a discriminative viewpoint instance discovery method is also proposed to discard the viewpoint instances unfavorable for action characterization. The wide-range experiments on five data sets demonstrate that our proposition can significantly enhance the performance of cross-view 3-D action recognition. And, it is also applicable to cross-view 3-D object recognition. The source code is available at https://github.com/3huo/ActionView.
Yancheng Wang 0002, Yang Xiao 0007, Zhiguo Cao 0001, Zhenjun Zhang, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.7
2022 Deep Multimodal Transfer Learning for Cross-Modal Retrieval
abstract
Cross-modal retrieval (CMR) enables flexible retrieval experience across different modalities (e.g., texts versus images), which maximally benefits us from the abundance of multimedia data. Existing deep CMR approaches commonly require a large amount of labeled data for training to achieve high performance. However, it is time-consuming and expensive to annotate the multimedia data manually. Thus, how to transfer valuable knowledge from existing annotated data to new data, especially from the known categories to new categories, becomes attractive for real-world applications. To achieve this end, we propose a deep multimodal transfer learning (DMTL) approach to transfer the knowledge from the previously labeled categories (source domain) to improve the retrieval performance on the unlabeled new categories (target domain). Specifically, we employ a joint learning paradigm to transfer knowledge by assigning a pseudolabel to each target sample. During training, the pseudolabel is iteratively updated and passed through our model in a self-supervised manner. At the same time, to reduce the domain discrepancy of different modalities, we construct multiple modality-specific neural networks to learn a shared semantic space for different modalities by enforcing the compactness of homoinstance samples and the scatters of heteroinstance samples. Our method is remarkably different from most of the existing transfer learning approaches. To be specific, previous works usually assume that the source domain and the target domain have the same label set. In contrast, our method considers a more challenging multimodal learning situation where the label sets of the two domains are different or even disjoint. Experimental studies on four widely used benchmarks validate the effectiveness of the proposed method in multimodal transfer learning and demonstrate its superior performance in CMR compared with 11 state-of-the-art methods.
Liangli Zhen, Peng Hu 0002, Xi Peng 0001, Rick Siow Mong Goh, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.5
2022 Utility Optimal Thread Assignment and Resource Allocation in Multi-Server Systems
abstract
Achieving high performance in many multi-server systems (e.g., web hosting center, cloud) requires finding a good assignment of worker threads to servers and also effectively allocating each server’s resources to its assigned threads. The assignment and allocation components of this problem have been studied extensively but largely separately in the literature. In this paper, we introduce theassign and allocate (AA)problem, which seeks to simultaneously find an assignment and allocation that maximizes the total utility of the threads. Assigning and allocating the threads together can result in substantially better overall utility than performing the steps separately, as is traditionally done. We model each thread by a utility function giving its performance as a function of its assigned resources. We first prove that the AA problem is NP-hard. We then present a$2 (\sqrt {2}-1) > 0.828$factor approximation algorithm for concave utility functions, which runs in$O(mn^{2} + n (\log mC)^{2})$time for$n$threads and$m$servers with$C$amount of resources each. We also give a faster algorithm with the same approximation ratio and$O(n (\log mC)^{2})$time complexity. We then extend the problem to two more general settings. First, we consider threads with nonconcave utility functions, and give a 1/2 factor approximation algorithm. Next, we give an algorithm for threads using multiple types of resources, and show the algorithm achieves good empirical performance. We conduct extensive experiments to test the performance of our algorithms on threads with both synthetic and realistic utility functions, and find that they achieve over 92% of the optimal utility on average. We also compare our algorithms with a number of practical heuristics, and find that our algorithms achieve up to 9 times higher total utility.
Pan Lai, Rui Fan 0004, Xiao Zhang 0006, Wei Zhang 0082, Fang Liu 0009, Joey Tianyi Zhou
IEEE/ACM Trans. Netw.6
2022 Deep Learning for Latent Events Forecasting in Content Caching Networks
abstract
A novel Twitter context aided content caching (TAC) framework is proposed for enhancing the caching efficiency by taking advantage of the legibility and massive volume of Twitter data. For the purpose of promoting the caching efficiency, three machine learning models are proposed to predict latent events and events popularity, utilizing collected Twitter data with geo-tags and geographic information of the adjacent base stations (BSs). Firstly, we propose a latent Dirichlet allocation (LDA) model for latent events forecasting because of the superiority of LDA model in natural language processing (NLP). Then, we conceive long short-term memory (LSTM) with skip-gram embedding approach and LSTM with continuous skip-gram-Geo-aware embedding approach for the events popularity forecasting. Furthermore, we associate the predict latent events and the popularity of the events with the caching strategy. Lastly, we propose a non-orthogonal multiple access (NOMA) based content transmission scheme. Extensive practical experiments demonstrate that: 1) the proposed TAC framework outperforms conventional caching framework and is capable of being employed in practical applications thanks to the associating ability with public interests; 2) the proposed LDA approach conserves superiority for natural language processing (NLP) in Twitter data; 3) the perplexity of the proposed skip-gram based LSTM is lower compared with conventional LDA approach; and 4) evaluation of the model demonstrates that the hit rates of tweets of the model vary from 50% to 65% and the hit rate of the caching contents is up to approximately 75% with smaller caching space compared to conventional algorithms. Simulation results also shows that the proposed NOMA-enabled caching scheme outperforms conventional least frequently used (LFU) scheme by 25%.
Zhong Yang 0001, Yuanwei Liu, Yue Chen 0002, Joey Tianyi Zhou
IEEE Trans. Wirel. Commun.4
2021 Contrastive Clustering
abstract
In this paper, we propose an online clustering method called Contrastive Clustering (CC) which explicitly performs the instance- and cluster-level contrastive learning. To be specific, for a given dataset, the positive and negative instance pairs are constructed through data augmentations and then projected into a feature space. Therein, the instance- and cluster-level contrastive learning are respectively conducted in the row and column space by maximizing the similarities of positive pairs while minimizing those of negative ones. Our key observation is that the rows of the feature matrix could be regarded as soft labels of instances, and accordingly the columns could be further regarded as cluster representations. By simultaneously optimizing the instance- and cluster-level contrastive loss, the model jointly learns representations and cluster assignments in an end-to-end manner. Besides, the proposed method could timely compute the cluster assignment for each individual, even when the data is presented in streams. Extensive experimental results show that CC remarkably outperforms 17 competitive clustering methods on six challenging image benchmarks. In particular, CC achieves an NMI of 0.705 (0.431) on the CIFAR-10 (CIFAR-100) dataset, which is an up to 19% (39%) performance improvement compared with the best baseline. The code is available at https://github.com/XLearning-SCU/2021-AAAI-CC.
Yunfan Li 0003, Peng Hu 0002, Zitao Liu 0001, Dezhong Peng, Joey Tianyi Zhou, Xi Peng 0001
AAAI5
2021 PointBA: Towards Backdoor Attacks in 3D Point Cloud
abstract
3D deep learning has been increasingly more popular for a variety of tasks including many safety-critical applications. However, recently several works raise the security issues of 3D deep models. Although most of them consider adversarial attacks, we identify that backdoor attack is indeed a more serious threat to 3D deep learning systems but remains unexplored. We present the backdoor attacks in 3D point cloud with a unified framework that exploits the unique properties of 3D data and networks. In particular, we design two attack approaches on point cloud: the poison-label backdoor attack (PointPBA) and the clean- label backdoor attack (PointCBA). The first one is straight-forward and effective in practice, while the latter is more sophisticated assuming there are certain data inspections. The attack algorithms are mainly motivated and developed by 1) the recent discovery of 3D adversarial samples suggesting the vulnerability of deep models under spatial transformation; 2) the proposed feature disentanglement technique that manipulates the feature of the data through optimization methods and its potential to embed a new task. Extensive experiments show the efficacy of the PointPBA with over 95% success rate across various 3D datasets and models, and the more stealthy PointCBA with around 50% success rate. Our proposed backdoor attack in 3D point cloud is expected to perform as a baseline for improving the robustness of 3D deep models.
Zekun Tong, Yabang Zhao, Andrew Lim 0001, Joey Tianyi Zhou
ICCV7
2021 Trusted Multi-View Classification
Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou
ICLR4
2021 Trustworthy Multimodal Regression with Mixture of Normal-inverse Gamma Distributions
abstract
Multimodal regression is a fundamental task, which integrates the information from different sources to improve the performance of follow-up applications. However, existing methods mainly focus on improving the performance and often ignore the confidence of prediction for diverse situations. In this study, we are devoted to trustworthy multimodal regression which is critical in cost-sensitive domains. To this end, we introduce a novel Mixture of Normal-Inverse Gamma distributions (MoNIG) algorithm, which efficiently estimates uncertainty in principle for adaptive integration of different modalities and produces a trustworthy regression result. Our model can be dynamically aware of uncertainty for each modality, and also robust for corrupted modalities. Furthermore, the proposed MoNIG ensures explicitly representation of (modality-specific/global) epistemic and aleatoric uncertainties, respectively. Experimental results on both synthetic and different real-world data demonstrate the effectiveness and trustworthiness of our method on various multimodal regression tasks (e.g., temperature prediction for superconductivity, relative location prediction for CT slices, and multimodal sentiment analysis).
Huan Ma 0006, Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu
NeurIPS5
2021 Video Corpus Moment Retrieval with Contrastive Learning
abstract
Given a collection of untrimmed and unsegmented videos, video corpus moment retrieval (VCMR) is to retrieve a temporal moment (i.e., a fraction of a video) that semantically corresponds to a given text query. As video and text are from two distinct feature spaces, there are two general approaches to address VCMR: (i) to separately encode each modality representations, then align the two modality representations for query processing, and (ii) to adopt fine-grained cross-modal interaction to learn multi-modal representations for query processing. While the second approach often leads to better retrieval accuracy, the first approach is far more efficient. In this paper, we propose a Retrieval and Localization Network with Contrastive Learning (ReLoCLNet) for VCMR. We adopt the first approach and introduce two contrastive learning objectives to refine video encoder and text encoder to learn video and text representations separately but with better alignment for VCMR. The video contrastive learning (VideoCL) is to maximize mutual information between query and candidate video at video-level. The frame contrastive learning (FrameCL) aims to highlight the moment region corresponds to the query at frame-level, within a video. Experimental results show that, although ReLoCLNet encodes text and video separately for efficiency, its retrieval accuracy is comparable with baselines adopting cross-modal interaction learning.
Hao Zhang 0048, Aixin Sun, Guoshun Nan, Liangli Zhen, Joey Tianyi Zhou, Rick Siow Mong Goh
SIGIR6
2021 You Only Look Yourself: Unsupervised and Untrained Single Image Dehazing Neural Network
Boyun Li, Yuanbiao Gou, Shuhang Gu, Zitao Liu 0001, Joey Tianyi Zhou, Xi Peng 0001
Int. J. Comput. Vis.5
2021 LPQ++: A discriminative blur-insensitive textural descriptor with spatial-channel interaction
Yang Xiao 0007, Zhiguo Cao 0001, Zhiwen Fang, Joey Tianyi Zhou
Inf. Sci.6
2021 Ordered or Orderless: A Revisit for Video Based Person Re-Identification
abstract
Is recurrent network really necessary for learning a good visual representation for video based person re-identification (VPRe-id)? In this paper, we first show that the common practice of employing recurrent neural networks (RNNs) to aggregate temporal-spatial features may not be optimal. Specifically, with a diagnostic analysis, we show that the recurrent structure may not be effective learn temporal dependencies than what we expected and implicitly yields an orderless representation. Based on this observation, we then present a simple yet surprisingly powerful approach for VPRe-id, where we treat VPRe-id as an efficient orderless ensemble of image based person re-identification problem. More specifically, we divide videos into individual images and re-identify person with ensemble of image based rankers. Under the i.i.d. assumption, we provide an error bound that sheds light upon how could we improve VPRe-id. Our work also presents a promising way to bridge the gap between video and image based person re-identification. Comprehensive experimental evaluations demonstrate that the proposed solution achieves state-of-the-art performances on multiple widely used datasets (iLIDS-VID, PRID 2011, and MARS).
Le Zhang 0001, Zenglin Shi, Joey Tianyi Zhou, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Zeng Zeng, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Nonlinear Regression via Deep Negative Correlation Learning
abstract
Nonlinear regression has been extensively employed in many computer vision problems (e.g., crowd counting, age estimation, affective computing). Under the umbrella of deep learning, two common solutions exist i) transforming nonlinear regression to a robust loss function which is jointly optimizable with the deep convolutional network, and ii) utilizing ensemble of deep networks. Although some improved performance is achieved, the former may be lacking due to the intrinsic limitation of choosing a single hypothesis and the latter may suffer from much larger computational complexity. To cope with those issues, we propose to regress via an efficient "divide and conquer" manner. The core of our approach is the generalization of negative correlation learning that has been shown, both theoretically and empirically, to work well for non-deep regression problems. Without extra parameters, the proposed method controls the bias-variance-covariance trade-off systematically and usually yields a deep regression ensemble where each base model is both "accurate" and "diversified." Moreover, we show that each sub-problem in the proposed method has less Rademacher Complexity and thus is easier to optimize. Extensive experiments on several diverse and challenging tasks including crowd counting, personality analysis, age estimation, and image super-resolution demonstrate the superiority over challenging baselines as well as the versatility of the proposed method. The source code and trained models are available on our project page: https://mmcheng.net/dncl/.
Le Zhang 0001, Zenglin Shi, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Joey Tianyi Zhou, Guoyan Zheng, Zeng Zeng
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Correction to "Nonlinear Regression via Deep Negative Correlation Learning"
abstract
Reports on changes to the author information presented in the above named paper.
Le Zhang 0001, Zenglin Shi, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Joey Tianyi Zhou, Guoyan Zheng, Zeng Zeng
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Dual Adversarial Transfer for Sequence Labeling
abstract
We propose a new architecture for addressing sequence labeling, termed Dual Adversarial Transfer Network (DATNet). Specifically, the proposed DATNet includes two variants, i.e., DATNet-F and DATNet-P, which are proposed to explore effective feature fusion between high and low resource. To address the noisy and imbalanced training data, we propose a novel Generalized Resource-Adversarial Discriminator (GRAD) and adopt adversarial training to boost model generalization. We investigate the effects of different components of DATNet across different domains and languages, and show that significant improvement can be obtained especially for low-resource data. Without augmenting any additional hand-crafted features, we achieve state-of-the-art performances on CoNLL, Twitter, PTB-WSJ, OntoNotes and Universal Dependencies with three popular sequence labeling tasks, i.e., Named entity recognition (NER), Part-of-Speech (POS) Tagging and Chunking.
Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Xi Peng 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Single-Image Dehazing via Compositional Adversarial Network
abstract
Single-image dehazing has been an important topic given the commonly occurred image degradation caused by adverse atmosphere aerosols. The key to haze removal relies on an accurate estimation of global air-light and the transmission map. Most existing methods estimate these two parameters using separate pipelines which reduces the efficiency and accumulates errors, thus leading to a suboptimal approximation, hurting the model interpretability, and degrading the performance. To address these issues, this article introduces a novel generative adversarial network (GAN) for single-image dehazing. The network consists of a novel compositional generator and a novel deeply supervised discriminator. The compositional generator is a densely connected network, which combines fine-scale and coarse-scale information. Benefiting from the new generator, our method can directly learn the physical parameters from data and recover clean images from hazy ones in an end-to-end manner. The proposed discriminator is deeply supervised, which enforces that the output of the generator to look similar to the clean images from low-level details to high-level structures. To the best of our knowledge, this is the first end-to-end generative adversarial model for image dehazing, which simultaneously outputs clean images, transmission maps, and air-lights. Extensive experiments show that our method remarkably outperforms the state-of-the-art methods. Furthermore, to facilitate future research, we create the HazeCOCO dataset which is currently the largest dataset for single-image dehazing.
Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Zhao Kang 0001, Shijian Lu, Zhiwen Fang, Liyuan Li, Joo-Hwee Lim
IEEE Trans. Cybern.4
2021 Deep Spectral Representation Learning From Multi-View Data
abstract
Multi-view representation learning (MvRL) aims to learn a consensus representation from diverse sources or domains to facilitate downstream tasks such as clustering, retrieval, and classification. Due to the limited representative capacity of the adopted shallow models, most existing MvRL methods may yield unsatisfactory results, especially when the labels of data are unavailable. To enjoy the representative capacity of deep learning, this paper proposes a novel multi-view unsupervised representation learning method, termed as Multi-view Laplacian Network (MvLNet), which could be the first deep version of the multi-view spectral representation learning method. Note that, such an attempt is nontrivial because simply combining Laplacian embedding (i.e., spectral representation) with neural networks will lead to trivial solutions. To solve this problem, MvLNet enforces an orthogonal constraint and reformulates it as a layer with the help of Cholesky decomposition. The orthogonal layer is stacked on the embedding network so that a common space could be learned for consensus representation. Compared with numerous recent-proposed approaches, extensive experiments on seven challenging datasets demonstrate the effectiveness of our method in three multi-view tasks including clustering, recognition, and retrieval. The source code could be found at www.pengxi.me.
Zhenyu Huang 0005, Joey Tianyi Zhou, Hongyuan Zhu 0002, Changqing Zhang 0002, Jiancheng Lv 0001, Xi Peng 0001
IEEE Trans. Image Process.2
2021 Triply Complementary Priors for Image Restoration
abstract
Recent works that utilized deep models have achieved superior results in various image restoration (IR) applications. Such approach is typically supervised, which requires a corpus of training images with distributions similar to the images to be recovered. On the other hand, the shallow methods, which are usually unsupervised remain promising performance in many inverse problems, e.g., image deblurring and image compressive sensing (CS), as they can effectively leverage nonlocal self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various artifacts due to naive patch aggregation in addition to the slow speed. Using either approach alone usually limits performance and generalizability in IR tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely, internal and external, shallow and deep, and non-local and local priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for IR. Following this, a simple yet effective algorithm is developed to solve the proposed H-PnP based IR problems. Extensive experimental results on several representative IR tasks, including image deblurring, image CS and image deblocking, demonstrate that the proposed H-PnP algorithm achieves favorable performance compared to many popular or state-of-the-art IR methods in terms of both objective and visual perception.
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Ce Zhu
IEEE Trans. Image Process.4
2021 Multi-Encoder Towards Effective Anomaly Detection in Videos
abstract
Given normal training samples, anomaly detection in videos can be regarded as a challenging problem of identifying unexpected events. The state-of-the-art approaches generally resort to the autoencoder model by using a single encoder to capture the motion and content patterns jointly. Nevertheless, due to the lack of accurate labels of normal and abnormal samples, how to detect anomalies is decided by the subjective understanding of models. It infers that different models will prefer to mine different patterns according to the characteristics of models. We call this problem as a pattern bias problem. To alleviate this problem, a novel Multi-Encoder Single-Decoder network, termed as MESDnet, is proposed in the spirit of encoding motion and content cues individually with multiple encoders. MESDnet is of end-to-end learning ability and real-time running speed. Particularly, the differences between adjacent frames and the raw frames are used as the motion and content sources, respectively. Then, a decoder takes charge of detecting anomalies in the way of observing reconstructing error towards the video frames by using the multi-stream encoded motion and content features simultaneously. The experiments on the CUHK Avenue dataset, the UCSD Pedestrian dataset, and the ShanghaiTech Campus dataset verify the effectiveness of MESDnet.
Zhiwen Fang, Joey Tianyi Zhou, Yang Xiao 0007, Yanan Li 0006, Feng Yang 0012
IEEE Trans. Multim.2
2021 Introduction to Big Multimodal Multimedia Data with Deep Analytics
abstract
introduction Share on Introduction to Big Multimodal Multimedia Data with Deep Analytics Authors: Yang Wang Hefei University of Technology Hefei University of TechnologyView Profile , Meng Fang Tecent AI Tecent AIView Profile , Joey Tianyi Zhou A-Star A-StarView Profile , Tingting Mu The University of Manchester The University of ManchesterView Profile , Dacheng Tao The UBTECH Sydney Artificial Intelligence Centre, the University of Sydney, Australia The UBTECH Sydney Artificial Intelligence Centre, the University of Sydney, AustraliaView Profile Authors Info & Claims ACM Transactions on Multimedia Computing, Communications, and ApplicationsVolume 17Issue 1sJanuary 2021 Article No.: 1pp 1–3https://doi.org/10.1145/3447530Online:31 March 2021Publication History 0citation212DownloadsMetricsTotal Citations0Total Downloads212Last 12 Months93Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Yang Wang 0023, Joey Tianyi Zhou, Tingting Mu, Dacheng Tao
ACM Trans. Multim. Comput. Commun. Appl.3
2021 Concurrent Processing Cluster Design to Empower Simultaneous Prediction for Hundreds of Vessels' Trajectories in Near Real-Time
abstract
The automatic identification system (AIS) plays a vital role in maritime traffic surveillance. AIS is designed for remotely tracking vessels but nowadays also becomes a useful data source to enable vessels' trajectory prediction so as to facilitate early alert of potential collision risks. Recent studies focus on improving prediction accuracy through the machine learning and knowledge-based technologies but the computational cost also greatly increases due to the model complexity of these methodologies. It becomes a practical challenge to realize near real-time (NRT) trajectory prediction of a large volume of ships. For risk alert application, forecasting timeliness is one of the key design considerations. In this paper, we propose a concurrent processing cluster solution to empower advanced trajectory forecasting for hundreds of vessels in NRT. The proposed solution relies on the properly determined frameworks by customizing the desired features for an integrated solution. Meanwhile, a novel task-based load balancing strategy with newly defined metrics are proposed, which aims to reduce the makespan of jobs and outperforms the existing load balancing algorithms. A practicable cluster system has been successfully implemented, serving as a step toward unlocking the power of advanced maritime traffic forecasting technologies and enabling the benefit from the latest progress on the methodological innovation.
Xiuju Fu, Wanbing Zhang, Joey Tianyi Zhou, Rick Siow Mong Goh
IEEE Trans. Syst. Man Cybern. Syst.5
2020 Unsupervised Domain Adaptation on Reading Comprehension
abstract
Reading comprehension (RC) has been studied in a variety of datasets with the boosted performance brought by deep neural networks. However, the generalization capability of these models across different domains remains unclear. To alleviate the problem, we investigate unsupervised domain adaptation on RC, wherein a model is trained on the labeled source domain and to be applied to the target domain with only unlabeled samples. We first show that even with the powerful BERT contextual representation, a model can not generalize well from one domain to another. To solve this, we provide a novel conditional adversarial self-training method (CASe). Specifically, our approach leverages a BERT model fine-tuned on the source dataset along with the confidence filtering to generate reliable pseudo-labeled samples in the target domain for self-training. On the other hand, it further reduces domain distribution discrepancy through conditional adversarial learning across domains. Extensive experiments show our approach achieves comparable performance to supervised models on multiple large-scale benchmark datasets.
Yu Cao 0014, Baosheng Yu, Joey Tianyi Zhou
AAAI4
2020 Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
abstract
Machine learning algorithms are often vulnerable to adversarial examples that have imperceptible alterations from the original counterparts but can fool the state-of-the-art models. It is helpful to evaluate or even improve the robustness of these models by exposing the maliciously crafted adversarial examples. In this paper, we present TextFooler, a simple but strong baseline to generate adversarial text. By applying it to two fundamental natural language tasks, text classification and textual entailment, we successfully attacked three target models, including the powerful pre-trained BERT, and the widely used convolutional and recurrent neural networks. We demonstrate three advantages of this framework: (1) effective—it outperforms previous attacks by success rate and perturbation rate, (2) utility-preserving—it preserves semantic content, grammaticality, and correct types classified by humans, and (3) efficient—it generates adversarial text with computational complexity linear to the text length.1
Di Jin 0005, Zhijing Jin 0001, Joey Tianyi Zhou, Peter Szolovits
AAAI3
2020 Hooks in the Headline: Learning to Generate Headlines with Controlled Styles
abstract
Current summarization systems only produce plain, factual headlines, but do not meet the practical needs of creating memorable titles to increase exposure.We propose a new task, Stylistic Headline Generation (SHG), to enrich the headlines with three style options (humor, romance and clickbait), in order to attract more readers.With no style-specific article-headline pair (only a standard headline summarization dataset and mono-style corpora), our method TitleStylist generates style-specific headlines by combining the summarization and reconstruction tasks into a multitasking framework.We also introduced a novel parameter sharing scheme to further disentangle the style from the text.Through both automatic and human evaluation, we demonstrate that TitleStylist can generate relevant, fluent headlines with three target styles: humor, romance, and clickbait.The attraction score of our model generated headlines surpasses that of the state-ofthe-art summarization model by 9.68%, and even outperforms human-written references. 1
Di Jin 0005, Zhijing Jin 0001, Joey Tianyi Zhou, Lisa Orii, Peter Szolovits
ACL3
2020 Multi-source Meta Transfer for Low Resource Multiple-Choice Question Answering
abstract
Multiple-choice question answering (MCQA) is one of the most challenging tasks in machine reading comprehension since it requires more advanced reading comprehension skills such as logical reasoning, summarization, and arithmetic operations.Unfortunately, most existing MCQA datasets are small in size, which increases the difficulty of model learning and generalization.To address this challenge, we propose a multi-source meta transfer (MMT) for low-resource MCQA.In this framework, we first extend meta learning by incorporating multiple training sources to learn a generalized feature representation across domains.To bridge the distribution gap between training sources and the target, we further introduce the meta transfer that can be integrated into the multi-source meta training.More importantly, the proposed MMT is independent of backbone language models.Extensive experiments demonstrate the superiority of MMT over state-of-the-arts, and continuous improvements can be achieved on different backbone networks on both supervised and unsupervised domain adaptation settings.
Ming Yan 0007, Hao Zhang 0048, Di Jin 0005, Joey Tianyi Zhou
ACL4
2020 Span-based Localizing Network for Natural Language Video Localization
abstract
Given an untrimmed video and a text query, natural language video localization (NLVL) is to locate a matching span from the video that semantically corresponds to the query.Existing solutions formulate NLVL either as a ranking task and apply multimodal matching architecture, or as a regression task to directly regress the target video span.In this work, we address NLVL task with a span-based QA approach by treating the input video as text passage.We propose a video span localizing network (VSLNet), on top of the standard span-based QA framework, to address NLVL.The proposed VSLNet tackles the differences between NLVL and span-based QA through a simple and yet effective query-guided highlighting (QGH) strategy.The QGH guides VSLNet to search for matching video span within a highlighted region.Through extensive experiments on three benchmark datasets, we show that the proposed VSLNet outperforms the state-of-the-art methods; and adopting span-based QA framework is a promising direction to solve NLVL. 1
Hao Zhang 0048, Aixin Sun, Joey Tianyi Zhou
ACL4
2020 3DV: 3D Dynamic Voxel for Action Recognition in Depth Video
abstract
For depth-based 3D action recognition, one essential issue is to represent 3D motion pattern effectively and efficiently. To this end, 3D dynamic voxel (3DV) is proposed as a novel 3D motion representation manner. With 3D space voxelization, the key idea of 3DV is to encode the 3D motion information within depth video into a regular voxel set (i.e., 3DV) compactly, via temporal rank pooling. Each available 3DV voxel intrinsically involves 3D spatial and motion feature for 3D action description. 3DV is then abstracted as a point set and input into PointNet++ for 3D action recognition, in the end-to-end learning way. The intuition for transferring 3DV into the point set form is that, PointNet++ is lightweight and effective for deep feature learning towards point set. Since 3DV may loose appearance clue, a multi-stream 3D action recognition manner is also proposed to learn motion and appearance feature jointly. To extract richer temporal order information of actions, we also split the depth video into temporal segments and encode this procedure in 3DV integrally. The extensive experiments on the well-established benchmark datasets (e.g., NTU RGB+D 120 and NTU RGB+D 60) demonstrate the superiority of our proposition. Impressively, we acquire the accuracy of 82.4% and 93.5% on NTU RGB+D 120 with the cross-subject and cross-setup test setting respectively. 3DV's code is available at https://github.com/3huo/3DV-Action.
Yancheng Wang 0002, Yang Xiao 0007, Fu Xiong, Wenxiang Jiang 0001, Zhiguo Cao 0001, Joey Tianyi Zhou, Junsong Yuan 0001
CVPR6
2020 Deep Reinforcement Learning for RIS-Aided Non-Orthogonal Multiple Access Downlink Networks
abstract
A novel reconfigurable intelligent surface (RIS) aided non-orthogonal multiple access (NOMA) downlink transmission framework is proposed. We formulate a long-term stochastic optimization problem that involves the optimization of phase shifting, aiming at maximizing the sum data rate of the mobile users (MUs) in NOMA downlink networks. For intelligently adjusting the phase shifting matrix of the access point (AP), we propose a deep deterministic policy gradient (DDPG) algorithm to collaboratively control multiple reflecting elements (REs) of the RIS. Extensive simulation results demonstrate that: 1) The proposed RIS-aided NOMA downlink framework achieves better sum data rate compared with orthogonal multiple access (OMA) networks. 2) The proposed DDPG algorithm is capable of learning a dynamic resource allocation policy, while conventional optimization approaches can not. 3) Compared with increasing the transmit power of the AP, increasing the number of reflecting elements (REs) is a more efficiency method to improve the sum data rate.
Zhong Yang 0001, Yuanwei Liu, Yue Chen 0002, Joey Tianyi Zhou
GLOBECOM4
2020 Adaptive Precision Training for Resource Constrained Devices
abstract
Learn in-situ is a growing trend for Edge AI. Training deep neural network (DNN) on edge devices is challenging because both energy and memory are constrained. Low precision training helps to reduce the energy cost of a single training iteration, but that does not necessarily translate to energy savings for the whole training process, because low precision could slows down the convergence rate. One evidence is that most works for low precision training keep an fp32 copy of the model during training, which in turn imposes memory requirements on edge devices. In this work we propose Adaptive Precision Training. It is able to save both total training energy cost and memory usage at the same time. We use model of the same precision for both forward and backward pass in order to reduce memory usage for training. Through evaluating the progress of training, APT allocates layer-wise precision dynamically so that the model learns quicker for longer time. APT provides an application specific hyper-parameter for users to play trade-off between training energy cost, memory usage and accuracy. Experiment shows that APT achieves more than 50% saving on training energy and memory usage with limited accuracy loss. 20% more savings of training energy and memory usage can be achieved in return for a 1% sacrifice in accuracy loss.
Tian Huang, Tao Luo 0014, Joey Tianyi Zhou
ICDCS3
2020 The Power Of Triply Complementary Priors For Image Compressive Sensing
abstract
Recent works that utilized deep models have achieved superior results in various image restoration applications. Such approach is typically supervised which requires a corpus of training images with distribution similar to the images to be recovered. On the other hand, the shallow methods which are usually unsupervised remain promising performance in many inverse problems, e.g., image compressive sensing (CS), as they can effectively leverage non-local self-similarity priors of natural images. However, most of such methods are patch-based leading to the restored images with various ringing artifacts due to naive patch aggregation. Using either approach alone usually limits performance and generalizability in image restoration tasks. In this paper, we propose a joint low-rank and deep (LRD) image model, which contains a pair of triply complementary priors, namely external and internal, deep and shallow, and local and nonlocal priors. We then propose a novel hybrid plug-and-play (H-PnP) framework based on the LRD model for image CS. To make the optimization tractable, a simple yet effective algorithm is proposed to solve the proposed H-PnP based image CS problem. Extensive experimental results demonstrate that the proposed H-PnP algorithm significantly outperforms the state-of-the-art techniques for image CS recovery such as SCSNet and WNNM.
Zhiyuan Zha, Xin Yuan 0002, Joey Tianyi Zhou, Jiantao Zhou 0001, Bihan Wen, Ce Zhu
ICIP3
2020 Query-efficient Meta Attack to Deep Neural Networks
Jiawei Du 0002, Hu Zhang 0005, Joey Tianyi Zhou, Yi Yang 0001, Jiashi Feng
ICLR3
2020 Speaker and Phoneme-Aware Speech Bandwidth Extension with Residual Dual-Path Network
abstract
Speech bandwidth extension aims to generate a wideband signal from a narrowband (low-band) input by predicting the missing high-frequency components. It is believed that the general knowledge about the speaker and phonetic content strengthens the prediction. In this paper, we propose to augment the low-band acoustic features with i-vector and phonetic posteriorgram (PPG), which represent speaker and phonetic content of the speech, respectively. We also propose a residual dual-path network (RDPN) as the core module to process the augmented features, which fully utilizes the utterance-level temporal continuity information and avoids gradient vanishing. Experiments show that the proposed method achieves 20.2% and 7.0% relative improvements over the best baseline in terms of log-spectral distortion (LSD) and signal-to-noise ratio (SNR), respectively. Furthermore, our method is 16 times more compact than the best baseline in terms of the number of parameters.
Nana Hou, Chenglin Xu, Van Tung Pham, Joey Tianyi Zhou, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH4
2020 Multi-Task Learning for End-to-End Noise-Robust Bandwidth Extension
abstract
Bandwidth extension aims to reconstruct wideband speech signals from narrowband inputs to improve perceptual quality. Prior studies mostly perform bandwidth extension under the assumption that the narrowband signals are clean without noise. The use of such extension techniques is greatly limited in practice when signals are corrupted by noise. To alleviate such problem, we propose an end-to-end time-domain framework for noise-robust bandwidth extension, that jointly optimizes a mask-based speech enhancement and an ideal bandwidth extension module with multi-task learning. The proposed framework avoids decomposing the signals into magnitude and phase spectra, therefore, requires no phase estimation. Experimental results show that the proposed method achieves 14.3% and 15.8% relative improvements over the best baseline in terms of perceptual evaluation of speech quality (PESQ) and log-spectral distortion (LSD), respectively. Furthermore, our method is 3 times more compact than the best baseline in terms of the number of parameters.
Nana Hou, Chenglin Xu, Joey Tianyi Zhou, Chng Eng Siong, Haizhou Li 0001
INTERSPEECH3
2020 Partially View-aligned Clustering
abstract
In this paper, we study one challenging issue in multi-view data clustering. To be specific, for two data matrices $\mathbf{X}^{(1)}$ and $\mathbf{X}^{(2)}$ corresponding to two views, we do not assume that $\mathbf{X}^{(1)}$ and $\mathbf{X}^{(2)}$ are fully aligned in row-wise. Instead, we assume that only a small portion of the matrices has established the correspondence in advance. Such a partially view-aligned problem (PVP) could lead to the intensive labor of capturing or establishing the aligned multi-view data, which has less been touched so far to the best of our knowledge. To solve this practical and challenging problem, we propose a novel multi-view clustering method termed partially view-aligned clustering (PVC). To be specific, PVC proposes to use a differentiable surrogate of the non-differentiable Hungarian algorithm and recasts it as a pluggable module. As a result, the category-level correspondence of the unaligned data could be established in a latent space learned by a neural network, while learning a common space across different views using the ``aligned'' data. Extensive experimental results show promising results of our method in clustering partially view-aligned data.
Zhenyu Huang 0005, Peng Hu 0002, Joey Tianyi Zhou, Jiancheng Lv 0001, Xi Peng 0001
NeurIPS3
2020 Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based Games
abstract
We study reinforcement learning (RL) for text-based games, which are interactive simulations in the context of natural language. While different methods have been developed to represent the environment information and language actions, existing RL agents are not empowered with any reasoning capabilities to deal with textual games. In this work, we aim to conduct explicit reasoning with knowledge graphs for decision making, so that the actions of an agent are generated and supported by an interpretable inference procedure. We propose a stacked hierarchical attention mechanism to construct an explicit representation of the reasoning process by exploiting the structure of the knowledge graph. We extensively evaluate our method on a number of man-made benchmark games, and the experimental results demonstrate that our method performs better than existing text-based agents.
Yunqiu Xu, Ling Chen 0006, Yali Du 0001, Joey Tianyi Zhou, Chengqi Zhang
NeurIPS5
2020 Multi-graph fusion for multi-view spectral clustering
Zhao Kang 0001, Guoxin Shi, Shudong Huang, Wenyu Chen 0001, Xiaorong Pu, Joey Tianyi Zhou, Zenglin Xu
Knowl. Based Syst.6
2020 Partition level multiview subspace clustering
Zhao Kang 0001, Xinjia Zhao, Chong Peng 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Xi Peng 0001, Wenyu Chen 0001, Zenglin Xu
Neural Networks5
2020 Attention-Driven Loss for Anomaly Detection in Video Surveillance
abstract
Recent video anomaly detection methods focus on reconstructing or predicting frames. Under this umbrella, the long-standing inter-class data-imbalance problem resorts to the imbalance between foreground and stationary background objects in video anomaly detection and this has been less investigated by existing solutions. Naively optimizing the reconstructing loss yields a biased optimization towards background reconstruction rather than the objects of interest in the foreground. To solve this, we proposed a simple yet effective solution, termed attention-driven loss to alleviate the foreground-background imbalance problem in anomaly detection. Specifically, we compute a single mask map that summarizes the frame evolution of moving foreground regions and suppresses the background in the training video clips. After that, we construct an attention map through the combination of the mask map and background to give different weights to the foreground and background region respectively. The proposed attention-driven loss is independent of backbone networks and can be easily augmented in most existing anomaly detection models. Augmented with attention-driven loss, the model is able to achieve AUC 86.0% on Avenue, 83.9% on Ped1, 96% on Ped2 datasets. Extensive experimental results and ablation studies further validate the effectiveness of our model.
Joey Tianyi Zhou, Le Zhang 0001, Zhiwen Fang, Jiawei Du 0002, Xi Peng 0001, Yang Xiao 0007
IEEE Trans. Circuits Syst. Video Technol.1
2020 Towards Real-Time Eyeblink Detection in the Wild: Dataset, Theory and Practices
abstract
Effective and real-time eyeblink detection is of wide-range applications, such as deception detection, drive fatigue detection, face anti-spoofing. Despite previous efforts, most of existing focus on addressing the eyeblink detection problem under constrained indoor conditions with relative consistent subject and environment setup. Nevertheless, towards practical applications, eyeblink detection in the wild is highly preferred, and of greater challenges. In this paper, we shed the light to this research topic. A labelled eyeblink in the wild dataset (i.e., HUST-LEBW) of 673 eyeblink video samples (i.e., 381 positives, and 292 negatives) is first established. These samples are captured from the unconstrained movies, with the dramatic variation on face attribute, head pose, illumination condition, imaging configuration, etc. Then, we formulate eyeblink detection task as a binary spatial-temporal pattern recognition problem. After locating and tracking human eyes using SeetaFace engine and KCF (Kernelized Correlation Filters) tracker respectively, a modified LSTM model able to capture the multi-scale temporal information is proposed to verify eyeblink. A feature extraction approach that reveals the appearance and motion characteristics simultaneously is also proposed. The experiments on HUST-LEBW reveal the superiority and efficiency of our approach. The comparisons with the existing state-of-the-art methods validate the advantages of our manner for eyeblink detection in the wild.
Guilei Hu, Yang Xiao 0007, Zhiguo Cao 0001, Lubin Meng, Zhiwen Fang, Joey Tianyi Zhou, Junsong Yuan 0001
IEEE Trans. Inf. Forensics Secur.6
2020 Zero-Shot Image Dehazing
abstract
In this paper, we study two less-touched challenging problems in single image dehazing neural networks, namely, how to remove haze from a given image in an unsupervised and zeroshot manner. To the ends, we propose a novel method based on the idea of layer disentanglement by viewing a hazy image as the entanglement of several "simpler" layers, i.e., a hazy-free image layer, transmission map layer, and atmospheric light layer. The major advantages of the proposed ZID are two-fold. First, it is an unsupervised method that does not use any clean images including hazy-clean pairs as the ground-truth. Second, ZID is a "zero-shot" method, which just uses the observed single hazy image to perform learning and inference. In other words, it does not follow the conventional paradigm of training deep model on a large scale dataset. These two advantages enable our method to avoid the labor-intensive data collection and the domain shift issue of using the synthetic hazy images to address the real-world images. Extensive comparisons show the promising performance of our method compared with 15 approaches in the qualitative and quantitive evaluations. The source code could be found at www.pengxi.me.
Boyun Li, Yuanbiao Gou, Zitao Liu 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Xi Peng 0001
IEEE Trans. Image Process.5
2020 Deep Subspace Clustering
abstract
In this article, we propose a deep extension of sparse subspace clustering, termed deep subspace clustering with L1-norm (DSC-L1). Regularized by the unit sphere distribution assumption for the learned deep features, DSC-L1 can infer a new data affinity matrix by simultaneously satisfying the sparsity principle of SSC and the nonlinearity given by neural networks. One of the appealing advantages brought by DSC-L1 is that when original real-world data do not meet the class-specific linear subspace distribution assumption, DSC-L1 can employ neural networks to make the assumption valid with its nonlinear transformations. Moreover, we prove that our neural network could sufficiently approximate the minimizer under mild conditions. To the best of our knowledge, this could be one of the first deep-learning-based subspace clustering methods. Extensive experiments are conducted on four real-world data sets to show that the proposed method is significantly superior to 17 existing methods for subspace clustering on handcrafted features and raw data.
Xi Peng 0001, Jiashi Feng, Joey Tianyi Zhou, Yingjie Lei, Shuicheng Yan
IEEE Trans. Neural Networks Learn. Syst.3
2020 Deep Clustering With Sample-Assignment Invariance Prior
abstract
Most popular clustering methods map raw image data into a projection space in which the clustering assignment is obtained with the vanilla k-means approach. In this article, we discovered a novel prior, namely, there exists a common invariance when assigning an image sample to clusters using different metrics. In short, different distance metrics will lead to similar soft clustering assignments on the manifold. Based on such a novel prior, we propose a novel clustering method by minimizing the discrepancy between pairwise sample assignments for each data point. To the best of our knowledge, this could be the first work to reveal the sample-assignment invariance prior based on the idea of treating labels as ideal representations. Furthermore, the proposed method is one of the first end-to-end clustering approaches, which jointly learns clustering assignment and representation. Extensive experimental results show that the proposed method is remarkably superior to 16 state-of-the-art clustering methods on five image data sets in terms of four evaluation metrics.
Xi Peng 0001, Hongyuan Zhu 0002, Jiashi Feng, Chunhua Shen, Haixian Zhang, Joey Tianyi Zhou
IEEE Trans. Neural Networks Learn. Syst.6
2020 RoSeq: Robust Sequence Labeling
abstract
In this paper, we mainly investigate two issues for sequence labeling, namely, label imbalance and noisy data that are commonly seen in the scenario of named entity recognition (NER) and are largely ignored in the existing works. To address these two issues, a new method termed robust sequence labeling (RoSeq) is proposed. Specifically, to handle the label imbalance issue, we first incorporate label statistics in a novel conditional random field (CRF) loss. In addition, we design an additional loss to reduce the weights of overwhelming easy tokens for augmenting the CRF loss. To address the noisy training data, we adopt an adversarial training strategy to improve model generalization. In experiments, the proposed RoSeq achieves the state-of-the-art performances on CoNLL and English Twitter NER-88.07% on CoNLL-2002 Dutch, 87.33% on CoNLL-2002 Spanish, 52.94% on WNUT-2016 Twitter, and 43.03% on WNUT-2017 Twitter without using the additional data.
Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Xi Peng 0001, Yang Xiao 0007, Zhiguo Cao 0001
IEEE Trans. Neural Networks Learn. Syst.1
2019 Inter-Class Angular Loss for Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have shown great power in various classification tasks and have achieved remarkable results in practical applications. However, the distinct learning difficulties in discriminating different pairs of classes are largely ignored by the existing networks. For instance, in CIFAR-10 dataset, distinguishing cats from dogs is usually harder than distinguishing horses from ships. By carefully studying the behavior of CNN models in the training process, we observe that the confusion level of two classes is strongly correlated with their angular separability in the feature space. That is, the larger the inter-class angle is, the lower the confusion will be. Based on this observation, we propose a novel loss function dubbed “Inter-Class Angular Loss” (ICAL), which explicitly models the class correlation and can be directly applied to many existing deep networks. By minimizing the proposed ICAL, the networks can effectively discriminate the examples in similar classes by enlarging the angle between their corresponding class vectors. Thorough experimental results on a series of vision and nonvision datasets confirm that ICAL critically improves the discriminative ability of various representative deep neural networks and generates superior performance to the original networks with conventional softmax loss.
Le Hui, Xiang Li 0041, Chen Gong 0002, Joey Tianyi Zhou, Jian Yang 0003
AAAI5
2019 Singe Image Rain Removal with Unpaired Information: A Differentiable Programming Perspective
abstract
Single image rain-streak removal is an extremely challenging problem due to the presence of non-uniform rain densities in images. Previous works solve this problem using various hand-designed priors or by explicitly mapping synthetic rain to paired clean image in a supervised way. In practice, however, the pre-defined priors are easily violated and the paired training data are hard to collect. To overcome these limitations, in this work, we propose RainRemoval-GAN (RRGAN), the first end-to-end adversarial model that generates realistic rain-free images using only unpaired supervision. Our approach alleviates the paired training constraints by introducing a physical-model which explicitly learns a recovered images and corresponding rain-streaks from the differentiable programming perspective. The proposed network consists of a novel multiscale attention memory generator and a novel multiscale deeply supervised discriminator. The multiscale attention memory generator uses a memory with attention mechanism to capture the latent rain streaks context at different stages to recover the clean images. The deeply supervised multiscale discriminator imposes constraints at the recovered output in terms of local details and global appearance to the clean image set. Together with the learned rainstreaks, a reconstruction constraint is employed to ensure the appearance consistent with the input image. Experimental results on public benchmark demonstrates our promising performance compared with nine state-of-the-art methods in terms of PSNR, SSIM, visual qualities and running time.
Hongyuan Zhu 0002, Xi Peng 0001, Joey Tianyi Zhou, Songfan Yang, Vijay Chanderasekh, Liyuan Li, Joo-Hwee Lim
AAAI3
2019 Dual Adversarial Neural Transfer for Low-Resource Named Entity Recognition
abstract
We propose a new neural transfer method termed Dual Adversarial Transfer Network (DATNet) for addressing low-resource Named Entity Recognition (NER).Specifically, two variants of DATNet, i.e., DATNet-F and DATNet-P, are investigated to explore effective feature fusion between high and low resource.To address the noisy and imbalanced training data, we propose a novel Generalized Resource-Adversarial Discriminator (GRAD).Additionally, adversarial training is adopted to boost model generalization.In experiments, we examine the effects of different components in DATNet across domains and languages, and show that significant improvement can be obtained especially for lowresource data, without augmenting any additional hand-crafted features and pre-trained language model.
Joey Tianyi Zhou, Hao Zhang 0048, Di Jin 0005, Hongyuan Zhu 0002, Rick Siow Mong Goh, Kenneth Kwok
ACL (1)1
2019 Reciprocal Multi-Layer Subspace Learning for Multi-View Clustering
abstract
Multi-view clustering is a long-standing important research topic, however, remains challenging when handling high-dimensional data and simultaneously exploring the consistency and complementarity of different views. In this work, we present a novel Reciprocal Multi-layer Subspace Learning (RMSL) algorithm for multi-view clustering, which is composed of two main components: Hierarchical Self-Representative Layers (HSRL), and Backward Encoding Networks (BEN). Specifically, HSRL constructs reciprocal multi-layer subspace representations linked with a latent representation to hierarchically recover the underlying low-dimensional subspaces in which the high-dimensional data lie; BEN explores complex relationships among different views and implicitly enforces the subspaces of all views to be consistent with each other and more separable. The latent representation flexibly encodes complementary information from multiple views and depicts data more comprehensively. Our model can be efficiently optimized by an alternating optimization scheme. Extensive experiments on benchmark datasets show the superiority of RMSL over other state-of-the-art clustering methods.
Ruihuang Li, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou, Qinghua Hu
ICCV5
2019 A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation From a Single Depth Image
abstract
For 3D hand and body pose estimation task in depth image, a novel anchor-based approach termed Anchor-to-Joint regression network (A2J) with the end-to-end learning ability is proposed. Within A2J, anchor points able to capture global-local spatial context information are densely set on depth image as local regressors for the joints. They contribute to predict the positions of the joints in ensemble way to enhance generalization ability. The proposed 3D articulated pose estimation paradigm is different from the state-of-the-art encoder-decoder based FCN, 3D CNN and point-set based manners. To discover informative anchor points towards certain joint, anchor proposal procedure is also proposed for A2J. Meanwhile 2D CNN (i.e., ResNet- 50) is used as backbone network to drive A2J, without using time-consuming 3D convolutional or deconvolutional layers. The experiments on 3 hand datasets and 2 body datasets verify A2J's superiority. Meanwhile, A2J is of high running speed around 100 FPS on single NVIDIA 1080Ti GPU.
Fu Xiong, Boshen Zhang, Yang Xiao 0007, Zhiguo Cao 0001, Taidong Yu, Joey Tianyi Zhou, Junsong Yuan 0001
ICCV6
2019 COMIC: Multi-view Clustering Without Parameter Selection
abstract
In this paper, we study two challenges in clustering analysis, namely, how to cluster multi-view data and how to perform clustering without parameter selection on cluster size. To this end, we propose a novel objective function to project raw data into one space in which the projection embraces the geometric consistency (GC) and the cluster assignment consistency (CAC). To be specific, the GC aims to learn a connection graph from a projection space wherein the data points are connected if and only if they belong to the same cluster. The CAC aims to minimize the discrepancy of pairwise connection graphs induced from different views based on the view-consensus assumption, i.e., different views could produce the same cluster assignment structure as they are different portraits of the same object. Thanks to the view-consensus derived from the connection graph, our method could achieve promising performance in learning view-specific representation and eliminating the heterogeneous gaps across different views. Furthermore, with the proposed objective, it could learn almost all parameters including the cluster number from data without labor-intensive parameter selection. Extensive experimental results show the promising performance achieved by our method on five datasets comparing with nine state-of-the-art multi-view clustering approaches.
Xi Peng 0001, Zhenyu Huang 0005, Jiancheng Lv 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou
ICML5
2019 Multi-view Spectral Clustering Network
abstract
Multi-view clustering aims to cluster data from diverse sources or domains, which has drawn considerable attention in recent years. In this paper, we propose a novel multi-view clustering method named multi-view spectral clustering network (MvSCN) which could be the first deep version of multi-view spectral clustering to the best of our knowledge. To deeply cluster multi-view data, MvSCN incorporates the local invariance within every single view and the consistency across different views into a novel objective function, where the local invariance is defined by a deep metric learning network rather than the Euclidean distance adopted by traditional approaches. In addition, we enforce and reformulate an orthogonal constraint as a novel layer stacked on an embedding network for two advantages, i.e. jointly optimizing the neural network and performing matrix decomposition and avoiding trivial solutions. Extensive experiments on four challenging datasets demonstrate the effectiveness of our method compared with 10 state-of-the-art approaches in terms of three evaluation metrics.
Zhenyu Huang 0005, Joey Tianyi Zhou, Xi Peng 0001, Changqing Zhang 0002, Hongyuan Zhu 0002, Jiancheng Lv 0001
IJCAI2
2019 CPM-Nets: Cross Partial Multi-View Networks
abstract
Despite multi-view learning progressed fast in past decades, it is still challenging due to the difficulty in modeling complex correlation among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Networks (CPM-Nets). In this framework, we first give a formal definition of completeness and versatility for multi-view representation and then theoretically prove the versatility of the latent representation learned from our algorithm. To achieve the completeness, the task of learning latent multi-view representation is specifically translated to degradation process through mimicking data transmitting, such that the optimal tradeoff between consistence and complementarity across different views could be achieved. In contrast with methods that either complete missing views or group samples according to view-missing patterns, our model fully exploits all samples and all views to produce structured representation for interpretability. Extensive experimental results validate the effectiveness of our algorithm over existing state-of-the-arts.
Changqing Zhang 0002, Zongbo Han, Yajie Cui, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu
NeurIPS5
2019 A deep learning framework for Hybrid Heterogeneous Transfer Learning
Joey Tianyi Zhou, Sinno Jialin Pan, Ivor W. Tsang
Artif. Intell.1
2019 Action recognition for depth video using multi-view dynamic images
Yang Xiao 0007, Jun Chen 0001, Yancheng Wang 0002, Zhiguo Cao 0001, Joey Tianyi Zhou, Xiang Bai
Inf. Sci.5
2019 Multi-class Heterogeneous Domain Adaptation
abstract
A crucial issue in heterogeneous domain adaptation (HDA) is the ability to learn a feature mapping between different types of features across domains. Inspired by language translation, a word translated from one language corresponds to only a few words in another language, we present an efficient method named Sparse Heterogeneous Feature Representation (SHFR) in this paper for multi-class HDA to learn a sparse feature transformation between domains with multiple classes. Specifically, we formulate the problem of learning the feature transformation as a compressed sensing problem by building multiple binary classifiers in the target domain as various measurement sensors, which are decomposed from the target multi-class classification problem. We show that the estimation error of the learned transformation decreases with the increasing number of binary classifiers. In other words, for adaptation across heterogeneous domains to be successful, it is necessary to construct a sufficient number of incoherent binary classifiers from the original multi-class classification problem. To achieve this, we propose to apply the error correcting output correcting (ECOC) scheme to generate incoherent classifiers. To speed up the learning of the feature transformation across domains, we apply an efficient batch-mode algorithm to solve the resultant nonnegative sparse recovery problem. Theoretically, we present a generalization error bound of our proposed HDA method under a multi-class setting. Lastly, we conduct extensive experiments on both synthetic and real-world datasets to demonstrate the superiority of our proposed method over existing state-of-the-art HDA methods in terms of prediction accuracy and training efficiency.
Joey Tianyi Zhou, Ivor W. Tsang, Sinno Jialin Pan, Mingkui Tan
J. Mach. Learn. Res.1
2019 N-ary decomposition for multi-class classification
Joey Tianyi Zhou, Ivor W. Tsang, Shen-Shyang Ho, Klaus-Robert Müller
Mach. Learn.1
2019 AnomalyNet: An Anomaly Detection Network for Video Surveillance
abstract
Sparse coding-based anomaly detection has shown promising performance, of which the keys are feature learning, sparse representation, and dictionary learning. In this paper, we propose a new neural network for anomaly detection (termed AnomalyNet) by deeply achieving feature learning, sparse representation, and dictionary learning in three joint neural processing blocks. Specifically, to learn better features, we design a motion fusion block accompanied by a feature transfer block to enjoy the advantages of eliminating noisy background, capturing motion, and alleviating data deficiency. Furthermore, to address some disadvantages (e.g., nonadaptive updating) of the existing sparse coding optimizers and embrace the merits of neural network (e.g., parallel computing), we design a novel recurrent neural network to learn sparse representation and dictionary by proposing an adaptive iterative hard-thresholding algorithm (adaptive ISTA) and reformulating the adaptive ISTA as a new long short-term memory (LSTM). To the best of our knowledge, this could be one of the first works to bridge the$\ell _{1}$ -solver and LSTM and may provide novel insight into understanding LSTM and model-based optimization (or named differentiable programming), as well as sparse coding-based anomaly detection. Extensive experiments show the state-of-the-art performance of our method in the abnormal events detection task.
Joey Tianyi Zhou, Jiawei Du 0002, Hongyuan Zhu 0002, Xi Peng 0001, Yong Liu 0026, Rick Siow Mong Goh
IEEE Trans. Inf. Forensics Secur.1
2019 Beyond Majority Voting: A Coarse-to-Fine Label Filtration for Heavily Noisy Labels
abstract
Crowdsourcing has become the most appealing way to provide a plethora of labels at a low cost. Nevertheless, labels from amateur workers are often noisy, which inevitably degenerates the robustness of subsequent learning models. To improve the label quality for subsequent use, majority voting (MV) is widely leveraged to aggregate crowdsourced labels due to its simplicity and scalability. However, when crowdsourced labels are "heavily" noisy (e.g., 40% of noisy labels), MV may not work well because of the fact "garbage (heavily noisy labels) in, garbage (full aggregated labels) out." This issue inspires us to think: if the ultimate target is to learn a robust model using noisy labels, why not provide partial aggregated labels and ensure that these labels are reliable enough for learning models? To solve this challenge by improving MV, we propose a coarse-to-fine label filtration model called double filter machine (DFM), which consists of a (majority) voting filter and a sparse filter serially. Specifically, the DFM refines crowdsourced labels from coarse filtering to fine filtering. In the stage of coarse filtering, the DFM aggregates crowdsourced labels by voting filter, which yields (quality-acceptable) full aggregated labels. In the stage of fine filtering, DFM further digs out a set of high-quality labels from full aggregated labels by sparse filter, since this filter can identify high-quality labels by the methodology of support selection. Based on the insight of compressed sensing, DFM recovers a ground-truth signal from heavily noisy data under a restricted isometry property. To sum up, the primary benefits of DFM are to keep the scalability by voting filter, while improve the robustness by sparse filter. We also derive theoretical guarantees for the convergence and recovery of DFM and reveal its complexity. We conduct comprehensive experiments on both the UCI simulated and the AMT crowdsourced datasets. Empirical results show that partial aggregated labels provided by DFM effectively improve the robustness of learning models.
Bo Han 0003, Ivor W. Tsang, Ling Chen 0006, Joey Tianyi Zhou, Celina Ping Yu
IEEE Trans. Neural Networks Learn. Syst.4
2019 Learning With Annotation of Various Degrees
abstract
In this paper, we study a new problem in the scenario of sequences labeling. To be exact, we consider that the training data are with annotation of various degrees, namely, fully labeled, unlabeled, and partially labeled sequences. The learning with fully un/labeled sequence refers to the standard setting in traditional un/supervised learning, and the proposed partially labeling specifies the subject that the element does not belong to. The partially labeled data are cheaper to obtain compared with the fully labeled data though it is less informative, especially when the tasks require a lot of domain knowledge. To solve such a practical challenge, we propose a novel deep conditional random field (CRF) model which utilizes an end-to-end learning manner to smoothly handle fully/un/partially labeled sequences within a unified framework. To the best of our knowledge, this could be one of the first works to utilize the partially labeled instance for sequence labeling, and the proposed algorithm unifies the deep learning and CRF in an end-to-end framework. Extensive experiments show that our method achieves state-of-the-art performance in two sequence labeling tasks on some popular data sets.
Joey Tianyi Zhou, Hao Zhang 0048, Chen Gong 0002, Xi Peng 0001, Zhiguo Cao 0001, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.1
2018 SC2Net: Sparse LSTMs for Sparse Coding
abstract
The iterative hard-thresholding algorithm (ISTA) is one of the most popular optimization solvers to achieve sparse codes. However, ISTA suffers from following problems: 1) ISTA employs non-adaptive updating strategy to learn the parameters on each dimension with a fixed learning rate. Such a strategy may lead to inferior performance due to the scarcity of diversity; 2) ISTA does not incorporate the historical information into the updating rules, and the historical information has been proven helpful to speed up the convergence. To address these challenging issues, we propose a novel formulation of ISTA (named as adaptive ISTA) by introducing a novel \textit{adaptive momentum vector}. To efficiently solve the proposed adaptive ISTA, we recast it as a recurrent neural network unit and show its connection with the well-known long short term memory (LSTM) model. With a new proposed unit, we present a neural network (termed SC2Net) to achieve sparse codes in an end-to-end manner. To the best of our knowledge, this is one of the first works to bridge the $\ell_1$-solver and LSTM, and may provide novel insights in understanding model-based optimization and LSTM. Extensive experiments show the effectiveness of our method on both unsupervised and supervised tasks.
Joey Tianyi Zhou, Kai Di, Jiawei Du 0002, Xi Peng 0001, Hao Yang 0033, Sinno Jialin Pan, Ivor W. Tsang, Yong Liu 0026, Zheng Qin 0004, Rick Siow Mong Goh
AAAI1
2018 A Learning-based Approach for Error Compensation of Industrial Manipulator with Hybrid Model
abstract
The industrial robot usually has high repeatability but relatively lower accuracy. Therefore, error compensation plays a pivotal role in many industrial robotic applications with high accuracy requirement. In this paper, we present a novel computational method that utilizes a hybrid model that consists of Local Product-Of-Exponential (POE) and Gaussian Process Regression (GPR) to compensate the positioning errors of the industrial robotic manipulator for high accuracy industrial robotic applications. Specifically in the proposed method, the Local POE calibration method is first applied to calibrate the robot forward kinematic model to reduce the geometric error. Then the GPR is applied to learn the inverse kinematic model to further compensate the residual error in task space. We also demonstrate the robustness and effectiveness of our proposed method by showing the reduction of norm pose error by up to 37.2%, compared to the existing methods with multiple datasets.
Joey Tianyi Zhou, Yong Liu 0026, Pey Yuen Tao, Guilin Yang
ICARCV2
2018 'Who Likes What and, Why?' Insights into Modeling Users' Personality Based on Image 'Likes'
abstract
The increased proliferation of data production technologies (e.g., cameras) and consumption avenues (e.g., social media) has led to images and videos being utilized by users to convey innate preferences and tastes. This has opened up the possibility of using multimedia as a source for user-modeling. This work attempts to model personality traits (based on the Five Factor Theory) of users using a collection of images they tag as `favorite' (or like) on Flickr. First, a set of semantic features are proposed to be used for representing different concepts in images which influence users to like them. The addition of the proposed features led to improvement over state-of-the-art by 12 percent. Second, a novel machine learning approach is developed to model users' personality based on the image features (resulting in upto 15 percent improvement). Third, efficacy of the semantic features and the modeling approach is shown in recommending images based on personality modeling. Using the modeling approach, recommendations are made regarding the factors that might influence users with different personality traits to like an image.
Sharath Chandra Guntuku, Joey Tianyi Zhou, Sujoy Roy, Weisi Lin, Ivor W. Tsang
IEEE Trans. Affect. Comput.2
2018 Structured AutoEncoders for Subspace Clustering
abstract
Existing subspace clustering methods typically employ shallow models to estimate underlying subspaces of unlabeled data points and cluster them into corresponding groups. However, due to the limited representative capacity of the employed shallow models, those methods may fail in handling realistic data without the linear subspace structure. To address this issue, we propose a novel subspace clustering approach by introducing a new deep model-Structured AutoEncoder (StructAE). The StructAE learns a set of explicit transformations to progressively map input data points into nonlinear latent spaces while preserving the local and global subspace structure. In particular, to preserve local structure, the StructAE learns representations for each data point by minimizing reconstruction error w.r.t. itself. To preserve global structure, the StructAE incorporates a prior structured information by encouraging the learned representation to preserve specified reconstruction patterns over the entire data set. To the best of our knowledge, StructAE is one of first deep subspace clustering approaches. Extensive experiments show that the proposed StructAE significantly outperforms 15 state-of-the-art subspace clustering approaches in terms of five evaluation metrics.
Xi Peng 0001, Jiashi Feng, Shijie Xiao, Weiyun Yau, Joey Tianyi Zhou, Songfan Yang
IEEE Trans. Image Process.5
2018 Transfer Hashing: From Shallow to Deep
abstract
One major assumption used in most existing hashing approaches is that the domain of interest (i.e., the target domain) could provide sufficient training data, either labeled or unlabeled. However, this assumption may be violated in practice. To address this so-called data sparsity issue in hashing, a new framework termed transfer hashing with privileged information (THPI) is proposed, which marriages hashing and transfer learning (TL). To show the efficacy of THPI, we propose three variants of the well-known iterative quantization (ITQ) as a showcase. The proposed methods, ITQ+, LapITQ+, and deep transfer hashing (DTH), solve the aforementioned data sparsity issue from different aspects. Specifically, ITQ+ is a shallow model, which makes ITQ achieve hashing in a TL manner. ITQ+ learns a new slack function from the source domain to approximate the quantization error on the target domain given by ITQ. To further improve the performance of ITQ+, LapITQ+ is proposed by embedding the geometric relationship of the source domain into the target domain. Moreover, DTH is proposed to show the generality of our framework by utilizing the powerful representative capacity of deep learning. To the best of our knowledge, this could be one of the first DTH works. Extensive experiments on several popular data sets demonstrate the effectiveness of our shallow and DTH approaches comparing with several state-of-the-art hashing approaches.
Joey Tianyi Zhou, Heng Zhao 0004, Xi Peng 0001, Zheng Qin 0004, Rick Siow Mong Goh
IEEE Trans. Neural Networks Learn. Syst.1
2017 MIML-FCN+: Multi-Instance Multi-Label Learning via Fully Convolutional Networks with Privileged Information
abstract
Multi-instance multi-label (MIML) learning has many interesting applications in computer visions, including multi-object recognition and automatic image tagging. In these applications, additional information such as bounding-boxes, image captions and descriptions is often available during training phrase, which is referred as privileged information (PI). However, as existing works on learning using PI only consider instance-level PI (privileged instances), they fail to make use of bag-level PI (privileged bags) available in MIML learning. Therefore, in this paper, we propose a two-stream fully convolutional network, named MIML-FCN+, unified by a novel PI loss to solve the problem of MIML learning with privileged bags. Compared to the previous works on PI, the proposed MIML-FCN+ utilizes the readily available privileged bags, instead of hard-to-obtain privileged instances, making the system more general and practical in real world applications. As the proposed PI loss is convex and SGD-compatible and the framework itself is a fully convolutional network, MIML FCN+ can be easily integrated with state-of-the-art deep learning networks. Moreover, the flexibility of convolutional layers allows us to exploit structured correlations among instances to facilitate more effective training and testing. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed MIML-FCN+, outperforming state-of-the-art methods in the application of multi-object recognition.
Hao Yang 0033, Joey Tianyi Zhou, Jianfei Cai 0001, Yew-Soon Ong
CVPR2
2016 Transfer Learning for Cross-Language Text Categorization through Active Correspondences Construction
abstract
Most existing heterogeneous transfer learning (HTL) methods for cross-language text classification rely on sufficient cross-domain instance correspondences to learn a mapping across heterogeneous feature spaces, and assume that such correspondences are given in advance. However, in practice, correspondences between domains are usually unknown. In this case, extensively manual efforts are required to establish accurate correspondences across multilingual documents based on their content and meta-information. In this paper, we present a general framework to integrate active learning to construct correspondences between heterogeneous domains for HTL, namely HTL through active correspondences construction (HTLA). Based on this framework, we develop a new HTL method. On top of the new HTL method, we further propose a strategy to actively construct correspondences between domains. Extensive experiments are conducted on various multilingual text classification tasks to verify the effectiveness of HTLA.
Joey Tianyi Zhou, Sinno Jialin Pan, Ivor W. Tsang, Shen-Shyang Ho
AAAI1
2016 Exploit Bounding Box Annotations for Multi-Label Object Recognition
abstract
Convolutional neural networks (CNNs) have shown great performance as general feature representations for object recognition applications. However, for multi-label images that contain multiple objects from different categories, scales and locations, global CNN features are not optimal. In this paper, we incorporate local information to enhance the feature discriminative power. In particular, we first extract object proposals from each image. With each image treated as a bag and object proposals extracted from it treated as instances, we transform the multi-label recognition problem into a multi-class multi-instance learning problem. Then, in addition to extracting the typical CNN feature representation from each proposal, we propose to make use of ground-truth bounding box annotations (strong labels) to add another level of local information by using nearest-neighbor relationships of local regions to form a multi-view pipeline. The proposed multi-view multiinstance framework utilizes both weak and strong labels effectively, and more importantly it has the generalization ability to even boost the performance of unseen categories by partial strong labels from other categories. Our framework is extensively compared with state-of-the-art handcrafted feature based methods and CNN based methods on two multi-label benchmark datasets. The experimental results validate the discriminative power and the generalization ability of the proposed framework. With strong labels, our framework is able to achieve state-of-the-art results in both datasets.
Hao Yang 0033, Joey Tianyi Zhou, Yu Zhang 0004, Bin-Bin Gao, Jianxin Wu 0001, Jianfei Cai 0001
CVPR2
2016 Improving Multi-label Learning with Missing Labels by Structured Semantic Correlations
Hao Yang 0033, Joey Tianyi Zhou, Jianfei Cai 0001
ECCV (1)2
2016 Transfer Hashing with Privileged Information
Joey Tianyi Zhou, Xinxing Xu, Sinno Jialin Pan, Ivor W. Tsang, Zheng Qin 0004, Rick Siow Mong Goh
IJCAI1
2016 Understanding Deep Representations Learned in Modeling Users Likes
abstract
Automatically understanding and discriminating different users' liking for an image is a challenging problem. This is because the relationship between image features (even semantic ones extracted by existing tools, viz., faces, objects, and so on) and users' likes is non-linear, influenced by several subtle factors. This paper presents a deep bi-modal knowledge representation of images based on their visual content and associated tags (text). A mapping step between the different levels of visual and textual representations allows for the transfer of semantic knowledge between the two modalities. Feature selection is applied before learning deep representation to identify the important features for a user to like an image. The proposed representation is shown to be effective in discriminating users based on images they like and also in recommending images that a given user likes, outperforming the state-of-the-art feature representations by ∼ 15 %-20%. Beyond this test-set performance, an attempt is made to qualitatively understand the representations learned by the deep architecture used to model user likes.
Sharath Chandra Guntuku, Joey Tianyi Zhou, Sujoy Roy, Weisi Lin, Ivor W. Tsang
IEEE Trans. Image Process.2
2014 Hybrid Heterogeneous Transfer Learning through Deep Learning
abstract
Most previous heterogeneous transfer learning methods learn a cross-domain feature mapping between heterogeneous feature spaces based on a few cross-domain instance-correspondences, and these corresponding instances are assumed to be representative in the source and target domains respectively. However, in many real-world scenarios, this assumption may not hold. As a result, the constructed feature mapping may not be precisely due to the bias issue of the correspondences in the target or (and) source domain(s). In this case, a classifier trained on the labeled transformed-source-domain data may not be useful for the target domain. In this paper, we present a new transfer learning framework called Hybrid Heterogeneous Transfer Learning (HHTL), which allows the corresponding instances across domains to be biased in either the source or target domain. Specifically, we propose a deep learning approach to learn a feature mapping between cross-domain heterogeneous features as well as a better feature representation for mapped data to reduce the bias issue caused by the cross-domain correspondences. Extensive experiments on several multilingual sentiment classification tasks verify the effectiveness of our proposed approach compared with some baseline methods.
Joey Tianyi Zhou, Sinno Jialin Pan, Ivor W. Tsang, Yan Yan 0006
AAAI1
2014 Deep Representations to Model User 'Likes'
Sharath Chandra Guntuku, Joey Tianyi Zhou, Sujoy Roy, Weisi Lin, Ivor W. Tsang
ACCV (1)2
2014 Heterogeneous Domain Adaptation for Multiple Classes
abstract
In this paper, we present an efficient Multi-class Heterogeneous Domain Adaptation (HDA) method, where data from the source and target domains are represented by heterogeneous features with different dimensions. Specifically, we propose to reconstruct a sparse feature transformation matrix to map the features of multiple classes from the source domain to the target domain. We cast this learning task as a compressed sensing problem, where each classifier can be deemed as a measurement sensor. Based on compressive sensing theory, the estimation error of the transformation matrix decreases with the increasing number of classifiers. Therefore, to guarantee the reconstruction performance, we construct sufficiently many binary classifiers based on the error correcting output correcting. Extensive experiments are conducted on both toy data and three real-world HDA applications to verify the superiority of our proposed method over existing state-of-the-art HDA methods in terms of prediction accuracy.
Joey Tianyi Zhou, Ivor W. Tsang, Sinno Jialin Pan, Mingkui Tan
AISTATS1
2014 Feature Disentangling Machine - A Novel Approach of Feature Selection and Disentangling in Facial Expression Analysis
Ping Liu 0004, Joey Tianyi Zhou, Ivor W. Tsang, Zibo Meng, Shizhong Han
ECCV (4)2